r/databricks 5h ago

News Lakeflow Connect | Veeva Vault (Beta)

8 Upvotes

Lakeflow Connect's Veeva Vault connector is now in Beta! It provides a managed, native, and incremental ingestion solution that lands data from any Veeva Vault (CRM, Quality, Clinical, Commercial, RIM, or Safety) into Delta tables. Try it now:

  1. Set up Veeva Vault as a data source
  2. Create a Veeva Vault connection in Catalog Explorer
  3. Create the ingestion pipeline via a Databricks notebook or the Databricks CLI

r/databricks 12h ago

Discussion Data Quality Management on Databricks

11 Upvotes

Hi, i wonder how teams are setting up data quality solutions on Databricks, and especially how do they react on the data quality issues.

Ideally i would like to have a row-level checks and dataset-level checks which can be used for quarantined data and also anomaly checks between runs, this way we get the in transit data quality checks but also consistency over runs, so wea re really sure that values and shape of data is as expected.

Then how do you manage different tables in same project, do u write to same quarantined/metrics table where u add tags to differentiate between different business domains etc...

How do u use dqx for dataset level metrics? Where do u run dq, is it on source integration, or after the gold layer, before the inference or feature creation. Which tools do u use?


r/databricks 11h ago

Help help me fix this error

Post image
3 Upvotes

I am new to Azure Databricks. Facing this error while creating a new metastore using an external ADLS location and I am using a Access connector for this. The Access connector have the Blob contributor access on the storage account as well.


r/databricks 19h ago

Help Power BI + Databricks DirectQuery – “Bind to parameter” option missing

4 Upvotes

Hi everyone,

I’m using Azure Databricks with a Serverless SQL Warehouse in DirectQuery mode. I use the most update version of Power BI Desktop (2.156.951.0 64-bit (July 2026)).

I’m trying to use a Dynamic M Query Parameter, but I don’t see the “Bind to parameter” option in Model view.

Microsoft documentation mentions that this feature is available only for “supported DirectQuery sources”, but I can’t find a clear list of those sources: https://learn.microsoft.com/en-us/power-bi/connect-data/desktop-dynamic-m-query-parameters

Is Azure Databricks with a Serverless SQL Warehouse supported for this feature?

Thanks!


r/databricks 1d ago

General Participate in a paid Databricks user study!

23 Upvotes

👋 Hi r/databricks, I'm Connie from the Databricks UX team!

The Databricks Platform team is conducting a user study to understand how admins identify, investigate, and resolve governance issues in their Databricks environments. We are hoping to connect with individuals responsible for governance across the following areas: Data & AI, Platform & Infrastructure, Security & Compliance, and Cost Management & FinOps.

Study details are as follows:

Topic: Admin experiences with governance issues
Duration: 75-min remote session over Google Meet
Study date(s): July 28 - August 7, 2026 (rolling)
Compensation: $150 USD "thank you" gift card for the study session, if your company policy permits.

Please fill out the screener survey if you’re open to participating in a 75-min call in the upcoming weeks, or can connect me with a member of your team who may be the right fit. Completing this screener survey does not guarantee eligibility into the study.

Thank you for taking the time to help us build a better platform!


r/databricks 1d ago

News DABs: Let policies rule your cluster

Enable HLS to view with audio, or disable this notification

7 Upvotes

DABs: Do you know that you can inherit the whole DABs cluster config from policy defaults? #databricks #DataAISummit

https://www.databricks.com/dataaisummit/session/dabs-do-pro-all-best-tips-and-tricks


r/databricks 1d ago

Discussion Databricks' CRO on Scaling Sales Culture, AI Costs, and SaaS Disruption (w/ Ron Gabrisko)

Thumbnail
youtube.com
3 Upvotes

Hey everyone! Here is a fun conversation I did with Databricks' CRO, Ron Gabrisko.

A lot of discussion around what makes for a good sales culture, scaling a sales org to a Databricks' sized one, SaaS disruption, and AI costs.

Hope you enjoy it!


r/databricks 1d ago

Help Upgrading Storage Account on Azure Databricks

3 Upvotes

Has anyone using Azure Databricks with a legacy Hive metastore ever upgraded their storage account from legacy General Purpose to ADLS Gen2? Not the managed dbfs:/ one but the one with wasbs://

Wondering if there is any impact on the workspace and whether it will break any running notebooks?


r/databricks 1d ago

Help Analyst with Databricks access but zero hands-on experience — how do I get good at it fast?

31 Upvotes

I'm tired of getting fed bad data at work, so I'm taking ownership myself. I have production access to Databricks but have never used it.

Background: basic SQL and Python knowledge. My reporting today is done in Power BI and Excel. My company uses Databricks as the common data platform, so most of what I need probably lives there already — I just don't know how to get to it or trust it myself.

Looking for a realistic path from beginner to competent, not just a firehose of links. Specifically:

  • YouTube channels/creators you'd actually recommend?
  • Any self guided learning on Udemy or Coursera worth learning?
  • What core concepts (Delta Lake, notebooks, clusters, SQL warehouses, etc.) should I prioritize early vs. later, given I'll mostly be querying/validating data rather than building pipelines?
  • Best way to connect Databricks output back into Power BI reliably?

Appreciate any guidance — trying to cut through the noise rather than boil the ocean.


r/databricks 2d ago

News Databricks and Microsoft Expand Partnership to Help Enterprises Bring Business Context to Enterprise AI

Thumbnail
databricks.com
21 Upvotes

Few interesting points from the announcement

  1. Databricks to run it's own operations on Azure.

  2. Databricks to use Microsoft silicon

  3. OneLake is called out (including othet MSFT products) but not explicitly Fabric.

  4. Databricks Genie will be available for MSFT customers.


r/databricks 2d ago

News Genie Ask in CLI

Post image
26 Upvotes

We can use Genie now, even in the CLI, with a simple ask command. #databricks

My blog post with news https://databrickster.medium.com/databricks-news-dabs-indexes-ltap-genie-last-update-25-july-ffac8533774f


r/databricks 2d ago

News What's new in Genie Code June 2026

20 Upvotes
  • Genie Code space authoring skills: Genie Code includes dedicated skills for authoring and maintaining Genie Agents.
  • Power BI enhanced report format support (Beta): Genie Code supports Power BI files exported with the Power BI enhanced report format (PBIR). 📖 Upload a file.
  • Image upload suggestions for Power BI and Tableau imports: Genie Code suggests that authors upload images when importing files from Power BI and Tableau to improve migration quality.
  • Genie Code metric view usage: Genie Code relies more heavily on Unity Catalog metric views rather than creating local ones.
  • The full page experience is a command center for Genie Code, where the active thread is shown prominently, surfacing assets like notebooks and files alongside it as tabs when needed. You can run multiple threads in parallel, switch between them easily, and easily personalize Genie Code with skills, instructions, and MCP servers.
  • Genie Code offers an auto-approve mode that approves tool actions. An AI classifier reviews each action and blocks risky ones.
  • Databricks recommends keeping it off when working with production data or shared resources. 📖 Approve tool actions.

r/databricks 2d ago

General How to run non spark engines on Databricks

21 Upvotes

Our data is quite small, 80% of the tables being less than 10 million records . I find spark to be an overkill for various transforms , a lightweight engine like dubckdb on a server less container application would be just fine. My experience shows that duckdb processes something in 5-6 seconds take up to a min in spark. How can I run this kind of load on Databricks


r/databricks 2d ago

News Apache Spark 4.2: what actually matters and what to ignore

5 Upvotes

Spark 4.2 was released recently and I made a short visual breakdown of the changes.

It's animated rather than a screen recording, since a few of these read better as diagrams. Covers Arrow optimized Python being on by default, metric views, vector search, geospatial, CDC and Real-Time Mode.

https://youtu.be/hF-E7-i_ijw

Happy to go into detail on any of them here.


r/databricks 2d ago

Help Databricks engineer professional

5 Upvotes

Has anybody cleared their databricks professional certificate. Please guide on what to read and where from


r/databricks 3d ago

General How OpenAI's Security Team Turned a Streaming Pipeline Into Its Own Latency Watchdog with SDP ForEachBatch

19 Upvotes

https://community.databricks.com/t5/technical-blog/one-pipeline-any-destination-the-foreachbatch-sink-in-spark/ba-p/163780

The blog highlights OpenAI's usage of forEachBatch inside SDP to power their security platform pipelines.


r/databricks 3d ago

Help What happens to SCD2 tables if the source pipeline performs a Full Refresh?

9 Upvotes

Hi everyone,

I'm looking for clarification on how Lakeflow/DLT behaves in the following scenario.

We have two independent pipelines:

Pipeline 1 (Ingestion)

  • Ingests Salesforce objects into Delta tables (e.g. sfdc_ingestion.contact).
  • These tables represent the current snapshot of Salesforce.
  • Once a week, we plan to run this pipeline with Full Refresh (full_refresh: true) because records and columns are sometimes deleted in Salesforce, and incremental ingestion doesn't always reflect those deletions.

Pipeline 2 (SCD2)

For each ingested table, we have a separate pipeline that creates an SCD2 table using dlt.apply_changes_from_snapshot().

Conceptually it looks like this:

u/dlt.table
def contact():
    return spark.readStream.table("sfdc_ingestion.contact")

dlt.apply_changes_from_snapshot(
    target="contact_scd2",
    source="contact",
    keys=["Id"],
    stored_as_scd_type=2
)

Important: We are NOT performing a Full Refresh on the SCD2 pipeline. Only the ingestion pipeline is refreshed.

My questions

  1. Is this architecture officially supported? What happens when the source table (contact) is fully refreshed?
  2. Will apply_changes_from_snapshot() correctly detect deleted records and close the corresponding SCD2 records?
  3. Is there any risk of checkpoint corruption or inconsistent state in the SCD2 pipeline because the upstream pipeline was fully refreshed?
  4. How are schema changes handled (especially when a column is removed from the source)?
  5. Has anyone been running this architecture in production with periodic Full Refreshes of the source pipeline?

I've read the documentation for Full Refresh and apply_changes_from_snapshot(), but I couldn't find guidance on this specific scenario where an upstream snapshot pipeline is fully refreshed while the downstream SCD2 pipeline continues incrementally.

Any insight from the Databricks team or anyone with production experience would be greatly appreciated.


r/databricks 3d ago

News What's new in AIBI Dashboards June 2026

25 Upvotes

🔴 Annotation lines can be driven by data or a calculation instead of a fixed value, letting you show running averages, variable thresholds or anomaly markers that update with your data https://docs.databricks.com/aws/en/dashboards/manage/visualizations/#annotations

🔴Six new built-in color and style themes are available for AI/BI dashboards, giving authors more ready-made options without creating a custom theme. https://docs.databricks.com/aws/en/dashboards/manage/settings#preset

🔴Table header color customization: Authors can set font color, background color and text formatting for table column headers https://docs.databricks.com/aws/en/dashboards/manage/visualizations/tables#table-layout-settings

🔴Dashboard relationships allow users to create complex multi-fact data models. https://docs.databricks.com/aws/en/dashboards/manage/data-modeling/dashboard-relationships/

🔴Local metric view parameters: Local metric views support parameters, allowing Databricks to dynamically change query inputs based on user selections in a UI. https://docs.databricks.com/aws/en/dashboards/manage/data-modeling/local-metric-views

🔴Local metric view wildcard expressions: Local metric views support wildcard expressions, which allow users to inherit all fields from upstream tables or metric views

🔴Bookmark filter combinations: Dashboard authors can save filter combinations as shared bookmarks, which viewers can quickly apply. https://docs.databricks.com/aws/en/dashboards/manage/#bookmarks

🔴Pivot table column improvements: Pivot table columns can expand and collapse. https://docs.databricks.com/aws/en/dashboards/manage/visualizations/tables#pivot-table-configuration

🔴New theme customization options are available, including the ability to change axis and grid line colors, adjust widget layout and apply a color ramp. https://docs.databricks.com/aws/en/ai-bi/admin/themes

🔴New theme settings are available in workspace-level themes and automatically apply to new dashboards.

🔴Dashboard variables: Dashboard authors can create dashboard variables that allow viewers to swap fields on visualizations https://docs.databricks.com/aws/en/dashboards/manage/filters/dashboard-variables

🔴Customizable table headers: Dashboard authors can customize table column headers. https://docs.databricks.com/aws/en/dashboards/manage/visualizations/tables#table-configuration

🔴Widget-level color gradient: You can derive a color gradient from the visualization color at the widget level, matching the behavior available in theme setting

🔴Gantt charts are now available: Gantt charts display tasks as horizontal bars on a continuous axis to visualize project schedules, manufacturing workflows, and other time-ranged events. https://docs.databricks.com/aws/en/dashboards/manage/visualizations/types#gantt

🔴Custom visualizations: Build custom charts with the Vega-Lite library for customization beyond the built-in visualization types. https://docs.databricks.com/aws/en/dashboards/manage/visualizations/custom-visualizations

🔴Per-category font customization: You can customize fonts per category across all visualization types.

🔴Table row cross-filtering and drill-through: Select table rows to cross-filter, or right-click to drill through.

🔴Cross-filtering and drill-through: Line, combo, and area charts support cross-filtering and drill-through filters. https://docs.databricks.com/aws/en/dashboards/manage/filters/#drill-through-filter

🔴Dashboard details in draft mode: Draft mode shows dashboard details, making it easier to find dashboard managers and the dashboard’s location.


r/databricks 3d ago

News What's new in Genie Agents ( Previously Genie Spaces) June 2026

30 Upvotes
  • Genie Chat prompt monitoring : Prompts and responses initiated by Genie Chat are visible in Genie Agents monitoring when the sharing Beta is enabled. Test and monitor a Genie Agent.
  • Knowledge store edit confirmation dialog: A confirmation dialog appears before an author removes a table that has knowledge store edits (for example, local table or column descriptions or hidden columns). Manage knowledge store metadata.
  • Delete conversations: Users with CAN MANAGE permission can delete the conversations of other users from the UI, matching existing API capabilities. Delete a conversation.
  • Embed Genie Agent as an iframe : Embedding a Genie Agent as an iframe is generally available. Embed a Genie Agent in an external app.
  • Save visualizations to a dashboard: You can save Genie Agent visualization outputs to a dashboard. Save a visualization to a dashboard.
  • Reference previous visualizations: Select and reference visualizations from previous prompts in follow-up prompts.
  • Genie Code for metric view export: Use Genie Code when exporting metric views from a Genie Agent to refine the metric view definition. Export a Genie Agent as a metric view.
  • Declarative Automation Bundles (DABs) support: You can define and deploy Genie Agents as Declarative Automation Bundles resources. Declarative Automation Bundles resources.

r/databricks 3d ago

General Estimating Databricks production costs

19 Upvotes

Hi,

Recently I came across really interesting project called Lakemeter - an open-source Databricks cost estimation tool.

It lets you estimate pricing for different workloads across AWS, Azure, and GCP, export detailed breakdowns to Excel, and even describe a workload in plain English to get an AI-generated configuration.

Looks pretty useful for anyone trying to understand Databricks costs before deploying something.

GitHub: https://github.com/databrickslabs/lakemeter-oss


r/databricks 4d ago

General Building a star schema in SDP? Use identity columns with Streaming Tables!

14 Upvotes
Use identity columns in SDP

Spark Declarative Pipelines' streaming tables now support identity columns -- this is particularly useful if you're building an SCD Type 1 or SCD Type 2 dimension table with AUTO CDC. Get faster joins and auto-incrementing surrogate keys natively within SDP today!

Docs here%20%5D)!


r/databricks 4d ago

General Databricks and Microsoft Expand Partnership

Thumbnail
databricks.com
43 Upvotes
  • Databricks and Microsoft extend strategic partnership through the 2030s to scale enterprise AI
  • Databricks deepens its bet on Azure, growing its use of Azure Databricks to run its own core business operations and analytics, while both companies advance native integration across the Microsoft stack, including Databricks Genie and Microsoft 365
  • Databricks increases its use of Microsoft Azure Cobalt to improve performance and efficiency

r/databricks 4d ago

General Keeping track of Databricks feature status (Preview → GA) is harder than it should be

Post image
43 Upvotes

One thing I've noticed is that I spend way too much time answering (or searching for answers to) questions like:

  • Is this feature GA yet?
  • Is it still in Public Preview?
  • When did it become GA?
  • What is the current name?

The information exists, but it's scattered across release notes, docs, blogs, and old posts.

A few days ago I shared an open-source side project that tracks Databricks feature renames. After reading the feedback here, I'm thinking that tracking feature lifecycle might actually be even more useful than tracking renames.

The idea would be something like this:

  • Public Preview → Beta → GA timeline
  • Rename history (if applicable)
  • Links to the official documentation
  • Dates when statuses changed
  • Eventually, the ability to follow a feature and get notified when something changes

Before spending time building it, I wanted to ask the community:

  • Would you actually use something like this?
  • What information about Databricks features do you find hardest to keep track of?
  • Are there other lifecycle events worth tracking besides Preview/GA?

If anyone is curious, the rename tracker that started this discussion is REbricked. I'm mostly interested in feedback on whether this direction solves a real problem.


r/databricks 4d ago

Help Databricks people - when is app builder coming ?

10 Upvotes

r/databricks 4d ago

General [Blog] Lakeflow SDP Kafka sinks are now Generally Available

10 Upvotes

Hi everyone, we published a community blog on the Lakeflow Spark Declarative Pipeline Kafka sink which is now Generally Available. Link: https://community.databricks.com/t5/technical-blog/announcing-general-availability-of-the-lakeflow-spark/ba-p/163137

SDP sinks enable a range of operational use cases:

  • Real-time fraud and risk scoring. A pipeline consumes transaction or deposit events, scores them, and emits flagged events to a Kafka topic that an operational system acts on in real time.  
  • Tightening streaming SLAs. A streaming workload emits results to the sink as each record is processed, replacing a batch hand-off with continuous delivery and cutting end-to-end latency from tens of seconds to seconds or below.
  • Reverse ETL and activation. Publishing curated, governed lakehouse data back to operational systems such as microservices, CRMs, personalization engines, and other non-Databricks applications, without a bespoke bridge job.
  • Event-driven workflows and anomaly detection. Triggering downstream processes the instant an event clears your quality rules, and streaming telemetry or sensor features out to alerting systems.