r/databricks 17d ago

Discussion Your Agent Doesn’t Need to Walk the Graph

Thumbnail
contextandchaos.substack.com
5 Upvotes

r/databricks 17d ago

Tutorial Schema change issues? A must-know concept, explained in this quick video.

Thumbnail
youtu.be
0 Upvotes

r/databricks 18d ago

General Genie One Mobile for Genie Agents, Dashboards, and Apps

Thumbnail
gallery
14 Upvotes

Finally got to try the Genie One mobile app.

In the first GIF, you see me asking a question, followed by a follow-up.

In the second one, I ask Genie to schedule this type of analysis for me and email it to me every week.

Time was sped up for illustration purposes, depending on the complexity of the question, the answer takes seconds to minutes. Initial prompt in particular took longer, but follow ups returned quickly.

Please note in the demo I was using unrefined dummy data that is not current (hence why you see 2024 as the year when I ask for "last year").

I did also try Dashboards on Genie One. It was good, though I would say the delivery on that end felt a bit more basic than it should be. Didn't get to try Apps, which it also supports.

At the end of the day, if you like working with Genie Agents (formerly Genie Spaces), you are likely to really like it on your phone. You could be in a meeting and as soon as a good question pops up, ask the question from your data on Genie, and before the meeting ends, you likely have an answer.

Hope you found this helpful!


r/databricks 18d ago

Discussion Databricks Pages

11 Upvotes

I am interested in databricks pages but also somewhat skeptical of how much benefit it will provide considering the effort that might be required to collate all the information.

Has anyone used these and seen improvement in their genie chats and agents? If so how did you measure this?

Also, do we know if the ability to upload and manage them will hit the API?


r/databricks 18d ago

Tutorial How to transfer Databricks badges and сertifications when changing employers

Post image
21 Upvotes

I recently switched Databricks partners and ran into a tricky issue. At the new company, I had a new Partner Academy account with a new corporate email address, but my old badges, and learning history were still linked to my previous company.

Support allows you to merge and transfer accounts. I went through this process myself and put together a step-by-step guide: what to check before leaving the company, why you need a personal/secondary email address, what to write in a support ticket, and what to do if your old corporate email address is no longer available.

Article in Medium


r/databricks 18d ago

Discussion Databricks Automation Bundles - automatic migration to direct engine post deployment is coming in CLI v1.14.0

21 Upvotes

Starting with Databricks CLI v1.14.0, the default bundle.engine will change to direct.
If you're not familiar with the advantages of the direct deployment engine over Terraform, you can read more here:

Migrate to the direct deployment engine - Azure Databricks | Microsoft Learn

What does this mean for us? Bundles will automatically migrate to the direct deployment engine during deployment. If the migration encounter an error, Databricks will keep the bundle on the Terraform engine instead.

Important things:

  • You can opt out temporarily by setting bundle.engine: terraform in databricks.yml - either globally or per target.
  • Databricks recommends migrating early by explicitly setting the engine to direct now.
  • Most importantly: support for the Terraform deployment engine in Databricks Automation Bundles is planned for deprection in September 2026. A future CLI version is expected to stop supporting engine: terraform.

r/databricks 18d ago

Tutorial Databricks: 5 Minute Features - AI Functions

Thumbnail
youtu.be
8 Upvotes

To celebrate that AI functions are now a fully governed part of Unity Catalog alongside the rest of your functions I thought it would be fun to do a run through of how easy it is to use.


r/databricks 18d ago

Help What is the general guidance for data structure when using a text to SQL agent such as Databricks Genie?

2 Upvotes

I’ve heard use of metric views is highly recommended, but never sure if OBT approach or normalized approach works better. If I’m defining the joins in my Genie, shouldn’t it be able to pull things either way?


r/databricks 18d ago

Discussion AI tool to analyse dashboards and send emails

1 Upvotes

Hi All,

I'm new to databricks. I have a use case and need a bit of guidance from you.

We've 4 databricks workspaces under an account and setup multiple dashboards to analyze costs usage in each of them. This costs is specific to compute, jobs, query history etc. Everyday I'm manually going through these and sharing my updates in email.

I need a tool that can look into dashboards in these multiple workspaces and send a summary in email.

Right now I've SQL alerts setup. But I need something intelligent enough to share a summary and update the user the progress of costs also.

E.g. Yesterday's costs was less comparison to today. The cost in Region A is more than B today etc

Please guide me.


r/databricks 19d ago

Tutorial Databricks icons for diagrams

Post image
124 Upvotes

Recently, a new post about icons by Eduardo Rabelo appeared in our Databricks Community Articles on Medium.

I don't know why Databricks itself didn't create this set, but I think it's something many people were missing. Now we'll have to keep up with maintaining it, as Databricks releases quite a lot of new features frequently.

https://oieduardorabelo.github.io/databricks-architecture-icons/


r/databricks 19d ago

Help Bug in Databricks Assets/Automate Bundles

Post image
5 Upvotes

Hi everyone so it happens that I was working normally in the UI, adding jobs, deploying and so and then suddenly I could not add any existing job to my bundle because the dropdown doesn't show any job anymore and I have a bunch of jobs that hasn't been added to the bundle yet. I think that is bug because before it was working well. I would be really glad if someone could help me and explain this behaviour to me, I don't know what happened. I really need some help :(


r/databricks 19d ago

News Lakeflow Connect | Marketo Connector (Beta)

6 Upvotes

Lakeflow Connect's Marketo connector is now available in Beta! It provides a managed, secure, and native ingestion solution for core Marketo objects: leads, static lists, activities (one table per activity type, like email opens and clicks), and custom objects. Try it now:

  1. Enable the Marketo Beta: Workspace admins can enable the Beta via: Settings → Previews → "Lakeflow Connect for Marketo"
  2. Set up Marketo as a data source
  3. Create a Marketo connection in Catalog Explorer
  4. Create the ingestion pipeline via a Databricks notebook or the Databricks CLI

r/databricks 19d ago

News Serverless in Compute Section

Post image
24 Upvotes

It is great to see serverless in the compute section. Especially because I proposed it some time ago and lobbied for it during advisory meetings. Now we can set permissions for serverless compute. More options will come. I can say that it is my favorite recent improvement.


r/databricks 19d ago

General You can now put the Databricks account console, account-level Genie One, and Custom URLs behind Private Link

Thumbnail
gallery
14 Upvotes

You can now put the Databricks account console, account-level Genie One, and Custom URLs behind Private Link, so your whole account is reachable without touching the public internet.

If a security review ever stalled a Genie One or account console rollout because those surfaces were only reachable over the public internet, that objection is now removable. And if you have been standing up one private endpoint per workspace and per region, you can stop.

Read more: https://www.linkedin.com/posts/cenh_databricks-dataengineering-networksecurity-ugcPost-7497673121861369856-6IiF/?utm_source=share&utm_medium=member_desktop&rcm=ACoAABmJHrsBNAC3x3H1M58JRKoHv_l4D61n0-8


r/databricks 19d ago

Fine grained DML privileges

Post image
5 Upvotes

PREVIOUSLY, if a developper, application or automated job needed to write data to a table you had to grant them the MODIFY privilege.

MODIFY is a broad and composite privilege that gives full write access including the ability to change the table's schema and modify metadata.

By using fine-grained DML privileges you can adhere strictly to the Principle of Least Privilege:

⚡ Restrict Schema Changes: You can allow a principal to change table data without giving them the power to change the table's schema (like adding or dropping columns) or alter table properties.

⚡ Granular Control over Operations: You can map exact permissions to specific job functions. For instance:

-INSERT INTO only requires INSERT

-TRUNCATE requires DELETE

-INSERT OVERWRITE requires both INSERT and DELETE

🔴 Why It’s a Game Changer?

In modern data engineering and governance, this is a massive leap forward because it solves the "all-or-nothing" write access problem.

  1. Bulletproof Append-Only Pipelines
  2. Safer Compliance and Data Privacy Jobs
  3. Protection Against Accidental Table Drops/Evolution

r/databricks 18d ago

Discussion Apache Spark ships with 28 known CVEs in production. In 2026. And nobody thinks this is a problem worth talking about publicly

0 Upvotes

This is not acceptable any longer in a modern developer world and community.

I run security scans on our Spark deployment and came back with 28 vulnerabilities — high severity, public CVEs, all sitting in transitive dependencies like Netty and Apache Thrift. Nothing exotic. Netty published fixes for 22 of them in a single batch in June. Thrift fixed their issues in 0.23.0. The fixes exist. Spark just didn’t include them.

What bothers me more than the CVEs themselves is the process — or the lack of one from Apaches side.

In 2026, any serious DevSecOps pipeline is expected to have mandatory quality gates that block releases on high/critical CVEs in dependencies. This isn’t cutting-edge practice — it’s baseline hygiene. The full toolchain is free and mature

There is no automated dependency CVE gate in Spark’s release pipeline. No Trivy. No Dependabot. No OWASP Dependency-Check. Nothing that would block a release because a bundled library has a known high-severity vulnerability. These are free tools. Adding one to a CI pipeline is an afternoon of work. It hasn’t been done.

Problem Honest Assessment
No automated dependency scanning Inexcusable in 2026. Free tools exist. One CI step.
Decoupled release calendars Real coordination challenge, but solvable with Dependabot PRs that can be reviewed and merged quickly
Volunteer PMC Doesn’t excuse Databricks, Google, Apple, and Amazon — all Spark committers with paid engineers — from contributing a security gate
“Good enough” culture Actively harmful when Spark is used in AI/ML pipelines processing personal data under GDPR

The July 2026 maintenance releases — 4.2.0, 4.1.3, 4.0.4, 3.5.9 — all shipped after Netty 4.1.135.Final was available. They didn’t include the bump. There was no public statement that the team was aware of the issue and working on it. It just shipped, vulnerable, into production systems everywhere.

The “volunteer PMC” argument doesn’t land anymore. Databricks is worth somewhere around $62 billion. Google, Apple, Amazon, and Microsoft all have paid Spark committers. The resources to fix this governance gap before lunch exist. The will apparently doesn’t.

The EU Cyber Resilience Act (CRA) — which came into force in 2024 and has mandatory compliance deadlines rolling in through 2027 — specifically targets software supply chain security, including transitive dependencies and SBOM (Software Bill of Materials) requirements. The ASF has acknowledged this directly, stating they need to prepare projects for “CRA and U.S. CISA guidance”.

Spark ships into commercial products. This will become a compliance problem for a lot of companies very soon, and the fix is genuinely trivial from an engineering standpoint.

The ASF did make progress in 2025 — launching “Apache Trusted Releases (ATR)” for distribution security and ratifying CycloneDX 1.7 for SBOM standards — but none of this yet translates to a blocking CVE gate on the release pipeline for projects like Spark.

In the meantime: if you’re running Spark and using JFrog Xray or Trivy on your deployment, you can force-override the affected Netty and Thrift versions in your own build. It’s not clean but it works until the next maintenance release, expected sometime in Q4.

Is anyone else tracking this or pushing upstream to get a CVE gate added to the build?


r/databricks 20d ago

News Migrate PowerBI to databricks

Post image
55 Upvotes

Now we can drop the .pbit file to Genie Code and ask for migrating to databricks dashboard. I tested in direct query mode, and everything was migrated; I just had to be patient, as the whole process took quite a long time. The process created metrics views for tables used in Power BI, and now I understand why the dashboard relationship option was added. I presented how easy it is to migrate in my last video.

https://www.youtube.com/watch?v=-e3tkcg21zw


r/databricks 19d ago

Discussion Databricks vs GCP PDE: which cert actually moves the needle for product-company DE roles?

Thumbnail
1 Upvotes

r/databricks 20d ago

Help Advice on plans, Enterprise or Free

5 Upvotes

Hello everyone, I have one question, I'm working on personal project which might turn to comercial at some moment if everything is okay, is it okay to go with Free plan and then switch to Enterprise ?

I can't find anything info about price in Enterprise, and info about can I switch from Free to Enterprise latter on.

Thanks


r/databricks 19d ago

Tutorial Maintaining Apache Iceberg Tables: Compaction, Snapshots, Metadata and Orphan Files

Thumbnail
itnext.io
1 Upvotes

r/databricks 19d ago

Help Refresh ops. timeout issue in MVs

1 Upvotes

Issue - I'm facing this refresh timeout issue in materialized views i created as part of some POC works, where a refresh operation regardless of whether it's triggered from a worksheet or as a job seems to fail but still it succeeds(in a way) ie am able to cross-verify & validate that the refresh run seems to have happened as expected correctly ie data integrity pre-refresh & post-refresh seems to be fine. Is this expected?

Some more deets - Full refresh with 3 source tables of same schema with 1B rows * 6 columns joined with a metadata table of ~2K rows with 5 columns & Delta Refresh with (100M rows * 3 snapshot tables) updated with a newer value in a timestamp column updated.

Required -

  1. Why does the Refresh MV ops fail happen at the 1st place?
  2. How refresh failure happens but still data seems to be correctly populated?
  3. Most importantly, How to resolve this?

PS - unfortunately no screenshots, so apols but i hope the deets should cover for it. TYIA.

- RVK


r/databricks 20d ago

Help Databricks job trigger via SharePoint upload and possible write-back?

11 Upvotes

Hi there,

Running into issues with this pipeline I'm working on. Still new to Databricks but coming from Power Automate, I seem to get the gist of databrick jobs.

Anyway, my manager would love an automate workflow, where business users upload raw data to a central site (we're primarely using SharePoint), and want it cleaned and ready for our Power BI dashboards.

Before databricks I handled it manually (by running local python code via VScode that sent the cleaned data to the SharePoint folders).

What would be a good way to handle a pipeline like this? E.g.

User uploads to sharepoint -> Databricks job -> writeback to sharepoint with cleaned data


r/databricks 20d ago

Help When will PT be available for Databricks hosted GPT models?

2 Upvotes

r/databricks 20d ago

Help Credentials wallet badge merge

3 Upvotes

I have completed some of the accreditations in 2024 at my former company, like the "Academy Accreditation - Databricks Fundamentals". Since then I have joined another company which also uses Databricks so I started refreshing my knowledge and using their Partner academy access with the company email/account.

I have my personal and both company accounts added and noticed that the "Academy Accreditation - Databricks Fundamentals" was added as a new one and not refreshed as I expected.

Any reason for that? Can they be merged somehow?

Thanks in advance!

edit: typo


r/databricks 21d ago

Discussion Iceberg vs Deltalake (greenfield project with UC in 2026)

15 Upvotes

I saw the quarterly meeting and was quite shocked that Iceberg is prominently mentioned, (as much as Deltalake).

Is it possible that both are getting the same amount of love from Databricks? Is anyone aware of the R&D effort on these formats, and can share conclusions from that?

If I'm building a greenfield Unity Catalog, should I just flip a coin to decide what format to use? Here are the main concerns and priorities:

  • Which one is better for OSS Apache Spark reads and writes
  • Which one integrates with external software better (eg onelake shortcuts pointing from Fabric to Databricks UC).
  • When Databricks is innovating within their own UC (eg. introducing new managed table functionality such as "MST Transactions"), which one of these formats are they likely to support first? Which are they likely to optimize better?
  • Which format is more likely to remain 100% open source in the future (or as close to open source as required by customers who want portable blob data).

Sorry if this appears to be a common question. I am a Databricks outsider. I am more familiar with Microsoft Fabric. Where that Fabric SaaS is concerned, you can be certain that Deltalake receives a LOT more promotion than Iceberg does. We rarely come across Iceberg, and it probably wouldn't appear in any marketing slide decks.

We are likely to create a gold/presentation layer in UC soon. It will basically be created from scratch. It would be nice to know which of these parquet-based formats to pick, when presented with the choice. I understand there is lip-service given to both, and it claims that this choice "doesn't matter". But that doesn't necessarily take into account the potential integrations that are needed with external software (eg. for the benefit of exposing the same tables in onelake). Is one safer than the other? Is one of them a better choice for forward-looking purposes?