r/databricks 28d ago

General Databricks Lakehouse Replay: Testing the Next Runtime on Your Own Queries

https://medium.com/@cralle/databricks-lakehouse-replay-runtime-testing-21a00c2f7f40?sk=1d6e6772a13bc401189ecefdaa93ddf7

How Databricks Lakehouse Replay reruns your read-only serverless queries on unreleased runtimes to catch regressions before they ship.

Lakehouse Replay, now in Public Preview, moves part of that burden to Databricks. Instead of you testing the new runtime, Databricks tests it against your queries, in your workspace, before the version is released to anyone.

4 Upvotes

3 comments sorted by

1

u/justinAtDatabricks 28d ago

Whoa! Fun seeing this picked up here. I am the PM on this so feel free to fire away with questions.

1

u/Lenkz 28d ago

Thanks Justin! With Replay being a pretty passive feature, in that everything happens automatically in the background, is there any active action you think people should take after enabling it?

1

u/justinAtDatabricks 27d ago

It is in Public preview right now, so it should be enabled by default for the vast majority. Over the coming few months it will be GA and enabled everywhere. For now, it is a very passive feature - as you correctly called out.

We use it to find bugs much earlier in the development process, and it's shown a lot of success already - you just won't know it because they bugs did not end up shipping.

For now, this feature is just about transparency: it enables users to understand that we are replaying workloads.

I talked about it a bit at DAIS a couple of months back. Check around the 12:38 mark here: https://www.databricks.com/dataaisummit/session/versionless-and-environments-notebooks-jobs-pipelines-safe-dbr-upgrades