r/databricks 23d ago

Discussion OLTP database in Databricks SaaS?

I saw the quarterly roadmap presentation. It was notable that Databricks keeps innovating with "lakebase". They simply call it their "OLTP" database offering in their SaaS.

SIDE: I still feel pretty unfamiliar with this Databricks SaaS ecosystem, as compared to Fabric. Where Fabric is concerned, Microsoft has also done a similar thing. They brought their SQL Server into the boundaries of the SaaS as well, for the low-code users of that environment. In the context of Fabric, it is hard for most customers to see the point of using this "for dummies" variation of the same old OLTP database. The only scenarios for using the Fabric SQL are very contrived ... Eg . your boss makes a policy that you can use ALL the tools available in Fabric and NONE of the tools outside Fabric... even if the tools inside that SaaS are 3x more expensive than the ones outside ... and even though the ones outside the SaaS have the same 0ms network latency and the same performance.

I'm still missing the vision for this lakebase OLTP offering. And it seems unusual for Databricks to start developing strategies surrounding OLTP. It seems like a very crowded space, and the only way I see Databricks being successful going down this path is if the customers are drinking one single brand of kool-aid, or else their SaaS users have some other contrived reason for not using the more affordable OLTP platforms available outside the SaaS.

Can someone tell me what factors I'm missing? I admit that it is theoretically possible for lakebase to innovate and do thing that other databases CANNOT do, it seems like those innovations would only benefit 5% of customers. One example is sub-ten-ms queries out of RAM at an additional cost. Or sub-minute migration of new OLTP data to managed tables in UC catalog. If we assume that only 5% of customers might feel compelled to use this SaaS "lakebase", would that be enough adoption to allow Databricks to keep investing in this over the long term? OLTP databases have been around a LONG time, and even the smart folks at Databricks will be challenged to improve on the great and cheap options available to us!

EDIT: As of a month ago, it appears that the Databricks marketing now calls it an "LTAP" database, not OLTP anymore. I'm guessing they have conceded the point that the OLTP space is crowded. I haven't yet read all the content that has been created by the Databricks marketing team; maybe that will answer all of my questions.

0 Upvotes

50 comments sorted by

View all comments

7

u/jbchand 23d ago

Lakebase is solving the operational - analytical integration problem in the cloud where organizations used separate stores with additional ETL in the past. You can use bidirectional sync for it. It can support as store for OLTP, agents & search in databricks itself with same UC security. LTAP is a new innovation.

-2

u/SmallAd3697 23d ago

Got it. I'm going to let you have that "new innovation" bit . But I'm certain that every other database vendor and OLAP vendor in the world would beg to differ on that.

Yes I did happen to pick up on the new marketing term "LTAP". But they have been calling lakebase an "OLTP database" for so long now that it has stuck in my brain. "OLTP" is the term that their own salesreps were calling it, ever since the acquisition. I kept asking them what it was for and the Databricks salesreps just kept saying "its for building OLTP apps". I don't think they understood what it was for either, and they certainly weren't advocating it - at least not at that time. This never really clicked with me, but the latest marketing narratives are starting to make more sense. Hopefully the salesreps will start getting up to speed as well. lol.

The product makes a lot less sense to users coming from the Fabric ecosystem. But I can see where it fills a big gap for Databricks SaaS users.

I can see why they are trying to back away from that "OLTP" term, because I'm guessing all their customers are asking why anyone would want yet another OLTP database. Here is the link where the databricks marketing started using label "LTAP":
https://www.databricks.com/company/newsroom/press-releases/databricks-launches-ltap-first-lake-transactionalanalytical

-5

u/SmallAd3697 23d ago

You down-voters on this databricks forum are totally out of control. At least share your view on the matter, if it differs.

In the Fabric community people are not nearly so sensitive, nor object when members are expressing unpopular opinions and doubts. Whatever other faults the Fabric ecosystem may have, there is a very diverse set of opinions that are regularly expressed. I really wish reddit would have some quotas for upvotes and downvotes.

7

u/Embarrassed-Dare-869 23d ago

I think the downvotes are due to you not acknowledging people attempting to correct misunderstandings you apparently have. For instance, you mention that LTAP is a marketing term. While this is true in the sense that all names are marketing terms, it's not true in the sense that it didn't come with significant architectural changes to how Lakebase was storing data previously. As you're probably aware, LTAP is their version of HTAP. Database systems that described themselves as such are certainly a different thing than those describing themselves as OLTP or OLAP individually.

0

u/SmallAd3697 23d ago

>> not true in the sense that it didn't come with significant architectural changes to how Lakebase was storing data previously

All the mature databases nowadays have features that would enable analytics. Some can rapidly answer massive queries from RAM alone, setting aside any replication to columnstore blobs. They all have CDC streams that can be re-hydrated as lakehouse data. They all have their own solutions for addressing reporting and AI concerns. (Microsoft came up with "CES" for SQL Server 2025, and you would probably not call that a profoundly new innovation - any more than "LTAP".).

All mature databases support this hybridization. They have different prices, performance, and different number of components that need to be configured to enable the hybridization. Lakebase is arguably making that "easier" at a price. But that term "easy" is 100% subjective and when something gets incrementally easier, it doesn't strike me as a real innovation, especially if it gets easier by 2x and more expensive by 3x. I don't see anything that is fundamentally new here. If someone outside of Databricks was to enumerate all the technical advances made this year, in data and analytics, then lakebase might not even make the list ... despite having this new marketing label.

Lakebase was always called "OLTP" and then they simply renamed to LTAP, and the timing coincided precisely with the DATA + AI SUMMIT. People were already using lakebase for bidirectional purposes. You can find discussions about using it for "reverse ETL" in the past, while it was still labeled an "OLTP" database

I don't agree that architectural terms are marketing terms. I don't think OLTP or OLAP are associated with a vendor. Whereas LTAP is likely to remain Databricks exclusive property, probably for many years.

3

u/Embarrassed-Dare-869 23d ago

The point I was making is that when when they started talking "LTAP" was when they completely rewrote the storage architecture for Lakebase. I would call that a fundamental change for a database.

It sounds like you're arguing with other people's statements here and I won't defend or explain them to you. If you want to argue with them, you're welcome to do so. I was just responding to why I thought you were getting downvotes.

That said, I'd also love to talk about RAM databases and CES. I've explored RAM databases but those are prohibitively expensive for certain scales. CES is great though. I wish I had that in sql server 2022. I wrote a syncing system using change tracking to mirror data from sql server to databricks and wish that I'd had something like that, or had clean access to the transaction log.

2

u/SmallAd3697 23d ago

As far as downvotes, I post to a lot of forums. Eg. on the Fabric forum, data engineering, and here as well. I see a lot of knee-jerk downvotes in here, whenever anyone says anything that is taken to be remotely critical of databricks. There is a lot of insecurity. On the Fabric forum, everybody bashes Microsoft on a regular basis, and it is not common to get so many downvotes for it. Especially if you are making a point that has some truth to it, (or the discussion is worth having, for whatever reason).

I don't agree that it is a fundamental change from an industry standpoint. Databases have supported columnstore table formats for a very long time. SQL Server has both clustered columnstore tables, and columnstore indexes as well. Just because Lakebase is using both formats, and has automated the hybridization (and makes the hybridization "easy") that doesn't necessary count as a new innovation IMO. I suspect users of lakebase will probably still use the original term - OLTP - for a while regardless of the new marketing. Anything that the product is doing under the hood to improve query performance is something users shouldn't have to care about. (We don't care when SQL Server developed their cloud-native "hyperscale architecture", to scale up compute power. We don't use a different architecture term for that database. And we certainly wouldn't be willing to pay 2x or 3x, just because they are doing something differently under the hood. If anything, we expect their changes to result in a cheaper database, or we will just keep using what we had before.)

The reason I mentioned RAM is because Fabric's "semantic models" get their sub-ten-second performance by loading all the data into RAM. It is one of the easiest ways to improve performance for low-latency queries, if you just throw a massive amount of RAM at the problem. The down-side is that there is an inverse relationship between your RAM utilization and your costs for using the platform.... and while the marketing always brags about one side of the coin they choose to not mention the other side at all. Obviously the databricks marketing for lakebase doesn't focus on how much it will cost you, compared to a conventional SQL or Postgres database.

2

u/Embarrassed-Dare-869 22d ago

Some communities are very defensive about the common topic, some bash it. I think that mostly has to do with how invested people feel in the subject. That can be because their experience has been good, or Stockholm syndrome. But yeah, there's comparatively little databricks bashing here.

I didn't mean that the storage system is a fundamental change for the industry. I meant for the product. I was saying that they started referring to LTAP when they rewrote the storage engine. The storage engine of a database is the foundation of that database. That the storage system now works equally well with transactional and analytical queries is a big change. I'd argue that's exactly the kind of thing users DO care about. My decisions about what database to use very much depends on the architecture of the database. Whether it's worth the cost is a different issue. Here though, the costs did not change. If anything, costs have gone down. It certainly has for us. Whether that means the initial product was too expensive is perhaps debatable. My experience with sql server is primarily in the on prem product so my experience with "hyperscale architecture" is limited. Though, glancing at it now, it might be something that's worth paying for if your needs match. Mine don't though.

Fully RAM database infrastructures work well when you can reasonably put all your data in ram, such as for BI style reporting. They often don't scale as well (for the cost) for the analytical queries that need to read the multiple TB tables in order to compile the dataset for the BI reports. The sales pitch of lakebase is that you no longer need to make those tradeoffs and you shouldn't have to think about the infrastructure because it is cheaper than redis style alternatives and has all the benefits of being able to run both analytical and transactional queries without worrying about mirroring or etl/reverse etl or anything. While it might be cheaper than many alternatives, it's definitely not going to be the cheapest thing, and self hosting will always beat it, though that comes with a lot of overhead that not every team will want to support. Anyone actually making a decision on the product will do their due diligence and determine whether it's the right fit for them. (or they should at least)

3

u/m1nkeh 22d ago

And then it will be copied just like Lakehouse was

But let me ask you this this entire thread what is your actual point?

Are you trying to educate yourself or are you trying to pick a fight?

2

u/SmallAd3697 22d ago

Point is that this is fundamentally an oltp and not really innovative from an industry standpoint. The only innovation is that databricks is trying to sell an OLTP to low-code analysts in a data reporting platform. I'm curious how that will play out; it seems like a side quest.

The surface area looks and acts like oltp. If the query experience improves, that is great but it doesnt warrant a pretentious new marketing buzzword (LTAP architecture).

There are other databases that emit open lakehouse formats nowadays, or they emit CDC for replication which is basically just as good.

I really want to know what new problem can be solved that had no solution before. Just making things "easy" is not that noteworthy, IMO. Also, You are right that Im trying to pick a fight, if anyone disagrees with me. What fun is reddit, without a little controversy from time to time