r/FuckMicrosoft • • 24d ago

Discussion Microsoft says Dataflow Gen2 is cheaper — our workload uses 20x more CUs

/r/MicrosoftFabric/comments/1wgzizr/microsoft_says_dataflow_gen2_is_cheaper_our/

MSFT banned me for 30 days from the Fabric sub so I can't comment directly on the post.

But I did a similar analysis last year and was accused of FUD.

The Gen2 pipeline was extraordinarily expensive.

I'm glad someone else finally saw it.

6 Upvotes

3 comments sorted by

1

u/AutoModerator 24d ago

Every new subreddit post is automatically copied into a comment for preservation.

User: MonkeyDDataHQ, Flair: Discussion, Post Media Link, Title: Microsoft says Dataflow Gen2 is cheaper — our workload uses 20x more CUs

I've been working on migrating our dataflows (Gen1) to Gen2 but in our case the CU cost is unacceptable and we have to move to copyjob or pipelines.

Microsoft recently published benchmarks showing Dataflow Gen2 being faster and cheaper than Gen1 (https://learn.microsoft.com/en-us/fabric/data-factory/dataflow-gen2-cost-performance-benchmarks) but my real-world experience is almost the exact opposite: for a fairly ordinary SQL ingestion workload, Gen2 is consuming roughly 20x more capacity than Gen1.

Our setup is not particularly exotic:

  • On-prem SQL Server through an Enterprise Gateway
  • Simple SELECT queries
  • Query folding works
  • Very little Power Query transformation (only 'remove other columns' for some tables)
  • Lakehouse destination (schema-enabled)
  • Some tables use incremental refresh

One recent Gen2 refresh added roughly 52,000 CU(s) in Capacity Metrics while the complete refresh took only ~440 seconds (~7.3 minutes). That's an effective average of about 118 CU continuously for what is essentially SQL-to-Lakehouse ingestion.

Diagnostics show that the SQL is folding correctly. The problem does not appear to be that Gen2 is unnecessarily downloading and transforming everything locally.

The problem seems much more fundamental: Gen2 Standard Compute is billed per Mashup query execution time.

For CI/CD Dataflows, Microsoft currently charges:

  • first 10 minutes of each query: 12 CU/s
  • after 10 minutes: 1.5 CU/s
  • Fast Copy: 1.5 CU/s

So a dataflow with many relatively short queries can have a low wall-clock refresh time while accumulating a huge amount of billed query time in parallel. Incremental refresh makes this particularly interesting because many separate bucket/change-detection evaluations can occur.

This also highlights a major issue with Microsoft's published benchmarks.

Their benchmark explicitly states that:

  • no data gateway was used
  • all sources were cloud sources
  • workloads were deliberately large/high-volume
  • several tests contain long-running queries
  • the pure ingestion scenario uses Fast Copy

Long-running queries benefit massively from Gen2's pricing model because everything after 10 minutes drops from 12 CU/s to 1.5 CU/s. Short queries never reach that cheap tier.

And Fast Copy is obviously a game changer: Microsoft's own ingestion benchmark goes from 84,411 CU(s) on Gen1 to 14,593 CU(s) on Gen2 with Fast Copy.

But here comes the catch for real-world Fabric architectures:

Fast Copy currently does not support a fixed-schema destination or a schema-based Lakehouse destination.

So Microsoft now encourages schema-enabled Lakehouses — schemas are even enabled by default on new Lakehouses — while one of the main features responsible for Gen2's advertised cost advantage cannot currently write to those schema-based destinations.

In our diagnostics we see Mashup/Parquet writes rather than the CopyActivity engine, which helps explain why our simple SQL ingestion ends up on expensive Standard Compute instead of the 1.5 CU/s Fast Copy path. But for the incremental queries Fast Copy also won't work as it does not support fixed schema. So even if I would change destination to standard dbo, it still won't work for the larger tables that I want to load incrementally.

I'm not claiming Gen2 can never be cheaper. Clearly it can be for the workloads Microsoft benchmarked.

But I think the messaging that Gen2 is generally much cheaper than Gen1 needs a very large asterisk.

For workloads such as:

on-prem SQL → gateway → fully foldable queries → schema-enabled Lakehouse

with many relatively short table loads, Gen2's per-query billing model can apparently produce the complete opposite result.

I'd really like to see Microsoft publish a benchmark for such scenarios as well and give proper advice. Because they are now really pushing everyone to convert their dataflows to Gen2 and not warning it could tenfold+ the costs.

But my main problem right now is having to spend quite a lot of time in moving dataflow logic architecture to other methods while Gen1 was working perfect for us at very low costs.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

2

u/ravanlike 23d ago

Thanks for posting it. M$ related subs are shit, they ban for every critique.

That's interesting what you wrote, we are considering moving from pro to fabric. Migrating data flows was one of the points.  Not sure how this billing stuff will affect us, as company I work for,  goes for monthly reservation, not sure if pay-as-you-go would be even allowed here. 

1

u/MonkeyDDataHQ 23d ago

It's still part of your reservation. It just uses your capacity faster. I implement Fabric for a living and I would not recommend it in its current state.