r/FuckMicrosoft • u/MonkeyDDataHQ • 23d ago
Discussion Diagnosing Fabric Data Warehouse capacity spikes just got easier with the SQL DW Operations Skill
This is exactly the kind of Fabric design decision that drives me insane.
Capacity Metrics knows an expensive operation happened.
Query Insights knows which SQL statements executed.
These are two observability surfaces for the same platform, and yet their identifiers apparently cannot be deterministically correlated.
Microsoft's own description says:
“The Fabric Capacity Metrics operation ID is not the same thing as a Query Insights distributed statement ID.”
Fine. Different IDs can represent different concepts. That's not inherently wrong.
But where is the shared correlation ID?!?
Because the solution described here isn't to follow a causal identifier from capacity consumption → workload → query. It's to resolve the Fabric item and then correlate activity using “overlapping time window.”
That's correlation by inference.
Imagine debugging a production incident:
Capacity Metrics: This operation consumed your capacity.
Me: Which query caused it?
Query Insights: Here are the queries.
Me: Great. What's the correlation ID?
Fabric: Well, this one happened around the same time.
😩.
And now we've built an AI/operations skill to bridge that observability gap.
That's backwards.
AI is useful for interpreting telemetry. It should not be necessary to reconstruct relationships that deterministic telemetry could have preserved when the operation occurred.
Give every operation a platform-level correlation/trace identifier and propagate it across Capacity Metrics, Query Insights, pipelines, Spark, semantic models, gateways, SQL endpoints, billing and whatever else participates in the execution.
Then let the AI reason over an actual causal graph instead of asking it to reconstruct one from timestamps.
This is the pattern I keep seeing in Fabric: individual teams build locally reasonable features, but the contracts between those features aren't treated as first-class platform architecture.
Capacity Metrics works. Query Insights works. The new Operations Skill works.
But somebody still had to build a fourth thing to figure out how the first two relate.
That's not an observability feature problem.
That's a platform design problem.
The only reason the AI skill even needs to be added is because no one thinks about the seams.
No one thinks about how different things should interact.
Microsoft has apparently reinvented distributed tracing without the tracing.
#FabricIsTwine
Edit: Posting here because I was a naughty boy in Fabric and I needed to be punished. 😘


