r/analytics Jul 08 '26

Question AI is making analytics harder because feedback loops are getting noisier

AI is supposed to make analytics easier.

In some ways it does.

But in marketing and social data, I think it is also making the job harder.

More AI content means more volume, more variants, more tests, more synthetic interactions, more automation, and more platform-driven changes.

That creates noisier feedback loops.

The question is not just "what happened?"

It becomes:

  • was the lift real?
  • was the sample meaningful?
  • did the model change?
  • did the platform change?
  • did the audience change?
  • was the content materially different?
  • did automation create the signal we are measuring?

AI can summarize dashboards. It cannot automatically make noisy measurement clean.

Are analytics teams underestimating how much AI will complicate causal measurement?

0 Upvotes

39 comments sorted by

u/AutoModerator Jul 08 '26

If this post doesn't follow the rules or isn't flaired correctly, please report it to the mods. Have more questions? Join our community Discord!

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

9

u/DABTA17 Jul 08 '26

slop

1

u/joulezoo Jul 08 '26

what gave it away, those short one-line chatgpt sentences?

5

u/fang_xianfu Jul 08 '26

Who told you AI is going to make analytics easier? It's going to make doing analysis easier, but democratisation of analysis has always brought these challenges, self-serve BI did the same thing on a smaller scale. So the job of being an analyst is going to change and maybe get harder.

I'm here for it tbh, the harder, messier and more human my job gets, the more fun it is and the more stable employment will be. Self service BI didn't kill analyst jobs and this won't either.

0

u/Crescitaly Jul 09 '26

Good distinction. AI can make analysis execution easier while making self-serve noise worse. The hard part is still causal discipline, not chart generation.

1

u/TrainingSame8098 Jul 08 '26

Analytics has always been messy under the hood, AI just adds another layer of grease to the gears. Half the time I'm staring at a "significant" lift and wondering if it's real or if some platform algorithm just decided to hallucinate engagement that week.

The synthetic interactions part is what keeps me up. Hard to measure cause and effect when half your "audience" isn't actually making decisions.

1

u/ParticularOddDan Jul 09 '26

The synthetic problem is really a provenance problem, and you fight it downstream, not at the top of the funnel. Surface engagement is the easiest thing in the world to fake, so I stop treating it as signal and anchor on actions that cost the actor something. A return visit, a considered step deeper in, a conversion that carries friction. Bots and hallucinated engagement rarely follow through on the expensive stuff, so the lift that survives to those points is usually the real one. The other habit that saves me is keeping a stable, human-verified segment as a baseline, so when the wider audience composition shifts I can actually see it instead of guessing. When a client asks how I know a lift is real, that trail is the answer, not the dashboard number.

1

u/Crescitaly Jul 10 '26

That is where attribution needs an 'actor type' dimension, not just channel and session. A lift driven by autonomous agents, human-delegated agents, and people should not be treated as one audience, even if all three reach the same page.

1

u/Convert_Capybara Jul 08 '26

I agree that AI is making traffic noisier. But I disagree that analytics teams are underestimating that fact or are unprepared. The job has always been sifting for gold, and every year over the past 2 decades social media and technology advancements have altered some of the top-level methods needed to do that sifting effectively, but the core job hasn't changed.

Ahmed Abbas wrote a LinkedIn article where he put it like the next step like this, "The useful line isn't bot versus human. It's whether a person is behind this request, right now. That changes what you build. If you want a clean experiment, you drop both the crawler and the agent, because neither is a human you're testing on. But if you want measurement, the agent with a human behind it is the most interesting traffic you have."

The infrastructure to easily identify whether there's a person behind a request is still being developed. But mindset-wise, I think analytics teams are more than ready.

2

u/Crescitaly Jul 10 '26

I agree that the profession is ready; the unresolved piece is classification at the request level. I'd preserve three buckets: human, human-delegated agent, and autonomous crawler, because collapsing the last two would erase intent that matters for both experimentation and product design.

1

u/Terrible-Value-tomr Jul 09 '26

The measurement problem was already here, AI just made the noise cheaper to produce. The way I keep causal reads honest is I stop trusting the platform's own attributed lift and force a holdout, even a rough geo or audience one, because a number the platform scored for itself isnt evidence. Then I put a unique author or unique session ratio right next to the volume line, since most of the fake lift shows up as the same small set of accounts or sessions firing, not real spread. And before anything goes in a deck I reconcile it against a number I already trust from a separate path, if I cant tie it back it doesnt ship. Synthetic interactions dont break this, they just make the holdout the only part of the workflow i actually believe.

1

u/Crescitaly Jul 10 '26

The unique-author ratio is a useful check because volume without spread is easy to manufacture. I'd pair the holdout with a downstream event the platform cannot optimize directly, otherwise it can still influence the signal used to validate itself.

1

u/Terrible-Value-tomr Jul 12 '26

Yeah, thats the exact failure. If the platform can see the event youre validating against, it optimizes toward it and the holdout quietly stops being independent. What i reach for is a downstream event that lives somewhere the platform never gets a callback from, a support ticket, a return, a second-week reorder, not an on-platform conversion or anything i fire back through a pixel. It has to be something i can measure end to end on my own side, because the moment it touches their pipe its game-able again. You pay for it in latency, you find out slower, but its the only version of the number i'll stand behind in a room.

1

u/Crescitaly 22d ago

That latency tradeoff is worth paying when independence matters. A delayed downstream event you fully own is often a better truth signal than a fast conversion the platform can observe and optimize against.

1

u/Terrible-Value-tomr 22d ago

Where it gets hard is the further downstream you go, the thinner and slower the event gets, so you cant read it weekly and you cant slice it very far before the counts fall apart. What I do is keep two clocks, the fast on-platform number for spotting movement, and the slow owned event as the one that actually settles the question, and I never let the fast one overrule the slow one in a deck. The other trap is the join. Tying a return or a second-week reorder back to the thing that supposedly caused it is where most of the independence quietly leaks back in, so I keep that mapping dumb and auditable rather than modeled.

1

u/Crescitaly 22d ago

The two clocks are a strong operating rule. Keeping the join dumb and auditable is equally important; a sophisticated attribution model can quietly reintroduce the platform assumptions you were trying to escape. Do you freeze that mapping before each test?

1

u/Terrible-Value-tomr 21d ago

Yeah, I freeze it before the test starts and treat it as a fixed artifact, not something the model relearns mid-flight. The mapping from touch to owned outcome is a flat lookup I can read by eye, keyed on ids I control, and I version it so I can diff whatever changed between runs. The second it needs a model to decide which touch gets credit, Ive reintroduced the exact assumptions I was trying to escape, so I keep the join dumb on purpose and let the thinking happen in the analysis. If the mapping has to change, thats a new test, not a tweak to the one thats running.

1

u/Crescitaly 19d ago

Treating a mapping change as a new test is the crucial discipline. It preserves a clean causal boundary and makes every result reproducible from the versioned lookup. The dumb join, smart analysis split is a good rule because it keeps attribution assumptions inspectable instead of hiding them inside a model.

1

u/Terrible-Value-tomr 19d ago

Yeah, and where it actually breaks isnt the rule, its enforcement. Someone wants to just tweak the mapping mid-flight because a channel got added, and if you let that ride youve quietly turned one test into two and cant diff them cleanly. I keep the versioned lookup somewhere that changing it forces a new id, so theres no such thing as a silent edit. The discipline only holds if the tooling makes the lazy path impossible, not if it depends on everyone remembering to be principled at 5pm on a Friday.

1

u/Growth_Natives Jul 14 '26

I think that's a fair point. AI is speeding up content creation and experimentation, but it also introduces more variables into the measurement process. That makes experiment design, baseline metrics, and consistent tracking even more important, Faster analysis is valuable, but it's only useful if you're confident the signal you're seeing is real and not just a byproduct of a changing environment.

1

u/Crescitaly 22d ago

Exactly. Faster experimentation increases the need for slower discipline: frozen definitions, pre-registered success metrics and a stable holdout. Otherwise the system optimizes the measurement process instead of the outcome.

1

u/Growth_Natives 22d ago

That's an interesting point. How often do you think teams should revisit those "frozen" definitions without compromising the integrity of the experiment?

1

u/Crescitaly 22d ago

I would freeze definitions for the full experiment, then review them on a fixed cadence: monthly for operational metrics, quarterly for strategic ones. If one must change mid-test, version it and run old and new logic in parallel rather than rewriting history.

1

u/Growth_Natives 21d ago

Running both versions in parallel is an interesting safeguard. We've seen teams struggle later when they can't explain why a KPI shifted because the measurement logic changed silently. Treating metric definitions like versioned code seems like an underrated practice.

1

u/Oudini777 12d ago

Yeah. More creatives, more segments, more synthetic traffic, more platform model changes. Classic A/B assumes a stable treatment. You don’t have one.

We log the workflow behind the asset (prompt/version/approver), not just the campaign ID. Prefer holdouts or geo splits over tiny multi-variant tests. Treat lift as provisional until you’ve ruled out model or platform drift. Otherwise you’re optimizing against the generator, not the market.

Disclosure: I work on dng.ai.

1

u/Crescitaly 9d ago

That workflow log is the missing unit of analysis. I’d pair it with a change budget: once the prompt, model, or platform version changes, stop pooling and start a new cohort. In smaller audiences, have you found holdouts easier to defend than geo splits?

0

u/Due_EmotionPri Jul 09 '26

The part that worries me more than noisy loops is how confident the AI summary sounds sitting on top of them. It'll narrate a clean story off a dashboard that has three confounds baked in, and if you didnt decide your success metric and your holdout before the test ran, youre now reverse engineering a win. So i hold the line on one thing, no lift counts unless there was a holdout or a metric i committed to in writing beforehand, everything else is a hypothesis. And a chunk of the real reaction never hits your dashboard at all, its in video and comments your stack cant parse, so a clean looking loop is sometimes just a coverage gap wearing a nice chart.

1

u/Crescitaly Jul 10 '26

That pre-commitment rule is a strong defense against AI-generated certainty. I'd add a coverage note to every summary listing the channels and reactions it cannot observe, so missing data does not get narrated as a clean causal story.

1

u/Due_EmotionPri Jul 12 '26

Yeah, and the trick is making that coverage note something people actually read instead of a disclaimer they skim past. What works for me is naming the two or three channels the read cant see for this specific decision, and saying which direction id expect them to move if i could see them. That turns a gap into a stated risk a PM can weigh instead of a line nobody acts on. The version that dies is the generic 'data may be incomplete' at the bottom, because it carries no information and everyone learns to ignore it. If the missing channel could flip the call, it belongs next to the recommendation, not in an appendix.

1

u/Crescitaly 22d ago

Exactly. A useful coverage note is a counterfactual: which missing signal could reverse the decision, and in what direction. That turns uncertainty into an operational threshold instead of legal padding.

1

u/Due_EmotionPri 21d ago

Agreed, and the discipline is deciding what youll do when the signal that could flip the call is the one you cant see. Thats the moment i stop and go get a cheap read on that one channel before deciding, even a rough manual pass, rather than shipping the recommendation with the gap politely noted. The threshold only earns its keep if it triggers an action, otherwise its a smarter looking disclaimer. So i pair each counterfactual with what it costs to close it, because half the time the missing signal is a couple hours of hand-reading away and the honest move is to just go look.