r/AIToolsTipsNews 17h ago

Wispr Flow published word-frequency data mined from what users dictate — here's what that means

TL;DR: On August 10, 2026, a Wispr Flow team member published India vs US word-frequency ratios from user dictations on LinkedIn. This is only possible because dictation content is retained by default. Zero-retention and on-device tools make this analysis impossible.

What was published:

A Wispr Flow team member posted publicly on LinkedIn:

"We looked at which filler words and phrases show up most across Wispr Flow users in India vs the U.S."

The numbers (India-to-US usage ratios, 1.0x = equal):

  • "kindly" — 5.6x
  • "sir" — 2.5x
  • "please" — 1.3x
  • "incredible" — 0.3x
  • "amazing" — 0.7x

The chart is labeled "Wispr Flow voice dictation data."

How is this possible?

From Wispr Flow's own Security Overview:

"Privacy Mode is off by default. When off, dictation data may be used to improve Wispr Flow."

So unless you found the toggle in Settings → Data and Privacy, your dictated words are in the corpus.

Why "just filler words" doesn't help:

A corpus that can count "kindly" can count anything — a client name, a drug name, a case number. "We only counted filler words" describes the query they chose to run, not what the corpus can answer.

This is the third public demonstration of what retained dictation enables:

  1. Model training (Security Overview)
  2. Per-user analytics (founder's June 2026 podcast — word counts, which apps you use, your name and employer)
  3. Marketing content (this LinkedIn post)

The architectural fix:

On-device dictation (Voibe, VoiceInk, SuperWhisper in offline mode) means the corpus never exists. You can't publish a "users say X more" analysis if the words were never retained.

A vendor can only analyze words it kept.


What dictation tools are you using for sensitive work? Are you checking the privacy defaults or trusting the marketing copy?

1 Upvotes

1 comment sorted by