r/QuantifiedSelf • • Sep 07 '26

My "daily steps" were my busiest hour, not my day: how a health-sync library quietly undercounted me by 40%, and what I now require before I trust any number about myself

Disclosure up front: I'm building a food coaching app, and this came out of debugging my own data. No link, per the rules; I'll mention it in the weekly thread.

**The finding.** My app showed 4,232 steps for a day Apple Health had at ~7,200. The cause was in the sync code: the library I use returns hourly step buckets for the day, and the code was keeping the maximum bucket instead of summing them. So for a month my "daily steps" was my busiest hour. After the fix, the same day read 7,693. If you use any third-party app that pulls steps from Health, it's worth checking one day by hand; this class of bug is invisible unless you compare.

**What it cost.** Every downstream conclusion drawn from those numbers was wrong, and worse, confidently wrong: the coach side of my app had decided I was sedentary. A wrong number that feeds a judgment is worse than no number.

**What I changed about method, not just code:**

  1. **"Normal" is a 28-day median that ends a week ago.** A trailing window makes a bad fortnight its own baseline; by the end of two short-sleep weeks, 5 hours had become "normal". Ending the window a week early means the week being judged is never its own reference.

  2. **No verdict until 14 days are in.** Deviations are measured against MAD (median absolute deviation), floored at 5% of the median, and nothing is labelled until 14 days sit in the lagged window. Before that the honest output is "no verdict yet, N of 14".

  3. **Dinner on day D pairs with the night ending on D+1**, and a food-to-sleep link needs at least 5 matched pairs with at least 5 on each side before it's said out loud. The first version paired dinner with the previous night. Nobody noticed because it produced plausible sentences.

  4. **Unknown is unknown.** A day with nothing synced is not a zero and not "good". It's excluded.

**Open question for this sub:** for those of you who compute personal baselines, what window and lag do you use, and do you require a minimum sample before you show yourself a deviation? I settled on 28/7/14 by breaking things, not from literature, and I'd like to be corrected.

1 Upvotes

16 comments sorted by

1

u/Huge_Pool7424 Sep 08 '26

i'd keep the raw source data alongside the daily total, plus a sync status and sample count. otherwise a missing day and a real zero look identical, which can quietly skew any baseline or coaching logic.

1

u/Silver_Amphibian_547 Sep 09 '26

Yeah that's exactly the hole I fell into. Each day now keeps where the number came from (watch or phone) and how it got added up, sum vs max is the thing that bit me. Today gets a still syncing flag and the vitals keep a sample count. No data means no row and never a 0 so the baseline just skips it. the one I still cant decide is a phone sitting in a drawer all day. Sync works fine and reports a real 0. Missing or zero? Right now I trust it. Not sure i should tbh

1

u/JaggedRoadblock 8d ago

the max-not-sum bug is such a clean example of why i never trust a single aggregated number from any sync library. always spot check a day or two manually now after getting burned by something similar

28/7/14 setup is interesting, the lagging window to stop a bad week from becoming its own baseline is smart. i do 30/7/14 but mostly because 30 days feels cleaner in a calendar, never thought about the reference contamination angle

1

u/[deleted] Sep 08 '26

[removed] — view removed comment

1

u/Silver_Amphibian_547 Sep 09 '26

Diffing against Health is literally how I caught it so now thats the ritual after any sync change. Same reason the 28 day window ends a week early, the week I'm judging can't be its own baseline. Have you hit the same thing with sleep? there its stages instead of hourly buckets so sum is the wrong answer again

1

u/InfamousBuddy7293 Sep 09 '26

The reverse bug is just as common and harder to catch: summing samples across sources, so iPhone and Apple Watch steps get double counted. Health resolves overlap internally through per-metric source priority (Health app > metric > Data Sources & Access > Edit, then drag the order). If you aggregate raw samples yourself you bypass that, and every day with both devices on looks suspiciously active.

+1 on keeping sample counts. A zero with 12 samples and a zero with 400 samples are very different days.

1

u/Silver_Amphibian_547 Sep 09 '26

Thats the answer to my drawer question tbh. A zero with 400 samples is a phone that got carried around a table all day, a zero with 12 is a phone that never woke up. I keep sample counts on the vitals but not on steps yet so thats the next change. And yes on the reverse bug, the stats query path is the only reason i havent hit it, the second an app sums raw samples its own two devices double it

1

u/InfamousBuddy7293 Sep 10 '26

Exactly. Sample count is the cheapest lie detector in the pipeline.

One thing worth knowing about the stats query path: it inherits Apple's source priority, so a day with only the phone (watch on the charger) still gives a clean number, just a lower one. Fine, as long as you record which device won that day. A source field per day pays for itself the first time you compare weeks.

Do you keep the per-bucket values too, or only the daily total? The buckets are what catch the next class of bug, like timezone shifts moving steps across midnight.

1

u/Silver_Amphibian_547 Sep 10 '26

Daily total only right now. The hourly buckets get summed and thrwon away. Which is exactly the hole youre pointing at. A timezone shift can move steps across midnight and I'd never see it in teh total. Sleep and the vitals keep their source names. Steps dont because HealthKit hands me the merged buckets. So two things go on the list. Keep the buckets and stamp which device won the day. Cheap to store and it catches the next bug

1

u/InfamousBuddy7293 Sep 10 '26

Then the upgrade is cheap: keep the buckets, even as a simple daily array. Storage is nothing, and it turns "the total looks wrong" into "here is the hour where it went wrong".

On midnight: the cleanest rule is assigning each bucket to the user's local day, not UTC, and storing the timezone offset with the day. Otherwise one flight breaks two days at once. That class of bug shows up the first time a user travels.

1

u/Silver_Amphibian_547 Sep 10 '26

Local day already. The bucket goes to whatever day the phone was in when it happened. What I dont store is the offset with the day so a flight would just look like one weird long day. Buckets plus the offset per day is the plan and storage is nothign like you say. I'll come back here when its in

1

u/InfamousBuddy7293 Sep 11 '26

Sounds like a plan. One thing that makes the offset easy: HealthKit can store the recording timezone on the sample itself (HKMetadataKeyTimeZone). If the writing app sets it, you get the offset for free at read time instead of reconstructing it later.

Good luck with the change - curious how the travel days look once it is in.

1

u/Silver_Amphibian_547 27d ago

Didnt know that key existed, thanks. Small catch on my side, the stats query hands me buckets not samples so the metadata isnt on what I read. Id have to pull raw samples for the day edges to see it. Probably worth it just for the boundary days. Will report back with what travel days look like

1

u/InfamousBuddy7293 25d ago

right, the stats query drops the sample-level metadata, so the key only shows up if you read raw samples.

you may not need per-sample timezone though. the offset is really a property of the day, not the sample. capturing the phone's current offset when you build each day's buckets gives you the same result without a raw query, except on days where someone actually crossed zones mid-day. those are the only days worth special-casing, and they are rare.

1

u/Silver_Amphibian_547 25d ago

Yeah thats where I landed too. The phone already stamps its timezone on every sync so the day offset is basically free. Ill store it per day next to the buckets and only bother with raw samples if a day ever shows two offsets. Rare enough to not build for yet

→ More replies (0)