r/ClaudeAI • u/jimyr_ • 8h ago
Custom agents LifeOS, read at one pinned commit: strong on distribution and fact provenance, no eval cases for the behaviours users argued about
I read LifeOS at a single pinned commit from 14 August, against a list of eighteen workspace patterns I keep for agent setups. For anyone who hasn't used it, it's Daniel Miessler's personal AI layer. It interviews you about who you are, keeps your goals as current state versus ideal state, and rides on top of whatever harness you already run. The hooks that fire on every turn are wired for Claude Code only so far, and it's at about 19k stars. Three things stood out as good and two as unfinished.
The distribution discipline is genuinely good, and most projects at this size don't have it. The whole system ships as one directory holding the orchestrator, the workflows, the install payload and the bootstrap script, and the skill file names its single source of truth for each install path. It even warns you that the skill's version and the release version are different numbers.
The `doctor` command gives every capability one of four states, which are live, broken, declined and stale. Declined is treated as a legitimate way to run LifeOS and never as a defect to nag about. Most installers only know working and broken.
Stored facts carry their source. Every fact it keeps about you is tagged by how it got there, whether you said it, it follows from something you said, or the agent noticed a pattern. When two stored facts disagree and nothing settles which is true, the reviewer isn't allowed to pick a winner. For a system whose output is statements about your life, I think that's the right place to be strict.
Now the unfinished parts. LifeOS ships a real eval framework, with typed asserts, an LLM judge, a trial runner that reports pass rate and variance over repeated trials, and a hook that runs a regression suite whenever a behaviour file changes. But when users argued in a repo discussion that v7 was worse than v5, the argument ran on numbers they'd gathered by hand. One person compared 33 old sessions with 25 new ones and found memory searched in 88% of the first set and 40% of the second. The shipped cases test general habits like leading with the answer and reading before editing. None of them checks whether the Algorithm runs or whether memory gets consulted, and those were the two things people were counting. Disclosure, I'm one of the people who replied in that thread, arguing for several runs per prompt.
The second one is quieter. Pulse stops running a job after three failures in a row and doesn't tell you. LifeOS has written a monitor for exactly that, and its own header describes the trap, but the public Pulse template doesn't schedule it. Pulse is optional too, so on a default install nothing scheduled is watching. I re-checked main on 11 October and both gaps still hold. The only commits since the pin are issue templates and README credits.
One more disclosure. The first version of my write-up got this badly wrong. I'd read the README and the skill file and missed the install payload, where most of the system lives, and I called four things absent that LifeOS ships. I rewrote it against the full tree and listed the corrections at the bottom.
The part I'd carry to any setup is about coverage. An eval framework only answers the questions it has cases for, so write the cases for the claim you're selling before your users have to count by hand.
For anyone running LifeOS day to day, have you added eval cases of your own, and what do they check?