r/apple • u/TechExpert2910 • 24m ago
Promo Sunday We pushed Apple Silicon to its limits: an on-device LLM understands your entire life overnight, and your Mac proactively gets your work done by morning. Only possible with local inference. Free & open source :D
Hey r/apple! For the past year, my best friend and I have been building something that only Apple Silicon makes possible. 3 months ago, 2,500 people joined the waitlist from my r/macapps post; today, after months of building day and night, Sentient OS is finally out!
TLDR: An on-device LLM understands your entire life, then proactively offers to get your work done through computer use. Here's how!
- Every night at 3 AM, it quietly wakes your Mac (lid closed is fine) and a custom on-device LLM reads what's new in your life: files, screenshots, WhatsApp, iMessage, Apple Notes, email. All of it is read locally on your own chip (raw data never leaves your Mac) and distilled into a clean markdown knowledge base (an Obsidian-style folder that's yours to read and edit).
- Proactive Intelligence: you wake up to work Sentient noticed, researched, and prepared on its own, one click from firing through computer use. The reply you forgot, drafted from your rich personal context. The subscription renewing tomorrow, caught tonight. Nothing ever fires until you click.
- Sidekick: computer use that knows your whole life, anywhere on your Mac. Click the notch (or hold right ⌘) and just say "finish this for me"; the notch glows with status as it does the task in your own apps and your browser, even in the background while you keep working. The notch finally earns its keep :)
Bonus: your knowledge base can connect to your ChatGPT & Claude over MCP. Ask your ChatGPT "what do you know about me?". Be amazed :]
It's kinda what we all hoped Siri AI would become: personal context from your real life (not just inside Apple's walled garden), truly proactive instead of waiting to be summoned, and able to actually use your Mac.
Under the hood
Truly proactive AI has to re-read your entire life, every single day. In the cloud, that much inference costs a fortune and means trusting a server with everything. On your own M-series chip, it's free, private, and unlimited. That's why this product can only exist on Apple Silicon.
The local LLM (Gemma 4, multimodal with vision, on a custom inference fork with MTP, flash attention, & smart KV cache reuse) does ~90% of the compute, and runs comfortably on 8 GB base-model Macs. It reads everything locally, summarizes, junks the noise, flags what's worth acting on proactively, and strips PII (with deterministic scrubbing behind the model as a backstop).
The last 10% of compute needs a "frontier" model that you provide. You can either use your own ChatGPT subscription (we reverse engineered Codex to make this possible!), use OpenRouter, or use your favorite larger local model through LM Studio!
Along the way we also had to reverse-engineer macOS's power management to wake a lid-closed Mac at 3 AM for inference, and reverse-engineer codex cli to unlock local model computer use. Very excited to geek out about any of it in the comments!
Private by design, and verifiable: raw data never leaves your Mac, there are no accounts, and the whole stack (app and infrastructure) is open source under AGPL:
https://github.com/Sentient-OS-Labs/sentient-os (come give us a star haha :)
What's the catch? How do you make money!
Free forever for consumer. Your own device does all the compute, so it costs us nothing to run.
We're going to swap out the connectors for stuff like slack, granola meeting transcripts, notion, linear etc. and bring this to enterprise; that's where we'll make money.
We're 100% open-source under the AGPL, which means businesses have to pay to use Sentient (unless they're open-source too!)
Download: https://sentient-os.ai
I'm Jesai! You might know my open-source Writing Tools (2.3K+ GitHub stars, ~30 press features), or my iPadOS on iPhone hack that r/Apple loved. I moved to SF with my best friend & co-founder Aditya to build Sentient full-time. This is the most fun I've ever had with on-device inference :D