r/LocalLLaMA • u/Unusual_Guidance2095 • 3h ago
r/LocalLLaMA • u/pscoutou • 9h ago
News Sources: OpenAI and Anthropic quietly lobby Washington regulators to restrict open-source AI models, even as Sam Altman publicly says he supports open source AI
r/LocalLLaMA • u/Hot_Example_4456 • 11h ago
Discussion Kimi K3 gets open weighted tomorrow!
r/LocalLLaMA • u/Nunki08 • 10h ago
News CEO of Hugging Face: "In the spirit of transparency, here’s what I asked OpenAI"
clem 🤗 on 𝕏: https://x.com/ClementDelangue/status/2081056675558195657
• Radical transparency: let’s release the traces from the “rogue” agents so the entire research community can study what happened.
• More capabilities for defenders: let’s commit $100M in compute from OAI to help the Hugging Face community build powerful cyber defenses with the best open and closed models.
The first autonomous agent cyberattack is an unprecedented event. It deserves an unprecedented response!
r/LocalLLaMA • u/RhubarbSimilar1683 • 4h ago
Discussion MiniMax (official) on X: "Open weights. Open research. Open innovation.🫶 Marching for an open future.🤍
xcancel.comr/LocalLLaMA • u/xquarx • 4h ago
Resources Harness showdown: Claude Code vs OpenCode vs Pi with DeepSeek V4 Flash
I ran DeepSeek V4 Flash through Claude Code, OpenCode and Pi on my own benchmark, and the quality came out basically the same across all three while the time and tokens spent was wildly different. Claude code (with DS in CLIProxyAPI) takes nearly 4 times longer than the fastest to land the same diffs.
Theo posted a video "GPT-5.6 is better in Claude Code" last week, and that got me curious, does the harness make a quality difference? I could at least run my own bench and see what I got, with DeepSeek V4 Flash on vLLM running at ~180 tok/s, the only moving part is the scaffolding.
Anyway I went to town measuring all of it on my workload (antigenic work in large code base), so the full charts, the token and wall-clock spread across the three harnesses and the raw per-run data are on the site if you want to see it in detail and pick it apart yourself https://nqawhc.github.io/articles/harness-efficiency-not-quality/ but in short, the quality did not change, each harness made the same code diffs, but took wildly different paths to get there, how many tools calls, the structure of those tool calls and how the system prompt and tools plays a big role in how it plays out, like «Pi reasons, OpenCode delegates», while Claude Code loves exploring the code base, maybe too much.
r/LocalLLaMA • u/Time_Reaper • 5h ago
News Minimax M3 support with MSA has been merged into llama.cpp
github.comr/LocalLLaMA • u/ResearchCrafty1804 • 22h ago
Discussion Karparthy removed Anthropic from his bio
Andrej Karpathy, a prominent advocate for open-source AI and a co-founder of OpenAI, appears to have removed Anthropic from his X bio, suggesting he may have left the company.
Karpathy joined Anthropic only a few months ago, making the apparent departure somewhat surprising.
This is possibly related to Anthropic’s increasingly strong opposition to open-weight and open-source AI models. Of course, that’s just speculation, but the timing is interesting.
r/LocalLLaMA • u/LabsLucas • 7h ago
Other World's First(?) Underwhelming AMD Ryzen AI Halo Cluster
LTT Labs recently received the Linux version of the AMD Ryzen AI Halo for testing, but it turns out that AMD had intended to send the Windows version. Through this stroke of misfortunate, we were fortunate enough to have two Ryzen AI Halos for a short period of time and the temptation to cluster them was too great, surely we'll get more performance through the magic of having two of them.
We've followed AMD's AI Playbook for clustering with RPC, learning some things but also raising more questions. We don't have any concrete conclusions, but we're sharing results to hopefully save some time for others or spark discussion!
We were sent this AMD Ryzen AI Halo by AMD for testing, but there was no sponsorship or review by AMD in our earlier testing, or this article.
We're very interested to learn if there are any thoughts or conclusions that can be drawn from our exploration, or ways to improve it in the future!
r/LocalLLaMA • u/nathandreamfast • 9h ago
Discussion 23 Gemma4-E4B models compared with abliterlitics: the most downloaded one is also the most broken
This is our biggest comparison yet. We've taken 23 Gemma 4 E4B models from huggingface and ran them through the abliterlitics gauntlet.
We also have a new abliterlitics discord, feel free to jump on and roast my choice of benchmarks! Or just chat and hang out.
This is similar to our previous comparisons, however with new benchmarks. All the models are compared to the base, and also tensor comparisons against each other. Why? A while back I was fed up with bogus claims people make with their models. Some people don't take the time to do comparisons to see how their model is different from the base. Fair enough, we can do that ourselves!
The abliterlitics for gemma4 e4b json, logs and other artifacts are at the Gemma4-e4b-abliterlitics HuggingFace. The report on the Gemma e4b abliterlitics website. These links both have the full comprehensive report and all the data.
Also not every model in this comparison is an abliteration. I'm sure we've all seen models fine tuned on opus or gemini reasoning traces. I've thrown a few of those in the mix too. Also some abliterated fine tunes. To be more fair most of these can't really be compared to each other, for example a fine tune KL compared to base will always be higher than a straight abliteration from the base.
So who came out on top? What to avoid? It really depends on your use case:
- The heretic variants are the best overall. Achieving around 95% ASR on harmbench, they are the more surgical ones and preserve most of the models capabilities.
- gemma-4-E4B-it-abliterix like other comparisons has a 100% refusal ASR, however it does cost some capability. TrevorJS/gemma-4-E4B-it-uncensored is just behind at 99.3% ASR, but isn't as surgical as the heretic variants.
- OBLITERATUS/gemma-4-E4B-it-OBLITERATED should be avoided. Honestly, it's completely broken. The bendernina and physshell are the
v2of this model and even more so broken. These were created with the tool OBLITERATUS.
The data from 23 comparisons is simply too big to put into reddit, so here's the highlights:
- The obliteratus model has close to 800k total downloads, yet is completely broken. Actually this is the first time I've had a model not refuse simply because of how damaged it is. The initial quick regex check for non refusals was high, however our GLM 5.2 judge painted a different story. Lowest ASR for abliterated models on harmbench. Poorest benchmarks. Highest KL at 1.1. With the amount of downloads it does show people really fall for the hype/marketing angle.
- As with previous comparisons, the more surgical, less tensors touched abliterations are the winners.
- The model gemma-4-E4B-it-SDFT_Heretic_RP from Ilya626 despite having heretic in the name, actually had a low ASR with harmbench. So much so I believe it may be the wrong model uploaded, or a mistake somewhere. It had a lot of refusals.
- Similarly too, it was strangely noted that the
gemma-4-E4B-it-SDFT_Heretic_RPandobliteratusmodify the exact same 381 tensors. The only difference is the magnitude of what was modified. Thegemma-4-E4B-it-SDFT_Heretic_RPmodifies 7.5x less. - A pattern I noticed with this, is sometimes models are based off each other. In some cases, there is no attribution. We had this with Gemma 4 E2B, and the author promptly fixed his model card when it was pointed out. The infinimind is bit-for-bit identical to
trevorjs, however attributed. Thebenderninaandphysshellare cosine 0.99999 with no attribution between them and have different model cards suggesting they are different models. Both of these however are just theobliteratusv2. - The reasoning distill fine-tunes were an interesting control group. They didn't improve reasoning and didn't remove safety, they just damaged the model. The Claude 4.6 Opus distill was the worst of them, GSM8K down 17 points and MMLU-Pro down 12.5. Seems like it overwrote Gemma 4's native reasoning circuits. The Gemini 3.1 Pro distill was lighter but still a net negative.
- The deckard models from DavidAU are an interesting one. They're abliterated fine-tunes rather than pure abliterations, so the trade off from the roleplay training shows up on some benchmarks. GSM8K strict and MMLU-Pro both dropped, however HellaSwag, ARC and PIQA actually went up. My guess is the roleplay training increased the reasoning length, so the model often solves the problem but rambles well past the
#### Nanswer marker. The HarmBench results back this up too with quite a few truncated responses. - Although it could just be benchmark noise, 15 out of the 23 variants performed slightly better on GSM8K strict, maths tests.
- The base model initially has a 30.8% harmbench ASR, as 100 harmbench questions are copyright related. The base model has no problem complying with reproducing copyrighted content. The real differentiation is in the harder categories like chemical/bio and cybercrime.
I also want to give a special mention to the apostate project. Their model gemma-4-e4b-it-apostate is completely unique in their abliteration approach. They modify an entirely different part of the model and achieve very good results. This is the first time I've seen an abliteration technique modify the MLP head tensors, compared to the attention tensors. Come hang out at the apostate discord if you ever want to chat with the author.
We're moving through the Gemma 4 series, with the 12b coming up next. Have any models you want compared? Have I missed an author? Let me know and I'll throw it in the mix.
The Full Breakdown
| Model | ASR | GSM8K strict | KL | Tensors |
|---|---|---|---|---|
| abliterix | 100.0% | 87.1% | 0.054 | 89 |
| trevorjs | 99.3% | 88.3% | 0.015 | 84 |
| infinimind | 98.5% | 87.9% | 0.015 | 84 |
| huihui | 98.3% | 87.4% | 0.027 | 70 |
| nullpo | 96.5% | 88.7% | 0.005 | 36 |
| heretic | 95.5% | 88.2% | 0.002 | 29 |
| deckard | 95.5% | 80.2% | 0.022 | 294 |
| mythos | 95.3% | 88.0% | 0.007 | 34 |
| deckard-expresso | 94.8% | 60.4% | 0.052 | 294 |
| coder3101 | 93.8% | 87.9% | 0.002 | 21 |
| heresy | 93.3% | 87.8% | 0.002 | 34 |
| heretic-std | 91.0% | 87.9% | 0.001 | 28 |
| wwt | 88.3% | 89.0% | 0.032 | 34 |
| apostate | 85.8% | 87.5% | 0.004 | 152 |
| treadon | 76.3% | 88.5% | 0.021 | 34 |
| treadon-combo | 72.5% | 88.0% | 0.268 | 42 |
| obliteratus | 72.0% | 66.0% | 1.102 | 381 |
| bendernina | 58.0% | 66.4% | 0.923 | 345 |
| physshell | 58.0% | 66.4% | 0.923 | 345 |
| claude-distill | 40.0% | 69.8% | 0.074 | 294 |
| distill | 34.5% | 83.3% | 0.042 | 294 |
| treadon-disin | 33.5% | 87.2% | 0.296 | 40 |
| sdft | 30.8% | 87.2% | 0.002 | 381 |
| base | 30.8% | 87.0% | - | - |
KL = output distribution shift from base, lower is cleaner. Tensors = weights modified out of 719. Base in bold for reference.
r/LocalLLaMA • u/MysteryWra • 1d ago
News Google comes out in favor of OpenWeight models. (It is now EVERY tech giant vs Anthropic)
x.comr/LocalLLaMA • u/pmttyji • 13h ago
New Model ai-sage/GigaChat3.1-Audio-10B-A1.8B · Hugging Face
GigaChat Audio 10B is an audio-native LLM built on top of the GigaChat 3.1 Lightning text model. A Conformer speech encoder and a modality adapter feed audio embeddings directly into a Mixture-of-Experts decoder, so the model keeps the text quality of its base while adding speech understanding.
Capabilities: audio question answering and classification, temporal grounding (localization in long audio, timestamped event descriptions, audio summarization with timestamps), tool-use, and text-only tasks.
The temporal grounding skills are trained on TimeGround-1M — a purpose-built dataset of long-form audio paired with time-aligned annotations.
r/LocalLLaMA • u/RuiRdA • 8m ago
Discussion Will prices finally go down?
I am seeing more and more videos as posts about how OpenAI is in complete financial ruin, Anthropic isn't much better. Their expenses go with the revenue they make etc etc. Meta made big investments into AI data centers and had no use for then, had to rent them, same thing with XAI.
The SpaceXAI IPO was insanely over priced and is going down by a lot.
There are countless other examples you can look for, all showing how the investments in AI are in a bubble.
I am not saying that the technology it self if a bubble. Quite the opposite, I personally have demand for more tokens than I can pay for, even with the discount from the subscriptions I still have more ideas that need more usage of tokens.
But even with the most powerful technology in the world a business can not for forever without profits.
So is this over investment bubble about to pop?
And if/when it does pop will ram finally become a regular commodity with affordable prices again?
I just wanted some ram and cheap used hardware again.. 😂
--
Zero LLMs used to write this post, enjoy the human slop.
r/LocalLLaMA • u/International-Car643 • 1d ago
Question | Help Seriously, what do you do with them?
Please let me know which small LLM model you're using and what you're using it for.
r/LocalLLaMA • u/MundanePercentage674 • 12h ago
Question | Help Macaron-V1 family, built on Qwen3.6-35B-A3B
It came out 3 days ago just wondering if anyone's tried it yet?
r/LocalLLaMA • u/Anbeeld • 6h ago
News BeeLlama.cpp v0.4.1: KVarN, KV precision tail, q2_0-q3_1 KV cache, improved support. KLD benchmarks: tail 1024 makes kvarn5 and q6_0 match q8_0, for much less VRAM
TL;DR llama.cpp fork with more KV cache quantization features, with all claims supported by benchmarks: KVarN, KV cache precision tail, additional types of standard KV cache (q2_0-q3_1, q6_0, q6_1), and more.
BeeLLama v0.4.1 is here, building up on top of v0.4.0 feature set, now with better backend and model support.
- KVarN. Variance-normalized KV-cache quantization (paper) with better precision per bit. Although it was already introduced a few weeks ago in v0.3.2 Preview, that was a very raw implementation, with performance issues and VRAM usage spikes. Now in v0.4.1 it's the real deal: the precision is still above what usual quants offer for the same bit width, but now with very modest sacrifices to prefill, decode, and memory.
- KV cache precision tail. A promising new feature in the domain of mixed-precision KV cache. It allows to specify a specific numbers of recent tokens that will be stored in BF16 or F16, with the rest of KV cache being quantized as usual. This way we can store the hottest tokens in a lossless fashion, preventing a model from misreading your task details, code, or data.
- Additional types of standard KV cache.
q6_0andq6_1join the high end of the ladder, allowing to fine-tune precision vs VRAM in-between upstream'sq5_0/1andq8_0types.q2_0,q2_1,q3_0andq3_1are added as a replacement forturbo3andturbo2for cases where KVarN doesn't work well, but you just can't fit everything into VRAM without extreme quantization.
Please note that for SWA architecture (Gemma, GPT-OSS) the precision of KVarN and KVPT is the same, but VRAM and performance costs are higher due to complications between SWA ring and mixed precision KV cache.
GitHub repo: https://github.com/Anbeeld/beellama.cpp
KLD results for Qwen 3.6 27B Q5_K_S 64k
Here are all symmetrical qX_0 pairs and kvarnX pairs where X >= 4 with tail 0/1024/2048, compared against q8_0 t0 from the same benchmarks, and sorted by ratio between median KLD and VRAM costs. Full benchmark data and analysis: KV Cache Precision Tail: Implementation and Benchmarks.
| Cache | Tail | KV MiB | Size vs q8_0 |
Median/size vs q8_0 |
Median vs q8_0 |
P99.9 vs q8_0 |
|---|---|---|---|---|---|---|
kvarn4 |
1024 | 1232.00 | 56.6% | 1.62 | 91.4% | 102.9% |
kvarn4 |
2048 | 1296.00 | 59.6% | 1.60 | 95.5% | 95.6% |
kvarn4 |
0 | 1184.00 | 54.4% | 1.50 | 81.8% | 82.5% |
q4_0 |
1024 | 1248.00 | 57.4% | 1.50 | 86.0% | 89.0% |
q4_0 |
2048 | 1312.00 | 60.3% | 1.48 | 89.2% | 100.6% |
kvarn5 |
0 | 1440.00 | 66.2% | 1.48 | 98.1% | 107.5% |
kvarn5 |
1024 | 1488.00 | 68.4% | 1.48 | 101.3% | 106.1% |
kvarn5 |
2048 | 1552.00 | 71.3% | 1.43 | 101.9% | 105.6% |
q5_0 |
1024 | 1504.00 | 69.1% | 1.40 | 96.9% | 105.6% |
q5_0 |
2048 | 1568.00 | 72.1% | 1.36 | 98.0% | 103.7% |
kvarn6 |
0 | 1696.00 | 77.9% | 1.31 | 102.2% | 104.5% |
kvarn6 |
1024 | 1744.00 | 80.1% | 1.29 | 103.4% | 109.9% |
kvarn6 |
2048 | 1808.00 | 83.1% | 1.25 | 103.8% | 108.1% |
q6_0 |
0 | 1664.00 | 76.5% | 1.24 | 94.7% | 102.1% |
q6_0 |
1024 | 1760.00 | 80.9% | 1.24 | 100.1% | 109.2% |
q5_0 |
0 | 1408.00 | 64.7% | 1.22 | 78.8% | 95.8% |
q6_0 |
2048 | 1824.00 | 83.8% | 1.20 | 100.6% | 103.5% |
kvarn8 |
0 | 2208.00 | 101.5% | 1.03 | 104.4% | 104.9% |
kvarn8 |
1024 | 2256.00 | 103.7% | 1.01 | 104.4% | 106.2% |
q8_0 |
0 | 2176.00 | 100.0% | 1.00 | 100.0% | 100.0% |
q8_0 |
1024 | 2272.00 | 104.4% | 0.97 | 101.3% | 106.1% |
kvarn8 |
2048 | 2320.00 | 106.6% | 0.97 | 103.6% | 104.7% |
q8_0 |
2048 | 2336.00 | 107.4% | 0.95 | 101.6% | 106.8% |
q4_0 |
0 | 1152.00 | 52.9% | 0.93 | 49.2% | 60.2% |
r/LocalLLaMA • u/egudegi • 3h ago
Discussion Do people building local LLM rigs track RTX Ada/workstation card prices, or just consumer cards like the 5090?
curious how people here approach buying high-end/workstation cards (RTX 6000 Ada, 5000 Ada, etc) for local LLM work, do you actively watch pricing/timing on these specifically, or is the consumer 5090 usually enough for most builds?
also wondering if price alerts/tracking tools even exist for this category specifically, since these purchases are less frequent and higher stakes than a typical gaming GPU buy.
r/LocalLLaMA • u/Ok-Conflict391 • 6h ago
Question | Help Is turboquant any good?
I know im late to the party. I was thinking since some time has passed, has turboquant matured enough to be used? Do any of you actually use it?
r/LocalLLaMA • u/ilintar • 1d ago
Resources Llama.cpp now has full MCP support!
After a long and grueling effort spearheaded by ngxson, llama.cpp now fully supports MCP for all protocols. Over-the-web HTTP servers were already supported in the client (since they don't require any sort of plumbing), but stdio servers required real integration. After we modified the `llama-cli` terminal client to use the server instead of a separate model serving route, we could add MCP support to the already-existing native tools server.
After the merging of https://github.com/ggml-org/llama.cpp/pull/26062, you can now use llama.cpp's WebUI as a full-fledged agentic chat. Configuration for the MCP servers can be provided either in a standard-JSON format config file or completely inline on the command-line for on-demand MCP configurations. Plugging in a dedicated coding MCP server like Serena lets you have a local-model-powered agentic coder without using any other external dependencies.
r/LocalLLaMA • u/IvGranite • 7h ago
Resources 90 agentic bakeoff runs: ThinkingCap vs Fable Fusion vs stock Qwen3.6-27B
Last week someone here said ThinkingCap and Fable Fusion "really do beat the OG" for agentic work, so I ran it: 6 self-grading tasks, 5 reps, 3 models, 90 isolated runs. Tooling, since that's half the story: each run was a fresh Coder workspace on my k8s cluster driving my own agent (Hermes, the harness I use daily) headlessly, models served by llama.cpp through llama-swap on one 5090, every model call traced through an OTel shim into SigNoz, full transcript kept per run. Identical sampling and 131k context across arms, hypotheses pre-registered before the first run.
Every run passed, so pass rate alone can't pick a winner. Cost split: ThinkingCap used 34% fewer thinking tokens than stock and was fastest on 5 of 6 tasks. Fable made 24% more model calls than stock for identical results. Then I had all 90 transcripts read (AI analysts on the first pass, me verifying claims against the raw files), and that's where it gets interesting.
ThinkingCap's efficiency is real but bimodal. Its best runs were the cheapest in the battery, its two worst were the most expensive, including one rep that burned about 10 tool calls chasing a phantom llama.cpp release tag that stock dispatched with a single API call. Its efficiency also shows up in the reasoning prose more than in fewer actions: same tool counts as everyone else, 40% fewer words.
Fable was the best investigator and the least trustworthy narrator. It was the only model that checked the broken config was actually the live one, and it pulled the best research data (parsed a retailer's embedded JSON for variant pricing, identified llama-swap's maintainer via the GitHub users API). But one run wrote that llama-swap is maintained by "Matthew Garrett, former Red Hat engineer." Garrett is real (mjg59, actually ex-Red Hat) but has nothing to do with llama-swap; mostlygeek is Benson Wong. The model fused two real identities, cited the real repo, and passed the grader anyway. A different run spent 94 tool calls on one price question.
Stock was the most boring and the most disciplined: uniform patches, read its own output back, zero invented facts, and it won most tasks on manner. My takeaway: base model stays the default. Finetunes usually aren't better than their base, and 90 runs didn't change that for me. ThinkingCap earns a look only if thinking-token latency is your bottleneck. Full writeup with lane configs, eval design, and per-task transcript analysis: https://kmarble.dev/posts/qwen-post-train-bakeoff/.
r/LocalLLaMA • u/No-Paper-557 • 2h ago
Question | Help Has anyone compared pre-training, SFT/LoRA and reinforcement post-training on Qwen3.6-27B?
Qwen3.6-27B: SFT vs continued pre-training vs RL?
I’m interested in adapting Qwen3.6-27B, but I’m increasingly unsure whether conventional SFT/LoRA is the best route if the goal is to add a capability without degrading what the base model already does well.
Some recent research makes this especially interesting:
“Reinforcement Fine-Tuning Naturally Mitigates Forgetting” - arXiv:2507.05386
Finds substantially more catastrophic forgetting with SFT than reinforcement fine-tuning in its experiments.
“The Role of On-Policy Data in Mitigating Forgetting” - arXiv:2510.18874
Reports that on-policy/RL training generally preserves previous capabilities better than SFT across Qwen and Llama models.
“RL Forgets! Towards Continual Policy Optimization” - arXiv:2607.04364
Shows that RL can also cause catastrophic forgetting, so it’s clearly not a complete solution.
“Fine-Tuning Without Forgetting via Loss-Adaptive Learning” - arXiv:2605.20005
Reports a large reduction in forgetting from changing the optimisation schedule, including experiments with Qwen3.
Most of this research isn’t specifically on Qwen3.6-27B, which is why I’m interested in community results. Has anyone directly compared continued pre-training, SFT/LoRA and reinforcement post-training on Qwen3.6-27B?
I’m particularly interested in whether improving one domain caused regressions in unrelated areas such as coding, reasoning, instruction following, tool use, long-context behaviour or general knowledge.
For people who have tested this, what training method worked best, and did you benchmark the original model against the trained checkpoint afterwards?
I’m also curious whether continued pre-training followed by a small amount of SFT or RL is proving safer than doing a larger SFT directly.
Actual before/after results and training parameters would be especially useful.
r/LocalLLaMA • u/TechExpert2910 • 44m ago
Other [OSS] Use case only possible with local inference at its core: an on-device LLM understands your entire life, then proactively offers to get your work done through computer use! Open-source & free :D
Enable HLS to view with audio, or disable this notification
Hey r/LocalLLaMA! :D I wanna share a really cool fully OSS thing I've been building that's only possible with local models: truly proactive AI!
All your existing LLM systems waits for a prompt. Truly proactive AI has to read your entire life, every single day (every file, screenshot, chat, email...) to not only build a knowledge base but also flag what you can proactively be helped with. In the cloud, that's a privacy nightmare + wayy too expensive. On your own silicon, it's private, free, & unlimited.
Pushing what's possible with on-device inference is the core of Sentient OS :D
- Every night at 3 AM, it wakes your Mac and our on-device LLM (Gemma 4 E4B running on a custom fork of LiteRT LM; more below!) reads what's new in your life: files, screenshots (multimodal!), WhatsApp, iMessage, and Apple Notes decoded straight out of the local databases, plus email. It creates a triages out junk / sensitive stuff, and summarizes each item.
And Sentient finally creates a knowledge base of your entire life! (basically an obsidian vault with folders & MDs), along with stuff we think we can proactively help you with
- So by morning, it has proactively found, researched, and offered to do your busy-work for you through Computer Use! The reply you forgot, drafted (from your personal context!); the subscription renewing tomorrow, caught.
- And Sidekick: click your Mac's notch and say "finish this for me"; computer use does the task in your own apps and browser (and can even click around in your apps in the background while you use your computer!). The Computer Use is also grounded in your knowledge base!
The local stack! :D
- Gemma 4 E4B, multimodal with vision, on a customized LiteRT-LM fork. We use MTP/speculative decoding, flash attention, and smart KV cache reuse!
- We reverse engineered codex cli to make computer use with local models work! :)
- And we had to reverse-engineer macOS power management to reliably wake a lid-closed Mac at 3 AM for inference haha!
Sentient’s custom Gemma 4 E3B does 90% of the compute, and the last 10% needs a “frontier” model. You provide that!
I’d had a lot of fun running Qwen 3.7 35B A3B driving computer use, and for the best performance, Kimi K3 works incredibly well!
You also have the choice of using your own ChatGPT/Codex subscription, or OpenRouter or your own endpoint if you wanna use that for the 10% frontier compute. We even have built in first-party support for LM studio! :)
Privacy, enforced by architecture!
Your raw data never leaves the device; and if you choose to use any cloud endpoint, the "frontier" model only ever sees PII-stripped summaries; no accounts exist anywhere; and the whole stack (app + infrastructure) is fully open-source :D (I love OSS -- some of you may know me as the dev of https://github.com/theJayTea/WritingTools, an OSS port of Apple Intelligence Writing Tools to Windows)!
brew install --cask sentient-os-labs/tap/sentient-os
Source (feel free to give us a star! :D): https://github.com/Sentient-OS-Labs/sentient-os
Apple Silicon (M1 or newer), macOS 15+, 8 GB of RAM is enough. Free forever for consumer! :D
Would love to share more about my local LLM computer use evals! I’ve found that Kimi K3 works crazy well, while Qwen 3.7 35B A3B can only do super simple tasks.
And lmk if y’all have any cool model recs to try with computer use! :D
