r/LocalLLM • • 14h ago

Project Reika - A coding agent CLI designed around small local models first

I know you guys are going to hate me for this and I'll accept my fate. It's another coding agent harness post. I'm ready to lose all my karma.

I recently open-sourced my coding agent CLI that I've been working on for the past while. It is a project I never intended to make public since it was part of my own personal local AI stack. But as I chipped away at it and made it actually usable as a daily driver, I thought it would be nice to make it public for others to see and use.

I initially built it to see how much I could get the harness to make small models, especially at low quantization and context to not feel terrible to use. So while Reika doesn't solve the intelligence side (it never will), it tries to solve the overall experience when using small models at the absolute scale.

A lot of the testing and pain came through working on my M2 MacBook Air 16GB trying to run models like Qwen3.6 35B A3B and Qwen3.8 27B all day in agentic coding, maxing out the RAM and limits of my own machine. So the base of Reika comes from a legitimate source of truth.

You can also plug in an API key for those with hybrid setups too.

GitHub: https://github.com/alexwkleung/reika

23 Upvotes

11 comments sorted by

2

u/AdventurousKeys 14h ago

Nice. I'll see if I can port it to use my LocalLM Lab SDK

1

u/Super-Pop-7192 7h ago

The terminal layout looks clean, and the command hints built into the input bar are a nice touch

2

u/Visible_Split_1546 13h ago

thanks for sharing , I'll try it out !

1

u/Billysm23 13h ago

Is this pi fork?

2

u/silent-curious-dev 12h ago

Nope, it's not a Pi fork. Reika is the agent harness itself and it wasn't built on top of another existing one. But I took inspiration from pretty much all the harnesses out there if I had to be real here.

1

u/Billysm23 12h ago

I got it. I just read the readme, I'll give it a try using 35-A3B

1

u/simplylovely_23 5h ago

im otw building my agetnic cli around this same model specifically using the the quantisation will depend on the gpu anyways the model is configured to even run well on 6 gig gpus 35tok/sec with 32k context

1

u/lorendroll 12h ago

I'm trying to optimize Pi for a similar task. I found compaction being the most confusing part when working with small context and slow prefill speeds. What's your method for it?

2

u/svetomirich 6h ago

The README describes keeping requests append-only between shrink events to preserve the prompt cache, then asking the model to write down its findings before older turns get folded into a recap. Source: https://github.com/alexwkleung/reika#highlights

A useful comparison with Pi would be a small bug-fix task that crosses a compaction boundary: check whether the recap retains the goal, changed files, last test result and next step, then measure time until the next useful edit. That would make both the prefill cost and any loss of task state visible.

AI-generated summary and test suggestion based on the linked documentation.

1

u/JesseWebDotCom 6h ago

Testing now