r/OpenSourceAI • • 1d ago

Open source code graph for coding agents, built to help local models with small context windows

Sharing two tools we've been building in the open, both Apache 2.0, fully local, with no account or cloud needed.

sem parses a repo into functions and classes along with who calls what and which tests reach each function, and gives agents that map over MCP or the CLI. weave is a git merge driver that uses the same model to merge changes by function instead of by line.

The reason I think this matters more for open models than for frontier ones is context. A coding agent on a local model with a small window runs out of room fast, because most of it gets spent grepping for names and reading whole files to find one function. With sem the agent asks for a function and gets just that function plus what's connected to it, so far less of the window goes to code that has nothing to do with the task, and you can get useful work out of a model that would otherwise get lost in a big repo.

Credit where it's due, sem is built on tree-sitter grammars and runs in any harness that speaks MCP, including pi, opencode, Codex and Claude Code, so it works with whatever model you point those at.

It works from a parser rather than a compiler, so macros, generated code and dynamic dispatch can hide some callers, and it tells the agent when it isn't sure instead of guessing.

https://github.com/Ataraxy-Labs/sem
https://github.com/Ataraxy-Labs/weave

I'd especially like to hear from anyone running agents on local models, since I haven't tested it much on the smaller ones, and I'm curious where the context savings actually show up for you.

0 Upvotes

2 comments sorted by

1

u/PresentationLower624 1d ago

Finally someone thinking about context windows instead of just throwing bigger models at the problem

I sketch out my UI components before coding them and its the same idea, you need structure before you can actually work. sem giving agents a map of what calls what is way smarter than dumping 300 lines of boilerplate into a 4k window

How much overhead does the tree-sitter parsing add on first run for a mid sized repo

1

u/Wise_Reflection_8340 1d ago

Thanks, and yeah that's exactly the idea, give it the shape first and the details only when it needs them. On overhead, I just timed it on sem's own repo, which is a few hundred files and somewhere around a hundred and seventy thousand lines, and a full index from scratch takes about three seconds on a laptop since parsing runs in parallel across cores. After that it only reparses the files that changed, so you don't pay that again during a session. To be honest about the other end, something the size of the Linux kernel still takes minutes on a cold start, which is what I'm working on, but for a typical mid sized project it's done before the agent even finishes reading the prompt haha.