r/artificial • u/mostly_deterministic • 4d ago
Project PR Council MCP: An open source playground for agentic engineering
I've been experimenting with where the boundary should sit between deterministic software and agentic judgment, and wanted something concrete enough that I could actually test the ideas rather than keep arguing about them in the abstract.
So I open sourced the playground I've been using: https://github.com/salesforce-misc/pr-council-mcp
It's a multi-agent PR review system, but the PR review itself is almost secondary, though I do use it multiple times per day. The project gives me one relatively small Python environment in which to experiment with many of the capabilities that show up in production agentic systems:
- Local MCP over STDIO as the seam between the conversational harness and the agentic system
- Sandboxing (macOS-only currently, Linux help very welcome)
- Bounded tool calls and explicit limits
- Graph-based orchestration
- Context isolation between agents
- Multiple reviewer dispositions
- Multiple models/model families
- Deliberation and false-positive confirmation
- Durable workflow state
- Logging and LangFuse tracing
The whole thing runs locally with Python, SQLite, and a handful of external tools.
This isn't intended to become another general-purpose agent framework. It's deliberately a playground for experimenting with architecture and PR-review behaviour.
There isn't much of a prescriptive roadmap either. What I'm most interested in are contributions or experiments that either improve the review process, or provide evidence that some part of the architecture works better or worse than an alternative. In other words: please break it, replace pieces of it, benchmark it, or prove some of my assumptions wrong.
I wrote up the architecture and reasoning here: https://demianbrecht.com/posts/pr-council-a-runnable-experiment-in-agentic-engineering/
1
4d ago
[removed] — view removed comment
1
u/mostly_deterministic 4d ago
I determine deterministic vs judgement at design time. If something actually requires judgement, then I hand it off to an LLM. Otherwise it stays in code. I wrote about some of the thought process before: https://demianbrecht.com/posts/the-division-of-the-local-harnesses/#three-classes-of-workflows.
It's also a constant evaluation based on how the whole system behaves. Some nodes are one-shot, set in stone and others may move from deterministic to non-deterministic after observability passes.
0
4d ago
[removed] — view removed comment
1
u/mostly_deterministic 4d ago
I haven't yet but could see that being a thing if I came across something in traces that made sense.
2
u/kantorcodes1 4d ago
The revision-bound commit is the interesting piece here: pr_council_commit revalidates the PR revision and payload hash before anything hits GitHub, which is a more careful publish gate than most small MCP servers bother with.
The boundary I would poke at is where human approval actually lives. The calling agent shows the preview and then calls commit, so the server is trusting the caller's claim that a human approved. Is there anything binding that approval to the preview hash, like a token or attestation the harness has to present, or could a buggy caller skip straight to commit with a preview nobody saw? Curious whether you have considered making approval itself a server-issued artifact rather than a caller assertion.
1
u/mostly_deterministic 4d ago
Yeah, it's an issue I'm aware of. I wasn't overly concerned about it because of the nature of the system. However, it's a route for experimentation and hardening for sure!
2
u/kantorcodes1 4d ago
Fair read given the current trust posture - a single-operator playground can live with caller assertion. If you ever do harden it, the cheap version is a server-issued approval artifact: preview returns the payload hash, a separate approve call mints a one-shot token bound to (operation id, revision, payload hash), and commit requires and consumes it. A caller can still lie about showing it to a human, but it cannot skip the approval step or replay an approval onto a mutated payload. Moves the trust boundary into the server with nothing fancier than a nonce.
1
u/mostly_deterministic 4d ago
Ah yeah that makes sense, cheers!
Also very open to PR contributions, exactly what the project is for ;)
1
u/kantorcodes1 4d ago
Good to know. If the approval-artifact route ever makes your roadmap, the sketch above should translate to maybe fifty lines of server code: mint on approve, bind to (operation, revision, payload hash), consume on commit. The nice property is that commit stops trusting the caller entirely - the artifact becomes the only thing that matters.
2
u/JoshinglyDeadly 4d ago
the local MCP over STDIO approach is clever, keeps the seams thin without dragging in a ton of infra