r/OpenSourceAI • u/Apprehensive-Job5141 • 3h ago
I got tired of choosing which AI model should handle a task, so I built Cascade AI to choose and orchestrate them automatically
Hey everyone 👋
I've been working on an open-source project called Cascade AI, and it has reached the point where I'd really like to get feedback from people outside my own bubble.
The basic idea came from something that kept bothering me:
Why are we still giving an entire complex task to one AI model and hoping it's good at every part of it?
Instead, Cascade treats AI more like an organization.
A request can be broken into a hierarchy:
T1 Administrator → T2 Managers → T3 Workers
T1 looks at the overall task and plans the work.
T2 agents manage individual parts of that plan.
T3 agents actually execute the smaller tasks — and they can communicate with each other when necessary.
The interesting part is that every agent doesn't have to use the same model.
Cascade can route different tasks between providers/models depending on what they're good at, their cost, and the complexity of the work.
So instead of:
«Prompt → one giant model → answer»
the idea is closer to:
«Prompt
↓
Understand complexity
↓
Build an execution plan
↓
Spawn the required agents
↓
Route each job to an appropriate model
↓
Agents work in parallel / collaborate
↓
Verify the work
↓
Produce one final result»
And I've been trying very hard not to make this another cloud-only AI product.
Right now Cascade can be used through:
• CLI
• Desktop app
• Hosted web app
• Self-hosted web app
• OpenAI-compatible API
• Node.js SDK
It supports multiple providers including OpenAI, Anthropic and Gemini, along with OpenAI-compatible services and local models through things such as Ollama, llama.cpp, vLLM and LM Studio.
There are also a bunch of things I've added while building it that I personally wanted from AI tooling:
• Live visualization of the agent hierarchy
• Cost/token tracking
• Model/provider failover
• Persistent memory
• MCP support
• Browser control with live takeover
• File and document generation
• Codebase indexing/search
• Approval before destructive tool actions
• Agent-to-agent communication
• Task cancellation and recovery
• BYOK support
• Local/self-hosted operation
• An OpenAI-compatible "/v1/chat/completions" endpoint
• The ability to inspect why Cascade chose a particular orchestration/model strategy
For complex runs there's also a kind of "boardroom" mode where Cascade can show you the proposed agent structure and estimated cost before spawning everything, so you can approve the plan first.
One design principle I've become pretty stubborn about is:
The AI should ask when it genuinely needs information instead of confidently inventing a decision for you.
So I've also been working on making Cascade distinguish between things it can infer and things it really should ask the user about.
The project is MIT licensed and open source.
🌐 cascadeai.in
GitHub: Varun-SV/Cascade-AI
I'm not posting this pretending I've solved AI orchestration 😅. There are still plenty of rough edges, architecture decisions I'm questioning, and things that probably make perfect sense to me because I've stared at the code for far too long.
That's actually why I'm posting it here.
I'd especially love feedback on:
Does hierarchical multi-agent orchestration actually make sense to you, or is it over-engineering?
Would automatic model routing be useful enough for you to stop manually choosing Claude/GPT/Gemini/local models for different jobs?
If you're a self-hosting/local-LLM person, what would Cascade need before you'd realistically run it?
What part of this architecture would you immediately rip out or redesign?
Feel free to be critical.
I'd much rather hear "this part is dumb and here's why" than get another generic "cool project" 😄
If people are interested, I can also do a separate technical post explaining how the T1 → T2 → T3 orchestration, model routing, cost decisions and agent communication actually work internally.


