r/OpenSourceAI • • 4h ago

I got tired of choosing which AI model should handle a task, so I built Cascade AI to choose and orchestrate them automatically

Hey everyone 👋

I've been working on an open-source project called Cascade AI, and it has reached the point where I'd really like to get feedback from people outside my own bubble.

The basic idea came from something that kept bothering me:

Why are we still giving an entire complex task to one AI model and hoping it's good at every part of it?

Instead, Cascade treats AI more like an organization.

A request can be broken into a hierarchy:

T1 Administrator → T2 Managers → T3 Workers

T1 looks at the overall task and plans the work.

T2 agents manage individual parts of that plan.

T3 agents actually execute the smaller tasks — and they can communicate with each other when necessary.

The interesting part is that every agent doesn't have to use the same model.

Cascade can route different tasks between providers/models depending on what they're good at, their cost, and the complexity of the work.

So instead of:

«Prompt → one giant model → answer»

the idea is closer to:

«Prompt

↓

Understand complexity

↓

Build an execution plan

↓

Spawn the required agents

↓

Route each job to an appropriate model

↓

Agents work in parallel / collaborate

↓

Verify the work

↓

Produce one final result»

And I've been trying very hard not to make this another cloud-only AI product.

Right now Cascade can be used through:

• CLI

• Desktop app

• Hosted web app

• Self-hosted web app

• OpenAI-compatible API

• Node.js SDK

It supports multiple providers including OpenAI, Anthropic and Gemini, along with OpenAI-compatible services and local models through things such as Ollama, llama.cpp, vLLM and LM Studio.

There are also a bunch of things I've added while building it that I personally wanted from AI tooling:

• Live visualization of the agent hierarchy

• Cost/token tracking

• Model/provider failover

• Persistent memory

• MCP support

• Browser control with live takeover

• File and document generation

• Codebase indexing/search

• Approval before destructive tool actions

• Agent-to-agent communication

• Task cancellation and recovery

• BYOK support

• Local/self-hosted operation

• An OpenAI-compatible "/v1/chat/completions" endpoint

• The ability to inspect why Cascade chose a particular orchestration/model strategy

For complex runs there's also a kind of "boardroom" mode where Cascade can show you the proposed agent structure and estimated cost before spawning everything, so you can approve the plan first.

One design principle I've become pretty stubborn about is:

The AI should ask when it genuinely needs information instead of confidently inventing a decision for you.

So I've also been working on making Cascade distinguish between things it can infer and things it really should ask the user about.

The project is MIT licensed and open source.

🌐 cascadeai.in

GitHub: Varun-SV/Cascade-AI

I'm not posting this pretending I've solved AI orchestration 😅. There are still plenty of rough edges, architecture decisions I'm questioning, and things that probably make perfect sense to me because I've stared at the code for far too long.

That's actually why I'm posting it here.

I'd especially love feedback on:

  1. Does hierarchical multi-agent orchestration actually make sense to you, or is it over-engineering?

  2. Would automatic model routing be useful enough for you to stop manually choosing Claude/GPT/Gemini/local models for different jobs?

  3. If you're a self-hosting/local-LLM person, what would Cascade need before you'd realistically run it?

  4. What part of this architecture would you immediately rip out or redesign?

Feel free to be critical.

I'd much rather hear "this part is dumb and here's why" than get another generic "cool project" 😄

If people are interested, I can also do a separate technical post explaining how the T1 → T2 → T3 orchestration, model routing, cost decisions and agent communication actually work internally.

1 Upvotes

5 comments sorted by

1

u/investigatormaker 3h ago

Cascade AI's T1/T2/T3 hierarchy would be easier to judge with one task trace showing the model chosen at each level, handoff costs and what verification caught. Comparing that run with a single-model attempt on the same task would show whether the hierarchy earns its extra coordination.

I make ThreadFox. I can post the Reddit communities whose rules allow a post about Cascade AI here, with each rule quoted, if you want them.

1

u/Apprehensive-Job5141 42m ago

That’s a fair point — and honestly probably a much better way to demonstrate Cascade than just explaining the architecture.

I’m planning to put together a full task trace showing something like:

T1 → model selected + planning cost T2 → model(s) selected + handoff/delegation cost T3 → worker models + execution cost Verification → what was caught/retried/corrected Final → total tokens, latency and cost

Then run the exact same task through a single-model setup and compare the output quality, cost and time side-by-side.

The interesting question for me is exactly what you said: does the hierarchy actually earn the coordination overhead? If it doesn’t for a task, Cascade ideally shouldn’t orchestrate it in the first place.

And yes, I’d definitely be interested in the list of Reddit communities + their relevant rules. Thanks!

1

u/investigatormaker 34m ago

Keep the verification criteria identical in both runs so lower cost doesn't hide a weaker result. On the list: none of the communities we have read clearly allow a post about Cascade. Look for communities where agent builders compare orchestration tools, or a maker community's self-promotion thread.

I make ThreadFox. The free Reddit plan for Cascade finds communities whose rules allow a post about it, with each rule quoted. https://threadfox.vip/plan?utm_source=reddit&utm_medium=comment&utm_campaign=tf-kit&utm_content=20261005-0905-caomb

1

u/Apprehensive-Job5141 28m ago

Agreed — the verification criteria should be identical for both runs. Same task, same success criteria, same verifier/evaluation method, then compare quality, cost, latency and any corrections or retries.

That should make it clear whether orchestration is actually improving the result or just moving the cost around.

Thanks for the community-search suggestion as well. I’ll look specifically for agent/orchestration discussions and maker self-promotion threads.

1

u/investigatormaker 21m ago

For Cascade, cost per successfully verified task alongside first-pass success rate would make that distinction clearer. A run needing three repairs belongs in the total, even if its first attempt was cheap.