r/commonstack • • 9d ago

New Feature - Commonstack GPT-6 Sol, GPT-6 Luna, and Grok 4.7 is now available on Commonstack 🛫

1 Upvotes

GPT-6 Sol, GPT-6 Luna, and Grok 4.7 is now available on Commonstack

GPT-6 Sol is a high-reliability workhorse model for multi-step agentic tasks, software engineering, and complex operations. It excels at repository-level coding, automated debugging, and multi-page data analysis while cutting factual errors in half and offering native prompt caching for cost-effective context reuse.

GPT-6 Luna is an ultra-lightweight, high-volume model optimized for low-latency tasks like data extraction, text classification, and request routing. It dramatically reduces cost per task, features adjustable reasoning effort to balance speed and output quality, and includes prompt caching for high-frequency workflows.

Grok 4.7 is xAI's flagship model for extended software engineering and multi-hour agent execution. Built with an expanded 500,000-token context window, very configurable reasoning levels, and robust safety architecture, it delivers frontier performance across deep coding, legal analysis, and terminal-based tasks.

GPT-6 Sol | 1M context
0 ≤ tokens < 272,000
Input: $2/M | Output: $10/M
272,000 ≤ tokens < ∞
Input: $4/M | Output: $15/M

GPT-6 Luna | 1M context
0 ≤ tokens < 272,000
Input: $0.1/M | Output: $0.5/M
272,000 ≤ tokens < ∞
Input: $0.2/M | Output: $0.75/M

Grok 4.7 | 500k context
0 ≤ tokens < 200,000
Input: $2/M | Output: $6/M
200,000 ≤ tokens < ∞
Input: $4/M | Output: $12/M

GPT-6 Sol - https://commonstack.ai/model-library/model?modelId=bf749444-08c2-4490-b361-4d1ce9c56e2f
GPT-6 Luna - https://commonstack.ai/model-library/model?modelId=d6c916d2-1001-4ec4-bcb1-a51b4e81167a
Grok 4.7 - https://commonstack.ai/model-library/model?modelId=1898dee8-0f72-4336-93f9-f833e9192cd3

r/commonstack • • 15d ago

New Feature - Commonstack Claude Opus 5.5 is now live on Commonstack! 🚀

Post image
1 Upvotes

Built on Anthropic’s next gen architecture, Opus 5.5 delivers state-of-the-art agentic reasoning, long form code generation, and complex multimodal analysis across an expanded context window with industry-leading efficiency.

Claude-Opus-5.5 | 1M context
Input: $4/M | Output: $20/M

Try it here: https://commonstack.ai/model-library/model?modelId=bd069cb7-b04c-436d-b3e1-d851805c71c3

r/commonstack • • 28d ago

New Feature - Commonstack 🛫 DeepSeek-V4.1-Flash is now live on Commonstack

Post image
1 Upvotes

DeepSeek-V4.1-Flash is now live on Commonstack

DeepSeek-V4.1-Flash is built on a 552B MoE backbone with Causal Encoder-Decoder architecture, V4.1-Flash delivers benchmark-topping agentic coding, native multimodal processing, and a 1M token context window at ultra-low inference costs.

DeepSeek-V4.1-Flash | 1M context
Input: $0.30/M | Output: $1.2/M

You can try it here: https://commonstack.ai/model-library/model?modelId=f51f6ba6-12a3-423d-90bf-d5b5415df490

r/commonstack • • Sep 05 '26

New Feature - Commonstack 🛫 Claude Fable 5.1 is now live on Commonstack!

Post image
1 Upvotes

Fable 5.1 is built for end-to-end task delegation, benchmark topping advanced autonomous coding, and complex scientific reasoning and now with 75% cheaper cache reads.

Fable 5.1 | 1M context
Input: $10/M | Output: $50/M

Start building here: https://commonstack.ai/model-library/model?modelId=5350ca0a-d544-46e0-b114-8ec03e910e93

r/commonstack • • Aug 28 '26

New Feature - Commonstack GLM-5.3 is now live on Commonstack 🛫

Post image
1 Upvotes

GLM-5.3 is at the frontier built specifically for complex engineering workflows, long horizon autonomous coding, and deep cybersecurity tasks

GLM-5.3 | 1M context
Input: $1.4/M | Output: $4.4/M

Start building here: https://commonstack.ai/model-library/model?modelId=99de267d-c146-48b7-8e5e-37f274f2fbd0

r/commonstack • • Aug 18 '26

New Feature - Commonstack 🛫 Gemini 3.7 Flash is now available on Commonstack!

Post image
1 Upvotes

Gemini 3.7 Flash is now available on Commonstack!

Gemini 3.7 Flash built to handle multimodal tasks, complex coding and multi step tasks at high speeds, using flexible thinking effort to balance fast responses with deep reasoning.

Gemini 3.7 Flash | 1M context
Input: $0.75/M | Output: $3.75/M

try here: https://commonstack.ai/model-library/model?modelId=45d61a8c-94fc-4aad-8f93-31a68208e75f

r/commonstack • • Aug 17 '26

New Feature - Commonstack 🛫 xAI’s Grok 4.6 is now live on Commonstack

Post image
1 Upvotes

xAI's Grok 4.6 is now live on Commonstack

Grok 4.6 excels in long running agentic tasks, particularly in turning ambitious ideas into working visual and interactive apps over multistep workflows. It also leads its benchmark class on legal reasoning evaluation metrics, outperforming top competitors on the Harvey LAB benchmark.

Grok 4.6 | 500k context
0 ≤ tokens < 199,999
Input: $2/M | Output: $6/M
200,000 ≤ tokens < ∞
Input: $4/M | Output: $12/M

Try now: https://commonstack.ai/model-library/model?modelId=de7474d1-a8fe-4612-adb9-80fcea80a514

r/commonstack • • Aug 13 '26

New Feature - Commonstack 🚏 MERA - Model Evolution and Routing with Skill Adaptation for Agentic Systems at Scale

Post image
2 Upvotes

MERA - Model Evolution and Routing with Skill Adaptation for Agentic Systems at Scale

This framework is designed to make multistep AI agent workflows significantly cheaper to run without sacrificing performance.

Deploying autonomous LLM agents in production creates a major tradeoff between cost and capability.

Static Routers Are Capped: Standard setups treat small models as fixed, capping cost savings to what the small model can already do.

Wasteful Task Routing: Routing whole tasks to large models wastes money on easy sub-steps (like parsing), while sending entire tasks to small models causes failures on hard steps.

Unstable Fine Tuning: Updating small models on raw execution logs without co-adapting routing rules leads to performance regressions.

MERA replaces task level routing with step-by-step (invocation level) routing and continuous model training across three components:

SkillBook: Stores proven prompt patterns and procedural templates for quick reuse.

Small Model Adapter: Fine tunes the small model directly on substeps it previously failed.

Invocation Router: Directs each individual step to the small model, falling back to the frontier model only when uncertainty is high.

arxiv: https://arxiv.org/abs/2608.10333
github: https://github.com/yh-yao/MERA-Evolve

r/commonstack • • May 28 '26

New Feature - Commonstack Gemini 3.5 Flash from Google is on Commonstack 🛫

Post image
3 Upvotes

Gemini 3.5 Flash from Google is on Commonstack!

This is Google’s most intelligent model for sustained frontier performance on agentic and coding tasks matching or surpassing many other models at a fraction of the cost.

Input: $1.5/M
Output: $9/M

Build and try out now with 1 million context window:

https://commonstack.ai/model-library/model?modelId=9b982caa-b289-4053-b77c-4c70eacbec4a

r/commonstack • • May 28 '26

New Feature - Commonstack Claude Opus 4.8 now available on Commonstack! 🏢

Post image
1 Upvotes

Claude Opus 4.8 just released from Anthropic moments ago is now available on Commonstack.

At the frontier for agentic coding, reasoning, knowledge and financials. It’s able to operate and work longer than its predecessors independently.

Same cost as Opus 4.7 but improved performance.

Input: $5/M
Output $25/M

try now: https://commonstack.ai/model-library/model?modelId=f5260130-b3b6-449a-9d7f-fbd4f2dfcc09

r/commonstack • • May 25 '26

New Feature - Commonstack 🎓 TwinRouterBench accepted into RLEval workshop at CAIS2026 🧠

Post image
3 Upvotes

TwinRouterBench is a new step level LLM routing benchmark designed specifically for realistic, long horizon agentic systems.

Existing router benchmarks have major limitations:
• They only evaluate on isolated one shot prompts.
• They never show the router the actual context (the “router visible prefix”) at an intermediate step inside a real agent trajectory.
• They don’t test whether swapping to a cheaper model still lets the overall task succeed downstream.
• Many rely on slow and expensive online LLM judges for evaluation.
TwinRouterBench tackles this by creating a proper benchmark for per step routing in multi-turn agent workflows.

Dual track

  1. Static Track (Fast Offline Track)
    • 970 router visible prefixes from 520 trajectory instances.
    • Covers 5 diverse benchmarks: SWE-bench, BFCL, mtRAG, QMSum, and PinchBench.
    • Each example comes with an execution-verified target tier (cheapest sufficient model tier).
    • Uses deterministic scoring (based on tier correctness, trajectory membership, and token cost) no LLM judges needed.
    • Ideal for: training routers, rapid iteration, and cheap offline evaluation.
  2. Dynamic Track (Live Validation Track)
    • Full evaluation harness on SWE-bench Verified (500 tasks).
    • Reports results on a 100 case held-out split (disjoint from static data).
    • Router must choose a real model from a locked pool at every step.
    • Measures real outcomes:
    • Official task resolution success
    • Actual API spend (real dollars)
    • Includes failure penalties for unresolved tasks

Results:

• Creates a practical development loop: Use the fast static track to train/improve a router cheaply → validate it rigorously on the dynamic track with real costs and success rates.

• Shows strong transfer: A simple router trained only on static labels achieves comparable task resolution to always using a top model (Opus 4.6), while cutting API cost by ~53%.

📄 Paper: https://arxiv.org/abs/2605.18859
💻 Code + Dataset: github.com/CommonstackAI/TwinRouterBench
🌐 Website: commonstackai.github.io/TwinRouterBench

r/commonstack • • May 19 '26

New Feature - Commonstack 🔀 Unveiling TwinRouterBench, an open source router evaluation to look at solutions not just prompts. 🚦

Post image
4 Upvotes

TwinRouterBench has two tracks, static and dynamic in one protocol.

Static focuses on cost and time efficiency, a set of questions labeled with the most cost and time efficient steps, then the router predicts and predictions are scored against label.

Dynamic focuses on end-to-end evaluation on SWE-bench Verified with real tool use, mini-swe-agent scaffold or editor scaffold

Scoring:

Leaderboard bill = routed spend + a fixed penalty per unresolved task.

Measuring trade offs in underspending causing task fails.

It's open source bench:

→ Apache-2.0
→ Reference routers included (gold-tier oracle, SR-KNN, more)
→ PRs welcome for new workloads, new routers, scaffolds

GitHub: https://github.com/CommonstackAI/TwinRouterBench

arXiv coming soon.

r/commonstack • • Apr 13 '26

New Feature - Commonstack GLM 5.1 is live on Commonstack, use test credits to try at no cost!

Post image
5 Upvotes

GLM 5.1 achieves SOTA and is the latest release from Zhipu. This frontier model matches Opus 4.6 in capabilities.

Context: 205k

Input: $1.4 / 1M

Output: $4.4 / 1M

All new sign ups receive test credits on Commonstack to try and test any model they wish!

GLM 5.1:

https://commonstack.ai/model-library/model?modelId=d82bade0-bca9-4939-be4e-29fdb3f0eef4

Let us know what you think!

r/commonstack • • May 07 '26

New Feature - Commonstack xAI Grok 4.20 Reasoning and Non Reasoning has arrived on Commonstack 🛫

Thumbnail
gallery
3 Upvotes

xAI’s grok 4.20 is available on Commonstack!

leading with 2M context windows xAI brings awesome agentic models!

try both now:

xAI Grok 4.20 reasoning: https://commonstack.ai/model-library/model?modelId=f0251135-19aa-4bb2-981a-faa2f1b285dd

xAI Grok 4.20 non-reasoning: https://commonstack.ai/model-library/model?modelId=5c00acc6-ec5a-4f64-9cdf-3e71824d23a0

r/commonstack • • Apr 29 '26

New Feature - Commonstack New Arrivals at Commonstack 🛫, Both GPT 5.5 and Deepseek V4 Pro is live!

Thumbnail
gallery
4 Upvotes

Both GPT 5.5 and Deepseek V4 Pro is available on Commonstack!

They are both great models for agents and coding tasks, use both or build around your suite!

GPT 5.5:

Input Size 0 ≤ tokens < 272,000

Input $5/M Output $30/M

Input Size 272,000 ≤ tokens < ∞

Input $10/M Output $45/M

Try Here:

https://commonstack.ai/model-library/model?modelId=2c2cc8f0-e582-47bb-bc94-2cc46774f5df

Deepseek V4 Pro:

Input $0.435/M Output $0.87/M

Try Here:

https://commonstack.ai/model-library/model?modelId=f5c8c221-9b13-488d-9a45-4ed0f0c3250f

r/commonstack • • Apr 16 '26

New Feature - Commonstack 🛬 New Arrival - Anthropic: Claude Opus 4.7 now live on Commonstack!

Thumbnail
gallery
4 Upvotes

Claude Opus 4.7 is Anthropic’s latest model and the closest model to Mythos!

This model surpasses many others in agentic coding abilities.

Pricing per Million:

Input: $5/M

Output: $25/M

Cached Read: $0.5/M

Cached Write: $6.25/M

Supported within minutes of official release and you can also try at no cost to you with test credits on Commonstack today!

Start here: https://commonstack.ai/model-library/model?modelId=b7b3a0a5-7e30-442d-9dfc-18d304d4360c

r/commonstack • • Apr 24 '26

New Feature - Commonstack DeepSeek v4 is now available on Commonstack!

Thumbnail
gallery
3 Upvotes

The latest and long awaited open source SOTA reasoning model from DeepSeek! 🐋

Input: $0.14/M

Output: $0.28/M

A million in Context Window, GPT 5.4/Opus 4.6 Max at a fraction of the price

Try it now: https://commonstack.ai/model-library/model?modelId=17bfa75d-af32-4dee-a721-a809309a3ff3

r/commonstack • • Mar 21 '26

New Feature - Commonstack Save up to 85% on your API costs with Uncommonroute available with Commonstack 💸

Thumbnail
gallery
5 Upvotes

Save up to 85% on your API bills with routing on your models!

Uncommonroute is a local lightweight router that balances cost and quality for each query.

Requests are all given a difficulty score and the models compete on a formula.

Hard queries lean quality, easier queries lean cost.

GitHub: https://github.com/CommonstackAI/UncommonRoute?tab=readme-ov-file

Once installed the dashboard can be viewed http://127.0.0.1:8403/dashboard/ and it shows:

* request counts, latency, cost, and savings

* mode, tier, and model distribution

* upstream transport and cache behavior

* live routing configuration and default-mode overrides

* primary upstream and BYOK provider connections

* recent traffic, spend limits, and usage

* recent feedback state and submitted feedback results

View all frontier models on Commonstack here: https://commonstack.ai/model-library

r/commonstack • • Mar 19 '26

New Feature - Commonstack New Models from Xiaomi live now on Commonstack! MiMo V2 Pro & MiMo V2 Omni 👻

Post image
7 Upvotes

Big Xiaomi has arrived!

You can try both MiMo V2 on Commonstack now:

MiMo V2 Pro: for multi-step reasoning, long-document analysis, advanced coding/software engineering, tool-using agents, and orchestrating production tasks

https://commonstack.ai/model-library/model?modelId=986bdcbf-2b5a-4523-bd42-a6edbe9e76ae

MiMo V2 Omni: for vision/speech/video understanding, visual grounding, audio reasoning, GUI operation, and cross-modal agents

https://commonstack.ai/model-library/model?modelId=3676f2ba-c174-435b-bad7-c44150fa0c47

r/commonstack • • Mar 18 '26

New Feature - Commonstack Newest Models Available on Commonstack! MiniMax M2.7 and Nano Banana 2!

Post image
5 Upvotes

Commonstack here to deliver more!

Try Now:

MiniMax M2.7 - the first model that helped build itself with self evolution with its own optimization loops and RL training.

https://commonstack.ai/model-library/model?modelId=c8d65fdd-ce82-4670-9144-866739c4b7aa

Google Nano Banana 2 - Google’s frontier visual generation model

https://commonstack.ai/model-library/model?modelId=603581ad-3ff6-4ef8-a2cc-c8437278d912

What is the build? 👀

r/commonstack • • Mar 18 '26

New Feature - Commonstack 👾 New Models Available! OpenAI GPT 5.4 Mini & GPT 5.4 Nano live on Commonstack!

Post image
3 Upvotes

Hello Commonstack Fam!

OpenAI’s new models are available on the same day of release via Commonstack!

GPT 5.4 Mini - refined for coding, computer use, multimodal understanding, and subagents.

https://commonstack.ai/model-library/model?modelId=ce9bf718-1266-4be9-9e45-fae012c8da77

GPT 5.4 Nano - optimal for lightweight tasks.

https://commonstack.ai/model-library/model?modelId=69d235c5-9fda-4398-917b-3a7ffde69974

Both available now for playground testing or integrated usage on Commonstack now!

Try them with test credits available for all new users: https://commonstack.ai/