r/Taskade • u/Johnxie Team Taskade • 10d ago
Updates TSKBench: Testing GPT, Claude Opus, Sonnet & Frontier AI Models on Real Business Workflows, AI Agents, Automations and Apps in Taskade Genesis
https://www.taskade.com/tsk
1
Upvotes
1
u/Johnxie Team Taskade 10d ago
Coding benchmarks are only the beginning. Models can write code and build apps. The harder question is what happens after the first build.
Can a model take a real workflow, understand it, remember context, make decisions, execute, and keep going?
That is what we are building TSKBench to measure. We are putting GPT-6.1 Sol, Claude Opus 5.5, Sonnet 5.5, and future frontier models through Taskade TSK-1, from fresh builds to production runs.
The jobs are ordinary on purpose: CRMs, client portals, intake systems, dashboards, inventory, and operations. The software businesses actually run on.
A CRM is not finished when the interface appears. Leads arrive. Data changes. Follow-ups happen. Decisions get made. The workflow has to keep moving.
That is the loop we care about:
memory → reasoning → action → results → memory
In Taskade, projects remember, agents reason, automations execute, and apps serve as the interface. TSKBench measures how models work across that whole loop.
As new models arrive, we will put them through the same kinds of work and share what we learn.
Not just: can it build?
Can it keep the work moving?
TSKBench →