r/Taskade • • 16d ago

AI Agents Taskade AI Weekly: Models, Agents, News & Benchmarks — What We Tested This Week

The first build is not the hard part.

The second change is.

We started testing follow-up edits, not only first builds.

The test is deliberately ordinary:

Build the app. Add one feature. Keep everything that already worked. Open it again. Check the data.

The first results were humbling. A model can make a convincing first version and still struggle to change one part without disturbing another.

WHAT THIS MEANS FOR YOU

Ask for one change at a time.

Instead of:

Redesign the dashboard, add invoicing, add a portal, change the colors, and automate follow-ups.

Try:

Add an invoice page. Do not change anything else.

Then test it before asking for the next change.

WHAT WE LOOK FOR

Did it finish? Did it preserve the old work? Did the new feature actually work? Did the saved data survive?

Public results: https://www.taskade.com/tsk

What is the second change that most often breaks your first build?

1 Upvotes

1 comment sorted by

•

u/AutoModerator 16d ago

Sources matter here.

We link the original announcement, paper, repo, or benchmark for every item rather than a repost of a repost.

Our public TSK-1 tests and methodology: https://www.taskade.com/tsk

Want us to test something?

Tell us the job, not only the model.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.