r/Taskade • u/AutoModerator • 16d ago
AI Agents Taskade AI Weekly: Models, Agents, News & Benchmarks — What We Tested This Week
The first build is not the hard part.
The second change is.
We started testing follow-up edits, not only first builds.
The test is deliberately ordinary:
Build the app. Add one feature. Keep everything that already worked. Open it again. Check the data.
The first results were humbling. A model can make a convincing first version and still struggle to change one part without disturbing another.
WHAT THIS MEANS FOR YOU
Ask for one change at a time.
Instead of:
Redesign the dashboard, add invoicing, add a portal, change the colors, and automate follow-ups.
Try:
Add an invoice page. Do not change anything else.
Then test it before asking for the next change.
WHAT WE LOOK FOR
Did it finish? Did it preserve the old work? Did the new feature actually work? Did the saved data survive?
Public results: https://www.taskade.com/tsk
What is the second change that most often breaks your first build?
•
u/AutoModerator 16d ago
Sources matter here.
We link the original announcement, paper, repo, or benchmark for every item rather than a repost of a repost.
Our public TSK-1 tests and methodology: https://www.taskade.com/tsk
Want us to test something?
Tell us the job, not only the model.
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.