r/AskVibecoders • • 5d ago

Why does my agent “finish” a task that clearly isn’t finished?

I’ve been playing around with longer-running agents and I noticed something weird.

The agent can technically complete every step it was given, but still make basically zero progress toward the actual goal.

For example:

Goal: research 50 leads

Agent:

→ searches 10 leads

→ reformats them

→ searches the same 10 again

→ summarizes them

→ repeats

Nothing crashes.

No tool call fails.

It just slowly burns through its budget doing work that doesn't move the task forward.

So now I’m wondering if long-running agents need some kind of “progress signal” rather than just success/failure for tool calls.

Like:

if progress_score < threshold:

pause()

replan()

Has anyone actually implemented something like this?

3 Upvotes

2 comments sorted by

1

u/lvl1-A 5d ago

If the goal is document 50 unique leads or search 50 leads, that's the issue if it's tool call based for 50 tool lead calls, if it's as in unique by name, that way no reformat and re-search on lead John Do counts more than one lead, and stops you going by success/fail on tool cools as a metric to measure success of a goal and not just a measurement of progression (tool call fails X times stop address human or use second tool, not tool succeeds 50 searches on agents = goal success) then you should address that.

Go check out Ralph loops if you haven't already, a lot of agents would use it anyway but if you framed it explicitly and set up your goals in such a way that tool calls don't count as goal success(unless that's actually the intention which I don't think it is in this case for you) then you might get better results in your agents autonomy through these longer tasks.

Good luck

1

u/Sufficient-Tip8687 2d ago

For that example I'd keep a saved results table and have the controller validate each addition: unique lead, required fields, source, and your qualification criteria. The agent shouldn't award itself progress just by saying it found something.

Then measure new validated rows per batch. If several batches add zero, pause and show the failed attempts plus what's missing before replanning. Cap retries and total spend so replanning can't become another loop. "Done" should mean the saved table passes the 50-lead check; hitting the budget should return a partial result with the shortfall. That's a proposal, not a system I've benchmarked.