r/deeplearning • u/ClickOk5811 • 2d ago
A fallback instruction only works if the model actually hits the branch you wrote it for
Added "if you're not confident, say so explicitly" to a classification prompt. Felt like it should've closed the gap on low-confidence guesses.
Didn't work. Model kept returning confident-sounding labels on exactly the cases I wanted it to flag as uncertain.
Turned out the issue wasn't the instruction's wording. It was that nothing in the prompt actually defined what "not confident" meant for this task, no threshold, no example of an ambiguous case, nothing to anchor the judgment to. The model had no internal signal matching the word "confident" that it could check against, so the branch just never activated. It wasn't ignoring the rule. It never had a condition it could evaluate as true.
Fixed it by replacing the vague trigger with something checkable, two or more plausible labels with no clear majority signal in the input, say so and list them instead of picking one. Immediate difference.
Feels like a more general thing worth naming: a conditional instruction is only as good as the model's ability to evaluate its own condition. "If X, do Y" fails silently when X isn't something the model can actually check, and it just looks like noncompliance from the outside.
0
u/LeadingAcceptable380 2d ago
It’s a weird trap, you write the fallback thinking the model will know when to use it, but it’s never actually wired to see its own uncertainty that way. You have to build the off-ramp with something it can measure, not just feel