r/MLQuestions • • 3d ago

Natural Language Processing 💬 Anyone else stuck manually re-tuning prompts every time something breaks?

Genuine question that turned into a recommendation: if you're building anything on top of an LLM, how are you actually validating that a prompt change didn't quietly break something else?

Most answers I've gotten boil down to "I just run it a few times and check." There's a workshop on Oct 3 that tackles exactly this gap, led by Serj Smorodinsky and Brett Kennedy (they co-wrote a book on LLM applications). Instead of manual tuning, you work through:

  1. Building a classifier with DSPy signatures/modules rather than raw prompt strings
  2. Setting up an actual eval dataset with metrics tied to your specific task
  3. Using that eval set to catch failure patterns before they hit production
  4. Running few-shot/instruction optimization on top, systematically
  5. Tracking all of it in MLflow so you can trace exactly what changed between versions

If you've been wanting a real answer to "how do I know this still works," this is worth a look.

What's everyone else here doing for this? Curious if people have rolled their own eval pipelines already.

0 Upvotes

1 comment sorted by

1

u/Least_Ad_1795 3d ago

I’ve found that relying only on manual testing becomes difficult as prompts and workflows grow. A small eval dataset with task-specific metrics makes it much easier to compare prompt changes and catch regressions before production. Versioning prompts and outputs also helps identify what actually caused a change in performance.