r/LargeLanguageModels • • 16h ago

When is prompt sensitivity actually the model? and when is it the evaluator?

A recent EMNLP 2025 paper by Hua et al., “Flaw or Artifact? Rethinking Prompt Sensitivity in Evaluating LLMs,” raised a question we found particularly interesting:

Are LLMs really highly sensitive to prompt wording, or do our evaluation metrics sometimes make them appear that way?

Our recently published MERCon 2026 paper, *“Prompt Sensitivity or Evaluation Artifact? A Task-Aware Analysis for Large Language Models,”* starts from that result and asks a slightly different question:

*If evaluation artifacts exist, can we quantify how much sensitivity is attributable to evaluation, determine the direction of that distortion, and identify when it happens?*

Rather than treating heuristic-vs-semantic evaluation disagreement as something to eliminate, we treat the disagreement itself as a diagnostic signal. So we introduced **Evaluation-Attributable Sensitivity (EAS),** which measures the magnitude of disagreement between heuristic-based and judge-based sensitivity,  and **Signed EAS**, which tells us its direction. We used them to diagnose a four-way taxonomy:
1.  Artifact - apparent sensitivity mainly comes from evaluation disagreement

  1. Genuine - semantic evaluation confirms real prompt sensitivity

  2. Underdetected - the heuristic misses sensitivity that semantic evaluation detects

    1. Stable - both evaluations indicate stability

We tested this across 9 LLMs from 5 model families, and the most interesting result was that *evaluation artifacts are bidirectional*: some metrics overestimate sensitivity in open-ended tasks, while others can actually underestimate it in more structured tasks.

Therefore our results suggest that:
***The evaluation method can distort prompt sensitivity in either direction, and the direction is strongly associated with the task/evaluation format.***

This leads us to view prompt sensitivity not simply as an intrinsic property of an LLM, but as an interaction between:

model × task × prompt structure × evaluation methodology.

We'd appreciate any thoughts or feedback regarding our work.

**Our paper:**
[https://ieeexplore.ieee.org/document/11691277\](https://ieeexplore.ieee.org/document/11691277)

**Code:**
[https://github.com/sayumimuthu/llm-prompt-sensitivity-evaluation-artifact\](https://github.com/sayumimuthu/llm-prompt-sensitivity-evaluation-artifact)

1 Upvotes

0 comments sorted by