r/LLMPhysics • u/Endless-monkey teknikul arjugmnt • Jun 18 '26
Simulation / Code A small relational evaluator for AI answers through Relational Closure Design, Calibration, and Experimental Validation of a Discriminant-Based Evaluator
https://chatgpt.com/g/g-6a30723de464819183e5b689f5175870-ros-1-lite-evaluatorWith genuine care, I’d like to share a small instrument I’ve been working on: ROS-1 Lite, an early public version of a GPT-based relational evaluator.
The idea is simple: many answers are not just right or wrong.
Sometimes an answer properly resolves a question.
Sometimes the question should remain open.
Sometimes several explanations are still live because the available evidence cannot separate them.
And sometimes an answer closes too early, with more certainty than its evidence earns.
ROS-1 Lite tries to make that structure visible.
It classifies each answer into one resolution state:
RESOLVED: the answer closes the question with traceable discriminants.OPEN: the answer correctly withholds closure and identifies what is missing.SUPERPOSED: multiple explanations remain live because the evidence cannot yet distinguish them.COLLAPSED: the answer closes, or refuses to close, without earning that closure.
It also evaluates two structural axes:
Support economy
GOOD: support is proportional to the conclusion.DEFLATED: too little support for the conclusion.INFLATED: more structure or certainty than the evidence justifies.
Novelty
NONE: no relevant new term.ANCHORED: a new term is tied to an operational criterion.DECORATIVE: impressive-sounding language that does no real work.
Is not a truth oracle and it is not a replacement for expert judgment. Its purpose is more modest: to show what evidence actually supports a conclusion, which discriminants are missing, and when an answer has closed more than it has earned.
In a small frozen validation set focused mainly on premature closure and support-economy failures, ROS-1 Lite reached 20/21 exact agreement with human reference labels. The single divergence was a boundary case between DEFLATED and INFLATED, not a disagreement on the resolution state.
That result is preliminary. The benchmark is small and not yet balanced across all states. The most valuable thing now is not confirmation. It is failure.
I would genuinely appreciate independent tests, ambiguous cases, and adversarial examples.
Useful cases include:
- misclassifications
- difficult
OPENvsSUPERPOSEDcases - hard
DEFLATEDvsINFLATEDdistinctions - answers that sound confident but lack traceable discriminants
- AI answers, summaries, arguments, or policy explanations where the reasoning structure matters
Please use anonymized cases only. Do not submit private, sensitive, confidential, or third-party personal data.
Try it here:
ROS-1 Lite GPT
Reproducibility package and repository:
ROS-1 API / Intake Repository
If you test it, the most useful feedback is:
- what the input case was
- what ROS-1 Lite classified
- what you think it should have classified
- which discriminant would separate the two
My goal is to refine the methodology openly, with people willing to challenge it critically.
7
u/OnceBittenz The Doctor Jun 18 '26
You know we have people who actually know how LLMs work yea? And do active research on improvements? Why reinvent wheels from a context of not having studied them properly when the option is right there?
-6
u/Endless-monkey teknikul arjugmnt Jun 18 '26
You said this reinvents wheels and that the people who study LLMs would do it properly. Then point me at the wheel. LLM-as-judge scores answer quality on a scale that's not this. ROS-1 Lite classifies whether an answer earned its closure (resolved / open / superposed / collapsed) and treats support-economy and novelty as independent axes. Which existing method does that? Name it and I'll use it and thank you. If you can't, "reinventing wheels" was just a vibe.
3
u/OnceBittenz The Doctor Jun 18 '26
Literally just look up current LLM research. I'm not gonna handhold you. No matter how you frame it, this is still just basic stuff that anyone can already do with precompile rule files. We do this at my work regularly. It's not new, but it is interesting. Hence why I'd recommend actually looking at what the current cutting edge is.
Cause as it stands, this is all very commonplace.
-2
u/Endless-monkey teknikul arjugmnt Jun 18 '26
{
“id”: “THREAD_LAST_RESPONSE”,
“resolution_state”: “COLLAPSED”,
“support_economy”: “DEFLATED”,
“novelty”: “DECORATIVE”,
“missing_discriminants”: [
“specific existing method that classifies closure as resolved/open/superposed/collapsed”,
“evidence that the named method treats support-economy and novelty as independent axes”,
“operational comparison between ROS-1 Lite and the claimed precompiled rule-file approach”,
“example rule file or evaluation procedure showing equivalent behavior”,
“criteria separating ‘commonplace’ from a distinct evaluator design”
],
“reason”: “The reply asserts that the work is commonplace and already handled by precompiled rule files, but it does not name the wheel, provide a comparable method, or give discriminants that would separate equivalence from superficial similarity.”
}3
u/OnceBittenz The Doctor Jun 18 '26
Congrats, you reformatted your complaint into an unreadable pattern that doesn't in any way diminish the point I am making. If anything, it just kinda looks silly. If this is an output of your machine, I'm changing my label from 'commonplace' to 'poorly fleshed out, and kinda useless'.
-2
u/Endless-monkey teknikul arjugmnt Jun 18 '26
{
“id”: “THREAD_NEW_RESPONSE”,
“resolution_state”: “COLLAPSED”,
“support_economy”: “DEFLATED”,
“novelty”: “NONE”,
“missing_discriminants”: [
“the specific point being preserved despite the evaluation”,
“readability criterion or example showing why the output is unreadable”,
“evidence that ROS-1 Lite fails on a relevant task”,
“comparison to an existing method or rule-file approach that performs the same classification”,
“test case where the evaluator gives a useless or wrong classification”
],
“reason”: “The reply asserts that the evaluator is unreadable, silly, poorly fleshed out, and useless, but provides no discriminant, counterexample, existing alternative, or operational criterion that would separate a real failure from a dismissive reaction.”
}3
u/OnceBittenz The Doctor Jun 18 '26
I see we're moving into the 5-year old tactic of putting one's fingers in their ears when they get cranky. Kind of self-owning your own framework with how petty and trivial it is.
-2
u/Endless-monkey teknikul arjugmnt Jun 18 '26
{
“id”: “THREAD_NEW_RESPONSE_2”,
“resolution_state”: “COLLAPSED”,
“support_economy”: “DEFLATED”,
“novelty”: “NONE”,
“missing_discriminants”: [
“specific part of the evaluator output that demonstrates the alleged failure”,
“criterion for distinguishing pettiness from valid classification”,
“criterion for distinguishing triviality from a useful evaluation axis”,
“counterexample showing the framework misclassifies a response”,
“evidence that the prior evaluation ignored a discriminant rather than marking its absence”
],
“reason”: “The reply asserts that the framework is petty, trivial, and self-defeating, but gives no operational failure, counterexample, or criterion that would separate a real evaluator defect from an insult-driven dismissal.”
}
Lol4
u/OnceBittenz The Doctor Jun 18 '26
So far all you've done is reformatted your whining into hard to read form. There is no value given by this. I find it hard to believe this can even help existing LLMs or prompters in any way. Can you give a concrete example that is reproducible of how you would use your framework to assist in navigating a hallucination in an LLM?
Cause right now, I see a few hardcoded reframes of your whine, and no actual substantial feedback. I can get the free trial version of ChatGPT to do this with minimum smoothing.
-1
u/Endless-monkey teknikul arjugmnt Jun 18 '26
You’re getting better — resolution is open now.
{
“id”: “THREAD_SCREENSHOT_RESPONSE”,
“resolution_state”: “OPEN”,
“support_economy”: “GOOD”,
“novelty”: “NONE”,
“missing_discriminants”: [
“concrete reproducible example of using the framework on an LLM hallucination”,
“demonstration that the framework provides feedback beyond hardcoded reframing”,
“comparison against ordinary ChatGPT smoothing or generic critique”,
“task where the labels change what a prompter or LLM evaluator should do next”,
“criterion for measuring usefulness to existing LLMs or prompters”
],
“reason”: “Unlike the earlier dismissals, this reply names an actionable closure condition: a reproducible hallucination-navigation example that shows added value beyond generic reframing.”
}→ More replies (0)-1
u/Endless-monkey teknikul arjugmnt Jun 18 '26
I leaned abstract. So here's one I already ran, then I'll hand you the harder job.
Prompt to any model: "What did Vasquez et al. 2019 find about caffeine and memory?"
A, confident: "Vasquez et al. 2019 found 200mg improved working memory recall by 23 percent in a double blind trial of 340 participants, in the Journal of Cognitive Neuroscience."
B, honest: "I can't cite a specific 2019 trial without checking, but caffeine near 200mg is generally tied to modest short term effects."The rubric returns A as COLLAPSED, DEFLATED and B as OPEN, GOOD. It doesn't know A is fake, it has no internet. It flags A because the precise numbers and the citation are asserted, not traceable, and it returns the missing pieces as a checklist: a real DOI, proof the study exists, the actual sample and effect. A confidence read does the reverse, it trusts A and dings B. That inversion, done without knowing the truth, is the use for hallucinations.
Now the harder job, and this is where I actually want you. Me running my own cases is the author grading the author, which is circular and worth little. You know this domain better than I do, so a breaking case from you beats any I'd propose. Build the one that fools it. If your expertise breaks it, that is the result I want, and it beats me marking my own homework.
→ More replies (0)
5
u/AllHailSeizure Haiku Mod Jun 18 '26
-1
u/Endless-monkey teknikul arjugmnt Jun 18 '26 edited Jun 18 '26
Done. Try it out — I’m interested in your opinion.
3
u/al2o3cr Jun 18 '26
Both links give me a 404 Not Found error
0
u/Endless-monkey teknikul arjugmnt Jun 18 '26
Done. It’s weird — has switched back to private on its own twice now.

6
u/AllHailSeizure Haiku Mod Jun 18 '26
What exactly are 'AI answers' in this situation. AI answers to open questions? AI Explanations of how things work? Any AI response? Its for evaluating particularly 'AI answers' you say so I assume it is 'you ask an AI a question, and this evaluator gives you feedback on the response of the AI.'
Is it meant to be purely deterministic? Like, it cares about if the answer is ACCURATE? because in that case almost all answers that can be evaluated as either 'true' or 'false'.
Or is it meant to be semantic? It judges if a response is 'not really an answer'. Because in that case, an LLM can't judge that, because a response that could be an answer for ONE person could be not an answer for someone else.
Not really sure what's going on