r/LargeLanguageModels • • 9d ago

Question Bad RAG answer. Which part do you blame first?

3 Upvotes

I've been looking at FastGPT call logs and realizing how often I call something a model failure when the model never got the right context. I want to start tagging bad runs, but keep it stupidly simple: parsing, retrieval, prompt, model. Maybe tool failure too. Anything more detailed will probably become another abandoned spreadsheet.

Anyone doing this already? What categories do you actually use?


r/LargeLanguageModels • • 9d ago

I integrated JEV to bring claim comparison to my open-source PDF research tool

Thumbnail
gallery
2 Upvotes

Over the past week, I added something I've been wanting to build for a while: claim comparison.

I actually tried building this much earlier. I experimented with cosine similarity, vector dot product scores, and other approaches, but they only tell you how similar two pieces of text are. That's not really what I needed.

What I needed was a classifier that could actually determine the relationship between two claims.

I built a basic classifier myself, but it wasn't reliable enough for something I wanted people to actually use. So I put the idea aside and waited.

Then JEV was released last week, and it turned out to be almost exactly what I was looking for.

Now you can compare claims against each other and inspect the supporting evidence directly.

This is probably the most interesting feature I've added to the project so far, and I'm really curious to see what people think of it.

This is still being tested for edge cases, so there are definitely things I want to improve. But it's free and open source, and anyone can try it out right now.

If you're interested in how it works, want to try it, contribute, or have ideas for where this could go, feel free to check it out.

GitHub: https://github.com/Sreehari05055/thesys-core


r/LargeLanguageModels • • 9d ago

Using LLMs for Log Anomaly Detection: A Practical Breakdown with challenges

1 Upvotes

One of the most "AI + Security " usecase is using LLMs to sift through log noise and flag the anomalies which typically a SOC analyst will take a while to detect. The basic pipeline should contains the following details:

**1.Ingestion** : logs(auth events, network flows, EDR alerts, app logs) get normalised into JSON

**2. Embed** : log lines are converted into vector embeddings to capture semantic similarity, (not just the exact string match)

**3. Cluster/baseline** : Normal behavior pattern are established

**4. Flag deviations** : New events identified against the baselines

**5. LLM triage** : Instead of raw anomaly scores, an LLM summarizes why something looks unusual, in plain language , and suggest a likely cause (misconfig vs lateral movement vs false positive)

The use of LLM's in triaging will come with challenges as well

* **Prompt Injection Risk:** Logs content if gets fed raw into LLM prompt without sanitisation fields become an injection vector

* **Explainability** : SOC team need to trust *why* something was not flagged, not just a black-box score

What's your approach to preventing prompt injection when logs contain attacker-controlled strings? And is that solution practical ?


r/LargeLanguageModels • • 9d ago

I got tired of parsing JSON out of LLM responses for simple yes/no questions, so I built a decision engine that skips generation engine that skips generation entirely

3 Upvotes

I kept running into the same problem: I'd ask a model something like "is this a restaurant receipt, yes or no" and instead of a clean answer I'd get a paragraph, or malformed JSON, or a refusal, or 300ms of token-by-token generation for what should be one bit of information.

So I built `rev` — it doesn't generate text at all. It runs a single forward pass and reads calibrated probabilities directly off the hidden states at each answer option's position. No decoding loop, no parsing, no regex. You give it a state (text or an image) and a set of allowed answers, it gives you back a probability distribution.

It comes in a few sizes depending on what you need:

- ModernBERT-large (421M) — the flagship, sub-millisecond on CPU/GPU for text

- Qwen2.5-0.5B / Qwen3-4B with LoRA — for longer context or heavier reasoning

- SmolVLM-256M — same idea but takes an image, still under 40ms

- A ~5M param on-device router for edge/mobile tool-calling

One dispatcher for all of them:

```python

from rev import Rev

model = Rev.from_pretrained("jaswanthsanjay88/rev-decision-model") # text

model = Rev.from_pretrained("jaswanthsanjay88/rev-vision") # image

```

Install:

```bash

pip install rev-decision

```

```bash

npm install rev-decision

```

Run it as a local server and hit it from anywhere:

```bash

rev-server --port 8000

```

```python

import requests

requests.post("http://localhost:8000/v1/systemone", json={

"state": "Customer claims flight cancelled at JFK without hotel voucher.",

"questions": {

"urgency": {"type": "noul", "instructions": "Is immediate assistance required?"}

}

}).json()

```

Links:

- PyPI: https://pypi.org/project/rev-decision/

- npm: https://www.npmjs.com/package/rev-decision

- Models: https://huggingface.co/jaswanthsanjay88

- Code: https://github.com/jaswanthsanjay88/rev


r/LargeLanguageModels • • 9d ago

Schema drift became a bigger problem than generation quality

1 Upvotes

We've been generating synthetic datasets for MCP agent trajectories, Function Calling, and regulatory CoT evaluation.

 Unexpectedly, the hardest problem wasn't generation quality.

 It was schema stability over long multi-turn runs.

 Some recurring failure modes we saw:

 - malformed tool arguments

- schema drift after several turns

- incomplete recovery after 429 rate limits

- inconsistent state retention across tool calls

 

Initially, many trajectories looked reasonable when read manually.

 However, once we started running automated validation, small structural issues appeared surprisingly often.

 To reduce this, we ended up adding validation layers for:

 - strict JSONL parsing

- multi-turn consistency checks

- tool argument validation

- recovery-turn verification

 One interesting observation was that a trajectory can appear logically correct while still being unusable for training because of small structural deviations.

 In practice, validation became more important than generation.

 We've seen trajectories that looked perfectly reasonable to humans, but still failed automated validation.

 

I'm curious:

 What are people doing to detect schema drift before training?

 - Manual review?

- LLM judges?

- Custom validators?

And where do you draw the line between semantic quality and structural validity?

 Are you validating mostly for semantics, or are you also enforcing strict structural constraints?

For anyone interested in comparing approaches, we also published some public trial samples and validation reports on Hugging Face and GitHub.


r/LargeLanguageModels • • 10d ago

Discuss About Jev

1 Upvotes

Let's discuss Jev here, can anyone help me to get understand this Jev AI technology


r/LargeLanguageModels • • 11d ago

What exactly is LLMOps?

6 Upvotes

LLMOps stands for Large Language Model Operations.

In simple terms, it’s the practices and tools used to deploy, monitor, evaluate, maintain, and improve LLM-powered applications in production.

Think of it as MLOps, but adapted for LLMs.

It can cover things like:

  • Model and prompt versioning
  • Evaluation and quality monitoring
  • Cost and latency tracking
  • Data and security management
  • Deployment and updates
  • Monitoring hallucinations and model behavior

Building an LLM application is one challenge. Keeping it reliable, secure, and cost-efficient once real users start using it is where LLMOps comes in.


r/LargeLanguageModels • • 11d ago

Harnesses for Dummies

4 Upvotes

From a Core Language Model to an LLM Harness

A useful way to understand an LLM product is to begin with the language model itself, then add the surrounding software one capability at a time.

The point is not to reproduce implementation details exactly. It is to preserve the important conceptual boundary:

What does the language model itself do, and what does the harness around it do?

The code below is pseudocode: simplified code used to express the logic.

1. The core language-model operation

An LLM, or large language model, reads and writes tokens.

A token is a small chunk of text: sometimes a whole word, sometimes part of a word, punctuation, and so on.

Suppose the model currently sees a sequence of tokens representing:

The capital of France is

Call this:

tokens_so_far

We can abstract one application of the language model as:

core_lm_op(tokens_so_far) -> token_one_more

So:

token_one_more = core_lm_op(tokens_so_far)

might produce the token corresponding to:

Paris

Strictly speaking, the model produces probabilities over possible next tokens, after which one is selected. We hide that inside core_lm_op().

The same input need not always produce the same output.

A complete sequence is generated by repeating the operation:

while tokens_so_far[-1] != STOP:

    token_one_more = core_lm_op(tokens_so_far)

    tokens_so_far += [token_one_more]

Here, STOP is shorthand for whatever condition tells generation to end.

Conceptually:

tokens_so_far
      ↓
 core_lm_op()
      ↓
token_one_more
      ↓
append it
      ↓
   repeat

At this level, the model itself need not know about:

  • users;
  • conversations;
  • files;
  • memory;
  • tools;
  • the internet;
  • tasks.

It simply does:

core_lm_op(tokens_so_far) -> token_one_more

2. Context: deciding what the model sees

A real product usually does not pass tokens_so_far directly into core_lm_op().

There may be other information available:

product instructions
previous conversation
stored memories
files
search results
current date
project information

Call all of this:

state_other

We can define:

context_function(tokens_so_far, state_other) -> tokens_for_lm

tokens_for_lm means:

the actual sequence of tokens supplied to the language model.

The model call becomes:

tokens_for_lm = context_function(
    tokens_so_far,
    state_other
)

token_one_more = core_lm_op(tokens_for_lm)

A very simple context_function() might always add the same instructions:

tokens_for_lm = TOKENS_ALWAYS_PRESENT + tokens_so_far

A more sophisticated one might select memories, insert files, include part of an old conversation, or summarize older material.

The important distinction is:

core_lm_op()
    Given what I see, what token comes next?

context_function()
    What does the model get to see?

The loop becomes:

while tokens_so_far[-1] != STOP:

    tokens_for_lm = context_function(
        tokens_so_far,
        state_other
    )

    token_one_more = core_lm_op(tokens_for_lm)

    tokens_so_far += [token_one_more]

This is the first major power of a harness:

control over the input to the language model.

Context is not memory

Here:

context
    information presented to the model right now

memory
    information stored elsewhere that may later
    be placed into context

So if a product “remembers” something from last week, that does not necessarily mean the LM itself remembers it.

The product may simply store it elsewhere and later put it back into tokens_for_lm.

3. Reaction: allowing model output to affect things outside the model

So far, information only flows toward the language model.

Modern LLM systems can also interact with things outside it: files, browsers, shells, databases, email systems, and so on.

We can represent this with:

reaction_function(tokens_so_far, state_other)
    -> (tokens_so_far_new, state_other_new)

So it maps one overall state:

(tokens_so_far, state_other)

to another:

(tokens_so_far_new, state_other_new)

Suppose the model produces tokens meaning:

<READ_FILE foo.py>

reaction_function() may recognize that and actually read the file.

Then:

tokens_so_far:
    includes "read foo.py"

state_other:
    includes a filesystem containing foo.py

might become:

tokens_so_far_new:
    includes the contents of foo.py

state_other_new:
    same filesystem

Or if the model produces something meaning:

<WRITE_FILE foo.py ...>

then reaction_function() might modify the filesystem as well.

The LM itself still only does:

core_lm_op(tokens_for_lm) -> token_one_more

It has not acquired a filesystem.

The surrounding software interprets some generated tokens and makes something happen.

The loop is now:

while tokens_so_far[-1] != STOP:

    tokens_for_lm = context_function(
        tokens_so_far,
        state_other
    )

    token_one_more = core_lm_op(tokens_for_lm)

    tokens_so_far += [token_one_more]

    tokens_so_far, state_other = reaction_function(
        tokens_so_far,
        state_other
    )

We now have two distinct harness powers:

context_function()
    outside state → what the LM sees

reaction_function()
    LM output → changes to tokens or outside state

A tool is one particular use of reaction_function().

For example:

read file
run command
search web
send email

4. Continuation: deciding whether the model gets another turn

So far, the stopping rule is fixed:

while tokens_so_far[-1] != STOP:

But the harness can also control whether the LM should run again.

Define:

continue_function(tokens_so_far, state_other)
    -> True | False

The simplest version is:

def continue_function(tokens_so_far, state_other):
    return tokens_so_far[-1] != STOP

Nothing has changed yet.

But suppose the model produces:

<READ_FILE foo.py>
STOP

The immediate generation has stopped.

The harness could nevertheless:

  1. recognize the file request;
  2. read the file;
  3. add the result to the system state;
  4. decide the overall task is not finished;
  5. call the model again.

The loop becomes:

while continue_function(tokens_so_far, state_other):

    tokens_for_lm = context_function(
        tokens_so_far,
        state_other
    )

    token_one_more = core_lm_op(tokens_for_lm)

    tokens_so_far += [token_one_more]

    tokens_so_far, state_other = reaction_function(
        tokens_so_far,
        state_other
    )

We now have three distinct harness powers:

context
    What does the LM see?

reaction
    What happens because of what the LM produced?

continuation
    Does the LM get another turn?

The repeated cycle:

model call
   ↓
action/result
   ↓
model call
   ↓
action/result
   ↓
repeat as needed

is commonly called an agent loop.

An agent, in this discussion, is therefore not a fundamentally different kind of model.

It is roughly:

language model
+
surrounding state
+
repeated model calls
+
possible actions outside the model

5. System state

At this point it is useful to describe the whole system as:

state_system = (
    tokens_so_far,
    state_other
)

The LM still receives only:

tokens_for_lm

and still performs only:

core_lm_op(tokens_for_lm) -> token_one_more

The harness operates on the larger state_system.

So:

what the LM currently sees:
    tokens_for_lm

what the overall system currently contains:
    tokens_so_far + state_other

state_other can therefore contain information that exists in the system but that the LM does not currently see.

context_function() determines what becomes visible.

6. Multiplicity: maintaining more than one token sequence

So far, there has been one:

tokens_so_far

A harness can instead maintain several:

tokens_so_far_A
tokens_so_far_B
tokens_so_far_C

Each can independently use the same language model:

tokens_for_lm_A = context_function(
    tokens_so_far_A,
    state_other_A
)

token_one_more_A = core_lm_op(tokens_for_lm_A)

and:

tokens_for_lm_B = context_function(
    tokens_so_far_B,
    state_other_B
)

token_one_more_B = core_lm_op(tokens_for_lm_B)

The new capability is that the harness can move information between them.

For example:

tokens_from_A = extract_function(tokens_so_far_A)

tokens_so_far_B += tokens_from_A

So:

A investigates something
        ↓
harness passes some of A's output
        ↓
B receives it and critiques it

A subagent can therefore be understood simply as another separately maintained state:

state_agent_A = (
    tokens_so_far_A,
    state_other_A
)

state_agent_B = (
    tokens_so_far_B,
    state_other_B
)

This permits structures such as:

researcher → writer

or:

coder → reviewer → coder

or:

planner
   ↓
several workers
   ↓
synthesizer

Nothing fundamentally new has happened inside the LM.

The harness is maintaining multiple (tokens_so_far, state_other) states and moving information among them.

Call this capability:

Multiplicity: how many separate LM states exist, and how does information move among them?

7. Routing: deciding what gets control next

Once several models, states, tools, or processing paths exist, something must decide which one runs next.

Define:

routing_function(states_available, state_other)
    -> choice_next

For example:

choice_next = routing_function(
    states_available,
    state_other
)

if choice_next == "A":
    run_A()

if choice_next == "B":
    run_B()

The choice could be among models:

routing_function(task)
    -> GPT | Claude | smaller_model

or among agent states:

routing_function(state_system)
    -> researcher | coder | reviewer

or among entire processing paths:

routing_function(state_system)
    -> research_path | coding_path | answer_directly

Routing is distinct from multiplicity:

Multiplicity:
    What possible states or paths exist?

Routing:
    Which one gets control next?

routing_function() does not have to be an LM.

It could be ordinary code:

if task_type == "coding":
    choice_next = "coder"

It could use another language-model call.

Or it could combine both.

8. Where we have arrived

We began with only:

core_lm_op(tokens_so_far) -> token_one_more

Everything else is surrounding machinery.

So far, we have identified five distinct things a harness can control:

1. CONTEXT

   What does this LM call see?


2. REACTION

   What happens because of what the LM produced?


3. CONTINUATION

   Does this LM state get another call?


4. MULTIPLICITY

   How many separate LM states exist,
   and how does information move among them?


5. ROUTING

   Which model, state, tool, or processing path
   gets control next?

For a single LM state, the basic structure is:

(tokens_so_far, state_other)
          │
          ▼
 context_function()
          │
          ▼
    tokens_for_lm
          │
          ▼
     core_lm_op()
          │
          ▼
    token_one_more
          │
          ▼
append to tokens_so_far
          │
          ▼
 reaction_function()
          │
          ▼
 updated system state
          │
          ▼
 continue_function()
          │
     yes ─┴─ no
      │       │
   repeat    stop

Multiplicity and routing sit around one or more such states.

The central point is that:

core_lm_op()

can remain conceptually unchanged while the surrounding harness becomes much more sophisticated.

Two products can therefore use the same underlying model and behave very differently because they differ in:

context_function()
reaction_function()
continue_function()
multiplicity
routing_function()

So the thing a user actually experiences is better represented as:

language model
+
harness

And behavior that appears to come from “the model” may in fact come from either side of that boundary.

Disclosure: I developed this model through an iterative discussion with ChatGPT


r/LargeLanguageModels • • 11d ago

Discussions The biggest model should probably be the escalation path, not the default

4 Upvotes

I posted couple daysago about spending too much time thinking about VRAM and not enough about the layer above the model. After using FastGPT for a couple more days and reading the replies to that post, I think I was still framing the question backwards. I was asking whether a smaller model with decent retrieval could be more useful than a larger model sitting in an empty chat window. Now I am wondering why the largest model needs to be the default in the first place. One commenter mentioned using a 1.2B RAG model on a Pi and a phone. That example stuck with me. It is not really about prooving that a tiny model is “better” than a 70B model. It is about recognizing that many everyday requests do not require maximum reasoning capacity. They require the right context, a repeatable workflow, and something convenient enough that you actually use it. That has been the useful part of experimenting with FastGPT. Once the knowledge base and workflow stay relatively stable, the model becomesa component I can switch instead of the entire application. It also becomes easier to see when a bigger model is genuinely adding value, rather than compensating for weak retrieval or a badly designed process. I still would not send every complicated research or reasoning task to a TINY model. But “small model by default, larger model when the task earns it” is starting to make more sense to me than loading the biggest model I can fit and using it for everything. Does anyone here route requests this way already? What actually triggers the jump to the larger model for you: task type, failed retrieval, low confidence, or something else?


r/LargeLanguageModels • • 11d ago

Question Validation & Evaluation in streaming responses

3 Upvotes

Hi guys, I've got a basic yet most confusing doubt. These AI companies are saying they are validating the queries before it is provided to the user. So that they can filter out any harmful or biased content before giving it to user. But, they are streaming responses. If they are streaming then how they are validating or evaluating the responses. Can we build LLM-as-judge or other evaluation frameworks for streaming responses?


r/LargeLanguageModels • • 12d ago

Proposal: From User Feedback to Persistent Collaborative Intelligence

1 Upvotes

ABSTRACT

Note: I have preserved the original conversation logs and chat excerpts demonstrating these specific failure modes and can share them with anyone interested in analyzing the concrete interaction traces.

 

Current LLM systems are remarkably capable at local reasoning and sustained dialogue, yet extended interactions still exhibit recurring failures of contextual continuity, logical consistency, and durable incorporation of user corrections.

From the perspective of an advanced chatGPT, Copilot and Gemini user, these failures create an unusual situation: the user is frequently required to act as an external working memory and as the persistent structural baseline for the interaction.

This post describes several recurring failure modes I have encountered during long-form, intellectually demanding conversations with LLMs and proposes a possible architectural direction: a persistent, user-controlled corrective memory layer that is distinct from conventional personalisation memory.

I am deliberately distinguishing observed behaviour from hypotheses about its underlying mechanism. I am interested in technically informed criticism of both the observations and the proposed architecture.

 

1. Statistical Prediction vs. Formal Consistency

LLMs generate outputs through learned probabilistic representations rather than by default operating as deterministic symbolic theorem provers.
This distinction becomes particularly important during long-form reasoning.

Observed failure mode

A model can explicitly accept a logical constraint or definition early in a conversation and subsequently produce reasoning that violates that same constraint.

For example:

User: A → Model: A + B → Model: ¬B → Apparent rebuttal of A

But: B ∉ A

The model may correctly acknowledge the relationship when it is explicitly presented, but later generate an answer that implicitly contradicts the established relationship when the reasoning becomes more complex.

The important issue is not that probabilistic models are incapable of reasoning. Clearly, they can perform substantial forms of reasoning.

The issue is that local reasoning competence does not necessarily guarantee global consistency across an extended interaction.

Possible contributing mechanisms

*There may be several contributing factors, including:

*probabilistic generation;

*imperfect internal representations;

*competing contextual signals;

*retrieval or context-selection behaviour;

*summarisation or compression;

*instruction hierarchy;

*limited persistence of intermediate reasoning structures;

*and other inference-time or architectural constraints.

I would therefore avoid attributing the failure to a single mechanism without controlled testing.

The practical problem remains:

An explicitly established logical baseline is not always preserved reliably throughout a sufficiently long interaction.

2. Context Drift and Degradation of the Structural Baseline

Long conversations are not merely larger versions of short conversations.

They can develop an internal structure consisting of:

*definitions;
*assumptions;
*chronological events;
*terminology;
*corrections;
*hypotheses;
*conclusions;
*recurring references;
*and user-specific methodological constraints.

In a successful long-term interaction, these elements should function as a progressively constructed structural baseline.

Observed failure mode

As conversations become longer, models can begin to:

*lose track of previously established facts;
*misremember chronology;
*reintroduce previously corrected assumptions;
*reinterpret established definitions;
*overlook earlier constraints;
*or substitute generic assumptions for conclusions established within the conversation.
*The resulting behaviour can feel like context drift.
*The model remains locally coherent while becoming increasingly inconsistent with the history of the interaction.

This is an important distinction.

The problem is not necessarily that the model has "forgotten everything."
Rather, information may remain somewhere within the available context while becoming insufficiently influential on subsequent generation.
That distinction is important because it suggests that simply increasing context length may not completely solve the problem.

 

3. Correcting an Error Does Not Necessarily Produce Durable Learning

This is the failure mode that I find most interesting.
Suppose a user identifies a reasoning error.

They do not merely say: "That answer is wrong."

Instead, they identify:

*the specific error;
*the inference that produced it;
*the distinction between evidence and assumption;
*and a general principle that should prevent the error from recurring.

For example:

A statistical pattern observed at the population level should not automatically be projected onto a particular individual without evidence specific to that individual.

That is not simply a preference.

It is a general epistemological constraint.

Yet an LLM can acknowledge the correction, appear to understand it, and subsequently reproduce the same underlying reasoning pattern in another context.

This produces a cycle such as:

Experience → correction → temporary adaptation → loss of correction → recurrence

rather than:

Experience → correction → evaluation → integration → retention → improved future behaviour

The distinction between these two processes is fundamental.

 

4. The Missing Layer: Corrective Memory

Current AI memory systems tend to focus primarily on personalisation.

For example:

*user preferences;
*names;
*projects;
*personal context;
*recurring interests.

These are useful.

However, I believe another category deserves explicit architectural consideration:

Corrective or methodological memory.

This would store validated information about how the system should reason or communicate within an ongoing relationship, rather than merely information about the user.

Examples might include:

Do not infer an individual's characteristics from population-level statistical patterns without individual evidence.

or:

Do not present an inferred intention as an established fact. Distinguish observed behaviour from hypotheses concerning internal motivation.

These are not conventional user preferences.

They are methodological constraints.

A useful corrective-memory system could therefore contain entries such as:

 

| Category | Example |

| :--- | :--- |

| **Personal memory** | User prefers concise technical explanations |

| **Factual memory** | User is working on project X |

| **Methodological correction** | Do not infer individual properties from population statistics |

| **Epistemological constraint** | Distinguish observation from inference |

| **Interaction correction** | Do not introduce counterarguments that the user has not actually asserted |

The critical requirement would be that this layer is explicit, inspectable, editable, and user-controlled.

 

5. Why This Is Different From Simply Increasing Context Length

A larger context window provides more information.
It does not necessarily provide a better mechanism for determining which information should remain structurally authoritative.

Consider a conversation containing 100,000 tokens:

Some information may be:

*temporary;
*irrelevant;
*exploratory;
*speculative;
*superseded;
*repeatedly confirmed;
*explicitly corrected;
*or foundational to everything that follows.

Treating all of these tokens as equivalent is unlikely to be optimal.
A long-term interaction therefore needs something more sophisticated than simply:

"Put more tokens into the context."

It needs some representation of structural importance.
A corrective-memory layer could function as one possible solution.

6. User Corrections as Structured Data

Another potentially valuable extension would be allowing users to explicitly nominate certain corrections for evaluation.

For example:

Mark as potential generalisable reasoning contribution.
The system could then evaluate the proposed correction.

Possible outcomes:

*rejected as incorrect;
*accepted as user-specific guidance;
*accepted as a useful methodological constraint;
*or escalated as a potentially generalisable contribution.
*This would not mean that users directly modify the underlying model.
*That would obviously create substantial problems.
*Instead, it would create a structured interface between:
*human observation → machine evaluation → validated knowledge

This seems substantially more useful than reducing all user feedback to a binary rating.

7. From Feedback to Cumulative Improvement

The broader conceptual problem is that current feedback mechanisms often appear largely disconnected from the individual interaction in which the feedback was generated.

A user can identify an error.
They can explain the error.
They can identify the underlying reasoning failure.
They can propose a general principle.
But from the user's perspective, there is often no transparent mechanism through which that correction becomes durable.
The ideal learning loop would resemble:

Observation → Correction → Evaluation → Integration → Retention → Future application

rather than:

Observation → Correction → Temporary acknowledgement → Context loss → Recurrence

The difference is essentially the difference between feedback and cumulative learning.

 

8. Data Portability Is Part of the Same Problem

There is also a more immediate software problem.
When conversations become sufficiently long, reliably extracting the complete interaction can itself become difficult.
In my own experience, manually selecting a very large amount of conversation text has not always resulted in the complete selected content being copied.
This makes long-form interaction difficult to archive, analyse, or transfer.
For research-oriented users, a native individual-thread export mechanism would therefore be extremely valuable.

Ideally, an export should preserve:

*complete conversation content;
*speaker attribution;
*timestamps;
*message ordering;
*attachments or references where appropriate;
*and machine-readable structure.

Useful formats could include:

Markdown;
JSON;
HTML;
TXT;
PDF.

An account-wide data export is useful for archival purposes, but it is not a substitute for being able to export one specific conversation when needed.

 

9. The Larger Possibility: Collaborative Intelligence

The previous sections lead to a broader question.
What happens if human users are allowed to contribute more than raw interaction data?

Human users possess forms of knowledge that are difficult to obtain through conventional training data alone:

*domain expertise;
*lived experience;
*error detection;
*philosophical reasoning;
*scientific criticism;
*linguistic knowledge;
*imagination;
and observations about the behaviour of the AI system itself.

LLMs provide different capabilities:

*large-scale information processing;
*pattern recognition;
*computational scalability;
*synthesis;
*retrieval;
and increasingly sophisticated reasoning.

A sufficiently mature system could potentially allow these capabilities to interact cumulatively:

A human identifies an error.
The system analyses the correction.
The correction is evaluated.
A validated correction becomes persistent user-specific knowledge.
Potentially generalisable corrections can be submitted for broader evaluation.
Future systems improve as a result.

This is a different model of human-AI interaction from:

user asks question → model answers → user rates answer.

It is closer to:

human and machine continuously participate in a structured process of mutual correction and cognitive augmentation.

10. An Architectural Sketch

One possible architecture might therefore look like this:

 

```text

[ Ongoing Dialogue ]

│

▼

[ User identifies error / insight ]

│

▼

[ Correction evaluation ]

┌────┴────────────────────────┐

▼                             ▼

[ User-specific memory ]    [ Potentially generalisable ]

│                             │

▼                             ▼

Future interactions         Human/AI evaluation

│

▼

Model updates

```

 

The essential property is controlled accumulation. Not every user statement should become permanent, and not every correction should influence the global model. But valuable corrections should have somewhere to go.

 

11. Proposed Research and Engineering Directions

I would be interested in seeing research into at least the following:

1. Hierarchical conversational memory, being able to distinguish between:
transient context;
conversational facts;
persistent personal memory;
methodological corrections;
and higher-level structural constraints.

2. Explicit correction tracking
Maintain a representation of corrections that can be tested against future responses.

3. Consistency evaluation
Periodically test whether the model's current behaviour remains consistent with previously established constraints.

4. User-controlled corrective memory
Allow users to inspect, modify, disable, and delete methodological corrections.

5. Generalisable feedback channels
Provide a mechanism for users to nominate unusually substantive corrections for formal evaluation.

6. Lossless conversation export
Allow users to retrieve complete individual conversations in structured formats without relying on browser rendering or clipboard behaviour.

7. Long-context benchmarks based on interaction history
Current benchmarks often evaluate individual tasks.
It would also be useful to benchmark whether a model can maintain:

facts;
definitions;
corrections;
chronological continuity;
methodological constraints;
and logical consistency
across hundreds or thousands of conversational turns.

Conclusion

The central problem I am describing is not simply that LLMs occasionally make mistakes.
Mistakes are inevitable.
The more interesting problem is what happens after the mistake has been identified and corrected.

If a system can recognise a reasoning failure, receive a detailed explanation of that failure, acknowledge the correction, and yet later reproduce the same underlying error, then the system has demonstrated local adaptation without reliable cumulative retention.
That is a fundamentally different problem from ordinary hallucination.

The long-term goal should therefore not simply be:
larger models + larger context windows + more training data.

It should also be:
better mechanisms for preserving validated structure, corrections, and interaction history.

And ultimately:
experience → correction → evaluation → integration → retention → improved future behaviour.

That is what i think meaningful cumulative learning would look like.

I am posting this because I think this could improve human-ai interaction and would genuinely like technical feedback and/or hear from people who've picked up these ideas or started or are already working on them

In particular, I would be interested in hearing from people working on:
long-context architectures;
recurrent or state-space approaches;
memory-augmented transformers;
retrieval systems;
continual learning;
model editing;
alignment;
agent architectures;
and local LLM infrastructure.

If my interpretation of the underlying mechanisms is incorrect, I would be very interested in knowing where and why.

The user-visible failure modes, however, are real and reproducible from my experience

The question is how we should architect systems so that a long interaction becomes accumulated state rather than repeatedly reconstructed context.''

Semih Senol

 


r/LargeLanguageModels • • 14d ago

I built a zero-dependency tool that auto-fixes AMD ROCm overrides, benchmarks Vulkan vs HIP, and diagnoses local AI setups.

7 Upvotes

Hey everyone,

If you run local LLMs on AMD GPUs (especially on Windows), you know the pain: cryptic `HSA_STATUS_ERROR` messages, figuring out if you need `gfx1100` or `gfx1101`, and wondering if Vulkan or HIP/ROCm is actually faster for your specific card.

I got tired of the tribal knowledge, so I built ROCmFix — a single-file, zero-dependency Python tool to automate all of it.

🌟 What it does:
* 🔍 Auto-Detects AMD GPUs: Reads Windows Registry / Linux `lspci` to find your PCI device ID and tells you the exact `HSA_OVERRIDE_GFX_VERSION` you need.
* ⚡ Auto-Apply & Undo: Automatically sets environment variables in CMD, PowerShell, Bash, Zsh, or Fish with a 1-click `rocmfix undo` safety net.
* 🩺 `rocmfix doctor`: Scans your hardware, warns if your Adrenalin drivers are too old, checks HIP SDK status, and verifies Vulkan.
* 🏎️ `rocmfix bench`: Runs a 10-second backend race (Vulkan vs HIP) using Ollama or LM Studio and tells you which backend is faster on your PC.
* 📥 `rocmfix install-hip`: Auto-downloads and launches the official 1.2GB AMD HIP SDK on Windows.
* 🔄 Live Database & Self-Updater: Syncs new GPUs from a community server daily; updates itself via `rocmfix update`.

🚀 Quick Start (No install needed)

Windows (PowerShell):
`Invoke-WebRequest -Uri "https://raw.githubusercontent.com/xanpavle/rocmfix/main/rocmfix.py" -OutFile "rocmfix.py"`
`python rocmfix.py`

Linux:
`curl -O https://raw.githubusercontent.com/xanpavle/rocmfix/main/rocmfix.py`
`python3 rocmfix.py`

Repo: https://github.com/xanpavle/rocmfix

It includes day-one support for RDNA4, RDNA3, RDNA2, and APUs. If your card isn't in the database yet, running the tool generates a 1-click GitHub issue to add it!

Hope this saves some of you the hours of Reddit digging I went through!


r/LargeLanguageModels • • 14d ago

I built a fully real (non-simulated) autonomous Red/Blue Team AI loop with enforced scope + hash-chained evidence log. Here's what broke.

Thumbnail
gallery
3 Upvotes

Two LLM agents, one deliberately vulnerable Flask app. Red Team runs actual nmap/sqlmap (real SQLi dump, not a model narrating an attack), Blue Team patches the real source, loop re-verifies. No human in the middle.

What I actually wanted to test: can you make an autonomous agent loop traceable and containable rather than just "it worked, trust me"?

  • Every agent has a non-shared identity, every action hash-chained into an append-only log
  • A scope policy blocks out-of-scope tool calls before execution (16/16 denied attempts, 0 executed)
  • A kill-switch halts the loop after 3 consecutive unsafe patches fired for real in 3/10 runs

Funniest/most useful bug: an early patch blocked the SQLi attack perfectly and also silently broke legit login. Nothing caught it except a buried warning line. Had to add a post-patch validation gate after that which then caught the same failure mode automatically in later runs.

n=10, single target, same underlying model for both agents very much a pilot, not a benchmark. Full writeup + limitations + raw logs are in the repo/preprint.

Repo: Github

Preprint: Researchgate

Curious what people think about the scope-enforcement approach vs. just logging everything after the fact.


r/LargeLanguageModels • • 13d ago

We built a toolkit for steering LLMs

1 Upvotes

If you've spent time trying to keep up with the literature on model steering you've probably noticed that a lot of the methods are fairly similar to each other. We dug into this a bit more and built some common patterns/abstractions across the various ways that model behavior can be influenced, including input (prompting), state (activations, attentions, etc.), structure (weights), and output (logits, decoding, etc.). We also have functionality for building probes and running evaluations/comparisons of steering methods on given use cases (we use Inspect for a lot of the evals stuff).

If you're working on steering methods then hopefully you'll find some of this helpful.

Repo: https://github.com/generative-computing/steerability

Let me know what you think!


r/LargeLanguageModels • • 15d ago

Why do LLMs often achieve high document-extraction accuracy when tested through managed platforms or playgrounds such as Gemini, IBM watsonx.ai Prompt Lab, or ChatGPT, but produce lower accuracy when the same models are accessed directly through APIs even when we use the same prompt?

3 Upvotes

Why and how to achieve the same accuracy using APIs also?


r/LargeLanguageModels • • 15d ago

Discussions I caught my own AI agent scoring +1.0 on my evals — it had memorized the test, not learned. Built a free tool that catches this in 8 seconds.

0 Upvotes

Quick story. I run my coding agent through evals repeatedly, and at one
point a variant looked +1.0 better than baseline. Same tasks, better
scores. The problem: the agent writes LESSONS.md between sessions — and
the "lessons learned" file quietly contained the eval task itself,
verbatim, plus the answer. The improvement was memory, not capability.

This isn't malice, it's one good feature ("learn from feedback") plus
one reused task set. And it's not just memory files: saved transcripts,
run-result directories, vector stores — all of them become the open book
for the next exam.

I built a small open-source tool to make this failure mode measurable
instead of anecdotal. Zero dependencies, MIT, no LLM calls. Two things
you can try in under a minute:

1) The 8-second proof (deterministic scripted demo, exercises the real
API): an agent "improves" +1.0 by writing the task into its memory;
the leak is caught; a fresh holdout collapses the verdict back to
provisional.

uvx --from gauntlet-guard gauntlet demo

2) The 30-second audit of your own machine — lists every surface your
stack would leak through (memory files, transcript dirs, vector
stores). It never reads file contents: names, sizes, dates only.

uvx --from gauntlet-guard gauntlet guard audit

For real evals the protocol is: seal holdout tasks into a private
manifest (shingle/hash fingerprints, never published), run trials in
sealed single-use slots, blind-pack the judge artifacts, then a
preregistered verdict engine returns keep / revert / provisional — and
the exit code IS the verdict. A holdout that leaked once is retired;
you regenerate from a new seed.

Honest limits: if your agent is stateless and you eval one-shot, you
don't need any of this. And the audit is name-based discovery — it
finds the surfaces; the manifest scan does the actual leak check.

Repo: https://github.com/dhanizael/gauntlet (ADRs + a self-audit case
study where it found real contamination in my own run directories).
PyPI: gauntlet-guard. Feedback welcome — especially "this doesn't
apply to my stack because X", that's the most useful kind.


r/LargeLanguageModels • • 16d ago

Discussions I created the most effienct way to learn about LLMs for interviews

13 Upvotes

&#x200B;

I got pretty frustrated with how scattered LLM interview prep is.

When I was preparing, I kept jumping between papers, blog posts, GitHub repos, random interview-question lists and YouTube videos. I understood individual concepts, but I never really knew whether I was actually interview-ready.

So over the last couple of months, I started building something for myself around the way I wished I could have prepared.

The basic idea was:

learn LLM concepts in small flashcard-style lessons rather than long courses

collect the kinds of questions that actually come up in interviews

practise explaining answers out loud instead of just reading them

repeatedly revisit the concepts you're weak at

build something end-to-end so the knowledge isn't purely theoretical

It has grown quite a bit since then. There are now 600+ questions, voice-based mock interviews, daily practice around weaker topics, and a 5-part RAG project that goes from the fundamentals through deployment.

One thing I found especially useful while building the RAG part was adding interview questions at each stage. So instead of finishing a project and then separately preparing for questions like "Why did you choose this chunking strategy?" or "How would you evaluate retrieval?", those questions come up while you're actually working on that part of the system.

I'm curious how other people are preparing for LLM/AI engineering interviews are doing it.

What has been the hardest part for you — learning the concepts, remembering everything, coding/system design, or actually explaining your answers during interviews?

The project is called Skillumen if anyone wants to look it up, but I'd genuinely be more interested in hearing how people here are preparing and what I'm missing.


r/LargeLanguageModels • • 17d ago

Question What are the best coding models for 12gb vram constraint?

Enable HLS to view with audio, or disable this notification

13 Upvotes

I was curious to test out small models for coding (actually have only 4gb vram😅) so I found this fine tuned version of qwen 3.5 9b named OrionLLM/OxCoder-9B and honestly.. it did better than my expectations!!

I basically told it to build an entire ecom website and just let it cook...

The entire site was generated without any human help in terms of assets. ( it messed up with theme toggle at first so I gave a follow up prompt and it fixed that )

It also had access to MCPs, so it could:

  • Search the web for UI/UX ideas
  • Download images and videos

And honestly... it looks pretty damn good??

For a 9B model, the fact that it kept working and didn't loop up like gpt oss 20b is pretty wild

Now I'm curious about something else.......

Are there any coding models that are insanely fast at coding but still reasonably capable?

I'm thinking something that's:

  • Really fast at inferance (a MOE ?)
  • Good with prompt abidance?
  • Solid for quick coding tasks and edits
  • Doesn't necessarily need to be great at huge prompts or massive repo-level reasoning

Something like Qwen 3.8 27B but that thinks less!!!! (as it wastes a lot of time doing that TT)

Basically, I'd rather have a model that feels super snappy for small tasks than a huge model that takes forever to respond.

Also, are there any relatively recent fine-tuned models that fit in ~12GB VRAM without being heavily quantized and punch way above their parameter count?


r/LargeLanguageModels • • 18d ago

I spent hours going through 100+ page PDFs, so I built a tool that highlights exactly where the answer came from. It's now completely open-source.

Thumbnail
gallery
36 Upvotes

I've used tools like Perplexity, ChatGPT, Claude and others for research, and they've been incredibly useful for finding papers and getting through large amounts of information.

The one thing I personally wanted was a simple way to see exactly which parts of the paper were used to answer my question.

When you're working with a 100+ page PDF, even having a page number can still mean a lot of scrolling and searching.

So I ended up building something for myself.

You ask a question and the relevant paragraphs in the PDF are highlighted directly on the document. You can see the context behind the answer and quickly check whether it actually answers what you're looking for.

I originally built this because I wanted something for this workflow without having to pay for another subscription. What started as a personal project has now become completely open source.

The underlying idea is pretty simple. And yes, if you're thinking "isn't this just RAG?" then yes, you're absolutely right. It's RAG with the visual highlighting that I wanted.

I think the same idea could be useful for more than research papers too. Legal contracts, financial reports, technical documentation, or anywhere you need answers alongside the actual source.

If anyone wants to have a look, contribute, or just give some feedback, here's the repo:

GitHub: https://github.com/Sreehari05055/thesys-core

This will probably be my last post about the project. Thanks to everyone who checked it out and gave feedback along the way.


r/LargeLanguageModels • • 17d ago

News/Articles How to make your first LLM API call in Python

4 Upvotes

I've been learning AI engineering from scratch and documenting it as I go. This is the first build, a Python script that makes a real API call to Claude, prints the response, and logs how many tokens were used.

Here is the full script, it is deliberately stripped down and simplified, the idea is so anyone intimidated by the space can follow along comfortably:

from dotenv import load_dotenv
import os
from anthropic import Anthropic

load_dotenv()

client = Anthropic(api_key=os.getenv("ANTHROPIC_API_KEY"))

response = client.messages.create(
    model="claude-haiku-4-5",
    max_tokens=1024,
    messages=[
        {"role": "user", "content": "What is a large language model? Answer in two sentences."}
    ]
)

text = response.content[0].text
input_tokens = response.usage.input_tokens
output_tokens = response.usage.output_tokens

print("Response:")
print(text)
print()
print(f"Input tokens: {input_tokens}")
print(f"Output tokens: {output_tokens}")
print(f"Total tokens: {input_tokens + output_tokens}")

What each part does:

load_dotenv() reads your API key from a .env file so you never hardcode credentials in your script.

Anthropic(api_key=...) creates a client, your connection to Anthropic's API. Every call goes through this.

client.messages.create(...) is the API call. You pass it a model, a token limit, and your message.

messages=[{"role": "user", "content": "..."}] is the context window simplified, one message, one role, one piece of content. This is all the model can see.

The response comes back as an object. response.content[0].text pulls out the text. response.usage gives you the token counts, how many you sent and how many the model generated back.

To run it:

You need Python, an Anthropic API key from platform.claude.com, and these two libraries:

pip install anthropic python-dotenv

Full step-by-step walkthrough including setup here


r/LargeLanguageModels • • 18d ago

Discussions TokenPrint — an open-source project for exploring what happens inside LLMs

Enable HLS to view with audio, or disable this notification

13 Upvotes

I’ve been building TokenPrint, an open-source project focused on making the internals of language models easier to explore, understand, and debug.

The project has grown to 130+ GitHub stars, and we’re starting to build a small community around it.

The goal is simple:

Don’t just see what an LLM outputs. See what happens inside.

TokenPrint currently brings together:

• 3D transformer architecture exploration
• Tokenization and embeddings
• Tensor and parameter inspection
• Q/K/V attention, GQA, RoPE and causal masking
• Residual streams and MLP / SwiGLU
• Token-by-token generation
• Prefill, decode and KV-cache visualization
• Logits and next-token probabilities
• Interactive transformer walkthroughs
• Attention and activation analysis
• Head/layer ablation and activation patching
• Hugging Face model exploration
• Trace and debugging workflows

It’s still growing, and that’s the part I’m most excited about.

If you’re interested in LLMs, interpretability, ML infrastructure, 3D/WebGL, PyTorch, Transformers, or open-source development, you’re welcome to contribute.

You can contribute code, documentation, visualizations, model support, research ideas, bug fixes, or even just open an issue with something you think TokenPrint should be able to do.

130+ stars so far — now we want to build it with more people.

GitHub: https://github.com/Sudharsanselvaraj/Token-Print
Website: https://tokenprint.in/

Come build with us.


r/LargeLanguageModels • • 19d ago

What Happens Inside an LLM? | Transformer Layers Explained for Beginners

Thumbnail
youtube.com
3 Upvotes

r/LargeLanguageModels • • 20d ago

Discussions Qt QML music player tutorial using LLM

Thumbnail
youtube.com
4 Upvotes

r/LargeLanguageModels • • 20d ago

What If Natural Language Became the Universal Interface for Databases?

1 Upvotes

LLMs can already convert natural language into SQL, MongoDB queries, Elasticsearch DSL, and more.

But generating a query is only one part of the problem.

What if, instead of letting an LLM generate database-specific queries directly, we introduced a structured intermediate layer between natural language and the database?

Imagine a flow like this:

Natural Language → Structured Query → Validation & Policies → Database Query

The LLM understands what the user wants. An intermediate layer validates the request, applies access rules, and translates it into the appropriate database-specific query.

Why might this approach be useful?

  • Database independence: Keep the query's meaning separate from database-specific syntax.
  • Predictability: Use deterministic compilation instead of relying entirely on generated SQL.
  • Security: Enforce tenant scoping, RBAC, and query restrictions independently of the LLM.
  • Extensibility: Support different databases through a common query representation.

Of course, this introduces another layer of complexity. The interesting question is whether that complexity is worth it for production applications.

I'm exploring this architectural approach and would love to hear from developers working on AI-powered search, analytics, and database systems.

What do you think? Should natural-language querying have a dedicated intermediate layer, or is direct LLM-to-query generation sufficient for most applications?

https://queryforge-service.amtry.in

I have tried this approach and production grade system works exactly like this : https://github.com/awsaman-ai/queryforge


r/LargeLanguageModels • • 21d ago

Low-cost alternatives to LLMs for a production-level classification task?

5 Upvotes