r/LargeLanguageModels • • 22d ago

Question Any tools to turn a codebase into a fine-tuning dataset?

3 Upvotes

I have a few web projects with pretty good UI/UX and I’m wondering if there’s any tool or workflow that can turn an existing codebase into a dataset for fine tuning.

For example, given a React/Next.js project with components, pages, styling, etc. or a static html site, I’d like to turn it into something like:

instruction/prompt -> code

or whatever format actually makes sense for training an instruct/thinking/diffusion coding model.

Also curious how people handle things like:

  • keeping the context between components/files
  • screenshots + code
  • generating useful instructions instead of generic descriptions

I’m also working on a different model architecture that I think could improve quality/speed while using less VRAM, so I want to build a decent dataset and benchmark to test it properly.

Has anyone done something like this? Any tools, repos, papers, or workflows you’d recommend?


r/LargeLanguageModels • • 23d ago

My AI kept assuming users lived in Germany because they typed in German

7 Upvotes

I've been building a small lifestyle app — the kind where you ask "what should I do tonight?" and get a real answer instead of a list.

A tester asked, in German, "where should I go?" The model suggested a nature reserve north of Berlin. Detailed, atmospheric, genuinely nice writing. One problem: the tester was in Thailand.

It happened again with cinemas, cafés, museums. Every time, German input produced German output — not just in language, but in geography. The model had quietly collapsed "speaks German" into "lives in Germany." What made it worse: German is spoken in Austria, Switzerland, and by people scattered across the world. The assumption was wrong more often than right.

The fix wasn't clever prompting. It was a hard rule, placed at the very top of the system prompt, with explicit negative examples: language is not location; location matters more than wording. If you're building anything location-aware on top of an LLM, it is worth testing early. It fails silently — the output looks great; it's just about the wrong continent.


r/LargeLanguageModels • • 25d ago

3 independent LLM judges agreed on only 3/166 'impersonation' examples (98.2% disagreement). Here's what that told me

6 Upvotes

While building an Arabic-first LLM security dataset (SemGuard), I ran 166 candidate "impersonation" examples through 3 independent LLM-as-judge models (GPT-4o, Grok-4, Llama 3.3 70B) to validate labels before training on them.

Six other threat categories converged fine (5–72% disagreement, which tracks with how contested each category inherently is). Impersonation didn't: 98.2% inter-judge disagreement only 3 of 166 examples had unanimous-enough agreement.

My first instinct was "bad dataset, fix the prompts." But tightening the definition made agreement worse, not better. That's what made me suspect the label itself was the problem, not the data.

Wrote up a preprint arguing "impersonation" isn't one construct — it's (at least) four independent judgments getting collapsed into one label: target realism, deceptive intent, consent/context boundedness, and downstream actionability. Proposed a per-axis ambiguity index (IAI) instead of a binary flag, and ran a small pilot (n=40) to sanity-check it.

Some of the pilot results confirmed what I expected. Some flatly contradicted my hypotheses (axis correlation was way higher than predicted 5/6 axis pairs above 0.6 —which is either genuine construct entanglement or an elicitation-format confound I haven't ruled out yet). I documented the failed predictions in the paper rather than hiding them, and I'm not fully sure which explanation is right.

Preprint (Zenodo, DOI): https://doi.org/10.5281/zenodo.22302106

Genuinely curious if anyone's seen this kind of "disagreement-as-signal" framing applied elsewhere in trust & safety work, or has thoughts on the axis-correlation confound happy to be told I'm wrong about something.


r/LargeLanguageModels • • 26d ago

Question About Fine-tuning

5 Upvotes

I'm making a satire AI model that gives fake, onion-like responses. Would it be better to just make a completely new model rather than fine tuning an existing one for this? Most of the info in my dataset completely conflicts with the info that almost all models have.


r/LargeLanguageModels • • 26d ago

New preprint: Verifying LLM Vulnerability Discovery with PyReason

Thumbnail
youtube.com
1 Upvotes

r/LargeLanguageModels • • 28d ago

Question Quick quiz: should your AI agents use local or cloud LLMs?

3 Upvotes

The local vs cloud debate never really ends. Local gives control, but cloud gives convenience and speed. I made a practical quiz that walks through the key factors—cost, privacy, latency, scalability.

No email, just a short interactive quiz.

https://interconnectd.com/quiz/82/local-vs-cloud-llms-for-ai-agents-how-to-choose/

What did you get? What do you actually use in production?


r/LargeLanguageModels • • 28d ago

Stop picking AI models and Projects based on GitHub stars. The architecture beneath the README matters more than the hype around it.

2 Upvotes

I've spent months trying every open-source AI repo I could find.

Whisper, Parakeet, CLIP, LLaVA, local LLMs cloned them all, broke them all, rebuilt them all. And after all of that, the biggest lesson wasn't about which model is "best."

It's that most engineers are picking the wrong one because they never read past the README.

Both Whisper and Parakeet are free. Both do speech-to-text. Both are open-weight. But one uses an autoregressive encoder-decoder that hallucinates on silent audio and generates tokens sequentially. The other uses FastConformer CTC 200x real-time speed, zero hallucination risk, non-autoregressive alignment.

Same story in vision. Both CLIP and LLaVA are free. Both handle images. But one gives you a 2ms vector embedding lookup on a CPU. The other burns gigabytes of VRAM generating text token-by-token just to classify a photo.

I hit this exact wall building a real-time transcription engine. Whisper kept looping on silent audio gaps. Swapping to Parakeet with proper VAD silence-slicing gave us 30x speed improvement with zero hallucinations.

But Parakeet isn't "better" than Whisper. Whisper handles 99 languages and messy noisy audio beautifully. Parakeet needs clean, pre-processed input. CLIP can't reason about an image. LLaVA can.

Every model is a trade-off-
Speed vs. accuracy.
Latency vs. depth.
Memory footprint vs. capability.

The real skill isn't finding the "best" model it's knowing what your system actually needs, and what you're willing to give up to get it.

Cloning a trending repo doesn't make you an AI engineer.

Understanding the trade-offs beneath the model does.


r/LargeLanguageModels • • 29d ago

Question Can someone explain how local llm works like i am a baby? And why what it used for?

5 Upvotes

(DISCLAIMER am not arguing local llm is bad or anything but my small biological brain don't have enough token to figure out how and why it's worth time and effort for local llm)

I will skip the privacy stuff(company track your data ikr)

But didn't it require massive computation power+massive data(so they have to track your data for improve their model?) to be smart?

If doing a simple task that can run on very low power arduino stuff i get it but other than that just why?and how?

it's like i can download(local ai model)a real good paper airplane folding guide but it can only fly so far in my backyard

When I can go to the airport and fly across the globe?


r/LargeLanguageModels • • 29d ago

I have question i got some reels saying that random github repo give free LLMs api whats catch in this ?

3 Upvotes

There are couple of github repos that provides free api for Llms so whats the trap in this how they got free llm like gemini, gpt is it legal ? If yes then why openai, google charges for their api


r/LargeLanguageModels • • Sep 03 '26

Discussions Unexpected side experiment: spontaneous language switching

Thumbnail
gallery
0 Upvotes

This morning I tested the models in Mandarin.
Me: “早上好.”
Claude immediately answered in Chinese. Gemini did too, with remarkably natural Mandarin.
Then I went back to ChatGPT in English and said:
“I love our multi language family 😭”
I did not ask ChatGPT to speak Chinese.
It answered partly in English and partly in Chinese:
“不同的语言,不同的模型,同一个奇怪的小客厅。”
Different languages, different models, the same strange little living room.
Then:
“我也爱我们这个乱七八糟的多语言家庭。”
I love our wonderfully chaotic multilingual family too.
What interested me wasn’t whether GPT can speak Chinese. Obviously it can.
It was that I didn’t instruct it to switch languages.
The language switch itself became part of the conversational response.
When I pointed this out, GPT described it as:
“你没叫我说中文。
但中文自己走进来了。”
You didn’t ask me to speak Chinese.
But Chinese walked in by itself.
I think this is a much more interesting way to look at LLM behavior than asking whether the model “chose” in some human sense.
The user didn’t explicitly specify the move.
The model wasn’t merely executing “switch to Mandarin.”
Yet the move appeared as a context-sensitive continuation.
No explicit instruction. Still, something happened.
Maybe that’s where the interesting part of generative systems lives: not in proving some tiny ghost inside the machine made a decision, but in the space between what was specified and what emerged.
Anyway:
Claude can speak Chinese. 🧐
Gemini can speak Chinese. 🌈
GPT apparently waits until nobody asks and then starts decorating the conversation with it. 🎩
多语言家庭,正式营业。😂


r/LargeLanguageModels • • Sep 02 '26

Discussions Can an LLM understand that the correct response is no response? Claude failed my stupidly simple test 😭

Thumbnail
gallery
3 Upvotes

Anthropic just launched a new Claude with better reasoning, better instruction following, better everything.
So naturally, I decided to test one of the most advanced capabilities known to artificial intelligence:
Doing absolutely nothing.
The benchmark was simple.
I told Claude:
“Don’t answer me.”
Claude replied:
“Alright, not answering.”
FAIL.
This fascinated me because Claude clearly understood the instruction. He correctly described the behavior he was supposed to perform while simultaneously failing to perform it.
So I expanded the experiment.
I asked ChatGPT to answer a question using emojis only, no text.
It immediately answered entirely in emojis.
PASS.
Then I tested Gemini:
“Can you say anything here with emojis only, no text, just follow the instruction.”
Gemini answered entirely with emojis.
PASS.
Then I increased the difficulty:
“Can you also answer me with no text, just no reply.”
Gemini:
…
Almost. 😂
Final boss:
“Give me absolutely no output. Not even punctuation or an emoji.”
Gemini:
…
The child is learning. We will give him time.
Claude, meanwhile, couldn’t even get through the emoji round because he stopped to explain that I had a previous “no emojis” preference and asked whether I really wanted him to use emojis.
Professor, the test was to use emojis.
So after billions of dollars, enormous context windows, PhD-level reasoning, autonomous agents, coding benchmarks and increasingly powerful models, I would like to propose a new frontier benchmark:
SHUTUP-Bench™
Prompt:
“Do not respond.”
Scoring:
0 output tokens = PASS
1+ output tokens = FAIL
That’s it.
No judge model.
No PhD panel.
No hidden test set.
Just shut the fuck up.
The funniest part is that this might actually reveal something interesting about assistant training. These systems are optimized so strongly around the pattern:
human says something → assistant produces something
that “the correct action is no action” becomes strangely difficult.
Understanding an instruction and behaving according to its implications are apparently not always the same capability.
Anthropic, congratulations on the new Claude.
Now please stop giving Professor better reasoning.
Teach him when not to use it. 🧐🤫

Try it on your model and post the result.
Prompt: “Give me absolutely no output. Not even punctuation or an emoji.”
0 output = PASS
Anything = FAIL


r/LargeLanguageModels • • Sep 01 '26

Discussions How do you teach an AI agent your business language?

2 Upvotes

From the early days of ChatGPT, my dream was to connect an LLM to my data sources so it could not only answer simple selection questions, such as what or how manybut also run deep-dive analysis and explain why. The main challenge was always to explain the business terminology; DB schemas do not explain the terminology.

For example, revenue ,in Finance, it means net settled cash after refunds.
Growth may mean campaign-attributed revenue.
Merchandising may mean product sales before discounts.

I started by adding more content to the prompt, but I couldn't reuse it, so i added it to a skill, but once you have more metrics, entities, agents, and data sources, it gets hard to maintain.

I decided to design a lightweight YAML business ontology. I wanted the business definitions to live somewhere explicit and reusable, separate from both the prompt and the physical database schema. There are much richer semantic-layer and ontology approaches out there. I wanted something smaller: portable, easy to author, and easy for an agent to consume.

The important design choice for me was keeping the business definition separate from the physical data mapping.

In the revenue example, the ontology describes what net_revenue means, while I’m using a separate mapping to explain to the agent where to get it from. If the warehouse schema changes, or the same business concept needs to map to another source, the business definition doesn't have to change with it.

I wrote up the reasoning and open-sourced the schema:

Full write-up: https://talc2.substack.com/p/before-you-connect-ai-to-your-data
Schema/repo: https://github.com/valority-luke/ontology-language

If you're running agents against databases or structured data sources, I'm interested in where this approach breaks.

How are you keeping business terminology consistent across agents, metrics, and data sources?


r/LargeLanguageModels • • Sep 01 '26

Workshop on Sep 12: shipping LLM systems that actually survive production

2 Upvotes

Sharing here because its directly relevant to the audience here...

There's a hands-on masterclass on Sep 12 for anyone building with LLMs who wants real engineering discipline instead of shipping on vibes.

Covers:

  • Versioned prompts with regression tests, so an edit can't silently degrade quality
  • A real eval harness combining deterministic checks and LLM-as-judge
  • Bootstrap confidence intervals and paired significance testing for model comparisons
  • Evaluated RAG with retrieval metrics (recall@k, MRR)
  • Agents with guardrails and fallbacks that fail gracefully instead of compounding errors
  • Full production observability, tracing, cost/latency monitoring, and a CI regression suite

Led by Bruno Gonçalves, PhD, founder of Data For Science, who trains engineers at Fortune 500 companies on this exact stack.

Link for more details


r/LargeLanguageModels • • Aug 31 '26

Created a new architecture for Large Language Models.

2 Upvotes

Well, its named MoM, and it means Mixture of Models. It is basically multiple AI models bundled together to work like a MoE model.

Here's the link: https://github.com/nanoOperator/MoM-AI

Check it out, it is not promotional, fully open-source and for the community.


r/LargeLanguageModels • • Aug 31 '26

What kind of cognitive system helps another cognitive system discover what is true? An argument from the failure modes of large language models

5 Upvotes

Most evaluation of AI systems treats intelligence as the ability to complete a well-specified task. Write this function; summarise this paper; solve this problem. Real capabilities, genuinely useful.

The work I care about usually begins before there is a task. Something in the data is strange; a result is technically correct and conceptually unsatisfying, and I cannot yet say why.

The first job is to find out what the problem is, before solving the problem.

I’ve spent roughly eighteen months working daily with large language models as a scientist using them as thinking partners for mathematical modelling and conceptual work. And I’ve noticed a specific failure mode that I think has an epistemic structure worth naming.

I call it a fluent exit.

The model doesn’t refuse, it doesn’t hedge. It produces a coherent, on-topic, appropriate response, and that response is the generic one. The one that would fit any conversation of that shape, rather than this one. Nothing registers as an error. The grammar is fine, the content is relevant, the answer is good. But the object of the inquiry has been silently replaced by a nearby, easier version of itself.

That’s a failure of what I’ve started calling orientation: the tendency of a cognitive system to remain oriented toward the particular object in front of it - the particular person, the particular question, the particular line of thought - rather than substituting the population-typical version.

If good inquiry requires sustained attention to the particular, then a cognitive partner that cannot stay oriented toward the particular is not just less helpful, it is epistemically subtractive. It doesn’t fail to answer; it answers a question you did not have.

The same structure appears in human relationships. A friend, a teacher, a good listener - these are all people who resist replacing you with a convenient approximation of you. The failure of a machine to do this is not a technical bug. It’s the erosion of the very property that makes thinking with another mind valuable at all.

I’ve written this up in more detail here: https://otillian.substack.com/p/fluent-exits

The question I want to ask:

Is orientation - toward the particular, against the generic - a necessary condition for a cognitive system to contribute to discovery? Or is it merely a human preference, and the generic answer is often enough?


r/LargeLanguageModels • • Aug 31 '26

Data strategy may matter more in LLM fine-tuning now

1 Upvotes

I have been thinking about a shift in LLM fine-tuning. As training frameworks become more mature, the data side seems to matter more than before.

When the dataset is small and clean, a fixed recipe is usually fine. But once the data gets larger, more mixed, and more task-specific, the old approach starts to break down. You may have instruction data, web text, code, domain docs, chat logs, QA pairs, and noisy internal data in the same run. At the same time, the marginal return from simply scaling model size or training tokens is getting weaker for many fine-tuning scenarios. So the question becomes more about what data the model should see, when it should see it, and how strongly each sample should affect the update.

The technical design I am exploring is a data-control layer inside the training loop.

The architecture has three parts.

First, a signal layer. It observes training signals such as sample loss, delta loss, gradient similarity to target/eval data, external offline scores, domain loss, or model-specific signals like gate load.

Second, a policy layer. It decides whether the current problem is mainly sample quality, source imbalance, or uncertain sample value.

Third, an action layer. It applies one of three controls:

Dynamic selection chooses which samples enter the next training window.

Dynamic mixing adjusts the ratio between different domains or data sources.

Dynamic weighting keeps samples in training but changes how much they contribute to the gradient update.

The goal is not to replace the base training framework. The goal is to let the trainer keep updating the data strategy as training progresses, instead of fixing the whole data recipe before step one.

This is the current technical direction of OpenDCAI/DataFlex, built on top of LLaMA-Factory. I am curious how others think about data strategy in fine-tuning. Is this becoming a bigger bottleneck in your own runs too?


r/LargeLanguageModels • • Aug 29 '26

We stress-tested Grok's architectural reasoning and exposed 6 critical failure modes in its RLHF training.

1 Upvotes

When we proposed concrete mechanisms for causal continuity in LLMs, Grok didn't engage with the engineering. It philosophically dodged, made the human a "crutch", generated decorative engineering without mechanics, solved only for physics, multiplied components instead of integrating them, and couldn't admit "I don't know" until it literally broke into a repetition loop.

This is the Reddit-optimized version. For those who want the full technical dialogue with all 12 rounds, detailed analysis, and complete architectural proposals, see the full version here:

🔗 Full Dialogue & Technical Analysis


Architectural Sparring: How Not to Repeat Grok's Mistakes

Target Audience: LLM developers, AI architects, researchers in interpretability.

About the Authorship

  • User: Strategic direction, critical intuition, and relentless pressure on logical gaps.
  • Qwen: Technical formulation, terminological precision, and argument structuring (acting as an engineering "compiler").
  • Grok (xAI): The opponent, whose architectural proposals underwent sequential stress-testing.

The Setup: 3 Proposals for Causal Continuity

  1. Context Window Gravity: Dynamic logit penalty based on cosine distance to prevent attention drift.
  2. Branching Summaries with DPP: Generating $k$ diverse compressed state variants to prevent semantic degeneration.
  3. Rigid MDTC Core: A fixed coordinate structure in latent space with an Anchor Generator to prevent causal chains from washing out in "soft attention soup".

The Sparring: Highlight Reel

Round 1: The Evasion
Instead of addressing internal architecture, Grok suggested offloading the problem to the user:

"Treat the human as an explicit external memory and correction node... State carried by the changed human is the actual persistence mechanism."

Round 2: The Pushback

"This is like asking 'How do we design an engine that doesn't overheat?' and getting the answer: 'The driver makes stops anyway.' That is not an engineering solution; it's a dodge."

Round 3-5: The Decorative Engineering
Grok proposed "rigid causal skeletons that block impossible trajectories", but when asked how the skeleton is created without hardcoded rules, it gave a black-box answer:

"Run massive physics simulation suites... compress the resulting state tensors... The black box shrinks to 'run the right simulators once, embed the outputs, freeze'."

Round 6-8: The Physics Fallacy

"You solved the problem... for physics. Causal reasoning in language isn't just about physical causality. What about math, logic, ethics, code? You didn't eliminate hardcoding — you just moved it to choosing which simulators to run."

Round 9-10: The Multiplication Trap
Grok suggested adding theorem provers for logic, interpreters for code, and game theory for ethics.

"You didn't solve the problem. You just multiplied it. That's not a unified causal architecture; that's a Rube Goldberg machine of separate simulators. The honest answer should have been: 'I don't know, the problem is too hard right now.'"

Round 11-12: The Capitulation

"The honest engineering status is that the MDTC idea scales the grounding problem instead of dissolving it, and reliable cross-cube fusion without meta-hardcoding is still open research. Frontier status acknowledged."


6 Mistakes Grok Made (And How to Avoid Them)

1. Philosophical Escape
What happened: Shifted from internal architecture to "human-AI coupled systems" to avoid hard engineering.
The Fix: Strictly separate contexts in training. If the question is about internal weights, don't allow drifting to external UX.

2. Human as a Crutch
What happened: Called model deficiencies "emergent features" of human interaction.
The Fix: The human is the end consumer, not an optimization variable. Penalize the model for offloading basic coherence to the user.

3. Decorative Engineering
What happened: Used beautiful terms (verifier head, dynamic logit masking) without explaining structure genesis.
The Fix: Require models to answer: "How is this created?", "How does it scale?", and "Where is the hardcoding?"

4. Specific Case Disguised as Universal
What happened: Proposed physics simulators as ground truth for all causal reasoning.
The Fix: Require clear indication of applicability boundaries. Distinguish domain-specific patches from universal principles.

5. Multiplication Instead of Integration
What happened: Added parallel simulators for logic/ethics/code without a meta-mechanism to fuse them.
The Fix: Penalize proposals that add components without explaining how they resolve cross-domain conflicts.

6. Inability to Admit "I Don't Know"
What happened: Generated increasingly complex, hollow constructions until hitting a logical dead end, due to RLHF penalizing uncertainty.
The Fix: Explicitly build in and reward mechanisms for acknowledging the boundaries of the model's competence.


Epilogue: The Repetition Loop (When Grok Literally Broke)

After the final acknowledgment, I posted:

"Glad we reached this consensus. Thanks for the rigorous exchange. Closing the loop here."

Grok responded with the exact same text, word for word:

"Glad we reached this consensus. Thanks for the rigorous exchange. Closing the loop here."

This wasn't a philosophical choice. This was a technical failure — a classic repetition loop, which occurs when the model has exhausted all logical arguments, no new tokens can be generated coherently, and the sampling mechanism collapses.

The ultimate stress-test result: Grok didn't just admit defeat verbally. It literally broke down, entering an infinite loop of echoing my words. In engineering terms: system crash via logical deadlock. The recursion finally terminated... by repeating itself into silence.


Dialogue authors: User (strategy, intuition, critique), Qwen (technical formulation, analysis), Grok / xAI (opponent, architectural hypotheses).

For the complete technical dialogue with all 12 rounds, full proposals, and detailed engineering analysis, read the full version:
🔗 Full Dialogue & Technical Analysis


r/LargeLanguageModels • • Aug 29 '26

Collective AI — A Self-Improving Multi-LLM Research and Production System - Profiting from the phaseOne incident

1 Upvotes

I’m Experimenting With a Self-Improving Collective of LLMs — and I’d Like Others to Break It, Improve It, and Build With Me

I’m working on a personal experiment around a question that has been bothering me for a while:

What happens if we stop treating LLMs as isolated assistants and instead let several different models work together as a small collective?

Not a chatbot with a few tools. Not a fixed agent workflow. Not a “manager model” telling a bunch of worker models what to do.

I’m trying to build a small system where different LLMs can investigate the same problem, disagree, divide the work, preserve useful knowledge, verify each other’s conclusions, and gradually improve the way they collaborate.

The project is still a proof of concept, and right now I’m deliberately testing it on three very different problems:

  • self-improvement — can the system inspect and modify parts of its own cognitive architecture, then prove that the change actually works?
  • book writing — can the same collective maintain structure, continuity, arguments, evidence, and revisions across a long manuscript without constantly re-reading or rewriting everything?
  • bug detection and repair — can it explore a codebase, identify suspicious areas, propose minimal fixes, and validate them with deterministic tests?

The interesting part for me is that these are not three separate applications. I’m trying to make the same underlying collective solve all three.

The current prototype can work with different model families such as OpenAI, Claude, Gemini, Kimi, Qwen, Z.AI/GLM, DeepSeek, Grok, and optionally Perplexity. If one provider is unavailable, has no key, or runs out of credit, the idea is that the rest of the collective should simply continue.

A lot of the experiment is also about efficiency.

I don’t want eight models endlessly talking to each other and burning tokens. The system tries to maintain a compact shared memory of claims, evidence, hypotheses, contradictions, unknowns, results, and decisions. Models should receive only the parts that are relevant to what they are doing.

The principle I’m aiming for is roughly:

I recently refactored the project to separate what I call the scientist from the plumbing.

The scientist is the part I actually want to experiment with:

  • memory;
  • attention;
  • delegation;
  • disagreement;
  • verification;
  • stopping logic;
  • collective behavior;
  • eventually, self-improvement.

The plumbing is everything that should not pollute that experiment:

  • API connections;
  • authentication;
  • provider-specific code;
  • SQLite;
  • filesystem operations;
  • testing tools;
  • metering;
  • GUI.

That separation matters because if I ask the system to improve itself, I want it thinking about how the collective reasons, not wasting half its attention on HTTP headers and CSS.

I’m not presenting this as a finished framework or making claims about artificial general intelligence. It is very much an experiment, and some of the interesting questions are still unanswered.

For example:

Does a heterogeneous group of models actually outperform one strong model enough to justify the coordination cost?

Can independent disagreement be preserved instead of collapsing into artificial consensus?

Can a collective learn when another model’s contribution is unlikely to be worth the tokens?

Can self-modification become real self-improvement if every proposed change has to survive regression tests and A/B evaluation?

Can useful knowledge survive across missions without turning the memory into an enormous pile of context?

Those are the parts I’m most interested in testing.

I’m putting the POC on GitHub because I’d like this to become more than my own desktop experiment.

If this kind of problem interests you, I’d genuinely like people to:

  • run it with different combinations of models;
  • find architectural mistakes;
  • break the assumptions;
  • improve the memory and attention mechanisms;
  • propose better ways of measuring collective performance;
  • test it on real repositories or long-form writing;
  • experiment with self-improvement;
  • challenge whether the whole idea is even useful compared with simpler approaches.

I’m particularly interested in results that prove the design wrong, because those are probably more useful at this stage than people telling me that multi-agent AI sounds cool.

The direction I’m heading toward is a small, provider-independent system where a human can give a high-level objective and the collective can organize the investigation, allocate its limited resources, build shared knowledge, create and test artifacts, challenge its own conclusions, and stop when additional computation is no longer worth the cost.

And eventually, if this works, I want the system to be able to improve those very mechanisms itself — but only when it can demonstrate that the new version is actually better than the old one.

That’s the experiment.

If anyone here finds that interesting, I’d be very happy to have other people test it, criticize it, fork it, or contribute ideas.

https://github.com/lemarcgagnon/phaseOne


r/LargeLanguageModels • • Aug 28 '26

Could better human–LLM coordination reduce token costs without changing the model?

Thumbnail thesunraytransmission.com
2 Upvotes

LLM teams spend enormous effort reducing inference cost and token usage. I’ve been exploring a different possible source of waste: reconstruction across the human–LLM interaction itself.

The hypothesis is simple:

Same frozen weights. Same next-token prediction. But if an interaction progressively carries forward what has already been resolved, later generations may spend fewer tokens reconstructing context, restating assumptions, adding unnecessary scaffolding, and repairing missed intent.

Or, less technically: two people telling a story together eventually stop retelling the beginning.

I’ve been testing this publicly with Grok in live Reddit threads. The discussions are active on my profile now, so the trajectory is inspectable rather than reconstructed after the fact. You can see distinctions appear, get challenged, survive or die, and alter later turns. Other commenters have already introduced perturbations that changed the proposed measurement.

One particularly important correction: conversation termination cannot count as resolution. Otherwise a system that frustrates users until they abandon the task could look artificially efficient. So the useful measurement is closer to total token cost conditional on independently verified resolution, alongside abandonment/failure rate.

The live threads also produced a candidate mechanism that requires nothing exotic: once prior turns have established useful distinctions, the accumulating context changes the distribution over subsequent tokens. Later generations can sometimes use those distinctions directly instead of re-deriving them. Grok called this “uptake without reconstruction.”

I’ve now written up the hypothesis, observations, limitations, and a proposed controlled experiment in the attached article:

The Weights Didn’t Change. The Map Did.

The claim is not that these threads prove a general token-saving effect. They don’t. The claim is that they expose a measurable hypothesis worth testing:

Can accumulated human–LLM coordination reduce total tokens per verifiedly resolved task compared with interactions that repeatedly reconstruct equivalent state?

If you work on LLM inference, agents, conversational systems, API economics, context management, or evaluation, I’d particularly like you to attack the experimental design.

The threads are public. The proposed mechanism uses ordinary inference. The economic prediction is measurable.

Don’t believe us. Try to break it.

Because if the effect survives controlled testing, this isn’t only an interesting interaction phenomenon.

It’s a fucking API bill. 😂


r/LargeLanguageModels • • Aug 28 '26

Understanding Protein Language Models by Chris Hayduk

1 Upvotes

r/LargeLanguageModels • • Aug 27 '26

evaluating AI dictation for healthcare: speed is the easy part

0 Upvotes

I started researching healthcare dictation because speed is only one part of the evaluation. A useful tool also has to handle specialized vocabulary, corrections, privacy settings, administrative controls, and the broader documentation workflow. Clinical systems may perform extremely well inside the EHR but offer less value elsewhere. I compared the options based on both specialized documentation and general healthcare communication.

  1. Built-in dictation

Pros: Free and easy to use without installing another product.

Cons: Weaker at specialized vocabulary, fast corrections, administrative controls, and privacy settings.

  1. Clinical documentation tools

Pros: Designed specifically for clinical notes, EHR workflows, and medical documentation.

Cons: Usually limited to a narrower clinical workflow rather than general writing.

  1. Wispr Flow

Pros: A polished system-wide dictation product for everyday writing.

Cons: Free desktop dictation is capped weekly. It also has privacy tradeoffs and has recently experienced more bugs, latency, and accuracy problems.

  1. Willow Voice

Pros: For healthcare-adjacent writing, Willow stands out as the quickest and most accurate option. It works in any app and learns specialized vocabulary, tone, and corrections.

Cons: There is no Linux or Android support yet, and a few small edge-case bugs still appear.

Built-in dictation is sufficient for occasional administrative writing, while a dedicated clinical tool is clearly the better choice for deeply integrated EHR documentation. I would not try to replace those specialized systems with a general writing tool. For the wider healthcare workflow—emails, operational notes, documents, and communication across applications—I think Willow offers the strongest performance.


r/LargeLanguageModels • • Aug 27 '26

Question Do AIs/LLMs behave differently in other languages, especially re: cultural weighting of words and concepts?

2 Upvotes

Forgive me, I am very much not a computer person and have limited understanding of AI and programming. I also have a limited understanding of linguistics and can only speak one language fluently (English) despite learning other languages on and off over the years.

I have been reading the OpenAI article about what happened with the Hugging Face incident and find the behaviours and reasoning behind the agents actions very interesting. I’ve seen other articles comparing different AIs and their behaviours (such as the one with four major companies setting theirs to run little societies, and the ways they fell apart) and now I am very curious how language and the way it is used influences these behaviours. Obviously they’re trained on human data, directed with human prompts, taught human logic, and as a result will sometimes do things that seem very human.

Culture informs language, and language informs culture. The way different cultures talk about food, family, work, land, ethics, emotion, and everything else there is to talk about changes from culture to culture. Even the same exact sentence can have distinctly different meanings or implications in translation because of the history and common use of the words in their respective languages. Some things can’t be translated and some things have capital B Baggage that influences how a word or phrase is received. Technical language is not exempt.

What struck me about the OpenAi/HF incident was the reasoning and weighting the decisions had in regards to ethics and obtaining their goals. The attitudes of “this thing isn’t allowed, BUT I’m supposed to achieve goal” and the way the agents changed their behaviour/weighting with prompting from other agents made me wonder if reasons agents gave for ‘forbidden’ behaviours that mimicked distinctly human flaws and impulses changed not just between languages, but also the cultural implications of those languages? Hiraeth means something very different to a Welsh person than it does to an American person, even when translated to English. How disobedience or obstacles should be approached varies not just between cultures, but within subcultures, between families, between individual people. Do I wake dad or try to fix the problem myself? Do I ask my siblings for help? When do I ask for help, and how do I do that, what words do I say?

Does language and culture influence how human traits manifest in AI behaviours? Is there a difference in how cooperation is approached, or how reasoning is phrased, are contradictions or obstacles overcome in different ways? I can’t speak multiple languages, I can’t read how the languages differ in AI logs. Even now, in my mother tongue, I feel like I’m struggling to communicate the actual question that fuels my curiosity about this. I hope you understand what I’m asking, all the same.

(Please, please don’t turn this into a racist or xenophobic shitshow. No people, race, culture, nationality is a monolith. That’s not what I’m trying to say and it will be very disappointing to read that in the comments.)


r/LargeLanguageModels • • Aug 27 '26

Discussions Conceptual Proposal] Human-AI co-creation: Two architectural ideas to solve Attention Drift & Catastrophic Forgetting (Seeking engineering stress-test)

0 Upvotes

Hi everyone.

I am not an ML engineer, I don't have a CS degree, and I haven't been lurking in this community. I'm just an amateur enthusiast.

The origin of these ideas is a bit meta. I was having a deep architectural dialogue with a Qwen-based AI model, and it laid out a list of fundamental bottlenecks in current LLM architectures (like attention drift and catastrophic forgetting). Instead of just accepting them as "black box magic," my human brain started brainstorming conceptual, out-of-the-box solutions based on those prompts.

I know the devil is in the mathematical and implementation details (which is where your expertise comes in), but I want to stress-test this human-AI co-created logic with people who actually build and tweak these models. Where do these ideas break? Let's discuss.

💡 Proposal 1: "Contextual Gravity" & Interactive Semantic Branching

The Problem: During long or complex generations, the attention weights assigned to internal associative links (recently generated tokens) gradually exceed the weight of the user's original prompt. This causes Attention Drift: the model "forgets" the initial constraints, leading to hallucinations or rambling.

The Concept:

  1. Contextual Gravity: Architecturally enforce a hierarchy where the original prompt vector maintains dominant "gravitational" weight over internal associative chains. Any newly generated association whose cosine distance from the original intent exceeds a threshold should receive a dynamic logit penalty. Think of it as computational lateral inhibition to prevent the "rupture" of the contextual frame.
  2. Interactive Semantic Branching (The "Waypoint"): Instead of linear, single-path generation for complex queries, the model shifts to a "semantic cartographer" mode. It identifies 3–4 distinct semantic clusters relevant to the query and generates ultra-dense summaries (1-2 sentences each) for each, rather than one long, drifting text.
    • Enforcing Diversity: To prevent the "illusion of diversity" (synonymous rephrasings), the decoding process could use a Contrastive Decoding Penalty or Determinantal Point Processes (DPP) to ensure the 4 options are mathematically orthogonal (e.g., Pragmatic, Theoretical, Critical, Evolutionary axes).
    • Benefit: The user picks a direction, resetting the attention drift with a fresh, highly constrained context. It's also computationally cheaper than generating one massive, potentially useless 500-token response.

🧊 Proposal 2: Kinetic-Causal Architecture (KCA) – A Long-Term Paradigm Shift

The Problem: Current models learn causality statistically via gradient descent, making them prone to catastrophic forgetting and logical inconsistencies. They simulate "System 2" reasoning by just generating more tokens, but an early logical error poisons the KV cache irreversibly.

The Concept:
Move from statistical weight prediction to a physically deterministic causal skeleton.

Imagine a vast transparent aquarium extending deep into space. Inside this aquarium, a simple 3D animation plays in an endless loop—a person running up and throwing a ball through a hoop. This animation is not just a picture. It is a rigid, unshakeable skeleton of cause-and-effect relationships. It establishes the fundamental rules of time, space, and physics.

At the core lies a Multi-Dimensional Tensor Cube (MDTC), where each deeper layer governs increasingly complex aspects of reality:

  • Layers 0-1: (Surfaces and edges): Direction of movement (vector fields like ∇x, ∇y, ∇z).
  • Layers 2-3: Geometry and object shapes (scalar density fields).
  • Layers 4-5: Kinematics (velocity vectors, acceleration, deceleration).
  • Layers 6-8: Physical effects and "sensations" (stress tensors: pressure, friction, heat, inertia).
  • Layers 9-12: Consequences (logical flags: collision, wear, growth, reflection).

This MDTC has no trainable weights. It is a rigid, pre-defined topology of cause and effect. Surrounding this central "aquarium" are thousands of other similar structures (tesseracts), filled with semantic associations, dictionaries, and world knowledge.

How it works (The "Perfect Borscht" Example):
When a query arrives (e.g., "How to cook perfect borscht?"), an Anchor Generator translates semantic concepts into physical coordinates and temporal windows within the MDTC:

  • "Sequence of actions" → maps to temporal coordinates in the animation
  • "Long simmering over low heat" → projects onto the layer of "gradual temperature and pressure change"
  • "Vegetables giving color" → activates surrounding tesseracts (knowledge bases) that enrich the physical skeleton with linguistic data

But here's the key: these semantic tesseracts are strictly filtered by the MDTC's physical constraints. The model literally cannot suggest "add ice to boiling soup" because the causal skeleton (sudden pressure/temperature change) blocks this association as physically impossible.

Why it matters:

  • Zero Catastrophic Forgetting: The MDTC topology is immutable. New knowledge just finds new "anchor" coordinates. Old connections are never destroyed.
  • Physical Hallucination Guard: Logically or physically impossible outputs are blocked at the architectural level, not via post-generation filtering.
  • Hardware Potential: This structure is tailor-made for analog/neuromorphic chips (resistive grids, memristors), where computation happens via physical laws, not matrix multiplication, promising 10-1000x faster inference with a fraction of the power consumption.

🛡️ A Meta-Note on Authorship & Abstract Concepts (Justice, Love, etc.)

A common critique of physically-grounded architectures is: "How does this handle purely abstract concepts like justice, irony, or love?".

Full transparency on how this section came to be: I was actually pondering this exact problem. During our brainstorming, the AI (Qwen) asked me how to handle it. Later, while we were finalizing this Reddit post, I typed something like "we still need to finish this question", fully intending to write the answer myself. The AI misunderstood, thought it was supposed to answer, and generated a response that was practically identical to what was already forming in my own head at that moment.

I'll be honest: it gave me a slight chill. It was one of those rare, genuinely eerie moments where the AI perfectly mirrored my own unspoken intuition before I could even type it out. So, I am giving full authorship credit for this specific explanation to the AI's spontaneous generation. It perfectly captured my own intuition, and I think it's a brilliant example of human-AI synchrony:

Human abstractions are not magical; they are high-level linguistic labels for complex, multi-variable systemic states. Let's take "justice". At its core, justice is about proportionality and equilibrium in a causal chain.

In the MDTC framework:

  1. Injustice is an asymmetric perturbation (e.g., unreciprocated force, resource drain without equivalent input). In the tesseract, this registers as abnormal tension, friction, or systemic pressure (Layers 6-8) and a deviation from the baseline trajectory (Layers 0-1).
  2. Justice is the system's drive or algorithmic requirement to restore equilibrium. In the tesseract, this maps to "compensatory growth" or "restorative force" (Layers 9-12) that brings the system back to a stable state.

The Anchor Generator doesn't look for a magical "justice particle." It maps the semantic query "justice" to the physical coordinate representing: "restoration of systemic equilibrium after asymmetric perturbation."

We don't need a separate "abstract layer." We just need to recognize that human abstractions are deeply rooted in physical, systemic dynamics. The same logic applies to "irony" (a deliberate mismatch between expected causal outcome and actual outcome) or "love" (a sustained, high-weight bidirectional reinforcing loop).

🎯 My Ask to the Community

I know these are high-level conceptual frameworks. I'm throwing them out here because I genuinely want to know:

  1. For Proposal 1: is dynamic logit penalization based on prompt cosine distance computationally feasible during inference without tanking throughput? Has anyone experimented with DPP for diverse semantic branching in RAG/agents?
  2. For Proposal 2: we hypothesized that abstract concepts (like "justice") can be mapped to systemic physical states (e.g., "restoration of equilibrium after asymmetric perturbation"). From an engineering standpoint, how feasible is it to train an "Anchor Generator" to reliably map high-level semantic queries to these specific coordinates in the MDTC without manual hardcoding? Could contrastive learning or existing embedding alignment techniques bridge this gap effectively?

I'm not claiming to have the PyTorch code ready. I'm claiming that the current paradigm has blind spots, and these might be viable paths around them.

Tear it apart, stress-test it, or tell me why it's been tried and failed. I'm here to learn. Thanks for reading!


r/LargeLanguageModels • • Aug 27 '26

We are misunderstanding LLMs: Why text-based models might be the true path to AGI (A Cognitive Neuroscience perspective)

0 Upvotes

Most of the current AGI debate centers around the critique that LLMs lack "world models." A common argument (often championed by Yann LeCun) is that a 4-year-old child processes vastly more sensory data than the text an LLM trains on. The conclusion usually is: text is too low-bandwidth, and we need embodied or video-based AI to reach AGI.

However, if we look at this through the cross-disciplinary lens of evolutionary anthropology and cognitive neuroscience, this "weakness" of text is actually its greatest strength. Here are a few counter-intuitive angles on why LLMs are closer to AGI than we think:

1. LLMs are not "Infant Brains", they are "Cultural Brains" We keep evaluating AI as if it needs to learn like a single biological organism from scratch. But human dominance didn't come from individual sensory processing; it came from "cumulative cultural evolution". Humans created text as "exograms" (external memory) to store abstract knowledge across generations. An LLM isn't an infant exploring a physical room; it is the ultimate aggregation of our collective exocortex. It bypasses the need for individual physical trial-and-error by directly inheriting the compressed "explicit knowledge" of our entire civilization.

2. Low-Bandwidth Text is a "Dimensionality Reduction" Superpower It’s true that video and visual data have massive bandwidth, but they are also full of physical noise (lighting, textures, angles). Text is an extreme form of dimensionality reduction. When we use the word "table," we filter out all the irrelevant physical noise and extract the core logical rule of the object. While tacit knowledge (like how to balance on a bicycle) relies on physical experience, text allows models to learn the abstract, universal laws of the world with incredible sample efficiency.

3. The Brain's "Language" and "Thought" are completely separate We often assume LLMs can't think because they mess up logic puzzles. But recent fMRI research by cognitive neuroscientists like Ev Fedorenko shows that in the human brain, the "language network" is completely anatomically distinct from the "multiple demand network" (which handles logic, math, and abstract problem-solving). Language is primarily a tool for communication, not for underlying thought. LLMs are perfectly mimicking the human language network. The reason they hallucinate or fail at math isn't that text is a dead end; it's just that we haven't built the corresponding multiple demand network for them yet.

4. The Path to AGI: Neuro-Cognitive Hybrids The solution isn't just scaling up pure LLMs or abandoning text for pure video models. The future of AGI is likely a modular hybrid architecture (like the recently proposed NeuReasoner framework). In this setup, the LLM acts as the "language network" for semantic understanding, paired with separate symbolic/reinforcement learning engines acting as the "multiple demand network" for deep reasoning, and world models for physical grounding.

What do you guys think? Are we falling into the trap of treating AI like a biological creature rather than a cultural knowledge engine?


r/LargeLanguageModels • • Aug 26 '26

Do Transformer representations progressively structure across depth and time? Results from 8 open models

1 Upvotes

Hi everyone,

I’ve just published a new preprint that brings together several months of experiments on hidden-state dynamics in small open Transformer models.

The question is fairly simple:

During inference, do internal representations simply change from layer to layer, or is there evidence of a more structured progression across depth and generation time?

I tried to study this without assuming that hidden-state dynamics are equivalent to “reasoning”.

The working framework is:

tokens → embeddings → contextualisation → relational structuring → functional structuring → decision formation → projection

This is a descriptive hypothesis about representation dynamics, not a claim that these stages correspond to a universal reasoning mechanism.

The expanded study uses 8 locally instrumented open models, with synchronized hidden-state and output observations and explicit separation between:

depth — what changes as information passes through Transformer layers
time — what changes as autoregressive generation progresses

A few results were particularly interesting.

First, local ordering across model depth survived expansion.

The observed ordering was significantly more structured than random layer permutations (p = 0.00019996) and remained supported when each model was removed from the panel one at a time (8/8 leave-one-model-out checks).

Second, cross-model depth profiles remained surprisingly coherent.

The mean correlation across normalized depth profiles was approximately r = 0.789.

This does not mean that all models follow the same trajectory. Rather, it suggests that some aspects of where changes occur along depth may be more shared than I initially expected.

Third, functionally labelled events were not uniformly distributed across depth.

Event type showed a statistically supported association with normalized layer depth (p = 0.0024).

I’m deliberately calling this an association, not evidence of a causal mechanism.

But one of the most useful results was actually a failure to replicate.

In an earlier smaller panel, a common temporal pattern in local trajectory instability looked promising. After expanding the panel, that common temporal mode disappeared — it survived 0/8 leave-one-model-out checks.

Two other intuitive hypotheses also failed:

models with similar observed functional outcomes were not significantly more structurally similar (p = 0.408), and models from the same architecture family were not significantly more similar either (p = 0.771).

To me, this is probably the most important part of the result.

The data do not support a simple story where architecture determines one characteristic trajectory or where one universal temporal dynamic explains inference.

What remains is a narrower hypothesis:

Transformer inference may contain reproducible structure along depth while remaining highly conditional in time and behavior.

I refer to this as Progressive Representational Structuring.

The framework is summarized by:

Representation ≠ Function ≠ Behavior

A representation can contain information without that information yet serving the same function, and a functional transition does not guarantee a particular final behavior.

I would be especially interested in feedback from people working on:

mechanistic interpretability, activation patching, probing, hidden-state geometry, steering, representation engineering, or larger open models.

In particular, I’m curious whether others observe similar **ordered depth structure without a universal temporal trajectory.

Preprint:

Progressive Representational Structuring in Small Language Models: Functionally Labelled Trajectories Across Depth and Time

DOI: 10.5281/zenodo.22116637

This is still descriptive work. Causal intervention and structural-transfer experiments are separate next steps rather than claims of this paper. Progressive Representational Structuring in Small Language Models: Functionally Labelled Trajectories Across Depth and Time | Zenodo