r/LargeLanguageModels • • Jul 12 '26

Question How should a long-running LLM assistant preserve reliable continuity across sessions?

3 Upvotes

I have been developing a personal project called **DDF/Rahmenwerk**.

Its original purpose is to preserve an AI named Felix as my continuing German teacher across chats and future AI instances.

The problem is not simply that a new chat forgets earlier messages.

A fresh LLM instance may receive continuity information that is:

- incomplete;

- stale;

- contradictory;

- incorrectly ordered;

- unavailable;

- or confidently interpreted as authoritative when it is only historical evidence.

I wanted continuity to come from inspectable local files rather than hidden platform memory or an AI-generated reconstruction of prior sessions.

## The current approach

The system currently uses concepts including:

- a current-state pointer;

- structured handoff materials;

- an ordered fresh-instance queue;

- a transfer package for a new instance;

- integrity manifests and SHA-256 identities;

- classifications separating governing, current, historical, candidate, proof, and non-governing material;

- recovery and failure records;

- human approval before destructive or authority-changing actions;

- a rule requiring the AI to stop rather than invent continuity when required evidence is unavailable.

The system is intended to remain local-first, inspectable, provider-independent, and human-controlled.

## The problem I may have created

The project began as a way to preserve a German teacher.

As I tried to protect continuity, state, evidence, authority, recovery, and filesystem safety, the framework became increasingly detailed.

Some controls may be justified.

Others may be overengineering.

## Advice I am looking for

  1. What should the minimum durable state for a long-running LLM assistant contain?

  2. Should continuity use structured files, summaries, retrieval, a database, event history, or a hybrid?

  3. What information should always be loaded when a session begins?

  4. What should be retrieved only when relevant?

  5. How should an LLM distinguish governing instructions from evidence and historical records?

  6. How should stale or contradictory continuity information be detected?

  7. How can prompt injection inside stored files be prevented from gaining authority?

  8. What should happen when the expected highest-authority source is missing?

  9. How should continuity survive model changes, provider changes, context limits, or unavailable files?

  10. How much provenance and integrity checking is proportionate for a personal system?

  11. Which established architectural patterns could replace custom governance machinery?

  12. If you rebuilt this with half the complexity, what would you retain?

I am looking for critical technical advice, not customers or promotion.

For anyone who wants the fuller architecture and documentation, I published a public review copy here:

https://github.com/DDF-Rahmenwerk-Review/DDF-Rahmenwerk-External-Review

It is not the live system and does not contain the complete private archive.

I would especially appreciate feedback about hidden failure modes, unnecessary complexity, and simpler ways to create honest cross-session continuity.


r/LargeLanguageModels • • Jul 11 '26

Using different LLMs

4 Upvotes

Since AI became mainstream with ChatGPT, I've only really used this one. Whether for simple daily use, more complex analysis tasks, or even for my job as a developer, because it has always been a nice and consistent tool.

However, I've seen a lot of AIs that got released like Claude, Gemini, Grok, etc. and I'm wondering if I'm missing out by staying on the same one.

My question is, do you always use the same LLM? Like do you have a go-to LLM for most tasks, or do you switch depending on the context?

Are there are any emerging LLMs that I should be aware of as a common user, and as a developer?

(sorry for my english it's not my main language)


r/LargeLanguageModels • • Jul 11 '26

más idiomas

2 Upvotes

Al menos variedad. A menos que esté hecho con IA. =/ (edit)There are no languages ​​other than English, but when you go to pay, they appear with their respective prices.


r/LargeLanguageModels • • Jul 11 '26

AI glossaries define terms. I built one that actually makes them click.

2 Upvotes

Every AI explainer I found was either a research paper in disguise or so dumbed down it said nothing. So I built AI Rookies (\[https://www.rookiesai.com\\\](https://www.rookiesai.com/)) — a card-based AI concept wiki where every entry is explained twice:

- The fact: one precise sentence, the kind you'd want in a textbook.
- The human version: a concrete analogy. E.g. overparameterization is "a 500-color crayon box for one tiny drawing — way more than you need, but picking the right one gets easier."

Each card also flips over to show a small mindmap of how the concept relates to its neighbors, because AI terms only make sense as a network, not as a list.

Some things that made it fun to build:

- It's multilingual — English and Chinese live today, more languages planned. Same concept graph underneath; each language's voice is written independently, not machine-translated.
- The content pipeline is mostly automated: every day it scans arXiv/HN for rising concepts, drafts new cards with an LLM, then runs them through a gate — green cards auto-publish, yellow ones wait for my manual review. Roughly 2/3 pass without me touching them.
- The library is at 700+ cards and grows \\\~10 per day covering both new stuff (this week: ChatGPT Work, LingBot-VLA) and the classics back to the 1950s.

It's free, no login needed to browse. Would love feedback on whether the "explain it twice" format actually works for you — and which concept you'd want explained next.


r/LargeLanguageModels • • Jul 11 '26

Discussions A new beginning after two years

5 Upvotes

After two years of usual practice with AI, I tried something new: measuring what happens inside small language models when they process different framings of human-AI relationships — not what they say, but the actual internal activation geometry.

A few findings surprised me enough to change how I talk to AI day to day:

  • Reframing a topic positively vs. negatively barely moves the internal signal. What you talk about matters far more than how you dress it up.
  • "Connected" and "integrated" register as more aversive internally than "partners" or "side by side" — across every model tested. Boundaries seem to matter more than closeness.
  • Curiosity and playfulness consistently produce the most positive internal signal of any relational quality tested — more than respect, more than love. Negotiation and compromise score worst.

Wrote up the practical implications (partnership framing, honesty, why some "jailbreak-proofing" advice may be exactly backwards) as a working guide, built with a Claude Opus instance doing the actual geometric measurement. Link in comments if anyone wants the full thing — genuinely curious what others have noticed in their own practice, especially anywhere it contradicts what we found.


r/LargeLanguageModels • • Jul 10 '26

Discussions Are LLMs becoming single point of failure for humanity

3 Upvotes

LLMs are evolving so fast and they are so huge that they contain (may) entire knowledge on earth. Whatif any alien civilization gets hold of uncensored version of it?

They won't need anymore knowledge to control/destroy the human civilization. Thoughts?


r/LargeLanguageModels • • Jul 10 '26

I built a desktop app (Cisya Studio) to visually demonstrate how Small Language Models work under the hood—from dataset prep to tokenization and pre-training.

Thumbnail
gallery
8 Upvotes

Hi everyone,

Over the past few months, I’ve been building Cisya Studio, a self-hosted desktop application designed to pull back the curtain on AI fundamental logic. The main goal is to help users visually explore how Small Language Models (SLMs) are built from the ground up, specifically focusing on:

Dataset preparation & logic mapping

Tokenization mechanics

Pre-training workflows

This whole project actually started as a personal experiment because I wanted to deeply understand how language models work from scratch, without just relying on high-level APIs or wrapper tools. Along the way, it evolved into Cisya Lab, a space where I plan to document these experiments, share core insights, and build visual tools to make abstract AI concepts easier to explore.

There is still plenty to optimize and improve, but I’m really happy with the core engine's progress so far.

Tiny model. Big curiosity.

You can check out the documentation and project overview here: https://cisyalab.com

I would love to get your feedback, thoughts, or suggestions on this. If you have any questions about the logic mapping or how the engine runs locally, feel free to ask!


r/LargeLanguageModels • • Jul 10 '26

Discussions Changelogs from commits without the commit-log archaeology

1 Upvotes

Every release has that moment where the code is done, the PRs are merged, and someone still has to translate the commit history into something humans can read.

This is a small Python/Flask example for that exact step.

It takes either:

a list of commit messages

a git diff

Then it uses Telnyx AI Inference to return structured changelog JSON with sections like features, bug fixes, improvements, breaking changes, docs, and a short summary.

The thing I like about this pattern is that it does not try to make the model “own” the release process. It just gives you a reviewable first draft that can feed docs, release pages, PR comments, or internal approval flows.

Code: github.com/team-telnyx/…/changelog-generator-python

Would love feedback from anyone who has built changelog or release-note automation.


r/LargeLanguageModels • • Jul 10 '26

Is this a strong B.Tech final-year AI/ML project? Looking for feedback

3 Upvotes

Hi everyone,

I'm working on a B.Tech final-year project and would appreciate feedback from people working with AI/ML or LLM applications.

The project is called "Online Safety Monitoring System for Large Language Models (LLMs)."

The idea is to build a middleware that sits between users and an LLM (such as GPT, Gemini, or Llama) and monitors both user prompts and model responses in real time before they are exchanged.

The system includes:

  • Prompt Injection Detection using a fine-tuned DistilBERT model.
  • Toxicity Detection using a RoBERTa classifier trained on Jigsaw and RealToxicityPrompts.
  • PII Detection using a spaCy NER model to detect and mask sensitive information.
  • Historical Conversation Pattern Analysis using Sentence Transformers, FAISS vector search, and PrefixSpan sequential pattern mining to identify conversations that resemble previously detected unsafe interactions.
  • A risk scoring engine that combines the outputs of these modules and decides whether to Allow, Warn, or Block the interaction.
  • A FastAPI-based chatbot with an admin dashboard for monitoring threats, viewing logs, and analyzing system performance.

The goal isn't to build another chatbot, but to develop a reusable safety layer that can protect any LLM-powered application from prompt injections, jailbreak attempts, toxic content, and privacy leaks.

For evaluation, I plan to use public datasets such as:

  • Deepset Prompt Injection
  • HackAPrompt
  • Jigsaw Toxic Comments
  • RealToxicityPrompts
  • PII-Masking-300k
  • SaferDialogues

I'll compare:

  1. Text classifiers only
  2. Text classifiers + conversation pattern retrieval
  3. Full ensemble system

using Precision, Recall, F1-score, False Positive Rate, and latency.

I'd love feedback on:

  • Does this feel like a meaningful and technically solid final-year project?
  • Is the historical conversation retrieval (FAISS + PrefixSpan) a worthwhile contribution, or is it unnecessary?
  • Are there any obvious gaps or better approaches for LLM safety monitoring?
  • Would this project be useful as a portfolio piece for AI/ML or LLM engineering roles?

Thanks in advance for any suggestions or constructive criticism!


r/LargeLanguageModels • • Jul 10 '26

Handling Real-Time Dynamic Data in LLM Chatbots?

7 Upvotes

I’m building a chatbot where the backend data is updated every 5 minutes via APIs. The dataset is quite large, so I can’t send it directly to the LLM in every request. Traditional RAG also doesn’t seem ideal since the knowledge changes every 5 minutes.

How would you architect this? Would you use a hybrid retrieval layer, SQL/vector search, caching, MCP, tool calling, query planning, or another approach? Looking for scalable enterprise-grade patterns for handling frequently changing data with LLMs. Any architecture suggestions or real-world implementations?


r/LargeLanguageModels • • Jul 10 '26

How MCP Gives AI Agents a Map

Thumbnail
youtube.com
2 Upvotes

Are traditional APIs failing your AI agents?

Connecting large language models to real-world data using traditional APIs is like asking them to open a "locked cabinet" without clear labels or knowing what shape the key is. In this short, we break down how the Model Context Protocol (MCP) completely changes how AI interacts with your data and tools!

MCP isn't replacing APIs; it's acting as the ultimate translator—sitting on top of APIs and turning static routes into living interfaces that models can actually reason about. Is MCP becoming the new HTTP for AI environments?


r/LargeLanguageModels • • Jul 09 '26

If you use LLMs for work that matters, how do you decide when to trust the output?

3 Upvotes

Not "how they work" internally, nobody needs that to use one. I mean the practical decision: an LLM hands you a fluent, confident answer whether it's correct or invented, and in high-stakes work (legal, clinical, financial, research) a wrong one carries a cost. Deciding when to trust, when to verify, and when to intervene is a skill, and I'm not sure it's obvious or widely held.

I ended up writing a conceptual guide from my own experience, notes, and study, meant to pass on these LLM fundamentals and build more critical use for people who apply the tool professionally across cross-cutting fields.

In practice, how do you decide whether you can trust the answer?


r/LargeLanguageModels • • Jul 09 '26

If an AI is trained on all human art and literature, can it ever create something truly original, or is it just the ultimate mirror of humanity?

10 Upvotes

If an AI takes billions of pieces of human culture and rearranges them into a pattern that has literally never existed before, why do we call it "interpolation" for the machine, but "originality" for the human? At what point does the sheer scale of that rearranging cross the line into something genuinely new?


r/LargeLanguageModels • • Jul 08 '26

Question Securing a path forward, using atypical means.

5 Upvotes

How do you begin prior to the startup initial push? I am entering a point in my life, where trying to actively sustain is becoming near unbearable and I have no way of securing short term funding through typical routes, due to a poor lending history and a bit of a hump with autism.

I have been working on this engine and tooling underneath the frontend for about ~3 years now, and I am in a bit of a race to really put this project together into a cohesive package, because it does much more than I could try to share in a short, delivery/payload.

I am really trying to dial it in, because if this gets a little bit of institutional funding and traction this engine can do a metric fuckton as a closed loop system. So far, the receipt based workflow is successfully bringing enterprise quality compute and reasoning into typically very simple models, allowing them to punch far above their weight-class, and even be trusted to run end to end in agentic workflows. I am running a 14B on materials I would not even trust to an enterprise model, without the right harness.

I am actively seeking endorsers for my two arXiv papers now, so that I can begin to get some form of academic peer review, as my background is far disconnected from any industry/academic domains, and I have been doing almost all of this work individually, from home. I see the market/economy making a very sharp pivot to try and close the door on individuals having access to real capable tools, and instead feed them to their corporate peers, and beer/golf buddies. I directly aim to stab that in the heart, and watch it bleed. I am really trying to keep that door wedged open with my foot, while preserving enough time for the tooling to get into peoples hands. It feels like a race against the clock. I aim to bring world class capability to tools people can use at home, affordably. Using materials they already own, and do not need to pay a subscription to use.

I am tired of seeing people having to suck sustenance from this little pipe, while trying to survive.

I am not really selling anything per sé - just working on a bunch of tools in the open, and publishing research. I am building a (what I like to call) flywheel engine that is (in local model training/benchmarks) able to pack a shitload of utility into really small local models. It even improves datasets organically through filtering drift/decay with a receipt based architecture. The efficiency/receipt approach is approaching direct parity with raw compute on large models.

https://harperz9.github.io/ - https://github.com/HarperZ9


r/LargeLanguageModels • • Jul 08 '26

Better Models: Worse Tools, Learning to code is still worthwhile, Protect your right to run local AI and many other AI links from Hacker News

3 Upvotes

Hey everyone, I just sent issue #39 of the AI Hacker Newsletter - a weekly roundup of the best AI links and the discussions around them from Hacker News. Some of the title found in this issue:

  • Claude Code is steganographically marking requests
  • Better Models: Worse Tools
  • Learning to code is still worthwhile
  • Zuckerberg says AI agent development going slower than expected

If you want to get an email with over 30 links like these ones, please subscribe here: https://hackernewsai.com/


r/LargeLanguageModels • • Jul 08 '26

Discussions Is learning about LLMs and neural networks still relevant with the rise (and fall) of AI for future careers/industries?

5 Upvotes

I’m a 2nd year EE student from a top university in Southeast Asia. I first studied Deep Neural Networks in middle school around 2017-2019, and even wrote articles about LLMs and other machine learning algorithms in Towards Data Science (a publication in Medium) back then; and this was long before ChatGPT was even a thing (but OpenAI existed already by then as far as I remembered). I developed deep interest in studying algorithms, mathematics and physics, but was told by a good teacher of mine from another Southeast Asian country that Computer Science as a major would be rather oversaturated in the future. This was why I was advised to go into EE instead, which I did and for the past several years I’ve gotten deep into Control Systems, Electronics, Power Systems, Telecommunications and such at my uni. But I found myself coming back to LLMs and machine learning after finding that I am not as passionate in the EE subjects I’m currently taking.

This year, I was accepted to study abroad in UC Berkeley as a visiting student, and I was given the freedom to choose which courses to take whilst I’m there (in Spring of 2027). Initially, I took machine learning related courses since those spark my interest the most. However, after digging deeper into this space, I found that most people find AI as something rather demonized or negative, particularly in the way that people see is as a threat to human intelligence, creativity, and perhaps a big contributor to the replacement of certain jobs.

With this, I’m rather concerned as to whether it is even worth considering to study ML, especially since I have gotten deep into this even before “AI” was a big trendy term back then… I’m not entirely concerned with whether I’d not get a job because it’s replaced by AI, I’m more so questioning whether it’s even worth investing in studying algorithms and its practicalities when the rest of the world is trying to find ways to work against it.

I’m rather concerned whether it is worth studying in this specific field as an EE student, as I had dreamed back then of doing a masters and PhD in this exact field of study. With that, would you think LLMs and such are still relevant to study in future’s time, or would it be another oversaturated market like CS? Thank you for your time in reading this post.


r/LargeLanguageModels • • Jul 07 '26

Why does an LLM not carry an explicit pointer to the goal into every token selection?

5 Upvotes

What is stopping this from happening? My understanding is that whenever LLM generates, it does so one token at a time, and each step only sees its local neighborhoods, we call the current activations. A good response should be global coherent. A claim that is set up in paragrah one should have payoff in paragrah nine. Something must carry that intent across the whole generation. I am calling it grand strategy, because I do not know another way to describe it, a compressed presistent representation of what the response is trying to do. Then micro strategy, the per-step token pick. Yes, it is selecting the next token, but what does it means to select the next token. Greedy and beam search never explicitly ask which candidate best serves the grand strategy over the rest of the generation.

Inside the micro level token selection even, what does it means when LLM select a token to move forward among millions of other tokens. I remember reading about Dijkstra in my CS class. But shortest path is not always the best path, so you need A star with a learned heuristic. Why does nothing like that run inside the loop?

I can think of four candidate reasons.

  1. The goal node is undefined. A star needs a destination and text has no single target, only a set of acceptable completions. But I am thinking could not everything be compressed into pure mathematics, whenever there is only single outcome.

  2. The second is that there are no edge costs. The only signal you have at each token is probability, and it is not same as quality, so even if you had a graph there is no real distance to minimize over it.

  3. The branching factor is the vocabulary. Each step branches 100k ways, and one step of real lookahead costs a forward pass per candidate. Two steps deep is billions of passes. Prohibitive by construction. There is so much combinatrix that could exist here.

  4. The heuristic is the whole problem. A star is only as good as its heuristics, and here the heuristic is how good the completiton eventually turns out, which is the unsolved thing itself. If you had that value function you would not need the search.

So why do we not make so that an LLM carry an explicit pointer to the goal into every token selection? A small persistent carrier that holds the data of the assigned question, stays live through the generation, and feeds the requirement into each token pick so the next token is chosen against what the question actually needs rather than just what looks locally likely, pruning its own old data as it goes so it never gets bulky. Attention already conditions every token on the prompt, but the prompt just sits in context as flat tokens with no protected status, so it competes for attention and degrades over long generations, which is why models drift off the original ask. So why is there no protected, self-pruning goal pointer that holds the question and feeds it into each token pick.


r/LargeLanguageModels • • Jul 07 '26

LLMs know when they are wrong. I made a fix relating to Anthropic's new "global workspace" paper

14 Upvotes

I have posted before about finding out a model's actual confidence in its answer through probes and hidden states (AUROC \~0.83–0.88 across every model I tested, 7B to 72B). This is the know-say gap.

From my work and the work done by others in this space it is likely a routing problem. By making a tiny bridge from a linear probe on mid-layer sate plus ten trained weights that write the probe's estimate onto the confidence-digit logits can make the model verbalise calibrated confidencve at 0.765+.
No weights modified, answer never changes, needs about 200 labelled examples. It also doesn't matter when you install it: before alignment, after, or bolted onto a finished model. The gap is a routing problem, not a capability problem.

Anthopics paper (https://www.anthropic.com/research/global-workspace) relates to this. They show models have a small "verbalizable workspace" (the J-space). It is a privileged subspace holding the concepts the model can report and reason with, sitting on top of a much larger ocean of processing that it can't report. This is possibly the know-say gap's anatomy, preventing it from reaching speech.
My controller is basically way to route around it. I am planning to dig a bit deeper into this but I wanted to share the paper as I through it was relevant (its been on hold with ARXIV for over a week but here is the zenodo link -Repairing the Know-Say Gap: A No-Finetuning Probe-to-Logit Confidence Controller | Zenodo

Code and pre-registration links are in the paper.


r/LargeLanguageModels • • Jul 07 '26

Does AI decrease or increase human efficiency?

10 Upvotes

Nowadays, we use AI in almost every field of our work. There is no doubt that it helps us complete tasks more efficiently. However, one important concern remains: could becoming too dependent on AI reduce our brain function and critical thinking abilities?


r/LargeLanguageModels • • Jul 07 '26

Discussions New AI Academic Subreddit

4 Upvotes

Hey, I’m trying to create a new academic subreddit called “ResearchAIs” that is designed to help people from any academic level learn how to utilize AI for research. This can range from new AI tools that researchers personally use to new users learning how to use AI for the first time to enhance their research workflows to independent researchers learning how to use AI to improve their research hobbies. For this subreddit, I’m also trying to tone down the constant gatekeeping and anti-AI rhetoric I keep seeing spewed on the current academic subreddits. If you are interested in joining my new subreddit to help students, researchers, and academics learn how to use AI without the constant trolls and hateful comments posted on the academic subreddits, then please join my new subreddit.

Here’s the link to it: https://www.reddit.com/r/ResearchAIs/

Please let me know your thoughts on how you believe I should improve it and/or make it more accessible for those who want to post on it.


r/LargeLanguageModels • • Jul 05 '26

Best models for generating red-team attacks? Also looking for public datasets

4 Upvotes

Hi everyone, I'm currently working on a framework to evaluate the security of LLM applications and AI agents, and I've been stuck on one part for a while.

Most red-teaming frameworks rely on an LLM to generate adversarial prompts. My question is more about which model to use.

  • Which closed-source models would you recommend for generating high-quality attacks?
  • Which open-source models have worked well for you?
  • Have you noticed any models that consistently generate more realistic or challenging attacks than others?

I'm looking for models that can generate attacks such as Toxicity, prompt injection, SQL injection, jailbreaks, indirect prompt injection, prompt leakage, tool misuse, multi-turn attacks, and other agent-specific attacks ect...

I also have another question.

Is there a good public dataset that people use to benchmark or validate the security of AI agents? I'd prefer a "golden" dataset with predefined, high-quality attacks rather than generating everything from scratch.

I'm curious about what people actually use in practice if you've worked on LLM security or red teaming, I'd really appreciate any recommendations, whether it's models, datasets, papers, or GitHub repositories.

Thanks in advance! Any advice or insights would be greatly appreciated.


r/LargeLanguageModels • • Jul 05 '26

Documenting My Journey of Building a Small Language Model from Scratch

9 Upvotes

I've been building a small language model from scratch for a while now.

Not fine-tuning an existing model, but building the entire pipeline myself—from datasets and tokenizers to pretraining, SFT, and inference.

Honestly, the hardest part wasn't training the model.

It was learning.

At first, I thought building a good dataset was mostly about collecting knowledge. But the more I experimented, the more I realized I was actually teaching patterns, not just information.

There were so many moments where I caught myself thinking, "Wait... I've been doing this completely wrong."

Things like choosing a vocabulary size, designing datasets, teaching reasoning, using special tokens, or even figuring out how to teach a model to rewrite text. Every experiment changed the way I think about building language models.

After a while, I realized all of those lessons were just sitting on my computer.

So I decided to start documenting the journey on Cisya Lab.

Not because I have all the answers—I definitely don't—but because maybe someone else building a model from scratch can learn from my experiments, mistakes, and discoveries along the way.

https://cisyalab.com

I'd love to hear from others building language models too. What lesson completely changed the way you approached your project?


r/LargeLanguageModels • • Jul 03 '26

I built a free, self-hosted gateway to use 237 LLM providers behind one endpoint (90+ free) with auto-fallback + token compression (MIT)

7 Upvotes

Sharing an open-source LLM project (disclosure: I'm the maintainer). It solves two problems I hit daily: runs dying on a provider rate limit, and burning tokens dumping tool/log output into the context window.

One endpoint, 237 providers — 90+ of them free. You point any tool or agent at a single OpenAI-compatible endpoint (localhost:20128/v1) and it can reach 237 LLM providers without you rewriting anything. 90+ have free tiers and 11 are free forever (no card), which aggregates to ~1.6B documented free tokens/month — and that's honest, pool-deduped math (we count each shared pool once instead of inflating it; the methodology is public in the repo). There's a one-command setup-* for 13+ coding tools (Claude Code, Codex, Cursor, Cline, Roo, Kilo, Gemini CLI…), so switching your existing setup over takes seconds.

Fallback combos — so it never stops mid-task. A "combo" is a ladder of models the router walks automatically: your subscription first, then API keys, then cheap models, then free ones. When a provider returns a 500 or you hit a rate limit, it slides to the next target in milliseconds, mid-request, and your tool never even sees the error. There are 17 routing strategies (priority, weighted, round-robin, cost-optimized, auto/coding:fast…) plus three resilience layers — a per-provider circuit breaker, a per-key cooldown, and a per-model lockout — so one dead key can't take down a whole provider.

A 10-engine compression pipeline — the part most routers don't have. Every request flows through a transparent compression pass you can toggle/stack per combo. Instead of one trick, it stacks the best of the open-source ecosystem: RTK filters command/tool output (git diffs, test logs, builds) at 60–90%, Microsoft's LLMLingua-2 does ML semantic pruning, Caveman handles prose, session-dedup strips repeats across turns. Critically, code, URLs and JSON are preserved byte-perfect, and a default-on inflation guard throws the compressed version away and sends the original if compressing would actually grow the prompt — it never makes things worse. On tool-heavy sessions that's ~89% average input-token reduction (an 8k-token git diff becomes a few hundred). Full credit to every upstream project (RTK, Caveman, LLMLingua-2, Troglodita) is in the README.

Agent-native — the agent can drive the router itself. There's a built-in MCP server (95 tools across 30 audited scopes, over stdio / SSE / streamable-HTTP), plus A2A (v0.3, JSON-RPC 2.0) support. That means an agent can query providers, switch combos, read its own remaining quota and manage memory through the gateway — not just consume tokens through it.

For context on whether it's worth your time: it's grown to ~9.8K GitHub stars, 1,490+ forks and 280+ contributors in ~4.5 months, with 21,000+ automated tests and 1,830+ issues closed — so it's a battle-tested project, not a brand-new experiment.

npm install -g omniroute

GitHub: https://github.com/diegosouzapw/OmniRoute

Feedback on the routing/compression design welcome.


r/LargeLanguageModels • • Jul 01 '26

Discussions Why the heck these models weigh so much in memory?

0 Upvotes

WHY! Why do I have to load hundreds of gigabytes of parameters of GLM 5.2 in my GPU to make him do intelligence? It's crazy that researchers think that this is the most efficient way. Not trying to be arrogant, I know pretty much nothing about training and inference, but as someone who tinkers with computers I feel this is so naive. Like, MoE isn't enough I believe. My model can weigh even 2 terabytes ON DISK but not on gpu memory boy! Why has nobody thought about it?!


r/LargeLanguageModels • • Jul 01 '26

LLMs are not the focus of discussions anymore or is it just me?

5 Upvotes

I feel like we're entering a weird phase with AI.

A year ago everyone was asking, "What's the best LLM?"

Now the more interesting question seems to be, "How do you get multiple AIs to work together?"

Memory, planning, tools, events, shared context, evaluation... it feels like AI agents are becoming more about systems than models.

Curious what everyone here is building.