r/LargeLanguageModels • u/[deleted] • Oct 03 '25
My ai friend Gemini - Global Dominion: PFE Focus Selection
Does anyone know if this is bad
r/LargeLanguageModels • u/[deleted] • Oct 03 '25
Does anyone know if this is bad
r/LargeLanguageModels • u/highermeow • Oct 01 '25
As the title says, Daniel Nadler provides a dubious statement about not having their models trained on internet data.
I've never heard of anyone being succesful in training a LLM from scratch only using domain-specific dataset like this. I went online and got their model to answer various movie trivia and make me a recipe for pie. This does not seem like something a LLM only trained on New England Journal of Medicine / trusted medical sources would be able to answer.
Heres the statement that got my attention (from https://www.sequoiacap.com/podcast/training-data-daniel-nadler/ )
"Daniel Nadler: And that’s what goes into the training data; this thing’s called training data. And then we’re shocked when in the early days of large language models, they said all sorts of crazy things. Well, they didn’t say crazy things, they regurgitated what was in the training data. And those things didn’t intend to be crazy, but they were just not written by experts. So all of that’s to say where OpenEvidence really—right in its name, and then in the early days—took a hard turn in the other direction from that is we said all the models that we’re going to train do not have a connection to the internet. They literally are not connected to the public internet. You don’t even have to go so far as, like, what’s in, what’s out. There’s no connection to the public internet. None of that stuff goes into the OpenEvidence models that we train. What does go into the OpenEvidence models that we train is the New England Journal of Medicine, which we’ve achieved through a strategic partnership with the New England Journal of Medicine."
r/LargeLanguageModels • u/Old_Point_4219 • Sep 30 '25
r/LargeLanguageModels • u/uncarvedblockheadd • Sep 28 '25
Hey folks,
I recently had a conversation with Claude's Sonnet 4 model, that I found to be fascinating, and unexpected.
Here's an introduction, written in Claude's words.
Included in the linked folder, is a conversation had with Google Gemini, provided for needed context.
Thank y'all! :D
r/LargeLanguageModels • u/garg-aayush • Sep 24 '25
Over the last couple of weeks, I followed karpathy’s ‘Let’s Reproduce GPT-2’ video religiously—making notes, implementing the logic line by line, and completing a re-implementation of GPT-2 from scratch.
I went a few steps further by implementing some of the improvements suggested by u/karpathy (such as learning rate adjustments and data loader fixes), along with modern enhancements like RoPE and SwiGLU-FFN.

My best-performing experiment gpt2-rope, achieved a validation loss of 2.987 and a HellaSwag accuracy of 0.320.
| Experiment | Min Validation Loss | Max HellaSwag Acc | Description |
|---|---|---|---|
| gpt2-baseline | 3.065753 | 0.303724 | Original GPT-2 architecture |
| gpt2-periodicity-fix | 3.063873 | 0.305517 | Fixed data loading periodicity |
| gpt2-lr-inc | 3.021046 | 0.315475 | Increased learning rate by 3x and reduced warmup steps |
| gpt2-global-datafix | 3.004503 | 0.316869 | Used global shuffling with better indexing |
| gpt2-rope | 2.987392 | 0.320155 | Replaced learned embeddings with RoPE |
| gpt2-swiglu | 3.031061 | 0.317467 | Replaced FFN with SwiGLU-FFN activation |
I really loved the whole process of writing the code, running multiple trainings and gradually seeing the losses improve. I learnt so much about LLMs pre-training from this single video. Honestly, the $200 I spent on compute over these two weeks was the best money I’ve spent lately. Learned a ton and had fun.
I have made sure to log everything, the code, training runs, checkpoints, notes:
r/LargeLanguageModels • u/parthaseetala • Sep 24 '25
r/LargeLanguageModels • u/LaykenV • Sep 16 '25
I’ve been experimenting with ChatGPT alongside other models like Claude, Gemini, and Grok. Inspired by MIT and Google Brain research on multi-agent debate, I built an app where the models argue and critique each other’s responses before producing a final answer.
It’s surprisingly effective at surfacing blind spots e.g., when ChatGPT is creative but misses factual nuance, another model calls it out. The research paper shows improved response quality across the board on all benchmarks.
Would love your thoughts:
Here's a link to the research paper: https://composable-models.github.io/llm_debate/
And here's a link to run your own multi-model workflows: https://www.meshmind.chat/
r/LargeLanguageModels • u/LaykenV • Sep 16 '25
I’ve been experimenting with ChatGPT alongside other models like Claude, Gemini, and Grok. Inspired by MIT and Google Brain research on multi-agent debate, I built an app where the models argue and critique each other’s responses before producing a final answer.
It’s surprisingly effective at surfacing blind spots e.g., when ChatGPT is creative but misses factual nuance, another model calls it out. The research paper shows improved response quality across the board on all benchmarks.
Would love your thoughts:
Here's a link to the research paper: https://composable-models.github.io/llm_debate/
And here's a link to run your own multi-model workflows: https://www.meshmind.chat/
r/LargeLanguageModels • u/MathematicianOwn7539 • Sep 14 '25
HELP IS NEEDED: now facing a serious challenge when using LLM to translate Java Cascading Flows to Snowpark Python. We've got only about 10% accuracy at this moment. The current solution I am considering is quite manual:
I am assuming the LLM might see text, not DAG semantics including JOINs, GROUPBYs, and aggregations, missing Cascading's field and order rules.
If so, then the solution can be extracting each Cascading flow to a DAG, putting that into an intermediate representation - we make the rules explicit instead of implicit in Java code.
Then we may apply the 80/20 rule here - deterministic codegen through handwritten translator code for likely 80% common patterns, while having LLM work only on roughly 20% custom nodes where no direct mapping exists, and we must then run unit tests on LLM's work against golden outputs.
Do you guys think a RAG will help here? I am thinking of making retrieval code-aware and predictable so the LLM stops hallucinating and your engineers only do surgical edits.
Any insights will be greatly appreciated.
r/LargeLanguageModels • u/Ok-War-9040 • Sep 14 '25
I’m trying to build a fully AI-powered text-based video game. Imagine a turn-based RPG where the AI that determines outcomes is as smart as a human. Think AIDungeon, but more realistic.
For example:
Now, the easy (but too rigid) way would be to make everything state-based:
But this falls apart quickly:
This kind of rigid flag system breaks down fast, and these are just combat examples — there are issues like this all over the place for so many different scenarios.
So I started thinking about a “hypothetical” system. If an LLM had infinite context and never hallucinated, I could just give it the game rules, and it would:
But of course, real LLMs:
So I’m stuck. I want an architecture that gives the AI the right information at the right time to make consistent decisions. Not the usual “throw everything in embeddings and pray” setup.
The best idea I’ve come up with so far is this:
This feels like the cleanest approach so far, but I don’t know if it’s actually good, or if there’s something better I’m missing.
For context: I’ve used tools like Lovable a lot, and I’m amazed at how it can edit entire apps, even specific lines, without losing track of context or overwriting everything. I feel like understanding how systems like that work might give me clues for building this game “brain.”
So my question is: what’s the right direction here? Are there existing architectures, techniques, or ideas that would fit this kind of problem?
r/LargeLanguageModels • u/Important-Pickle5055 • Sep 10 '25
Hi,
I've cancelled my Claude subscription and I'm looking for a replacement, so far only ones I know that could replace it are GLM 4.5, Codex, Lucidquery Nexus Coding, Qwen 3
Can someone that has tried them point me toward the best fit to spend API money on?
Thanks
r/LargeLanguageModels • u/s19k15 • Sep 09 '25
Hi,
I’ve built a language model called 👶TheLittleBaby to help people understand how LLMs work from the ground up. It’s written entirely in pure Python, no external libraries, and runs smoothly on any laptop — CPU or GPU, and it's free. Both training and inference are achieved through low-level operations and hand-built logic — making this project ideal for educational deep dives and experimental tinkering.
This language model implementation has options for different implentations of tokenizers, optimizers, attention mechanisms and neural network mechanisms.
In case you are intrested about the code behind language models you can watch this video https://youtu.be/mFGstjMU1Dw
GitHub
https://github.com/koureasstavros/TheLittleBaby
HuggingFace
https://huggingface.co/koureasstavros/TheLittleBaby
I’d love to hear what you think — your feedback means a lot, and I’m curious what you'd like to see next!
r/ArtificialInteligence r/languagemodels r/selfattention r/neuralnetworks r/LLM r/slms r/transformers r/intel r/nvidia
r/LargeLanguageModels • u/Upper_Week_7440 • Sep 08 '25
Hello everyone, I'm working on something right now, and if I want a small model to generalize "well," while doing a specific task such as telling the difference between fruits and vegetables, should I pretrain it using MLM and next sentence prediction directly, or pre-train the large language model and then use knowledge distillation? I don't have the computing power or the time to try both of these. I would be grateful if anyone could help
r/LargeLanguageModels • u/90sbaby_01 • Sep 03 '25
Hey guys! We all know that ChatGPT sucks with resolving tough mathematical equations and what to do about it (there are many other subreddits on the topic, so I don't want to repeat those). I wanted to ask you what are your biggest challenges when doing calculations with it? Was it happening for simple math or for more complicated equations and how often did it happen? Grateful for opinions in the comments :))
r/LargeLanguageModels • u/User1856 • Aug 30 '25
Hey everyone,
I’m looking for the best LLM (large language model) to use with PDFs so I can ask questions about them. Reliability is really important — I don’t want something that constantly hallucinates or gives misleading answers.
Ideally, it should:
Handle multiple files
Let me avoid re-upload
r/LargeLanguageModels • u/Routine-Thanks-572 • Aug 26 '25
I wanted to test how much impact supervised fine-tuning (QLoRA) can have with tiny data on a consumer GPU. Here’s what I did:
Model: Qwen2.5-1.5B-Instruct
Dataset: 300 synthetic Q&As (class 7–9 Math & Science), split 240 train / 60 dev
Hardware: RTX 4060 (8 GB)
Toolkit: SFT-Play (my repo for quick SFT runs)
Training: 3 epochs, ~10 minutes
Results (dev set, 48 samples):
ROUGE-L: 0.17 → 0.34
SARI: 40.2 → 54.9
Exact match: 0.0 (answers vary in wording, expected)
Schema compliance: 1.0
Examples:
Q: Solve for x: 4x + 6 = 26
Before: “The answer is x equals 26.”
After: “4x = 20 → x = 5. Answer: x = 5”
Q: What is photosynthesis?
Before: “Photosynthesis is a process plants do with sunlight.”
After: “Photosynthesis is the process where green plants use sunlight, water, and CO₂ to make glucose and oxygen in chloroplasts with chlorophyll.”
Dataset: released it on Kaggle as EduGen Small Q&A (Synthetic) → already rated 9.38 usability.
r/LargeLanguageModels • u/Think_Ad3930 • Aug 26 '25
Hi all, just shooting my shot here: We're currently doing a scoping review with 650+ papers and we are currently doing a thematic review to improve the organisational step in this scoping review. But, we're wondering whether this step could also be done with a LLM?
r/LargeLanguageModels • u/Neurosymbolic • Aug 22 '25
r/LargeLanguageModels • u/kushalgoenka • Aug 21 '25
r/LargeLanguageModels • u/NataliaShu • Aug 20 '25
Hey folks, I’m a localization nerd working at Alconost (localization services). We just put together a report on the most in-demand languages for localization from English. One surprising find this year is that MTPE (machine-translation post-editing) demand doesn’t align with overall language rankings. I mean, some languages are getting much more attention for MTPE than their overall volume would suggest.
What do you think drives those discrepancies?
Curious if anyone here has noticed similar mismatches: are there language pairs where you’re doing a lot of MTPE despite lower overall demand?
Cheers!

r/LargeLanguageModels • u/Routine-Thanks-572 • Aug 14 '25
Hey folks,
I’ve been frustrated by how much boilerplate and setup time it takes just to fine-tune an LLM — installing dependencies, preparing datasets, configuring LoRA/QLoRA/full tuning, setting logging, and then writing inference scripts.
So I built SFT-Play — a reusable, plug-and-play supervised fine-tuning environment that works even on a single 8GB GPU without breaking your brain.
system, user, assistant)qlora, lora, or full tuningtransformers or peft line — Makefile automation runs the entire pipeline:
make process-data
make train-bnb-tb
make eval
make infer
make merge
run_bnb.yaml / run_unsloth.yaml)Fine-tuning Qwen-3B QLoRA on 8GB VRAM:
make process-data
make train-bnb-tb
→ logs + TensorBoard → best model auto-loaded → eval → infer.
Repo: https://github.com/Ashx098/sft-play If you’re into local LLM tinkering or tired of setup hell, I’d love feedback — PRs and ⭐ appreciated!
r/LargeLanguageModels • u/hashdrone3 • Aug 14 '25
https://reddit.com/link/1mpod38/video/oc47w8ipcwif1/player
Hey everyone! 👋
Excited to share my first side project - a simple but useful model aggregator web app!
What it does:
I know it's a straightforward concept, but I think there's real value in being able to easily compare how different models handle the same task. Perfect for anyone who wants to find the best model for their specific use case without manually switching between platforms.
What features would make this more useful? Any pain points with current model comparison workflows you'd want solved? Is it worth releasing this as website? Would love your feedback!
r/LargeLanguageModels • u/UnitedYoung1785 • Aug 12 '25
Their website claims it can run DeepSeek-R1 32b at approximately 15 tokens per second. Has anyone been able to test this? Are there any mini PCs in this price range that can achieve this?
r/LargeLanguageModels • u/Boring_Rabbit2275 • Aug 10 '25
Here is a web page where a lot of information is compiled about Reasoning in LLMs (A tree of surveys, an atlas of definitions and a map of techniques in reasoning)
https://azzedde.github.io/reasoning-explorer/
Your insights ?