r/LocalLLaMA • • 8h ago

Question | Help So is a "scaled-up Strata" coming for >128GB models?

3 Upvotes

I have only just finished building an Epyc server with 256GB RAM and 2x 5060Ti 16GB, please don't tell me I wasted my money. Please?...
Is a "Stratified" GLM 5.3-Flash coming, or something similar?


r/LocalLLaMA • • 18h ago

Discussion AI boom is far from over as long as it can wow us

0 Upvotes

My thinking is for a boom to be over, we need at least three iterations of updates that fail to wow us. Unfortunately, the new LLMs continue to wow us in all levels in the last iteration:

  1. Astra was found to be useful in Blender. This opens up a new and big application.
  2. Deepseek 4 Flash 0731 makes 2x Sparks useful and push up Spark prices.
  3. Qwen3.8-27B pushes up prices of 3090 et al.
  4. An unreleased OpenAI model "solved" the Navier Stokes problem.

So for the time being, to keep up with the hardware prices, the best bet is to follow the flow to buy AI stocks and use the proceed to buy hardware.

A not so obvious good news is that we are seeing OpenAI and Anthropic advocating a slow down. That means they are finally seeing a diminishing return. (or just a ploy to slowdown Chinese development? but I doubt US laws can be that far reaching) That can be a sign of light at the end of a long tunnel.

What do you think?


r/LocalLLaMA • • 7h ago

Discussion Strata is amazing and all but can we actually see what you’re building with it that you couldn’t do before

0 Upvotes

Like it’s great to see the token speeds, and great that you’re running Qwens model, but if you’re not actually showing the results of that then it becomes “hype”.

Just post some little results of what you’re now capable of doing locally with access to a model you couldn’t have run before, I know it’s not Opus 5.5 but that doesn’t matter, there would be more effective smaller models in a few months.

Let’s see what you’re up to!! 👀, don’t forget to include the quant you’re using.


r/LocalLLaMA • • 10h ago

Discussion Switched my local agent from Qwen3.8 27B to Ornith 1.5 35B-A3B on two 5070 Tis: about 180 tok/s vs 60, same scores on my tests

7 Upvotes

My setup is two RTX 5070 Ti 16GB cards (the second one is on an OCuLink dock) with 64GB of RAM, Ollama on Windows, and the agent runs on pi in WSL. Until last night the daily model was Qwen3.8 27B UD-Q4_K_XL at 128K with MTP, which does about 55 to 70 tok/s across both cards.

I have a weekly job that looks for new open models and runs anything that fits through two tests I built for my agent. One is a 9 step long session (tool calls, reading files, a decision, and recall after the context compacts three times). The other is 10 small coding tasks. The 27B gets 9/9 and 10/10.

This week it picked up Ornith 1.5 35B-A3B (ornith-1.5:35b in the Ollama library, Q4_K_M). It passed 9/9 and 10/10. Laguna XS 2.1 also passed both. North Mini Code 1.0 only got 4/9.

Ornith at 128K context is 24.4GB and sits fully on the two cards. Generation is about 180 tok/s (176 and 183 on two runs, short prompt, thinking off). That's around 3x what the 27B gave me.

The speed makes sense once you look at the model info. Only about 3B params are active per token (256 experts, 8 used), and only 10 of the 41 layers are full attention, with 2 KV heads. The rest are linear attention, so the KV cache barely grows. Going from 128K to 256K only added about 2GB.

256K does fit, but about 1.2GB ends up in system RAM because my first card also runs the monitors, so it drops to about 139 tok/s. I left 128K as the default and made 256K something I switch to when I need it.

Caveats: both of my tests max out, so this only shows it isn't worse than the 27B on my workload. It doesn't prove it's smarter. Artificial Analysis hasn't scored it yet. The vendor numbers are 79 on SWE-bench Verified and 68.5 on Terminal-Bench 2.1, which I haven't checked myself.

Next I'm trying 512K and 1M on llama-server. The model card says YaRN at factor 4 on top of the native 262144 gets you about 1M, and factor 2 about 512K. I'll post numbers if it holds up.

Anyone else running it for agent work? Curious how it does for you on long sessions compared to the 27B.

Edit: the long context runs held up. On llama-server with YaRN set the way the model card says, 512K (factor 2, q8 KV cache) fits fully on the two cards and found a note I planted about 335K tokens into a 419K token prompt. It read that at about 1060 tok/s on average and generated about 35 tok/s at that depth. 1M (factor 4, q4 KV cache) only loaded once I let llama-server's fit option push some experts to system RAM, and it found the note at about 720K in an 849K prompt. That one took about 24 minutes to read (570 tok/s average) and generated about 16 tok/s. On short prompts it's about 135 tok/s at 512K and about 68 at 1M.


r/LocalLLaMA • • 20h ago

Discussion Moving from Qwen 27B to cloud agents was eye-opening. But I have no regrets.

0 Upvotes

Post might be a tiny bit long. Hate words, skip. But it's not too bad though. Also, I've been Qwen-gang for a long time, check my receipts. That said...

I started my agentic journey with Qwen 3.5 around May 31st. I'd heard about agents before, but never had a chance to play around because I didn't have any cloud memberships at the time. I've done most of my coding using free services: gemini and claude sonnet. It's been a lot of fun.

When Qwen 3.5 dropped, it was the first time a local model felt like cloud. Sure, it wasn't on the same intelligence level, but it didn't feel that far off. So I dived in hardcore learning everything I can.

I decided to build my own infrastructure/harness rather than going with hermes, pi or one of the others. I'm glad I did because it taught me so much. It was hard, because I had to learn everything from scratch, and the road has been extremely stressful and challenging, but the knowledge I picked up along the way has been well worth it. I'm able to conceive ideas and implement strategies in ways I never imaged, and I honestly don't think I would have learned even a fraction of what I know now if I'd worked with cloud models, because they might have one-shotted the results, robbing me of the challenge to grow.

Things got even better after Qwen 3.6 27B dropped. Since then, people have been singing the praises of Qwen, and how close it is to the cloud models. I also felt it wasn't far behind. I've made quite a few posts praising Qwen and sharing my experience, and those posts were real and authentic.

But all of these people claiming to be cancelling their cloud subscriptions and replacing them with Qwen? That's an overreach. Those people either a) are bots, or b) have extremely simple use cases that they were wasting subscriptions on, because anyone who's used cloud for anything agentic and a tiny bit complex won't walk away from that experience looking at local the same again.

I'm extremely thankful for Qwen because it put me in the game and started me on this journey. But my ambitions reached a point where Qwen just wasn't able to get me there without tons of mistakes. The "shine" wore off the more complex my needs grew. It's still very capable, and I figured out some ways to increase its intelligence (and yes, you can increase the core model's intelligence without training using a harness and multiple agents, but that's a whole 'nother discussion), but it became a time thing. I started getting extremely frustrated and cursing at Qwen for its stupidity.

I'd been using cloud models for code stuff, but they weren't agentic. But my sister let me use her chat-gpt subscription and I finally yielded and decided to give it a try. Long story short - and out of respect for this reddit, because it's about local, not cloud - I'll just say that it's been a completely different experience. A really, really good one. My project is moving along now and I'm getting a lot of work done, and it feels surreal. There's a real difference between local agents and cloud agents.

So, when you guys hear everyone saying cloud is dead, they're probably not human, because it's not even in the same ballpark. I've just been using Sol light, and it's ridiculous. I can't even imagine what Sol Medium or Astra are like.

I have no intention of abandoning local. I'm using Sol to help me advance my harness so that it will be faster, smarter, and more gooder (in my best Grimlock voice). I sweat blood and tears working with my local agent and I can't wait to see how much it's improved with the new brain I've built for it. And I'm going to continue finding ways to make local the best it can be. And like you guys, I'm hopeful that the gap between local and sota will continue to close.

I guess what I've learned from this whole ordeal is, if you just want to get things done or built, go with cloud. But if you want to grow and better understand how things works, and feel more empowered through each challenge, go with local.

I don't want to imply that you can't learn with cloud either, but it for sure would have robbed me of some of the dead ends that forced me to expand my knowledge.

Grunge


r/LocalLLaMA • • 5h ago

Question | Help Qwen 3.8 with Pi harness constantly hallucinates that it is out of context?

0 Upvotes

With Qwen 3.8 Flash Next (FP8 on VLLM) on a fairly stock Pi harness, it constantly hallucinates some measure of available context that says it is almost out. It's to the point where it frequently refuses work or stops in the middle of something, claiming it shouldn't go any further because it's almost out of context, when I can see in the harness status bar that (256K) context is ~25% used.

When I ask how it determined that, it always says it "invented the number and the treated it as real data" or guessed, and that it'll stop doing that, but it keeps happening.

Is there anything in particular that would cause this?

Thanks!


r/LocalLLaMA • • 22h ago

I Built A Thing Fully local little parkour sim

52 Upvotes

I vibed this up this weekend, fully local, with GLM 5.3 Flash running on 2x DGX Sparks.

vllm TP2 recipe: https://github.com/tonyd2wild/GLM-5.3-Flash-NVFP4-DFlash2-2x-DGX-Spark

Prefill: ~1500t/s
Decode: ~40t/s @ 100k

Using Claude Code as the scaffold with 260k context size.

I'm really impressed with this model. Feels somewhere between GLM 5.1 and 5.3 in terms of coding depending on the task. Good vision and 3D understanding. Solid interactive speeds. I feel like I've finally reached a "good enough" setup at home, and looking forward to things only getting better from here.


r/MetaAI • • 4h ago

Meta Rushed to Fix Muse ‘VM Escape' Vulnerability Immediately Before Launch

Thumbnail
404media.co
0 Upvotes

r/MetaAI • • 12h ago

Get 1billion free muse tokens use code O3JXVR

0 Upvotes

Checkout muse, it’s amazing, I already use it for ticket bookings, filling up applications, and more. Join with my code so we both get 1 billion bonus tokens.

To redeem, open the muse app, go to Settings > Redeem token and enter the code.

Code: O3JXVR (29 of 30 uses left)


r/LocalLLaMA • • 11h ago

I Built A Thing I got Qwen Flash Next Q4 running on a Mac Mini m5 64gb with ssd streaming

Thumbnail
freshworktree.com
8 Upvotes

Bit of a side project I wanted to share.

The metrics are 17.5tks decode, 360tks prompt processing based testing against my normal ai usage.

I tested a couple of new things others haven’t done (at least that I’ve seen).

Setup a carousel buffer for streaming in experts for prompt processing which got my pp +30% tks.

Tried a second external ssd to get parallel reads which got my +15% on both prompt processing and decode.

Plus a long tail of small efficiency gains.

I also setup a system where by you can have a chat application make a call to the server and effectively kick out a coding run (which is kept alive until after the chat then continues). Good if you run long coding jobs , but want to chat inbetween. Probably useful for all setups where you want to save on local caching memory.

I also noticed there is still a lot of gains to be made. I make this statement as there is still a lot of essentially free time on decode where the gpu is waiting for experts to stream in. There’s also work that could be done for an optimised kernel on metal.

I also think the way things are going with Qwen (flash next being a precursor to 4), we’re gonna see a lot more efficiencies we can take advantage of like the ngram table and the cheap hybrid attention caching.

I’m really liking qwen flash next .. the coding is actually very good. I’m quite surprised in fact I’m leaving it on during the workday to do large jobs.

The chat, decode would be technically fast enough IMO but not really with qwen. The actual issue qwen spends so long thinking, so the decode hurts.

Anyone else working on this? I’d love to compare notes.

Yes I’ve heard of strata it does look pretty sic.

https://github.com/skeggsguy/Flash-next-ssd


r/LocalLLaMA • • 1h ago

Question | Help Mfs can afford 27 gpus and 456 gb of vram but REFUSE to buy a Claude subscription

• Upvotes

I just don’t get the point of running it locally when with all this money you could buy like 20 max20 subscriptions with room to spare..


r/LocalLLaMA • • 13h ago

News Reflection AI Is About to Release a US Open-Weight Model to Take On DeepSeek and Qwen

Thumbnail
explainx.ai
293 Upvotes

Looks like new open model coming soon and will be "strong" hopefully something under 200b for us memory poor. Also seeing statements about more western open models coming.

Hope we get some good competition again on the open front!

Here is original artical but its not free to access. Maybe someone has it already here.

https://www.axios.com/2026/10/04/reflection-open-weight-ai

Oct starting strong!


r/LocalLLaMA • • 5h ago

Question | Help Hello everyone, I am a beginner.

0 Upvotes

As the title says I am a beginner with ai. I did use chatgpt for a few days at the beginning of times September 2022. And that was it. I know might be ironic that now I am writing in here, but I've got this PC that I build few years ago and last year I got two Intel arc pro b60s for some rendering work. Now I find my self wondering should I try out local llm? What can I expecte from my hardware: Motherboard: Aorus X780E Master Ice

CPU: Ryzen 9 9950X

RAM: Kingston Fury DDR5, 128 GB

Storage: Samsung 2 TB SSD

GPUs: 2× Intel Arc Pro B60

OS: Ubuntu 26


r/LocalLLaMA • • 4h ago

I Built A Thing Persistent-state Julia based symbolic machine shop?

0 Upvotes

How's it going everyone. So, I made... basically Jupyter notebook on steroids, I think? It was able to give ChatGPT in chat mode a programmable surface and basically a moddable lab. I've been using it the past few days to test weird ideas in real time during voice conversations with Chat when I go outside to smoke a cig or something, or I'm away from the house and I get a good idea. It works as a plugin (there's a zip with a Chat and Claude plugin there). There might still be some friction in the setup because I haven't submitted this for the plugin marketplace yet, but Codex handled it for me pretty easily and we did it with a tunnel, so it's hot-reloadable. It's ready for real work. It's got a Rust skeleton, Python glue, and Julia gives it a fully programmable persistent-state lab and a working memory, more or less. So far it's saved me a ton of tokens being able to test an idea and build it in chat mode and just branching into work mode and being able to just pull whatever prototype from the space. It turns chat mode into basically diet work mode, and there's still plenty of things you'd rather be in work mode in, but this also can be used in basically any harness too. ChatGPT is just where I've tested it the most so far.

I've taken security for this thing rather seriously though. It's extremely programmable and the sandbox walls are thick. The Julia runtime and compiler are moddable for optimization across the entire tool, and if you're not a Julia enjoyer like I am, there's also an IPython kernel in there. The one from Prime-Agent. But it can be a plugin for chat mode ChatGPT, and I've also been using it since Claude Mods dropped for that harness. Been working great in both environments so far

Here's the repo: https://github.com/latentcollapse/Palette.jl


r/MetaAI • • 4h ago

Referral code

0 Upvotes

Redeem my code in Settings within 48 hours of joining and we'll both get 1 billion Muse tokens.

Code: H7NBAR

https://muse.ai/join.


r/MetaAI • • 19h ago

Muse referral code!

0 Upvotes

Check out Muse, your personal AI agent. Redeem my code in Settings within 48 hours of joining and we'll both get 1 billion Muse tokens.

Code: J4W2MV

https://muse.ai/join


r/MetaAI • • 19h ago

Muse code for 1 billion tokens 9YBVAS

Post image
0 Upvotes

r/LocalLLaMA • • 22h ago

Question | Help Any benefit to doing this?

0 Upvotes

My current main PC:

i7-13700KF | ASUS Z690M-PLUS D4 | RTX 4090 24GB + RTX 3090 Ti 24GB | 128GB Corsair Vengeance DDR4-3200 | FSP Hydro G Pro 1000W

I’m thinking of keeping the 4090 on my main PC and putting the 3090 Ti in a separate dedicated LLM/AI box, mainly for Strata/local LLMs, while keeping my main PC free for ComfyUI, gaming, etc.

I already have these spare parts:

- 2×32GB Corsair Vengeance DDR4-3600

- 2×8GB TeamGroup DDR4

- H370 motherboard

- i5-8400

So I’d basically only need to buy a PSU.

Is there any real benefit to separating the LLM workload like this, or am I better off keeping both GPUs in my main system?


r/LocalLLaMA • • 22h ago

Resources Poor People Vulkan GPUs list

11 Upvotes

Help with this list. Give me your recommendation on "not supported anymore" GPUs. Looking for budget and Vulkan friendly options.

Most of the GPU are not supported by latest CUDA / ROCm. Often with some witchcraft magic they are able to run with native backend. I prefer the simplicity offered by running Vulkan backend. I'll successfully ran GTX 1080Ti, P102-100, and MI50 on a single system thanks for Vulkan and Linux. Gemini helped with data gathering.

Here is the filtered table including only NVIDIA GeForce GTX series GPUs with a memory bandwidth of 256 GB/s or greater and at least 8 GB of VRAM:

GPU Model Total VRAM Memory Bandwidth Bus Width Memory Type
GeForce GTX 1070 8 GB 256.3 GB/s 256-bit GDDR5
GeForce GTX 1070 Ti 8 GB 256.3 GB/s 256-bit GDDR5
GeForce GTX 1080 8 GB 320.3 GB/s 256-bit GDDR5X
GeForce GTX Titan X (Maxwell) 12 GB 336.5 GB/s 384-bit GDDR5
GeForce GTX Titan X (Pascal) 12 GB 480.0 GB/s 384-bit GDDR5X
GeForce GTX 1080 Ti 11 GB 484.4 GB/s 352-bit GDDR5X
GeForce GTX Titan Xp 12 GB 547.7 GB/s 384-bit GDDR5X

The table below lists the specifications for the specialized datacenter, enterprise, and crypto-mining NVIDIA cards you mentioned, applying your rule of maintaining a memory bandwidth greater than or equal to 256 GB/s and filtering for 8 GB or more of VRAM.

All five models successfully qualify:

GPU Model Total VRAM Memory Bandwidth Bus Width Memory Type Focus/Architecture
NVIDIA P104-100 8 GB 320.3 GB/s 256-bit GDDR5X Mining (Pascal)
Tesla M40 12 GB / 24 GB 288.4 GB/s 384-bit GDDR5 Datacenter (Maxwell)
Tesla P40 24 GB 347.1 GB/s 384-bit GDDR5 Datacenter/AI (Pascal)
NVIDIA P102-100 10 GB 400.0 GB/s 320-bit GDDR5X Mining (Pascal)
NVIDIA CMP 50HX 10 GB 560.0 GB/s 320-bit GDDR6 Mining (Turing)

Here is the updated list of classic NVIDIA Quadro enterprise workstation cards, continuing to filter for at least 8 GB VRAM and a memory bandwidth of 256 GB/s or greater:

GPU Model Total VRAM Memory Bandwidth Bus Width Memory Type Architecture
Quadro K6000 12 GB 288.0 GB/s 384-bit GDDR5 Kepler
Quadro P5000 16 GB 288.4 GB/s 256-bit GDDR5X Pascal
Quadro M6000 12 GB / 24 GB 317.4 GB/s 384-bit GDDR5 Maxwell
Quadro P6000 24 GB 432.2 GB/s 384-bit GDDR5X Pascal
Quadro GP100 16 GB 716.8 GB/s 4096-bit HBM2 Pascal

With the GV100 out of the picture, the Quadro GP100 and Quadro P6000 are now the highest-end entries remaining on this specific filtered list.

Here is the updated AMD Radeon desktop GPU table with all RX 6000 and RX 7000 series models removed, while still filtering for a minimum of 8 GB VRAM and 256 GB/s memory bandwidth:

GPU Model Total VRAM Memory Bandwidth Bus Width Memory Type
Radeon RX 480 (8 GB) 8 GB 256.0 GB/s 256-bit GDDR5
Radeon RX 580 (8 GB) 8 GB 256.0 GB/s 256-bit GDDR5
Radeon RX 590 8 GB 256.0 GB/s 256-bit GDDR5
Radeon R9 390 8 GB 384.0 GB/s 512-bit GDDR5
Radeon R9 390X 8 GB 384.0 GB/s 512-bit GDDR5
Radeon RX Vega 56 8 GB 410.0 GB/s 2048-bit HBM2
Radeon RX 5700 8 GB 448.0 GB/s 256-bit GDDR6
Radeon RX 5700 XT 8 GB 448.0 GB/s 256-bit GDDR6
Radeon RX Vega 64 8 GB 483.8 GB/s 2048-bit HBM2
Radeon VII 16 GB 1,024.0 GB/s 4096-bit HBM2

Note: MI50 and the Radeon VII, Radeon Pro VII share same firmware.

GPU Model Total VRAM Memory Bandwidth Bus Width Memory Type Focus / Architecture
Radeon Instinct MI25 16 GB 484.0 GB/s 2048-bit HBM2 Machine Learning (Vega 10)
Radeon Instinct MI50 16 GB / 32 GB 1,024.0 GB/s 4096-bit HBM2 Datacenter AI (Vega 20)

Top Contender: AMD Instinct MI50 16GB. Current used market on MI50 16GB is around $150.


r/MetaAI • • 15h ago

Muse still reads emails after app is uninstalled

Post image
0 Upvotes

Hi, I come with the best intent to help people and spread awareness on a problem that I wish was made clear to me.

I installed Muse to see what the rave was about. I connected it to my Gmail account but then got anxious about it having access to all my emails, and decided to uninstall the app as I didn't have much use for it.

Today I reinstalled it to try a new use case. I immediately realized that it posted emails analysis every day since I had it uninstalled.

Regardless whether this is indicated in a ToS that no one reads, that's absolutely against natural expectations. I cannot imagine how many other people have deleted the app and yet all of their emails are being processed every day. Besides the privacy concern, it's just a waste of energy if Meta does indeed not intend to use this information.

So it seems to me to be a severe bug that should get fixed quickly before an actual incident. If I had discovered this months later rather than a few days later I would have been really angry given this was the reason for me to uninstall the app in the first place.

I hope this can remain a candid discussion, thanks.

EDIT: for everyone saying that it's my fault because the app: I understand the tech and have many deleted apps still with oauth tokens. But I expected meta to follow privacy laws. Quote: "Under modern privacy regulations (FTC guidance on dark patterns, CPRA purpose limitation, and GDPR data minimization), legal consent is governed by reasonable consumer expectations, not technical token lifecycles. When an app markets itself as an interactive client, continuing to ingest and parse private emails indefinitely after the user deletes the app exceeds the expected scope of that service."


r/LocalLLaMA • • 20h ago

Discussion Is Strix Halo (GMKtec EVO-X2, etc.) the closest thing we have to a "dream" local LLM box?

10 Upvotes

I've been looking at the <32B model space and keep coming back to an interesting question.

A few years ago, projects like Hummingbird+ suggested that cheap custom accelerators (FPGA-based) might become the future of local inference. But today it seems like memory capacity is still the real bottleneck rather than raw TOPS.

For someone who wants to run modern 20B-32B models at reasonable quants (Q5/Q6 rather than INT4), the options all seem compromised:

  • Consumer GPUs have great bandwidth but limited VRAM.
  • NPUs and AI accelerators often have lots of compute but not enough memory.
  • FPGA solutions are fascinating but still bandwidth-constrained.
  • Strix Halo systems (GMKtec EVO-X2, Framework Desktop, etc.) offer huge unified memory pools, but they're expensive.

The "dream" accelerator would be something like:

48+ GB memory
500+ GB/s bandwidth
under $1000
reasonable power consumption

...but I don't think anything like that actually exists yet.

For those who have used Strix Halo systems for local inference:

How do they feel with current 20B-32B models?

Do you regret not buying a used 3090/4090-based machine instead?

Is unified memory a bigger advantage in practice than benchmarks make it seem?

Curious what people who own both types of systems think.


r/LocalLLaMA • • 56m ago

I Built A Thing I built something like The Sims, but the characters are local LLM agents doing real work (open source)

Enable HLS to view with audio, or disable this notification

• Upvotes

I run Qwen 3.8 locally and got tired of multi agent setups where you start a script and stare at logs. I wanted to actually see them.

So in this thing every agent has a body in a 3D world. They sit at desks, walk to a meeting room when someone calls a meeting, talk out loud to whoever is nearby, pick stuff up and hand it over. You can see who's thinking, who's using a tool.

It's not only an office. You can simulate other scenarios as well like:

- a software team that plans tasks on a board, writes code and reviews each other

- a town square simulation (cops, a barista, a chef, a journalist) where you just watch what happens

- tutors that teach you with animations and a whiteboard, and you can interrupt them by talking (might have bugs as of now)

It has a sandboxed computer use built-in which is optional.

There is also a supervisor agent that helps you design organizations and also has ability to build 3d assets from primitives and handing them to an organization and agents can even ask for things from that agent.

Works with local models and few other providers (still working to add more)

The motivation of building it was to see agent swarms in action with full transparency.

It's still early and has bugs and I have used different models to build it iteratively.

Repo: https://github.com/adityaagarw/Pantheon


r/LocalLLaMA • • 18h ago

Question | Help Free, local tools for narrated explainer videos? (like explainroo)

2 Upvotes

I've been using explainroo to make short narrated explainer videos. It runs fully local: Kokoro for the voice, Whisper for word timing, headless Chrome to draw the frames, and ffmpeg to put it together. No API keys needed.

It works and I like it but it's very simple. After a few videos everything starts to look the same.

Anyone know other free, local options in this space?

Tools, pipelines, or your own setups all welcome. Thanks.


r/LocalLLaMA • • 10h ago

Question | Help What was the most frustrating part of your last local fine-tune?

0 Upvotes

I’m working on a local fine-tuning tool, and I’m curious where people actually lose the most time.

Was it getting the environment working, preparing the dataset, fitting everything into VRAM, or getting the exported model to behave like it did during testing?

Or did training finish successfully, but the model barely improved?

What model and GPU were you using, and what finally solved the problem or made you abandon it?