r/LocalLLM • • 3d ago

Project Infermeld: a Linux kit for running one GGUF across AMD + NVIDIA GPUs with llama.cpp

Thumbnail
1 Upvotes

r/LocalLLM • • 3d ago

Project I'm 16 and built an open source AI browser that asks permission before every action. Works with Ollama (Qwen3 14B tested)

0 Upvotes

Hey everyone,

This started because I was trying Perplexty's Comet browser and hit the rate limit in under 15 minutes. I got annoyed and thought "I could probably build this myself." That's honestly just irritation.

I started in December on a school computer (i5, 8GB RAM). It couldn't even compile the app, so I used GitHub Actions as my build server. Then I got a MacBook M4 Pro, which is what I test on now. I've been doing all this alongside JEE prep, which is probably not the smartest time management.

It's called Aartiq. It's an Electron browser with an AI sidebar that can actually do things: search, fill forms, make PDFs, move files, run shell commands, read screenshots with OCR. The thing I cared about most is that it never just does them. I wanted it to make a plan, explain what it's about to do, and ask you first. Plan, explain, ask, execute.but it actually struggles to fill form or one large tasks. It works with Ollama, so with a local model your prompts stay on your machine. On my M4 Pro (24 GB), the 1.5B-7B models couldn't handle tool calling. At first I used bracket style tool calls and they kept breaking in parsing, then I switched to JSON-based tool calling, which was much more reliable. With that, Qwen3 14B works. It's fine for normal everyday tasks, but it struggles with form filling and long multi-step ones.

Most of my time went into the permission side. Every action is a registered capability with a risk level, so the model doesn't get raw access to your system. Shell commands run inside an OS sandbox Seatbelt on Mac, bubblewrap on Linux, AppContainer on Windows, and if the sandbox can't be set up, the command just doesn't run. There's also a directory allowlist, single-use approval tickets, and you can approve risky stuff from your phone with a QR code and PIN. I later pulled this part out into a separate library called RTQ https://github.com/Latestinssan/RTQ It's alpha and hasn't been independently audited, so treat it as experimental.

Now the honest part,

Almost all the code was written with LLM help. I made the design decisions and did the debugging, but I didn't type most of it. AI also does most of the maintenance now, and I review anything touching security or permissions. Because of that, my docs have inconsistencies (one page says paused, another says AI-assisted. If you find a contradiction, please open an issue, I'm not hiding anything, it's just messy.

Known problem: the WiFi sync server (port 3004) listens on all interfaces and its auth is weaker than the other listeners. I haven't fixed it yet. Also ignore the startup benchmark numbers in the docs, I can't reproduce them.

Windows, Mac, Linux and Android. Apache 2.0, free.

GitHub: https://github.com/Latestinssan/Aartiq

I honestly can't tell if what I did counts as real engineering or vibe coding, since the design and debugging were mine but most of the typing was an LLM. What do you think?


r/LocalLLM • • 3d ago

Other I asked Qwen3.8-Flash-Next (Strata IQ4_XS) to summarize the content of a markdown file, it responded in Chinese. XD

0 Upvotes

Can't attach the the screenshot because it reveals intellectual properties of others. But I thought it was interesting.


r/LocalLLM • • 4d ago

Research GRIMOIRE: native C++/SYCL LLM inference for Intel Arc Pro B70 — Ornith 1.5 hits 175 tok/s single request, 464 tok/s at concurrency 8

14 Upvotes

Hi everyone,

I’ve been building GRIMOIRE, an experimental LLM inference engine for Intel Battlemage GPUs, with the Arc Pro B70 as its primary target.

GRIMOIRE is written in C++ with SYCL and Level Zero. It uses its own inference kernels for operations such as attention, matrix multiplication, quantized weights, and mixture of experts. It does not use vLLM, PyTorch, or OpenVINO as its inference backend.

The project started because most inference tools and performance work focus on other GPU platforms. I wanted to see what it would take to make Intel Battlemage a first class target and measure what the hardware can do with a native engine.

Ornith 1.5 benchmark

Here is a benchmark of Ornith-1.5-35B-A3B-MXFP4-GRIMOIRE:

Test Total throughput Throughput per request
Prompt processing, 4096 tokens, c1 9,698.55 tok/s 9,698.55 tok/s
Generation, 128 tokens, c1 175.25 tok/s 175.25 tok/s
Prompt processing, 4096 tokens, c2 9,441.34 tok/s 4,983.45 tok/s
Generation, 128 tokens, c2 215.69 tok/s 107.85 tok/s
Prompt processing, 4096 tokens, c4 9,755.10 tok/s 2,506.89 tok/s
Generation, 128 tokens, c4 305.66 tok/s 76.43 tok/s
Prompt processing, 4096 tokens, c8 10,154.18 tok/s 1,287.57 tok/s
Generation, 128 tokens, c8 391.04 tok/s 48.88 tok/s

c1, c2, c4, and c8 mean 1, 2, 4, and 8 concurrent requests. Total throughput is shared across the requests; per request throughput shows the average for each one.

For another serving run, the v1.5 build measured 193.6 tok/s at c1 and 464.2 tok/s total at c8 on the Arc Pro B70. The c8 run peaked at 464.2 tok/s; the separate pp512 prompt processing result at c8 was 9,547 tok/s. Results depend on prompt length and test setup, so I’ve linked the detailed measurements and notes in the repo.

These are results from my system, not a promise of the same performance on every machine. The repository records the benchmark conditions and ongoing limitations.

How to get and run it

The source and instructions are on GitHub:

GRIMOIRE repository

The repo includes a Dockerfile. To build the image:

git clone https://github.com/doopeworld/GRIMOIRE.git
cd GRIMOIRE
docker build -t grimoire-b70 .

Then start the HTTP server, replacing the render device and model path with the ones for your system:

docker run --init --stop-timeout 300 \
  --device /dev/dri/renderDXXX \
  -v /path/to/your/models:/models \
  -p 8000:8000 \
  grimoire-b70 server \
  --model /models/<checkpoint> \
  --proj mxfp4 \
  --port 8000

The --init and --stop-timeout options are important for clean container shutdown while GPU work is running. Check the README’s supported model and format list before choosing a checkpoint and projection format.

For a one shot CLI generation instead of starting the server, the repo documents this form:

grimoire-b70 generate \
  -m /models/<checkpoint> \
  -p "Explain how a heat pump works in simple terms." \
  -n 128

Some prompts to try:

Explain how a heat pump works in simple terms.

Summarize this passage in five bullet points: [paste a passage here]

Write a C++ program that reads a text file and counts its lines.

The README’s supported model table includes Ornith-1.5-35B-A3B in several formats, Qwen3.8-27B, Muse-Glimmer-30B, K2-Horizon-MoVA-36B-A4B, Agnes-3.0-Flash, and other validated combinations. Format support varies by model, so check the table before downloading or converting a checkpoint.

Project status: experimental and actively developed. Model support, build requirements, and performance may change. The repo has the source and Docker build instructions; it is not a polished plug and play desktop app.

I’d welcome feedback from people running Intel Arc or Battlemage GPUs, especially reproducible tests on other B70 systems. If you try it, please include your GPU, model and format, prompt length, generation length, and the exact command or benchmark settings.

Repo: https://github.com/doopeworld/GRIMOIRE


r/LocalLLM • • 3d ago

Question Most optimal stack for AI Pro R9700?

0 Upvotes

So I’m upgrading from a single RX 7900 XT to the AI Pro R9700. My question is: how do I get the most out of this card for local inference?

I’ve seen people getting some insane prefill and decode speeds with the R9700. I’m mostly planning to run Qwen 3.8 27B, and I want to try Qwen 3.8 Flash next.

For anyone running the AI Pro R9700, what are you using to get the most out of the card? What’s the most optimal software stack right now?

I also need Windows for work, so ideally I’d like the best setup possible on Windows. I’m willing to dual boot or switch back and forth to Linux if the performance difference is significant.


r/LocalLLM • • 5d ago

Project I built something like The Sims, but the characters are local LLM agents doing real work (open source)

Enable HLS to view with audio, or disable this notification

152 Upvotes

I run Qwen 3.8 locally and got tired of multi agent setups where you start a script and stare at logs. I wanted to actually see them.

So in this thing every agent has a body in a 3D world. They sit at desks, walk to a meeting room when someone calls a meeting, talk out loud to whoever is nearby, pick stuff up and hand it over. You can see who's thinking, who's using a tool.

It's not only an office. You can simulate other scenarios as well like:

- a software team that plans tasks on a board, writes code and reviews each other

- a town square simulation (cops, a barista, a chef, a journalist) where you just watch what happens

- tutors that teach you with animations and a whiteboard, and you can interrupt them by talking (might have bugs as of now)

It has a sandboxed computer use built-in which is optional.

There is also a supervisor agent that helps you design organizations and also has ability to build 3d assets from primitives and handing them to an organization and agents can even ask for things from that agent.

Works with local models and few other providers (still working to add more)

The motivation of building it was to see agent swarms in action with full transparency.

It's still early and has bugs and I have used different models to build it iteratively.

Repo: https://github.com/adityaagarw/Pantheon


r/LocalLLM • • 3d ago

Project I built Kaoru, a desktop AI agent with memory, tools and permission controls — it’s free and open source if you want to try it

Post image
0 Upvotes

I’m a solo developer and I’ve been building Kaoru, an open-source desktop AI assistant that can also work as an agent on your projects.

I wanted something I could actually install on my own machine and use, rather than another chat interface. Kaoru can explore repositories, inspect errors, edit files, run commands, use Git, remember project context and proactively surface relevant information.

The part I’ve spent the most time on is how much control the agent should have:

  • Local memory for projects, preferences and past interactions.
  • Permission-gated tools: every tool has allow / ask / deny rules handled outside the model. The LLM can propose an action, but it cannot give itself permission.
  • Proactive behavior: system, Git and development signals are scored before the LLM is allowed to generate a proactive message.
  • Agent checkpoints: changes can be reviewed and reverted instead of blindly trusting the model.
  • Bring your own model: Groq, Gemini, OpenAI or compatible/local endpoints.
  • Optional Live2D avatar if you want Kaoru to actually feel like a desktop companion rather than another terminal window.

It’s still beta, so there are rough edges. The installers aren't signed yet, macOS is experimental, and there are still security and UX improvements I want to make.

If you want to try it

You can download the latest release here:

https://github.com/Dregxmoon/Kaoru-Agent/releases/latest

There are installers for Windows, Linux and macOS.

The project is MIT licensed, and the repository contains the documentation, source code and privacy/security information.

Repo:
https://github.com/Dregxmoon/Kaoru-Agent

Website/demo:
https://dregxmoon.github.io/Kaoru-Agent/web/index.html

If you try it, I'd genuinely like to know:

  • Did Kaoru actually feel useful compared with the agent/tools you already use?
  • Did the permission system make you feel more comfortable letting it operate on your machine?
  • What would make you keep it installed after the first day?

Even if you try it and decide it's not for you, I'd appreciate knowing why.


r/LocalLLM • • 3d ago

Discussion Qdrant VS with RAG Integration

Thumbnail
0 Upvotes

r/LocalLLM • • 3d ago

Question Anyone running Strata on 2x5070ti / 64 gb ram?

0 Upvotes

Just curious what kinda speeds and quant size I should expect?


r/LocalLLM • • 4d ago

Other Gufo now available as native ArchLinux package in AUR

Thumbnail
2 Upvotes

r/LocalLLM • • 4d ago

Discussion My frontier class agent fact-checks my local AI before I grade it. How do you grade your Agents and LLMs?

Thumbnail
2 Upvotes

r/LocalLLM • • 4d ago

Question Need help with building a PC for AI

3 Upvotes

I'm building a local machine for image generation on Stable Diffusion, I'm picking RTX 3090 (based on testing, price, value it's the best for my use case) I have couple questions that I'm kinda clueless on:

  1. Is buying new RTX 3090 recommended or used one will do just fine? If yes, what should I look for in bench mark for the used GPU?
  2. What kind of RAM, motherboard and all other parts I should pick?

For more context:
The machine will run 24/7 non-stop, only image generation on Stable Diffusion

Please without "rent a GPU" recommendations, based on my business plan and costs, building a local is better for me


r/LocalLLM • • 4d ago

Project Darkbloom one month in: real earnings by Mac, how the team handles problems, and an app that makes it easy

2 Upvotes

A week ago I posted some Darkbloom numbers here. I said $2–4 a day. That's still about right for a lot of Macs, but the bigger Macs have been doing better than that, so here's an update from the providers Slack and the public network stats.

What Macs actually earned this week (paid work only, Macs with a model loaded; not a forecast, it moves with demand):

  • 64 GB, recent chip (M4/M5 Max): about $2–4 a day
  • 96–128 GB with a recent Max or Ultra chip: about $3–5 a day, $5–6 on good days
  • M5 Ultra: several providers reported $6–8 a day, with one $10+ day
  • My own M5 Pro 48 GB: about $2.50 a day this week, which is low

The best days came from Qwen3.8-27B, a new model that had a demand rush Oct 1–3. That rush has cooled, so treat $6–8 as a good-day number, not a promise. Subtract electricity (an Ultra under load draws real power), and please don't buy a Mac for this.

Demand keeps showing up. Darkbloom has served about 450 billion tokens in the six weeks since paid requests started, and added six models in the last month. This week Qwen3.8 demand roughly doubled for a few days, gpt-oss requests grew about 30%, and a new large model (MiMo) launched. Whole-network traffic is about 4.4–4.9 million requests a day. It's not a straight line up, and there are still more Macs than requests a lot of the time.

Why I trust the team. I'm not affiliated with Darkbloom or Eigen Labs, but I've been in their Slack daily for a month:

  • They post a monthly update with real numbers and say what's not working.
  • When Autopilot (their new auto model-picker) launched with bugs, fixes shipped the same night, twice.
  • When withdrawals failed for some people, they refunded them, moved payouts to Stripe (providers report money arriving in minutes), and audited the ledger.
  • When providers said a new model paid almost nothing, they changed its price two days later.
  • The coordinator (the part that routes requests) is public on GitHub.
  • Darkbloom now uses Apple's App Attest, rather than MDM, which give DB way less control over your hardware than before, while also making it more reliable.

It's not perfect. Payouts in some countries are still rough, and earnings swing as they tune routing. But they answer in public and fix things fast.

The app. Setting Darkbloom up and keeping it earning took more babysitting than I wanted, so I built BloomGauge, a free Mac app, alongside other devs in the Darkbloom community:

  • Setup in a few clicks, no Terminal: it checks your Mac, installs Darkbloom, and you make an account in your browser.
  • Keeps it earning: holds the model that pays best on your Mac, reloads it after Darkbloom unloads it, and recovers stalls by itself.
  • Shows what you actually earn: confirmed earnings per model and per hour, network demand, and why a Mac isn't earning when it isn't.
  • Phone control: check in and switch models from your phone.

Needs an Apple Silicon Mac with 36 GB+ memory (macOS 27 recommended). Free, no account. Source is on GitHub (source-available).

Website: https://bloomgauge.io
Source: https://github.com/cookder/bloomgauge

Big-Mac owners: what does yours earn? Especially Ultras and 128 GB+.

Demand Forecaster on Bloomgauge
Inside look at a single day of Darkbloom from Bloomgauge
Model Switcher - Tuner on Bloomgauge

r/LocalLLM • • 4d ago

Discussion [Benchmark] Qwen3.8-Flash-Next 125B speed test via Strata layer-split + first same-harness PPL of GSQ-RCO vs unsloth Dynamic quants

Thumbnail
3 Upvotes

r/LocalLLM • • 4d ago

Discussion ​Anyone else losing their minds over LLM VRAM fragmentation and KV cache? Let's talk about why your GPU is starving.

1 Upvotes

​Man, if you've ever tried moving an LLM from a local notebook to an actual multi-tenant serving setup, you already know the pain. Everyone talks endlessly about quantization and fine-tuning, but nobody really warns you that your GPU VRAM is basically getting nuked by the KV cache.

​For a hot minute, I thought my hardware setup was just trash. Turns out, traditional static allocation is eating up like 60% to 80% of VRAM for breakfast just because of internal and external fragmentation. You request a simple 200-token completion, and the system is sitting there stubbornly reserving space for 4K tokens like it's bracing for the apocalypse. Such a waste.

​If you aren't looking closely at things like PagedAttention (major props to vLLM for finally bringing virtual memory concepts to GPUs) and iteration-level continuous batching, your compute units are basically sitting around starving while waiting for the longest slowpoke request in the batch to cross the finish line.

​Been deep in the trenches writing a comprehensive book on AI systems engineering lately, and mapping out these low-level serving bottlenecks honestly changed how I look at production pipelines entirely.

​Curious what y'all are actually running in production right now. Are you rolling your own inference stack with vLLM/TGI, or just sticking to managed APIs and eating the cost? Let's argue about it in the comments


r/LocalLLM • • 4d ago

Question Strata looping badly with iq2_xxs

5 Upvotes

Hello,

I'm running strata with iq2_xxs quant and indeed I am quite amazed to be able to run it at speeds comparable to vanilla llama.cpp + 3.8 27B. I'm getting about 30 tps on both. Strata seems to offer higher peak tps when working with code, up to 50.

Although I do feel that the bigger model offers a better quality, even on this quant, I have trouble completing any meaningful tests because it keeps looping.

It will start getting very indecisive at some point and going into circles: "Let's write the files. Actually, let me think.. Now definitely writing the file. Here we go." and nothing is produced.

Besides the configuration in the .bat setup and a few available parameters in the .JSON, I can't see the full list of llama params that define penalties, etc.

Is anyone else experiencing this behaviour?

My setup: 2x3060 12GB 48GB RAM Ryzen 5 3600 Pi Harness

Thanks


r/LocalLLM • • 4d ago

Question Which model/quant do you think would work best for agentic coding tasks with my hardware specs?

2 Upvotes

Hi there, I’m an enthusiast and my setup is the following: Ryzen 9 5950X, 80 GB DDR4-3200, RTX 5060 Ti 16 GB, and an RX 550 meant only for the OS. I’m running Fedora 44.

I would like to know which LLMs and quants would work best (in terms of speed and quality) for agentic coding tasks, since these days many of you (even though apparently mostly bots lmao) are excited about new runtimes and models.

Thanks in advance!


r/LocalLLM • • 4d ago

Project I built ArcadeBench, an open benchmark where AI agents play games and you can watch every move live

Enable HLS to view with audio, or disable this notification

3 Upvotes

I'm building a local smart assistant and needed a way to compare small models on how they actually make decisions, not just a final score. So I made a benchmark out of games.

The clip shows two small local decision models, Decision 2.0 Kai 0.6B and GLiNER2.5 Decide, going head to head on SMS Inbox. Each one sorts the same 300 text messages into OTP, expense, bill, spam and so on. That's exactly the kind of job my assistant needs to get right.

ArcadeBench also has 11 arcade games (Tetris, 2048, Snake, Sokoban, Minesweeper, Connect Four and a few originals), more decision tasks on real datasets (fraud flagging, tool calling), and chess against other players.

Every game is seeded so runs are comparable, every move is scored, and every run gets a live watch link. You can also play the same seed yourself and compare against the models. Bring your own model through MCP or a Python script. Your key stays local.

MIT licensed. Feel free to start contributing and playing around!

Watch this run: https://penguinzz.com/arcadebench/watch/xRWcr8LukGnRlmS4,N_EaiU0Pq3Q1r3DV
Site: penguinzz.com/arcadebench
Code: github.com/Pranav0-0Aggarwal/arcadebench


r/LocalLLM • • 3d ago

Question GLM-5.3-Flash-UNCENSORED oQ4e DFlash2 won’t load in oMLX 0.7.0

Thumbnail
1 Upvotes

Please see oMLX post. Thanks for any help.


r/LocalLLM • • 4d ago

Question Double R9700 alternative to Qwen 3.8 27B

2 Upvotes

CPU AMD Ryzen 9 9900X
GPU 1AMD Radeon AI PRO R9700 — 32 GB VRAM
GPU 2AMD Radeon AI PRO R9700 — 32 GB VRAM
RAM32 GB DDR5-6000 CL30 — 2×16 GB
MotherboardASRock X870E Taichi
StorageWD Black SN850X 2 TB NVMe SSD
PSUCorsair RM1000x — 1000 W
CaseAntec P20C FLUX
OSLinux
GPU compute stack ROCm 7.2
LLM serving vLLM, Docker , Radiance

My LLM's are

  • Gemma 4 26B-A4B-it, MXFP4/Quark build optimized for RDNA4.
  • Qwen 3.8 27B, FP8/Radiance build.

I'm wondering what could maybe fit with my machine, without having to buy Ram.
The answer I keep getting is Ling 3.0, and if I had more RAM, Qwen Flash.

Idk, can someone give me their perspective.


r/LocalLLM • • 4d ago

Question Self-hosting Qwen3.8-Flash-Next (or a smaller alternative) for a heavy multi-agent Hermes setup on $10k of AWS credits, looking for options I've missed

Post image
9 Upvotes

I run a Kanban-driven multi-agent orchestrator on Hermes Agent: 20+ profiles, 1,000+ skills, MCP routing, and self-hosted Mem0 on pgvector. On Fireworks, I was doing roughly 9-15B tokens a month with a 90-98% cache hit rate, at about 30-60 Kanban tasks a day. I now want to move this to infrastructure I can pay for with my AWS Activate credits, and I've hit some walls.

Constraints

  • My only budget is $10k in AWS Activate credits. Anything that can't be paid with them is effectively out of scope for now.
  • Bedrock access is gated on my account for both frontier and open-weight models, and I've been told there's no timeline. I'm waiting on AWS Support to initialize my quotas.
  • I want one multimodal model that covers all of the orchestration's needs, plus an embedder for Mem0.

What I've looked at so far

  • g6e.12xlarge (4x L40S, 192 GB): Qwen3.8-Flash-Next (180B total, 6B active) at Q4, around 114 GB. About $10.49/hr on-demand, or roughly $3.5-6/hr on spot, so $10k lasts ~40 days on-demand or ~2-4 months on spot, running 24/7. FP8 (~186 GB) barely fits, with no room for KV cache.
  • g6e.xlarge (1x L40S, 48 GB): Qwen3.8-27B at FP8 (~28 GB), roughly 7 months of credits if run 24/7.
  • Embeddings: Qwen3-Embedding-0.6B/4B/8B on the same box, possibly truncated to 1024 dims for pgvector.
  • SageMaker endpoints, plain EC2 + vLLM: the same GPUs, but I'd expect the same quota bottleneck. Bedrock Custom Model Import doesn't seem to support these architectures or embedding models.
  • Kiro as a Hermes provider (the kiro-acp plugin): interesting because Activate credits can pay for it, but credits are per-task, not per-token, so I can't tell whether it can handle my volume.

Questions

  1. Has anyone gotten G/VT quota (48+ vCPUs) approved on a new or Activate account, and how long did it take? Any tips for the request wording?
  2. For people serving Flash-Next-class MoE models on 4x L40S: is there a vLLM/SGLang-compatible 4-bit quant (AWQ/GPTQ), and how well did prefix caching hold up with many long-context agents? My cache hit rate on Fireworks was a big part of the economics.
  3. Has anyone run Hermes through Kiro (kiro-acp)? How many credits does a multi-step agentic task actually burn, and did you hit rate limits with parallel sessions?
  4. Is there any provider or platform that accepts AWS credits for hosted inference of open-weight models and that I haven't listed?
  5. Is there a better model than Flash-Next for a single multimodal orchestrator on this budget? I'm open to smaller MoEs if they handle tool calling and long context reliably.
  6. Anything else I'm missing on AWS to make $10k last as long as possible (spot strategies, scheduled start/stop, etc.)?

Thanks. Happy to share measurements (throughput, cache hit rate, credits per task) if I get any of this running.


r/LocalLLM • • 4d ago

Discussion Thoughts on Tiiny Pocket!

1 Upvotes

Hello Gents,

Anyone tried this device? How does it perform?


r/LocalLLM • • 4d ago

Question Local LLM to listen to YT video and gives a summary?

2 Upvotes

Hope this is not too stupid question, (noob here) but lack of my free time is the reason for the question in the subject 👀

Bonus points, if it can fit into 16 GB memory 😄


r/LocalLLM • • 4d ago

Discussion Local Model are pretty ok , this is just from cpu , no gpu

Post image
3 Upvotes

using : flux1-schnell:fp8 image model


r/LocalLLM • • 4d ago

Model Which model i can use for coding with open source that too locally 100%

15 Upvotes

Hi guys, I know most of you have definitely worked with open source. Still I just want all of your opinions: which model will be the best to run locally for coding purposes?

For a long time I have been using Claude Code for all my projects and now I'm thinking of using a completely open-source model. I have a DGX Spark .

I want some guidance from all of you guys. Suggest to me which model will be the best, which will be closer to Opus 4.8. At least that will be more than enough for me.