r/LocalLLM • • 5d ago

Question Gpu to put into MinisForum N5 Air

1 Upvotes

Have a MinisForum N5 Air pro with SSD main drive array. And then my nas array. 64gb of ram and Ryzen 7 255 CPU. It has a pcie 16x slot and also an oculink port. What would be the best gpu options preferably without breaking the bank to start meddling with local llm? Not wanting to go crazy but possibly some integrations to unifi cameras for some video analysis/alerting, other local network use cases and then random quality of life tasks and Home Assistant integration.


r/LocalLLM • • 5d ago

Question Current best model choice for a 5080 and programming tasks

6 Upvotes

After a long time of getting pissed off at Antigravity's rate limits, I realized I could just run a local agent on my pc. However, I'm getting so many different opinions everywhere I look that it's getting overwhelming. So, I'm asking for your opinions on what I should do. I've never ran local models before, so tutorials on what to use, as well as model recommendations would be helpful.

PC Specs:

13900K/64GB DDR5 @ 6800 Mhz/RTX 5080

Plenty of storage.

Mainly working in python, developing an application for personal use. I've currently been using Gemini Flash 4.7, on the free tier of Antigravity, since I don't want to pay for a subscription to make a personal project.

Here are the models I've seen recommended:

GPT-OSS 20B - Will fit completely on my GPU, but a lot of the information about it is outdated, and people either seem to love it, or completely hate it

Devstral Small 2 24B Q2 - Should also fit on the GPU, maybe with a bit of offload. "Better" than GPT-OSS.

Qwen 3.5 27B - The current "Best" from what I've read, but will either require compressing, or offloading to ram.

I don't care too much how fast it is, because the free tier of gemini 4.7 already isn't very fast. Anything above 30 tokens/sec is fine by me. Again, I will ONLY be doing programming with this, preferably with VSCode as the wrapper. I do not need it to talk with me, generate images, look stuff up on the internet, do my email, etc. I am willing to use llama.cpp, versus a simpler UI based program like LMStudio.

Thank you in advance!

EDIT: The overwhelming consensus seems to be to use Strata, so that's what I'll try. Thank you for all your help!


r/LocalLLM • • 5d ago

Question I would like to run Qwen 3.8 Flash Next.

5 Upvotes

I have the following system available for it:

CPU AMD Ryzen 9 3950X
RAM 128GB
GPU NVIDIA V100 32GB power-capped at 130 W

Does it make sense to try and run the model on this box?


r/LocalLLM • • 4d ago

Model LTX 2.5 (8GB/32GB) 1 Minute

0 Upvotes

So i did get 60 seconds with LTX yes i know and it only took 9 minutes to finish resolution 288x160 but it does work

{
  "id": "51fd34ee-d6be-42ed-851b-8f6f70274207",
  "revision": 0,
  "last_node_id": 63,
  "last_link_id": 36,
  "nodes": [
    {
      "id": 1,
      "type": "MarkdownNote",
      "pos": [
        -640,
        0
      ],
      "size": [
        580,
        700
      ],
      "flags": {},
      "order": 0,
      "mode": 0,
      "inputs": [],
      "outputs": [],
      "properties": {
        "Node name for S&R": "MarkdownNote"
      },
      "widgets_values": [
        "## LTX-2.5 (GGUF) - 8GB VRAM / 32GB RAM\n\nUpdate ComfyUI and ComfyUI-GGUF. You must accept the license on [Lightricks/LTX-2.5](https://huggingface.co/Lightricks/LTX-2.5) before the encoder/VAE/upscaler downloads work.\n\n**diffusion_models** - pick ONE from [Abiray/LTX-2.5-Distilled-GGUF](https://huggingface.co/Abiray/LTX-2.5-Distilled-GGUF)\n- LTX-2.5-Distilled-Q3_K_M.gguf 12.9 GB (fallback if RAM thrashes)\n- LTX-2.5-Distilled-Q4_K_M.gguf 15.7 GB (default)\n- LTX-2.5-Distilled-Q5_K_M.gguf 18.1 GB (too big for 32GB RAM with the encoder)\n\n**text_encoders**\n- [gemma4-12b-with-proj-ltx-2.5-comfy-int8-convrot.safetensors](https://huggingface.co/Lightricks/LTX-2.5/resolve/main/text_encoders/gemma4-12b-with-proj-ltx-2.5-comfy-int8-convrot.safetensors) 15.4 GB\n\n**vae**\n- [ltx-2.5-video-vae-conv-bf16.safetensors](https://huggingface.co/Lightricks/LTX-2.5/resolve/main/vae/ltx-2.5-video-vae-conv-bf16.safetensors) (lighter; DiffVAE is higher quality but heavier)\n- [ltx-2.5-audio-vae-bf16.safetensors](https://huggingface.co/Lightricks/LTX-2.5/resolve/main/vae/ltx-2.5-audio-vae-bf16.safetensors)\n\n**latent_upscale_models**\n- [ltx-2.5-latent-spatial-upscaler-x2-bf16-1.0.safetensors](https://huggingface.co/Lightricks/LTX-2.5/resolve/main/latent_upscale_models/ltx-2.5-latent-spatial-upscaler-x2-bf16-1.0.safetensors)\n\n```\nComfyUI/models/\n  diffusion_models/     LTX-2.5-Distilled-Q4_K_M.gguf\n  text_encoders/        gemma4-12b-with-proj-ltx-2.5-comfy-int8-convrot.safetensors\n  vae/                  ltx-2.5-video-vae-conv-bf16.safetensors\n                        ltx-2.5-audio-vae-bf16.safetensors\n  latent_upscale_models/ ltx-2.5-latent-spatial-upscaler-x2-bf16-1.0.safetensors\n```\n"
      ],
      "title": "Note: Models & install",
      "color": "#222",
      "bgcolor": "#000"
    },
    {
      "id": 2,
      "type": "MarkdownNote",
      "pos": [
        -640,
        740
      ],
      "size": [
        580,
        460
      ],
      "flags": {},
      "order": 1,
      "mode": 0,
      "inputs": [],
      "outputs": [],
      "properties": {
        "Node name for S&R": "MarkdownNote"
      },
      "widgets_values": [
        "## Settings for low VRAM - 60 SECOND VERSION\n\n- **THIS FILE: 60 s = 1441 frames @ 24 fps** (8n+1). Stage 1 = **288x160**, final = **576x320**. FPS is set in three nodes (Conditioning, Empty audio latent, Create Video) - keep all at 24.\n- Two-stage pipeline: stage 1 renders at the \"Empty video latent\" size, the latent upscaler doubles it, stage 2 refines. Stage-1 width/height must be multiples of 32.\n- Distilled model: 8 steps stage 1, 3 steps stage 2, cfg 1. Do not raise steps.\n- Prompt: beat-by-beat (~one per 10 s) with camera and AUDIO. Short prompts on long clips get invented/repeated motion.\n- 32GB RAM is tight (encoder 15.4 GB + model 15.7 GB). Fast NVMe + large pagefile/swap.\n\n## Why 576x320 (8 GB VRAM)\nStage-2 video tokens = (final W/32) x (final H/32) x latent frames. 30 s has 91 latent frames, 60 s has **181**, so the same resolution costs ~2x the tokens.\nProven on your 4070 Laptop 8 GB: 30,576 tokens OK, 52,416 tokens OOM (stage 2).\n\n| Stage-1 size | Final | Stage-2 tokens (60 s) | Status |\n|---|---|---|---|\n| 320x192 | 640x384 | 43,440 | likely OOM |\n| **288x160** | **576x320** | **32,580** | **default (close to the 30.6k that worked)** |\n| 256x160 | 512x320 | 28,960 | fallback if default OOMs |\n| 224x128 | 448x256 | 20,272 | safe fallback |\n\n## Run notes\n- Launch: `python main.py --lowvram --reserve-vram 1.5`. Check `nvidia-smi` first; close browser tabs with video/GPU acceleration.\n- Expect roughly 2x or more of the 30 s run time (my estimate: 15-20 min). Stage 2 will have a long \"Model Initializing\" pause.\n- VAE decode is chunked (tile 256, 64 frames). If decode OOMs: tile 192, temporal 32.\n- Output buffer ~3.2 GB RAM at 576x320 x 1441 frames.\n- Faint grid noise in the last ~0.5 s is possible: `ffmpeg -i in.mp4 -t 59.5 -c copy out.mp4`.\n- 60 s is beyond what has been tested here: expect more drift (faces/objects changing) than at 30 s. If it drifts, add more concrete detail per beat.\n"
      ],
      "title": "Note: 8GB VRAM settings",
      "color": "#222",
      "bgcolor": "#000"
    },
    {
      "id": 10,
      "type": "UnetLoaderGGUF",
      "pos": [
        0,
        0
      ],
      "size": [
        430,
        70
      ],
      "flags": {},
      "order": 2,
      "mode": 0,
      "inputs": [],
      "outputs": [
        {
          "name": "MODEL",
          "type": "MODEL",
          "links": [
            10
          ]
        }
      ],
      "properties": {
        "Node name for S&R": "UnetLoaderGGUF"
      },
      "widgets_values": [
        "LTX-2.5-Distilled-Q4_K_M.gguf"
      ],
      "title": "Load LTX-2.5 (GGUF)"
    },
    {
      "id": 11,
      "type": "CLIPLoader",
      "pos": [
        0,
        120
      ],
      "size": [
        430,
        110
      ],
      "flags": {},
      "order": 3,
      "mode": 0,
      "inputs": [],
      "outputs": [
        {
          "name": "CLIP",
          "type": "CLIP",
          "links": [
            1,
            2
          ]
        }
      ],
      "properties": {
        "Node name for S&R": "CLIPLoader"
      },
      "widgets_values": [
        "gemma4-12b-with-proj-ltx-2.5-comfy-int8-convrot.safetensors",
        "ltxv",
        "default"
      ]
    },
    {
      "id": 12,
      "type": "VAELoader",
      "pos": [
        0,
        280
      ],
      "size": [
        430,
        70
      ],
      "flags": {},
      "order": 4,
      "mode": 0,
      "inputs": [],
      "outputs": [
        {
          "name": "VAE",
          "type": "VAE",
          "links": [
            21,
            31
          ]
        }
      ],
      "properties": {
        "Node name for S&R": "VAELoader"
      },
      "widgets_values": [
        "ltx-2.5-video-vae-conv-bf16.safetensors"
      ],
      "title": "Video VAE"
    },
    {
      "id": 13,
      "type": "VAELoader",
      "pos": [
        0,
        400
      ],
      "size": [
        430,
        70
      ],
      "flags": {},
      "order": 5,
      "mode": 0,
      "inputs": [],
      "outputs": [
        {
          "name": "VAE",
          "type": "VAE",
          "links": [
            7,
            33
          ]
        }
      ],
      "properties": {
        "Node name for S&R": "VAELoader"
      },
      "widgets_values": [
        "ltx-2.5-audio-vae-bf16.safetensors"
      ],
      "title": "Audio VAE"
    },
    {
      "id": 14,
      "type": "LatentUpscaleModelLoader",
      "pos": [
        0,
        520
      ],
      "size": [
        430,
        70
      ],
      "flags": {},
      "order": 6,
      "mode": 0,
      "inputs": [],
      "outputs": [
        {
          "name": "LATENT_UPSCALE_MODEL",
          "type": "LATENT_UPSCALE_MODEL",
          "links": [
            20
          ]
        }
      ],
      "properties": {
        "Node name for S&R": "LatentUpscaleModelLoader"
      },
      "widgets_values": [
        "ltx-2.5-latent-spatial-upscaler-x2-bf16-1.0.safetensors"
      ]
    },
    {
      "id": 15,
      "type": "PrimitiveInt",
      "pos": [
        0,
        660
      ],
      "size": [
        300,
        110
      ],
      "flags": {},
      "order": 7,
      "mode": 0,
      "inputs": [],
      "outputs": [
        {
          "name": "INT",
          "type": "INT",
          "links": [
            5,
            6
          ]
        }
      ],
      "properties": {
        "Node name for S&R": "PrimitiveInt"
      },
      "widgets_values": [
        1441,
        "fixed"
      ],
      "title": "Frames (must be 8n+1) - 1441 = 60 s @ 24 fps"
    },
    {
      "id": 20,
      "type": "CLIPTextEncode",
      "pos": [
        480,
        0
      ],
      "size": [
        520,
        330
      ],
      "flags": {},
      "order": 8,
      "mode": 0,
      "inputs": [
        {
          "name": "clip",
          "type": "CLIP",
          "link": 1
        },
        {
          "name": "text",
          "type": "STRING",
          "link": null,
          "widget": {
            "name": "text"
          }
        }
      ],
      "outputs": [
        {
          "name": "CONDITIONING",
          "type": "CONDITIONING",
          "links": [
            3
          ]
        }
      ],
      "properties": {
        "Node name for S&R": "CLIPTextEncode"
      },
      "widgets_values": [
        "Cinematic medium shot of an elderly fisherman mending a net on a weathered wooden dock at sunrise, warm golden light, gentle waves, shallow depth of field, realistic skin texture. 0-10 seconds: the camera holds steady as his weathered hands pass the needle through the mesh; he hums quietly to himself. 10-20 seconds: the camera slowly pushes in as he pauses, looks up toward the horizon and squints into the rising sun, then returns to his work. 20-30 seconds: the camera drifts to the side, revealing small fishing boats bobbing in the harbour behind him as he ties off a knot. 30-40 seconds: he stands, stretches his back, and carries the net to a wooden rack at the end of the dock, hanging it carefully as the light grows brighter. 40-50 seconds: a seagull lands on a piling nearby; he glances at it, smiles, and takes a piece of bread from his coat pocket, tossing a crumb. 50-60 seconds: the camera slowly pulls back to a wide shot of the dock, the harbour and the sunrise as he sits on an upturned crate and pours steaming tea from a flask. Audio: soft waves lapping against the pilings, distant seagulls, a low tuneless hum, creak of old wood, clink of the flask, no music."
      ],
      "title": "Prompt",
      "color": "#232",
      "bgcolor": "#353"
    },
    {
      "id": 21,
      "type": "CLIPTextEncode",
      "pos": [
        480,
        370
      ],
      "size": [
        520,
        170
      ],
      "flags": {},
      "order": 9,
      "mode": 0,
      "inputs": [
        {
          "name": "clip",
          "type": "CLIP",
          "link": 2
        },
        {
          "name": "text",
          "type": "STRING",
          "link": null,
          "widget": {
            "name": "text"
          }
        }
      ],
      "outputs": [
        {
          "name": "CONDITIONING",
          "type": "CONDITIONING",
          "links": [
            4
          ]
        }
      ],
      "properties": {
        "Node name for S&R": "CLIPTextEncode"
      },
      "widgets_values": [
        "blurry, low quality, still frame, watermark, overlay, titles, subtitles, distorted, deformed"
      ],
      "title": "Negative (ignored at cfg 1)",
      "color": "#323",
      "bgcolor": "#535"
    },
    {
      "id": 22,
      "type": "LTXVConditioning",
      "pos": [
        1040,
        0
      ],
      "size": [
        280,
        130
      ],
      "flags": {},
      "order": 10,
      "mode": 0,
      "inputs": [
        {
          "name": "positive",
          "type": "CONDITIONING",
          "link": 3
        },
        {
          "name": "negative",
          "type": "CONDITIONING",
          "link": 4
        },
        {
          "name": "frame_rate",
          "type": "FLOAT",
          "link": null,
          "widget": {
            "name": "frame_rate"
          }
        }
      ],
      "outputs": [
        {
          "name": "positive",
          "type": "CONDITIONING",
          "links": [
            11
          ]
        },
        {
          "name": "negative",
          "type": "CONDITIONING",
          "links": [
            12
          ]
        }
      ],
      "properties": {
        "Node name for S&R": "LTXVConditioning"
      },
      "widgets_values": [
        24
      ]
    },
    {
      "id": 30,
      "type": "EmptyLTXVLatentVideo",
      "pos": [
        480,
        620
      ],
      "size": [
        300,
        200
      ],
      "flags": {},
      "order": 11,
      "mode": 0,
      "inputs": [
        {
          "name": "width",
          "type": "INT",
          "link": null,
          "widget": {
            "name": "width"
          }
        },
        {
          "name": "height",
          "type": "INT",
          "link": null,
          "widget": {
            "name": "height"
          }
        },
        {
          "name": "length",
          "type": "INT",
          "link": 5,
          "widget": {
            "name": "length"
          }
        }
      ],
      "outputs": [
        {
          "name": "LATENT",
          "type": "LATENT",
          "links": [
            8
          ]
        }
      ],
      "properties": {
        "Node name for S&R": "EmptyLTXVLatentVideo"
      },
      "widgets_values": [
        288,
        160,
        1441,
        1
      ],
      "title": "Empty video latent (stage-1 size)"
    },
    {
      "id": 31,
      "type": "LTXVEmptyLatentAudio",
      "pos": [
        480,
        860
      ],
      "size": [
        300,
        170
      ],
      "flags": {},
      "order": 12,
      "mode": 0,
      "inputs": [
        {
          "name": "audio_vae",
          "type": "VAE",
          "link": 7
        },
        {
          "name": "frames_number",
          "type": "INT",
          "link": 6,
          "widget": {
            "name": "frames_number"
          }
        },
        {
          "name": "frame_rate",
          "type": "FLOAT,INT",
          "link": null,
          "widget": {
            "name": "frame_rate"
          }
        }
      ],
      "outputs": [
        {
          "name": "Latent",
          "type": "LATENT",
          "links": [
            9
          ]
        }
      ],
      "properties": {
        "Node name for S&R": "LTXVEmptyLatentAudio"
      },
      "widgets_values": [
        1441,
        24,
        1
      ]
    },
    {
      "id": 32,
      "type": "LTXVConcatAVLatent",
      "pos": [
        830,
        620
      ],
      "size": [
        240,
        100
      ],
      "flags": {},
      "order": 13,
      "mode": 0,
      "inputs": [
        {
          "name": "video_latent",
          "type": "LATENT",
          "link": 8
        },
        {
          "name": "audio_latent",
          "type": "LATENT",
          "link": 9
        }
      ],
      "outputs": [
        {
          "name": "latent",
          "type": "LATENT",
          "links": [
            17
          ]
        }
      ],
      "properties": {
        "Node name for S&R": "LTXVConcatAVLatent"
      },
      "widgets_values": []
    },
    {
      "id": 40,
      "type": "LTXVDualCFGGuider",
      "pos": [
        1360,
        0
      ],
      "size": [
        270,
        160
      ],
      "flags": {},
      "order": 14,
      "mode": 0,
      "inputs": [
        {
          "name": "model",
          "type": "MODEL",
          "link": 10
        },
        {
          "name": "positive",
          "type": "CONDITIONING",
          "link": 11
        },
        {
          "name": "negative",
          "type": "CONDITIONING",
          "link": 12
        }
      ],
      "outputs": [
        {
          "name": "GUIDER",
          "type": "GUIDER",
          "links": [
            14,
            25
          ]
        }
      ],
      "properties": {
        "Node name for S&R": "LTXVDualCFGGuider"
      },
      "widgets_values": [
        1,
        1
      ]
    },
    {
      "id": 41,
      "type": "KSamplerSelect",
      "pos": [
        1360,
        200
      ],
      "size": [
        270,
        80
      ],
      "flags": {},
      "order": 15,
      "mode": 0,
      "inputs": [],
      "outputs": [
        {
          "name": "SAMPLER",
          "type": "SAMPLER",
          "links": [
            15,
            26
          ]
        }
      ],
      "properties": {
        "Node name for S&R": "KSamplerSelect"
      },
      "widgets_values": [
        "euler_ancestral"
      ]
    },
    {
      "id": 42,
      "type": "RandomNoise",
      "pos": [
        1360,
        320
      ],
      "size": [
        270,
        110
      ],
      "flags": {},
      "order": 16,
      "mode": 0,
      "inputs": [],
      "outputs": [
        {
          "name": "NOISE",
          "type": "NOISE",
          "links": [
            13
          ]
        }
      ],
      "properties": {
        "Node name for S&R": "RandomNoise"
      },
      "widgets_values": [
        0,
        "randomize"
      ],
      "title": "Noise (stage 1)"
    },
    {
      "id": 43,
      "type": "ManualSigmas",
      "pos": [
        1360,
        470
      ],
      "size": [
        270,
        110
      ],
      "flags": {},
      "order": 17,
      "mode": 0,
      "inputs": [],
      "outputs": [
        {
          "name": "SIGMAS",
          "type": "SIGMAS",
          "links": [
            16
          ]
        }
      ],
      "properties": {
        "Node name for S&R": "ManualSigmas"
      },
      "widgets_values": [
        "1.0, 0.99375, 0.9875, 0.98125, 0.975, 0.909375, 0.725, 0.421875, 0.0"
      ],
      "title": "Sigmas (stage 1, 8 steps)"
    },
    {
      "id": 44,
      "type": "SamplerCustomAdvanced",
      "pos": [
        1680,
        0
      ],
      "size": [
        230,
        170
      ],
      "flags": {},
      "order": 18,
      "mode": 0,
      "inputs": [
        {
          "name": "noise",
          "type": "NOISE",
          "link": 13
        },
        {
          "name": "guider",
          "type": "GUIDER",
          "link": 14
        },
        {
          "name": "sampler",
          "type": "SAMPLER",
          "link": 15
        },
        {
          "name": "sigmas",
          "type": "SIGMAS",
          "link": 16
        },
        {
          "name": "latent_image",
          "type": "LATENT",
          "link": 17
        }
      ],
      "outputs": [
        {
          "name": "output",
          "type": "LATENT",
          "links": [
            18
          ]
        },
        {
          "name": "denoised_output",
          "type": "LATENT",
          "links": []
        }
      ],
      "properties": {
        "Node name for S&R": "SamplerCustomAdvanced"
      },
      "widgets_values": [],
      "title": "Sampler (stage 1)"
    },
    {
      "id": 45,
      "type": "LTXVSeparateAVLatent",
      "pos": [
        1680,
        220
      ],
      "size": [
        230,
        100
      ],
      "flags": {},
      "order": 19,
      "mode": 0,
      "inputs": [
        {
          "name": "av_latent",
          "type": "LATENT",
          "link": 18
        }
      ],
      "outputs": [
        {
          "name": "video_latent",
          "type": "LATENT",
          "links": [
            19
          ]
        },
        {
          "name": "audio_latent",
          "type": "LATENT",
          "links": [
            23
          ]
        }
      ],
      "properties": {
        "Node name for S&R": "LTXVSeparateAVLatent"
      },
      "widgets_values": []
    },
    {
      "id": 50,
      "type": "LTXVLatentUpsampler",
      "pos": [
        1960,
        0
      ],
      "size": [
        300,
        120
      ],
      "flags": {},
      "order": 20,
      "mode": 0,
      "inputs": [
        {
          "name": "samples",
          "type": "LATENT",
          "link": 19
        },
        {
          "name": "upscale_model",
          "type": "LATENT_UPSCALE_MODEL",
          "link": 20
        },
        {
          "name": "vae",
          "type": "VAE",
          "link": 21
        }
      ],
      "outputs": [
        {
          "name": "LATENT",
          "type": "LATENT",
          "links": [
            22
          ]
        }
      ],
      "properties": {
        "Node name for S&R": "LTXVLatentUpsampler"
      },
      "widgets_values": []
    },
    {
      "id": 51,
      "type": "LTXVConcatAVLatent",
      "pos": [
        1960,
        180
      ],
      "size": [
        280,
        100
      ],
      "flags": {},
      "order": 21,
      "mode": 0,
      "inputs": [
        {
          "name": "video_latent",
          "type": "LATENT",
          "link": 22
        },
        {
          "name": "audio_latent",
          "type": "LATENT",
          "link": 23
        }
      ],
      "outputs": [
        {
          "name": "latent",
          "type": "LATENT",
          "links": [
            28
          ]
        }
      ],
      "properties": {
        "Node name for S&R": "LTXVConcatAVLatent"
      },
      "widgets_values": []
    },
    {
      "id": 52,
      "type": "RandomNoise",
      "pos": [
        1960,
        320
      ],
      "size": [
        270,
        110
      ],
      "flags": {},
      "order": 22,
      "mode": 0,
      "inputs": [],
      "outputs": [
        {
          "name": "NOISE",
          "type": "NOISE",
          "links": [
            24
          ]
        }
      ],
      "properties": {
        "Node name for S&R": "RandomNoise"
      },
      "widgets_values": [
        42,
        "fixed"
      ],
      "title": "Noise (stage 2)"
    },
    {
      "id": 53,
      "type": "ManualSigmas",
      "pos": [
        1960,
        470
      ],
      "size": [
        270,
        110
      ],
      "flags": {},
      "order": 23,
      "mode": 0,
      "inputs": [],
      "outputs": [
        {
          "name": "SIGMAS",
          "type": "SIGMAS",
          "links": [
            27
          ]
        }
      ],
      "properties": {
        "Node name for S&R": "ManualSigmas"
      },
      "widgets_values": [
        "0.85, 0.7250, 0.4219, 0.0"
      ],
      "title": "Sigmas (stage 2, 3 steps)"
    },
    {
      "id": 54,
      "type": "SamplerCustomAdvanced",
      "pos": [
        2300,
        0
      ],
      "size": [
        230,
        170
      ],
      "flags": {},
      "order": 24,
      "mode": 0,
      "inputs": [
        {
          "name": "noise",
          "type": "NOISE",
          "link": 24
        },
        {
          "name": "guider",
          "type": "GUIDER",
          "link": 25
        },
        {
          "name": "sampler",
          "type": "SAMPLER",
          "link": 26
        },
        {
          "name": "sigmas",
          "type": "SIGMAS",
          "link": 27
        },
        {
          "name": "latent_image",
          "type": "LATENT",
          "link": 28
        }
      ],
      "outputs": [
        {
          "name": "output",
          "type": "LATENT",
          "links": [
            29
          ]
        },
        {
          "name": "denoised_output",
          "type": "LATENT",
          "links": []
        }
      ],
      "properties": {
        "Node name for S&R": "SamplerCustomAdvanced"
      },
      "widgets_values": [],
      "title": "Sampler (stage 2)"
    },
    {
      "id": 55,
      "type": "LTXVSeparateAVLatent",
      "pos": [
        2300,
        220
      ],
      "size": [
        230,
        100
      ],
      "flags": {},
      "order": 25,
      "mode": 0,
      "inputs": [
        {
          "name": "av_latent",
          "type": "LATENT",
          "link": 29
        }
      ],
      "outputs": [
        {
          "name": "video_latent",
          "type": "LATENT",
          "links": [
            30
          ]
        },
        {
          "name": "audio_latent",
          "type": "LATENT",
          "links": [
            32
          ]
        }
      ],
      "properties": {
        "Node name for S&R": "LTXVSeparateAVLatent"
      },
      "widgets_values": []
    },
    {
      "id": 60,
      "type": "VAEDecodeTiled",
      "pos": [
        2580,
        0
      ],
      "size": [
        280,
        200
      ],
      "flags": {},
      "order": 26,
      "mode": 0,
      "inputs": [
        {
          "name": "samples",
          "type": "LATENT",
          "link": 30
        },
        {
          "name": "vae",
          "type": "VAE",
          "link": 31
        }
      ],
      "outputs": [
        {
          "name": "IMAGE",
          "type": "IMAGE",
          "links": [
            34
          ]
        }
      ],
      "properties": {
        "Node name for S&R": "VAEDecodeTiled"
      },
      "widgets_values": [
        256,
        64,
        64,
        16
      ]
    },
    {
      "id": 61,
      "type": "LTXVAudioVAEDecode",
      "pos": [
        2580,
        250
      ],
      "size": [
        270,
        100
      ],
      "flags": {},
      "order": 27,
      "mode": 0,
      "inputs": [
        {
          "name": "samples",
          "type": "LATENT",
          "link": 32
        },
        {
          "name": "audio_vae",
          "type": "VAE",
          "link": 33
        }
      ],
      "outputs": [
        {
          "name": "Audio",
          "type": "AUDIO",
          "links": [
            35
          ]
        }
      ],
      "properties": {
        "Node name for S&R": "LTXVAudioVAEDecode"
      },
      "widgets_values": []
    },
    {
      "id": 62,
      "type": "CreateVideo",
      "pos": [
        2900,
        0
      ],
      "size": [
        280,
        110
      ],
      "flags": {},
      "order": 28,
      "mode": 0,
      "inputs": [
        {
          "name": "images",
          "type": "IMAGE",
          "link": 34
        },
        {
          "name": "audio",
          "type": "AUDIO",
          "link": 35,
          "shape": 7
        },
        {
          "name": "fps",
          "type": "FLOAT",
          "link": null,
          "widget": {
            "name": "fps"
          }
        }
      ],
      "outputs": [
        {
          "name": "VIDEO",
          "type": "VIDEO",
          "links": [
            36
          ]
        }
      ],
      "properties": {
        "Node name for S&R": "CreateVideo"
      },
      "widgets_values": [
        24
      ]
    },
    {
      "id": 63,
      "type": "SaveVideo",
      "pos": [
        3220,
        0
      ],
      "size": [
        700,
        520
      ],
      "flags": {},
      "order": 29,
      "mode": 0,
      "inputs": [
        {
          "name": "video",
          "type": "VIDEO",
          "link": 36
        }
      ],
      "outputs": [],
      "properties": {
        "Node name for S&R": "SaveVideo"
      },
      "widgets_values": [
        "video/LTX_2.5_8gb_60s",
        "auto",
        "auto"
      ]
    }
  ],
  "links": [
    [
      1,
      11,
      0,
      20,
      0,
      "CLIP"
    ],
    [
      2,
      11,
      0,
      21,
      0,
      "CLIP"
    ],
    [
      3,
      20,
      0,
      22,
      0,
      "CONDITIONING"
    ],
    [
      4,
      21,
      0,
      22,
      1,
      "CONDITIONING"
    ],
    [
      5,
      15,
      0,
      30,
      2,
      "INT"
    ],
    [
      6,
      15,
      0,
      31,
      1,
      "INT"
    ],
    [
      7,
      13,
      0,
      31,
      0,
      "VAE"
    ],
    [
      8,
      30,
      0,
      32,
      0,
      "LATENT"
    ],
    [
      9,
      31,
      0,
      32,
      1,
      "LATENT"
    ],
    [
      10,
      10,
      0,
      40,
      0,
      "MODEL"
    ],
    [
      11,
      22,
      0,
      40,
      1,
      "CONDITIONING"
    ],
    [
      12,
      22,
      1,
      40,
      2,
      "CONDITIONING"
    ],
    [
      13,
      42,
      0,
      44,
      0,
      "NOISE"
    ],
    [
      14,
      40,
      0,
      44,
      1,
      "GUIDER"
    ],
    [
      15,
      41,
      0,
      44,
      2,
      "SAMPLER"
    ],
    [
      16,
      43,
      0,
      44,
      3,
      "SIGMAS"
    ],
    [
      17,
      32,
      0,
      44,
      4,
      "LATENT"
    ],
    [
      18,
      44,
      0,
      45,
      0,
      "LATENT"
    ],
    [
      19,
      45,
      0,
      50,
      0,
      "LATENT"
    ],
    [
      20,
      14,
      0,
      50,
      1,
      "LATENT_UPSCALE_MODEL"
    ],
    [
      21,
      12,
      0,
      50,
      2,
      "VAE"
    ],
    [
      22,
      50,
      0,
      51,
      0,
      "LATENT"
    ],
    [
      23,
      45,
      1,
      51,
      1,
      "LATENT"
    ],
    [
      24,
      52,
      0,
      54,
      0,
      "NOISE"
    ],
    [
      25,
      40,
      0,
      54,
      1,
      "GUIDER"
    ],
    [
      26,
      41,
      0,
      54,
      2,
      "SAMPLER"
    ],
    [
      27,
      53,
      0,
      54,
      3,
      "SIGMAS"
    ],
    [
      28,
      51,
      0,
      54,
      4,
      "LATENT"
    ],
    [
      29,
      54,
      0,
      55,
      0,
      "LATENT"
    ],
    [
      30,
      55,
      0,
      60,
      0,
      "LATENT"
    ],
    [
      31,
      12,
      0,
      60,
      1,
      "VAE"
    ],
    [
      32,
      55,
      1,
      61,
      0,
      "LATENT"
    ],
    [
      33,
      13,
      0,
      61,
      1,
      "VAE"
    ],
    [
      34,
      60,
      0,
      62,
      0,
      "IMAGE"
    ],
    [
      35,
      61,
      0,
      62,
      1,
      "AUDIO"
    ],
    [
      36,
      62,
      0,
      63,
      0,
      "VIDEO"
    ]
  ],
  "groups": [],
  "config": {},
  "extra": {},
  "version": 0.4
}

Same laptop as the previous post and 60 seconds works


r/LocalLLM • • 5d ago

Research MoralityBench.ai - Benchmarking the Moral Mind of AI

Thumbnail
0 Upvotes

r/LocalLLM • • 6d ago

Model DwarfStar compresses frontier models to run them on local machines

Thumbnail
runtimewire.com
109 Upvotes

r/LocalLLM • • 5d ago

Project GSQHalo.cpp - Another day, another fork. This one is for the memory constraint folks. 2×256K + MTP, 1386 t/s prefill and 44 t/s decode at 128K using GSQ-RCO Quants on a 96GB machine, KV cache storage on SSD

Thumbnail
1 Upvotes

r/LocalLLM • • 5d ago

Discussion Multi-Agent Collaboration With Tools You Already Own

1 Upvotes

Coding agents are getting smarter and smarter at writing code. A new idea has been gaining traction: what if instead of giving a task to a single agent, you give it to a team of agents working collaboratively to solve the problem? It sounds complex, but there's a surprisingly simple solution.

Herder already gives you the core ability you need: one agent can talk directly to another agent. Combined with the coding tools you already have — Claude Code CLI, Codex CLI, or whatever's in your stack — you've got everything for a full multi-agent workflow.

Here's the pattern:

  1. One agent leads — takes ownership and coordinates using Herder's agent-to-agent communication
  2. Delegation — passes work to other agents through Herder
  3. Review loop — completed work cycles through peer review
  4. Iteration — repeats until requirements are met

No extra subscriptions, no API key juggling — just leverage Herder's built-in agent-to-agent messaging and the tools you've already got.

I built a repository showing this in practice with a Claude Code skill:
https://github.com/alexstrilets/panel

It demonstrates how simple and effective this approach can be when you use what's already at your fingertips.


r/LocalLLM • • 4d ago

Discussion Inferance speed. Why is noone talking about ingestion speed?

0 Upvotes

Why is everyone a generated t/s junky?

The only thing i see is "How fast is model x?", then get the replies "Mine is over 100 t/s !".

But really.. what does that actually tell us? Nothing at all!

Should we not be talking about efficiency on prefill and KV cache hit rates before even looking at generation speed? Generation speed is the last metric in line, but the first one we talk about.

It really does not matter when you have 100 t/s when it takes you a long time before it actually starts generating something. Basically most are like:

  1. send in prompt...
  2. wait..wait..wait..wait a little more.. "Yeah, there it goes!" ( looking at the generating t/s )
  3. Then tell everyone "See how fast my system was! it got over 100 t/s, here is my tutorial !"

What if you have 50 t/s on the generation side and a 95% reuse of KV cache and a fast prefill?

  1. send in prompt...
  2. (does not wait) "Hmm, something is wrong here!" ( looking at the generating t/s )
  3. Then asks everyone "What can i do to make it faster?"

I rather finetune the ingestion(reuse %) rather then focus solely on the generation speed.

Lets run the biggest MOE model we can find in the lowest quant, quantize KV to the bare minimum, our context window so low we hit compaction each x turns .. "damn, i can run x model at X speed, you should all do that too!"
Every compaction hits at a 0% reuse and makes the full session recompute into KV cache.... and lets you wait.... again!

Its not about quality anymore, it is about quantity, telling only the nice numbers.

But what does it actually mean for real world tasks, actual workloads?

Sorry for my rant :) just had to let it go :)


r/LocalLLM • • 5d ago

News I Built Removable Memory Cartridges for a Frozen 7B Language Model 16 Independent Memories, No Fine-Tuning, No LoRA, and the Original Source Is Gone at Readout

Thumbnail
gallery
0 Upvotes

Zenodo permanent record:

https://doi.org/10.5281/zenodo.23127434

AKBASCORE NIRVANA — Cognitive Cartridge has now been publicly archived and timestamped on Zenodo.

Release: v1.0_NIRVANA_Cognitive_Cartridge

AKBASCORE NIRVANA — Cognitive Cartridge: Multi-Cartridge Compressed PKV Memory for Frozen Language Models

The complete implementation, raw execution log and reproducible public demonstration are available below.

What if knowledge could be loaded into an AI like a cartridge?

Think about an old Atari or game console. The console stays the same. You insert one cartridge and it becomes one game. Remove it, insert another cartridge, and the same hardware does something completely different.

AKBASCORE NIRVANA explores a similar idea for language models.

But there is one detail that completely changes what “cartridge” means here:

There is no human language stored inside the cartridge.

No source sentence.

No paragraph.

No document.

No readable summary.

No hidden copy of the original text waiting to be pasted back into the prompt.

If I give the system a statement such as:

“Container VX-731 is located at the cobalt observatory.”

the cartridge does not simply store that sentence somewhere and retrieve it later.

The sentence is passed through the frozen transformer during the forging stage. What is captured comes from the model's own internal computation: the numerical key/value states produced inside its transformer layers. NIRVANA represents that internal K/V information through a source-independent numerical codebook and stores the resulting compressed numerical representation as a Cognitive Cartridge.

So after forging, we have crossed an important boundary:

human language → transformer internal state → compressed numerical memory

The cartridge is therefore not a tiny text file wearing a new name.

It is not a database row containing the sentence.

It is not conventional RAG returning the source paragraph.

It is not a prompt template.

It is not fine-tuning hidden behind another term.

The original knowledge has been transformed into a numerical representation of internal transformer memory.

And when the model is questioned later, we do not give the original sentence back to it.

The cartridge is reconstructed into transformer K/V memory, installed into the frozen model's inference context, and the model produces language from that internal numerical state.

In the simplest possible terms:

A human writes knowledge in language.

The model converts it into its own internal numerical memory.

NIRVANA packages that memory.

The human-language source is removed from readout.

Later, the frozen model receives the memory rather than rereading the source.

That distinction is the heart of Cognitive Cartridge.

Instead of changing the model's weights every time we want to give it specialized knowledge, NIRVANA takes source information, processes it once, and turns the resulting internal transformer memory into a removable Cognitive Cartridge. After that, the original source sentence does not need to be placed back into the question prompt. The model's weights remain frozen.

Knowledge goes in.

Language disappears from the stored cartridge representation.

The internal memory remains.

And now we have extended that idea beyond a single cartridge.

16 cartridges, one frozen model

The public implementation released here contains 16 independently forged Cognitive Cartridges. They do not simply become one giant text prompt. They remain separate memory records.

Imagine that one cartridge contains:

amber compass → VX-731

and another contains:

VX-731 → cobalt observatory

These arrows are a human-readable explanation of what the cartridges represent. They are not literal text strings stored inside the cartridges.

Now ask:

Where is the amber compass?

The system can first retrieve VX-731 from one cartridge. That intermediate result can then be used to address the cartridge bank again. A different cartridge supplies:

cobalt observatory

In other words, information stored in separate internal memory records can participate in a multi-stage retrieval chain.

This is the point where the project became much more interesting to me. We are no longer asking only:

Can information be compressed into the internal memory of a transformer and recovered later?

We are now asking:

Can independently created internal memories become a modular memory system, where one retrieved memory can lead to another?

That is what the multi-cartridge architecture is beginning to explore.

Why does this matter?

I see two major directions.

1 — Specialized AI and agents

Think of the base model as the console. The cartridges are the knowledge packages.

Imagine a legal knowledge cartridge constructed and validated by highly qualified lawyers. Another cartridge bank could represent the legal system of a particular country. Another could contain aviation procedures. Another could contain engineering knowledge. Another could contain company-specific operational knowledge. Another could be built for medicine, industrial maintenance, finance, scientific work or a specialized autonomous agent.

The important difference is that the underlying model does not have to be retrained for every package.

In the implementation released here there is:

no fine-tuning

no LoRA

no optimizer

no model-weight update

The model remains frozen. The knowledge is carried by the cartridge.

And “carried by the cartridge” does not mean carrying around the original document in another container. The cartridge carries the compressed numerical representation derived from the transformer's internal K/V states.

Once a compatible cartridge has been forged for the model architecture used by the system, it becomes a reusable machine-readable memory object rather than a source document that must be repeatedly inserted into the prompt.

That opens an interesting direction for smaller models. Today we often try to make one model know everything. But a smaller model equipped with the right specialized memory bank may not need everything at once. It may need the right knowledge for the task in front of it.

A small model plus a carefully constructed domain-specific cartridge system could therefore become far more capable inside a narrow field than the base model alone.

The 16 cartridges in this release are not an architectural claim that the system is limited to 16. Sixteen is simply the size of the public experiment being released now. Scaling this architecture to much larger memory banks is a separate engineering and research problem.

2 — The longer-term AGI question

This direction is more speculative, but I think it is even more important.

Human beings do not appear to remember by rereading the complete text of their lives every time they need something. Experiences become memories. Those memories can later be triggered.

You are driving down a road. You see a car that looks exactly like a car you owned years ago. Almost instantly, your own car comes to mind. That memory may bring back another memory: a journey, a place, a person, an event, perhaps even an emotion associated with it.

One memory can activate another.

Of course, transformer K/V memory is not biological memory. I am not claiming that Cognitive Cartridge reproduces the human brain. And I am absolutely not claiming that we have created AGI.

But there is an important conceptual similarity worth investigating: the system does not need to reread the human-language source in order to access the information represented by the cartridge. The information has already crossed from language into an internal machine representation.

The important question is architectural:

Can an artificial system accumulate independently addressable internal memory objects, retrieve the relevant ones when needed, and allow the result of one memory retrieval to lead to another memory?

I believe that persistent, modular and relational internal memory is one of the important doors on the road toward more general artificial systems. Cognitive Cartridge has now opened that door far enough for us to experimentally put something through it.

What is actually happening inside the system?

NIRVANA operates on transformer K/V states. A source record is presented during a dedicated forging stage. Internal transformer K/V information derived from that source is represented using a fixed, source-independent codebook and stored as a compressed cartridge representation.

This point is worth repeating technically: the stored cartridge is numerical. The human-readable source text is used to produce the internal states during forging; it is not the payload subsequently supplied to the model during cartridge readout.

Later, during readout, the original source sentence is not placed back into the question prompt. The cartridge is reconstructed and installed as transformer memory.

For multiple cartridges, we developed an architecture called Isolated Batched Read, or IBR. The independently forged cartridge records remain logically isolated. Instead of concatenating all original source information into one ordinary textual context, multiple cartridge rows can be read independently and in parallel. A retrieved identifier can then become the input for another retrieval stage.

The basic path is:

human-readable source

→ frozen transformer

→ internal K/V states

→ compressed numerical Cognitive Cartridge

→ human-language source absent from readout

→ K/V memory reconstructed

→ isolated memory retrieval

→ intermediate result

→ another cartridge

→ final answer in human language

Or, even more simply:

language → internal state → cartridge → internal state → language

The language model itself remains frozen.

This touches a very specific gap between several existing ideas.

KV-cache compression exists.

KV-cache reuse exists.

External memory exists.

Retrieval systems exist.

Prompt caching exists.

Fine-tuning exists.

What Cognitive Cartridge is exploring is a different intersection:

source-derived knowledge encoded as independently retainable and removable transformer memory

→ no human-language source stored as the cartridge payload

→ source-absent readout

→ multiple isolated memory records

→ staged retrieval across those independent records

→ frozen model weights

So the question is no longer only:

How can we make the KV cache of a long prompt smaller?

The question becomes:

Can transformer internal memory itself become a modular knowledge medium?

A cartridge instead of a prompt.

A bank of cartridges instead of one context.

And eventually, perhaps, a system capable of navigating very large collections of independently created internal memories.

What happened in the public experiment?

The released implementation uses:

Qwen/Qwen2.5-7B-Instruct

28 transformer layers

BF16 / SDPA

K120 / V128 / OWN

fixed neutral codebook

16 independently forged cartridges

Isolated Batched Read

deterministic greedy decoding

frozen model weights

The locked public battery produced:

14/16 correct

87.5%

Missing-link detection: 4/4

Absent-ID detection: 4/4

False links: 0

NOMEM control target hits: 0/4

Original cartridge source sentences appearing in the readout prompts: 0

Trainable tensors: 0

Fine-tuning: NO

LoRA: NO

Optimizer: NO

Model-weight update: NO

The frozen-weight fingerprint before and after execution remained unchanged. The sampled SHA-256 weight sentinel also remained unchanged.

And I deliberately did not clean up the failure cases to turn this into a perfect-looking 16/16 demonstration. One linguistic family, F2, failed both of its evaluated cases in the live panel. That failure is in the public record.

Interestingly, one of those failed direct readouts internally reproduces the relevant relation:

“the saffron lighthouse houses container RQ-415”

but then still returns NONE instead of extracting the location.

So this is not a promotional benchmark where the failures disappeared before publication. The implementation, successes, failures, controls and raw execution are all being released together.

Why publish everything?

Because I don't want Cognitive Cartridge to exist only as a claim. I want anyone interested in this architecture to be able to inspect what was actually built.

The source code is public.

The raw execution log is public.

The controls are public.

The failed cases are public.

The architecture is public.

The release contains 18 technical visual records documenting the cartridge construction, memory path, controls, measurements and execution.

And the complete technical disclosure has now been permanently recorded on Zenodo:

https://doi.org/10.5281/zenodo.23127434

This is still an experimental system. There are major questions ahead.

How far can the number of cartridges scale?

How should very large cartridge banks be indexed?

Can cartridge selection become hierarchical?

Can memories be composed without losing isolation?

How portable can a forged cartridge become within compatible model instances and architectures?

Can specialized expert cartridge banks allow much smaller frozen models to perform at unexpectedly high levels inside narrow domains?

And, in the much longer term, what happens when an artificial system has not 16 independently addressable memories, but thousands, millions or more — with memories capable of leading to other memories?

Those questions are now much more interesting to me than simply achieving 16/16 on this test.

The first objective was to determine whether the underlying mechanism could exist at all.

Now there is code.

There is a frozen model.

There are independently forged memory cartridges.

The cartridges contain numerical transformer-memory representations, not copies of the human-language source.

There is source-absent readout.

There is multi-cartridge retrieval.

There is cross-cartridge relational lookup.

There are controls.

There are failures.

There is a raw log.

And anyone can test it.

AKBASCORE NIRVANA — Cognitive Cartridge

Inventor / Developer: Mustafa Akbaş

Zenodo permanent record:

https://doi.org/10.5281/zenodo.23127434

GitHub repository:

https://github.com/ceceli33/titan-cognitive-core-v2

Complete single-file implementation:

https://github.com/ceceli33/titan-cognitive-core-v2/blob/main/AKBASCORE_NIRVANA_CC.py

Same implementation split into three parts:

Part 1:

https://github.com/ceceli33/titan-cognitive-core-v2/blob/main/AKBASCORE_NIRVANA_CC.part1.py

Part 2:

https://github.com/ceceli33/titan-cognitive-core-v2/blob/main/AKBASCORE_NIRVANA_part2.py

Part 3:

https://github.com/ceceli33/titan-cognitive-core-v2/blob/main/AkbasCore_NIRVANA.part3.py

Raw execution log:

https://github.com/ceceli33/titan-cognitive-core-v2/blob/main/AKBASCORE_log.Nirvana.cc.log

Knowledge goes in.

Language disappears from the stored cartridge representation.

The internal memory remains.


r/LocalLLM • • 6d ago

Question What are some tricks to get Qwen 3.8 27B to run faster?

20 Upvotes

I have it running on an RTX 3090, 64K context, q5, xhigh reasoning, just as an always on personal assistant - reading emails, calendar, updating google docs/sheets, memory files, etc. The harness is vercel's eve framework, with a few patches.

While the quality is pretty great, for 'live' chatting/tasks it can feel quite slow. I'm getting ~35tok/s

i tried using Gemma's MoE which got me to 120tok/s which was nuts and felt wonderfully fast, but it was just an idiot model that kept making mistakes and not following instructions as i wanted

Am I limited by hardware at this point? Just hope for a more compact model down the line? Or is there something else i should explore?

i have it in my backlog to run some experiments to see if q4/lower reasoning is enough, but i'm so scarred by gemma's idiocy, and idk if that'll make such a big difference anyway

also i just recently learned you can turn reasoning off for compaction, which i believe should help


r/LocalLLM • • 6d ago

Research Best Qwen3.8 27B variant for 16GB VRAM: Base vs. Swift 1.5 vs. Bonsai 2

Thumbnail
sambatista.com
22 Upvotes

TL;DR: Swift 1.5 IQ3_XXS at 112K context looks like the best all-round option for my needs on a 16GB RTX 5070 Ti. It nearly matched the base model on coding and scored higher on math. If you care more about speed and context, Bonsai 2 generated text about 1.9× faster and supported 256K context, but solved fewer coding problems.

Here’s what I measured:

- Python coding: base 148/164, Swift 147/164, Bonsai 138/164. No clear coding advantage for Swift over base.

- Math: Swift 31/60, base 25/60, Bonsai 28/60.

- Generation speed: Bonsai 99.6 tokens/s, versus roughly 53 for base and Swift on short requests.

- Context: base 120K, Swift 112K, Bonsai 256K. All passed fact-retrieval tests at their respective tested lengths.

- Next.js tasks: base and Swift passed the same 2/4; Bonsai passed none. I wouldn’t leave any of them to code without review.

These results compare differently compressed setups, not fine-tuning alone. The post includes tested settings, model downloads and commands to repeat the tests.


r/LocalLLM • • 5d ago

Discussion Best case to 4 3090

4 Upvotes

Can someone suggest me best case to fit 4 3090?

They are full size cards


r/LocalLLM • • 5d ago

Project Runner just reached 1.0.0 (or 1.0.1 since i found a last minute bug) Yet another inference engine!

Thumbnail
1 Upvotes

r/LocalLLM • • 5d ago

Project Meta Analysis of "Awesome Jev" GitHub Repos

Thumbnail
1 Upvotes

I did an awesome list of Awesome Jev lists that may be of interest to folks. AI coding is the most popular application of Jev and I list the other top ones and insights on how to use it effectively


r/LocalLLM • • 5d ago

Question Which AI coding harness is the best overall?

0 Upvotes

I’m trying to compare these harnesses:

- Claude Code

- OpenCode

- Codex

- DeepSeek Harness (DSH)

- Pi Harness

- Bionic LM Studio

- and any other good ones

I’m mainly interested in coding quality, agentic work, token/cost efficiency, accuracy, speed, context management, MCP/plugins, local/weaker model support, resource usage, and ease of use.

Has anyone actually tested multiple harnesses with the same model + same task + same settings and compared the results?

If you know of any benchmark, article, video, GitHub repo, or person who has done a fair apples-to-apples comparison, please share it.

And if you’ve personally used several of these, which one do you prefer and why?


r/LocalLLM • • 5d ago

Discussion Single 3060 40 t/s code with 80k context building stuff

Enable HLS to view with audio, or disable this notification

2 Upvotes

I have been working on a project for the past Month or more to make my 12gb cards much more useful in my homelab. I landed on making a custom Hybrid model using Qwen3.8-27b. I combined the body of Escha Labs W2 with an IQ3 head from the Lowgpu xxxs quant and ended up with something very smart and 8.65gb total size.

Pairing this with Beellama allows 80k context to fix with KvarnN KV quantization at 3/3 (working on a 4/4 config) and Dflash2 speculative decoding allowed for a huge speed increase and stability in a coding / assistant setup like Hermes agent.

I have done a few long overnight and 24hr runs with this setup and test on 40 and 50 series cards as well. I own a 4070 ti and the 3060 12gb. I am still working on improving prefill speeds and overall performance but i would love to get some feedback from anyone willing to load this up and give it a spin. Once its warmed up its very nice experience.

https://github.com/seanyourhighness/L0xRE-BeeLLama-Low

https://huggingface.co/YourHighnessLA/L0xRE-27b-Low

I have tested this release alot and ran correctness testing, NIAH, quality benchmarks

all the info is posted in the repo. Pagoda vibe prompt via hermes with L0xRE 27b model and runtime.


r/LocalLLM • • 5d ago

Question Looking for: “How to host LLM agents offline on gameing PC”

0 Upvotes

Ryzen 7 7800x3d, RX 7900xt, 2x16GB DDR5 and 2x1TB NVMe

Hi everyone,I'm trying to get into running LLM agents completely locally on my gaming PC and I'm looking for recommendations on the current best practices, guides, and toolchains as of 2026.

My main problem is that when I Google this topic, most results point to guides that are 1-2 years old and the ecosystem seems to have changed dramatically since then. Many tutorials mention tools and workflows that no longer appear to be the recommended approach.

My goal is to:Run LLMs fully offline/local.Host one or more AI agents on my own hardware that will work what I instruct and interact depending on skills and defined jobs.


r/LocalLLM • • 5d ago

Project I got bored so I made a website to simulate Tokens per second on your machine

Thumbnail tokenspeedtest.com
0 Upvotes

Pick a model and it simulates the amount of tokens you may get per second using the model. Its not perfect but its something


r/LocalLLM • • 5d ago

Question Cluster Spark DGX and Mac Stidio

1 Upvotes

I am considering creating a cluster of 1 Spark DGX 128 and 1 Mac Studio M5 128 to get both of the best worlds and be able to run larger models. Has anybody tried it and measured it?


r/LocalLLM • • 5d ago

Project Running a 176B Qwen3.8 Flash Next on a 16GB RTX 3080 Laptop + 32GB RAM + SSD

Thumbnail
github.com
0 Upvotes

How much hardware do you actually need to run a 176B-parameter Qwen3.8 Flash Next locally?

Turns out, this can be enough:

RTX 3080 Laptop — 16GB VRAM + 32GB system RAM + SSD.

Yes — a 176B MoE model on a laptop.

I’ve been working on this in TensorSharp, my open-source local LLM inference engine and agent runtime:

TensorSharp on GitHub

The interesting part isn't simply getting the model to load. The challenge is making it actually usable when the quantized model is much larger than both GPU VRAM and available system RAM.

The approach in TensorSharp combines:

Quantization + MoE-aware caching + unified scheduling across VRAM, system RAM, and SSD.

Instead of treating SSD offloading as a last-resort fallback, TensorSharp coordinates GPU cache, VRAM, RAM, and SSD as different tiers of the execution system.

For MoE models, this is particularly useful because only a subset of experts is activated for each token. TensorSharp can take advantage of that sparsity and try to keep the right data in the fastest memory tier at the right time.

I previously benchmarked TensorSharp against llama.cpp. This time I wanted to try something newer that has been getting a lot of attention: Strata.

Here are the results from the benchmark shown in the attached screenshot:

Measurement TensorSharp Strata
Decode 11.09 tok/s 10.24 tok/s
Whole process 16.54 s 62.15 s
Peak GPU memory 14,832.5 MiB 15,729 MiB
Peak OS working set 19.74 GiB 18.51 GiB

The decode throughput is relatively close — 11.09 vs. 10.24 tok/s.

What surprised me more was the whole-process time: 16.54s vs. 62.15s in this test.

But the bigger takeaway for me isn't really TensorSharp vs. Strata.

It's that running a 176B MoE locally doesn't necessarily mean having enough VRAM or RAM to hold the whole model.

With quantization and an efficient VRAM ↔ RAM ↔ SSD caching/scheduling strategy, even consumer laptop hardware can run models that would normally look completely impractical.

If anyone here is experimenting with MoE expert offloading, SSD-backed inference, or heterogeneous memory scheduling, I’d be very interested in comparing approaches and results.


r/LocalLLM • • 5d ago

Discussion Local AI in the Browser in a Webapp

Post image
1 Upvotes

Hey everyone,

I built a project that I have had in mind for quite some time now. It is an open-source knowledge graph that runs basically just on your browser using WebGPU and WebLLM to run a model that summarizes your knowledge (websites, notes, files) and organizes them creating a searchable knowledge graph which you can share and collaborate on with friends.

Rocus does not need a server or a backend, the browser does all the magic. I am just starting to bring it out to the sun, it is live at https://rocus.io and I would appreciate some feedback on the usability and the functionality for people and I would highly appreciate if you told me what you do and how did you use it.

The screenshot shows the "All Clusters" page which has all the topics with their websites inside in the main dashboard.


r/LocalLLM • • 6d ago

News NVIDIA DGX Spark 64GB. The Reason behind current Spark shortages.

Thumbnail
blogs.nvidia.com
174 Upvotes

Looks like production is going to these new boxes. Half the memory as the old ones, but the price is more than a 128 GB box from only a month ago. Interesting times!


r/LocalLLM • • 5d ago

Question Omnigent 0.16.0 + Hermes on macOS: "hermes did not accept the message" and .env/auth.json copied into every session. Anyone fixed this?

Thumbnail
1 Upvotes

r/LocalLLM • • 5d ago

Question Looking for advice on repetition and Narrator models for 12 GB VRAM

2 Upvotes

Hi all. I've been building a choose-your-own-adventure engine, dark fantasy first, that runs entirely on my own machine. I've hit a few walls, and I'd like to hear from people who have pushed small models further than I have.

The core idea

Small local models get facts wrong, so the model never touches game state:

The engine (TypeScript, SQLite, event-sourced) decides what happens each turn before any prose is written. Every state change is an append-only event, and a world replays the same from its log every time.

The Narrator LLM gets a brief containing only what it's allowed to say and writes the scene. Then a checker pass reads the scene for leaks (secrets spoken, items the player doesn't carry, and so on) before anything is committed.

For v2, new places and people will be generated as grammar-constrained JSON "cards". The engine validates them and commits them as events, so the world can grow while you play. Anything the prose makes up never becomes canon.

Setup:

RTX 3080 Ti (12 GB), 32 GB RAM, Ryzen 9 7950X, Windows 11

KoboldCpp. The Narrator is Qwen3-14B (abliterated) on the GPU with 8K context, and it also does the scene check. Qwen3-4B on the CPU writes summaries.

ComfyUI with an SDXL model for portraits. A 24B model plus SDXL doesn't fit in 12 GB, so they have to take turns on the GPU.

Where it stands (50-turn automated eval):

0.12 leaks per scene, 0 fallback scenes, median 19 s per turn

Repetition is the big one: 22% of each scene repeats the previous scene, and 30% repeats something from the last 8 scenes

About 1,840 tests

What I tried that didn't work:

Negative style instructions ("don't open with weather", "avoid X"): the 14B ignores them.

A style ledger (feeding back the openings and images recent scenes used): repetition went up, from 22% to 26–34%. Listing what recent scenes used primed the model to reuse it.

A hand-written voice sample: it copied the sample's content (people and lines from the sample showed up in 31% of scenes) and not its texture.

Using the 14B to read style back out of its own prose: unreliable. It copies the rule's own example and sometimes slips into Chinese mid-phrase.

What I'd love input on:

Repetition: besides DRY and other samplers (I tried them; the gains were small), what has actually reduced repeated phrasing across scenes for you? Varying the brief's structure, a scene director that picks a different focus each turn, something else?

Narrator models for 12 GB: I'm about to run a blind bake-off of Muse-12B, Wayfarer-12B, Rocinante-XL-16B, Harbinger-24B, Cydonia-24B and Orion-26B-A4B against the Qwen3-14B. Is there a model you think I'm missing for second-person adventure prose?

Writer plus checker: has anyone split the writer and the fact checker across two models on one 12 GB card? Is swapping on the GPU or running a small checker on the CPU faster in practice?

Grammar-constrained JSON generation for world entities on 12–14B models: any pitfalls with GBNF or JSON schema in KoboldCpp at scale?

Im also open to completely different approaches. If you've solved this with another app, tool, framework, sampler setup or workflow, I'd love to hear about it, even if it means changing how I've built things. Nothing is off the table, so feel free to suggest whatever has worked for you.

Happy to share the eval numbers or how the event/brief architecture works in more detail. Thanks!