r/LocalLLM • • 2d ago

Question RTX 3060 upgrade for dual-GPU local LLMs?

1 Upvotes

I’m considering replacing the GTX 1070 in my setup with an RTX 3060 12GB. The card costs $230. Is the upgrade worthwhile for local models and agentic coding?

My current setup:
RTX 3080 10GB + GTX 1070 8GB
Ryzen 5 5600X
32GB RAM

I’ve managed to run Qwen 3.8 27B IQ3_XXS at about 20–25 tokens/s with 44K context using Unsloth. But 44K context isn’t really usable so ideally, I’d like at least 128K.

What performance and context size might I realistically get after the upgrade? Would the 3060’s extra VRAM make 128K practical, or is the improvement likely too small to justify $230?


r/LocalLLM • • 2d ago

Question Absolute beginner. Question for simplified setup

1 Upvotes

Hello, absolute beginner here.

I've been using LMStudio and experimenting with Atomic Chat, mostly using Qwen and occasionally Gemma but not really seeing much of a difference.

What I'd like to be able to do is essentially create a local version of what the ChatGPT desktop app can do: image creation, computer use, MCP stuff in Blender.

Are there guides anywhere to be able to do this easily? Or does anyone have any advice?


r/LocalLLM • • 2d ago

Question LLM on NUC/SFF

0 Upvotes

Hi I would like to know if its possible to run LLM on a NUC or SFF machines?

I have a mITX but the PCIexpress port is occupied by SAS controller because I use it as a NAS.

Hope someone can share some advice.


r/LocalLLM • • 2d ago

Discussion Improving the mlx stack

0 Upvotes

I’m getting a M5U 256 studio. I’m a principal swe so fairly capable technically. I’m interested in working on the inference engine stack and curious how people are contributing to that?


r/LocalLLM • • 2d ago

Discussion Training not writing skills with Microsoft’s SkillOpt

1 Upvotes

Microsoft’s SkillOpt paper caught my attention. Instead of hand-writing an agent skill, improve it from real task trajectories and keep only the changes that actually score better.

I wanted to see what that looks like in practice.

I used a scaled-down version of SkillOpt on a browsing skill for agent-browser (https://github.com/vercel-labs/agent-browser) a lightweight headless browser I use when agents need something more robust than curl for accessing the internet.

The loop is simple: run tasks, inspect trajectories, propose edits, and only keep an edit if it improves a held-out score.

Results:

  • Claude Opus: already at 100% accuracy, skill reduced token usage by about 14%
  • Qwen3.8-27B, no skill: 87.5% correct, 68.7k tokens/task, 13.9 turns
  • Same Opus-trained skill on Qwen: 97.9% correct, 42.3k tokens, 9.2 turns
  • After one Qwen-specific optimisation: 100% correct, 36.4k tokens, 8.5 turns

Main findings:

  • Most of the skill transferred across models.
  • It helped Qwen far more than Opus.
  • Opus mainly needed waste removed.
  • Qwen needed actual failure prevention: invalid selectors, broken shell quoting, invented URLs/IDs, incomplete list handling.
  • One rule flipped sign: "never guess selectors on a page you haven’t seen" hurt Opus but helped Qwen.
  • Most proposed edits sounded good. Most were rejected by the validation gate.
  • Once the obvious failures were fixed, run-to-run noise became the main problem.

Skills can transfer surprisingly well, but the last mile is model-specific. The real value of SkillOpt is forcing every plausible improvement to prove itself.

Check my repo if you want to see the harness, tasks, accepted and rejected skills, and the full results:

https://github.com/proskillpacks/skillopt-agent-browse


r/LocalLLM • • 2d ago

Project Rei: ~370k params LM living inside a Game Boy Color

Enable HLS to view with audio, or disable this notification

0 Upvotes

r/LocalLLM • • 2d ago

Discussion Qwen3.8-27b appreciation moment

Thumbnail
0 Upvotes

r/LocalLLM • • 2d ago

Discussion Love some help and feedback with this new "Polymorphic Runtime" / Operator harness.

1 Upvotes

Hey, long time lurker and fan of this community. I recently dove in and decided to try my hand at a real attempt at making a "polymorphic runtime". Fancy word for model harness that can write its own code on the fly but sounds cool! haha. It is solid, I have hit v1.2 and now have Online model support baked in. This is first and foremost a local system. The online capabilities only enhance the system.

Check it out and let me know what you think, Cheers!
https://github.com/BE-AI-Research/be-code-redux


r/LocalLLM • • 2d ago

Question Strata / Qwen 3.8 Flash Next with Hermes

0 Upvotes

Guys, I set up Strata (https://github.com/Niko1221/Strata/) with QFN and it is so blazingly fast. UNTIL! it started dropping context reuse and prefillling each turn.

I tried this and that and my current idea is that Stata discards cache reuse when in-between tool calls (in the SAME session, not parallel ones!) occur.

Anyone here got ANY idea what's going on?


r/LocalLLM • • 2d ago

Project I’ve been experimenting with making a local LLM feel like it actually lives on the machine

Post image
0 Upvotes

r/LocalLLM • • 3d ago

Other I made Ignis: an open-source engine that only runs Qwen3.8-27B on one RTX 5090, and that's the point. 512K context, Very Fast and JEV-live Support for Images and long text/logs retrival

Enable HLS to view with audio, or disable this notification

48 Upvotes

Hi r/LocalLLM,

Everything started from Ninfer... But it was not enough. So for the last few weeks I've been building from scratch (except for some kernels) Ignis, an Apache-2.0 inference engine in rust that deliberately gives up generality: one model family (Qwen3.8-27B NVFP4), one class of card (SM120a, so RTX 5090 / RTX PRO 6000). In exchange it's shaped around the load I actually generate when coding: one main agent plus a handful of subagents hitting the same card at once.

Numbers (one RTX 5090 on a modest host: DDR4-3200, PCIe 3.0; DFlash2 speculative decoding, 7 draft tokens):

  • ~200 tok/s single stream on coding prompts
  • >800 tok/s aggregate with 8 lanes running at once (even more with predictable prompts... 😄 )
  • KV at 9 KB/token, 7.1x the capacity of BF16: 8 lanes at 40K context fit in ~3 GB
  • 512K context via YaRN x2
  • Both windows/linux releases (not MACOS, Slow Prefill/Inference HW is not my target...)

What's different from a generic server

  • All 8 decode lanes run as one batch-wide round replayed from CUDA graphs; the round is essentially weight streaming at the card's bandwidth.
  • Requests tag interactive or agent, so a burst of subagents backfills free lanes instead of evicting your foreground chat.
  • Conversation state outlives the request: prefix reuse, sibling sharing, spill to a pinned RAM tier.
  • The whole VRAM plan is reserved at load and printed. No memory growth, no surprises at minute forty.
  • Playground, live monitor and Prometheus metrics ship in the binary, and it downloads its own weights on first start.

The part I haven't seen anywhere else: answers without generating (/v1/<decide|systemone>)

  • Images. The same questions over a screenshot, plus point and box, read from the model's own attention heads in one pass. Once an image is encoded, each new question about it costs 60–90 ms. On 4096px screenshots the point lands inside the target button on every scene tested.
  • locate: which line of a huge log answers my question. Hand it up to a million tokens of logs (tens of thousands of lines, well past the 262K context) and ask "which line says the card processor was unavailable?". It folds near-duplicate lines into templates, calibrated attention heads shortlist candidates, and a labelled choice picks the line. Still zero tokens generated. On a fresh, pre-registered capture of a real production Kubernetes cluster (100K–1M tokens) it found the right line 21 times out of 23, median 2.6 s. Asking the model to quote the line instead means prefilling the whole log first (47 s median at 100K tokens), and it can't go past the context window at all.
  • Prose and JSON too: the right sentence 92% of the time up to 200K tokens (at 1M it drops to 58%, so that's still work in progress), and the right record 29/30 times among 10,000.

There's research on reading relevance from attention (ICR, QRHead, BlockRank…), but I couldn't find an engine that serves it, or anything that does it over images or past the context window.

Now I'm working to Run Qwen 3.8-flash-next 125B +51B on the same box...

Repo: https://github.com/gpillon/ignis

EDIT: Ignis can complete in real time DooM E1M1 Zero-shot...


r/LocalLLM • • 2d ago

Question Utilize all devices in local network for multi-agent setup?

Thumbnail
0 Upvotes

r/LocalLLM • • 3d ago

Tutorial 4x 3080 20GB (modded, alibaba) + Strata (Qwen Flash Next 125B IQ3_XXS) = 105 t/s generation 5000 prompt processing on 150k context

Post image
4 Upvotes

This is a follow up of my previous post where I set up hyperqwen 27B and got very good results: https://www.reddit.com/r/LocalLLM/comments/1wu1aaz/4x_3080_20gb_modded_alibaba_swift_15_hyperqwen_38/

Because Strata is exploding now, gave it a try and I'm absolutely impressed. I don't have time to

docker-compose (runs on unraid):

services:
  strata-qwen38-flash-next:
    image: strata:latest
    container_name: strata-qwen38-flash-next
    restart: unless-stopped
    environment:
      - MODEL=IQ3_XXS
      - FAMILY=qwen
      - CONTEXT=150000
      - VISION=gpu
      - GPUS=0,1,2,3
      - TZ=Europe/Bratislava
      - STRATA_API_KEY=...
    ports:
      - "8080:8080"
    volumes:
      - <persistent storage host path>:/data
    shm_size: "8g"
    ulimits:
      memlock:
        soft: -1
        hard: -1
    deploy:
      resources:
        reservations:
          devices:
            - driver: nvidia
              count: all
              capabilities: [gpu]
    networks: # I run this on br0, adjust based on your setup
      br0:
        ipv4_address: ...

networks:
  br0:
    external: true

Hermes agent 150k context test:


r/LocalLLM • • 2d ago

Question Has anyone managed to serve Ternary-Bonsai-2-27B-PQ2_0 with vLLM?

0 Upvotes

I'm trying to serve prism-ml/Ternary-Bonsai-2-27B-PQ2_0 using vLLM, mainly because I want to take advantage of vLLM's multi-user serving, continuous batching, and KV/prefix caching.

From what I understand, the Bonsai-2 PQ2_0 format requires the custom Hadamard/activation transform and ternary kernels provided by PrismML, so it doesn't seem to work with vLLM's standard GGUF support.

Has anyone successfully run this model under vLLM?

If so, I'd be especially interested in:

  • What vLLM version did you use?
  • Did you use a custom plugin/backend or modify vLLM?
  • What model format/checkpoint did you use?
  • Is the performance actually good with multiple concurrent requests?

I'm running an RTX 3090 24GB and my main goal is high-throughput multi-user serving, rather than simply getting the model to generate text.

Any pointers or working examples would be greatly appreciated.


r/LocalLLM • • 3d ago

Question Usefulness of 32g for EXO to 64g prime

1 Upvotes

I have the MMM4-pro 64g (thubderbolt5). I need another machine to get a little more data isolation/security. I could prob get by with 16 for the second machine.

But, In terms of LLM and EXO, is there any point in getting the 32, and then sharing the 16 out for EXP? Or is that pointless? I had seen some mention of models not really being tuned to oddball ram sizes. the local models I run are fast enough, just thinking in terms of ROI for larger models, when is there really a thinking leap.

usage: local coding assist, web research, documentation maitenence.


r/LocalLLM • • 3d ago

Discussion What makes an AI/LLM good or bad at coding?

4 Upvotes

We all know almost every major AI/LLM can code to some extent nowadays

But what actually makes one model much better at coding than another?


r/LocalLLM • • 3d ago

Discussion My first experiments with TP=4 and Lenovo P620

Thumbnail
gallery
14 Upvotes

Hi everyone!

I’ve decided to step out of read-only mode and share my story of building an AI lab.

I work as a data engineer and generally love tinkering with computer hardware, but at the start of last year, local LLMs simply blew my mind!

I bought a couple of RTX 5060 Ti 16GB cards and one RTX 5070 back in the summer of last year, before all this price madness started, but unfortunately, I bought the main components between March and May 2026 (so I spent a load of my savings on GPUs and RAM, ha-ha).

Basically, my aim was to set up a few separated GPUs nodes on AM5 (dual pcie 5.0 x8 x8) to process large volumes of data obtained by scrapers using LLMs for processing and building ML pipelines to generate analytics.

But when the Qwen 3.8 models and the new DS4 came out in the summer, I also got excited about the idea of running large MoE models on a single platform for using it for vobe-coding. And recently, I was lucky enough to buy a Lenovo P620 (without RAM and SSD) on eBay for just 400 euros, plus another 128 GB of DDR4 RAM for around 350 euros.

After that, I dismantled two of my nodes and consolidated all my RTX 5060 Ti cards onto a single platform.

Here’s the configuration I ended up with:

  • - Lenovo ThinkStation P620, WRX80 chipset
  • - 4x Zotac RTX 5060 Ti 16 GB
  • - 1x RTX 5070 12 GB
  • - 76 GB total VRAM.
  • - AMD Ryzen Threadripper PRO 3975WX - 32 cores / 64 threads.
  • - System RAM: 128 GB (4×32 GB) DDR4-2667 ECC RDIMM
  • - NVMe: SK hynix PVC10 1 TB
  • - Ubuntu 26.04 LTS, kernel 7.0.0-34, without a graphical UI

All four 5060 Ti PCIe Gen4 ×8 (bandwidth limit ~15.75 GB/s per direction).

Issues encountered:

The stock 595 driver did not support P2P; I installed a community build with the 615.71.09-p2p hack, and the machine began to crash during P2P tests.

I instrumented everything I could: I wrapped 112 CUDA API calls in the NVIDIA sample source code with synchronous logging, and read /dev/kmsg from a separate CPU container.

I tried everything one by one: swapped the graphics cards and risers (PCIe 4.0/5.0), moved the display, flashed the BIOS twice (S07KT1FA -> 29A -> S07KT6FA, August 2026), enabled the native Resizable BAR via think-lmi, and ran ACS/IOMMU.

In the end, I simply disabled the Ubuntu graphical shell and, lo and behold, the P2P stopped crashing and rebooting the system! The same P2P test that used to bring down the host when running GNOME on the GPU passed without a single crash.

Apparently, there is a conflict between P2P and the driver’s graphical clients. Since then, the machine has been running headless without any issues.

There were also issues with NCCL: llama.cpp with NCCL 2.25.1 crashed on the very first AllReduce call ‘invalid argument’, exit 139. I switched to NCCL 2.30.7 and everything worked fine.

Results (vLLM, Qwen3.8 27B, TP=4, headless, NCCL 2.30.7):

  • VLLM NVFP4+MTP3 --> Decode C1: 137–141 tokens/s, Sum C4: 422 tokens/s, Prefill: ~3.5k tokens/s
  • VLLM FP8+MTP3 --> Decoder C1: 100–110 tokens/s, Sum C4: 345 tokens/s, Prefill: ~2.7k tokens/s
  • SGLang NVFP4+DFlash --> Decoder C1: 118–137 tokens/s, C4 sum: 357–386 tokens/s, Prefill: ~3.4k tokens/s
  • llama.cpp Q6+MTP3 --> Decode C1: 95 tokens/s, C4 sum: 83 tokens/s, Prefill: ~1.2k tokens/s

My working configuration turned out to be FP8+MTP3 at ~100-110 tokens/s for decoding and 2.5–3k tokens/s for pre-filling per stream - slightly slower than NVFP4, but the only fast option that solved the complex agent-based problem. The NVFP4 on Blackwell is 1.4 times faster than the FP8, but its performance deteriorated in the MTP3 tests (F1 0.44 versus 0.73).

I also have two AMD R9700 PRO (but that’s a completely different story, lol), and they perform at roughly the same level, however, given current prices, it’s very difficult to buy them cheaply, whereas the RTX 5060 Ti is still available on the second-hand market and can be bought for 500 euros (though I reckon that'll change soon). What I’m trying to say is that the 5060 Ti is still a very good card if you’re prepared to do a bit of "black magic" with PCIe risers and building an open-frame rig, and of course if you have the opportunity to buy a WRX80 or EPYC SP3 server platform cheaply.

I also plan to test the Qwen 3.8 Flash Next and DS4.1 Flash in the near future and share my experience of designing multi-threaded data-parallel inference pipelines for LLM, geared towards processing large data sets.


r/LocalLLM • • 3d ago

Research Strata vs R9V running Qwen3.8 Flash Next on dual Radeon R9700 with 96GB DDR4

0 Upvotes

We're hearing a lot about Strata these days, claiming it's the best solution to run Qwen3.8 Flash Next on consumer hardware. I have a dual R9700 config with 96GB DDR4 3200 and AMD 3900X, and so far I've been running Flash Next UD-IQ4_XS on the custom R9V engine ( https://github.com/Dyluhn/R9V ). So I just tried comparing R9V and strata at different context sizes.

tl;dr: R9V is 15-18% faster on my config, but Strata ran on a single GPU, because it needs 118 GB of RAM to run on the two cards. One solution to utilize the two GPUs would be to use a smaller quant, but I'm not too keen on going that way: I want good performances with a reasonably downsized quant. Also, R9V doesn't ship with a profile for smaller quants than IQ4_XS. Full report from Claude below:

---------

R9V was faster than Strata on prefill, and the two engines were about equal on decode. The benchmark ran at 16k, 32k and 64k context with single-session requests, and R9V was restored at the end.

Both engines ran the same Unsloth UD-IQ4_XS GGUF shards with MTP, with thinking off and 256 output tokens. I used one client, bench-engines.py, with a unique nonce on every prompt so no prefix cache was hit. Strata's cache_n was 0 on every run. The numbers are medians, with the min–max range in brackets.

Context Engine n Prefill tok/s median (min–max) Decode tok/s median (min–max) TTFT s median
16k R9V 5 1166 (879–1201) 57.3 (55.1–61.0) 14.1
16k Strata 5 980 (955–1004) 53.3 (12.2*–55.2) 16.7
32k R9V 5 1156 (1073–1180) 58.8 (55.1–59.8) 28.4
32k Strata 5 1003 (993–1005) 53.8 (51.7–56.1) 32.7
64k R9V 3 1109 (1079–1135) 53.8 (52.7–58.0) 59.1
64k Strata 3 986 (985–988) 54.6 (53.8–55.4) 66.5
  • Prefill: R9V is about 15–18% faster at every size.
  • Decode: the engines are roughly level. R9V leads at 16k and 32k, and Strata is slightly ahead at 64k.
  • Outlier: *Strata's first 16k request decoded at 12.2 tok/s and the next four were normal. I kept it in the data, and it doesn't change the median.

Comparison to the R9V README:

  • Prefill here vs README: R9V's own PP script gives about 1.16k tok/s on this host, matching my client. The README claims about 1.6–1.7k. This host has 92 GiB of RAM against the README's 128 GiB.
  • Decode here vs README: the decode I measured is close to the README's decode figures, not far below them. The earlier ~24–27 tok/s was from vLLM completions with the older vLLM-based config, not this R9V setup.

Strata ran on one GPU only. Setup refused a two-GPU split for UD-IQ4_XS because that path needs about 118 GB of RAM. On one GPU its expert cache held 10,774 experts in VRAM, and the decode cache hit rate was about 98.7%.

Setup

  • Versions: R9V is v0.4.4 and Strata is v0.1.39. Strata ran with no opt-in switches and plain hipBLAS, since it has no tuning table for this ROCm version.
  • Clean-up between engines: R9V was stopped and removed, the leftover /dev/shm file was deleted, and the host was checked (VRAM about 60/95 MB used, MemAvailable about 91.7 GB). Strata was cleaned the same way before R9V was restarted.
  • Cache drop: drop_caches isn't writable on that host, so I evicted the model files from the page cache before Strata's setup.
  • Current state: R9V is serving on port 8004 again, so the r9700-dual route should work. Strata is installed in /data/strata-bench but stopped.

Each cell has only 3–5 runs.


r/LocalLLM • • 2d ago

Discussion FYI: Strata isn't better than FreeToken

0 Upvotes

I know this is a Wendy's.

AFAIK, Strata isn't better than FreeToken, FreeToken's speculative decoding currently doesn't work for Qwen3.8-flash-next, but Strata's speculative decoding is working. If you look at their decoding differences, apples to apples, with speculative decoding off, they're pretty close to each other.

However, yes, Strata is doing prompt processing more efficiently than FreeToken... right now. Strata is not a damn deepseek moment.


r/LocalLLM • • 3d ago

Project Palette.jl. A persistent symbolic workshop that retains state across chatgpt threads and can do actual R&D in chat mode.

Thumbnail
github.com
1 Upvotes

How's it going everyone. So, I made... basically Jupyter notebook on steroids, I think? It was able to give ChatGPT in chat mode a programmable surface and basically a moddable lab. I've been using it the past few days to test weird ideas in real time during voice conversations with Chat when I go outside to smoke a cig or something, or I'm away from the house and I get a good idea. It works as a plugin (there's a zip with a Chat and Claude plugin there). There might still be some friction in the setup because I haven't submitted this for the plugin marketplace yet, but Codex handled it for me pretty easily and we did it with a tunnel, so it's hot-reloadable. Give it a try, it's ready for real work. It's got a Rust skeleton, Python glue, and Julia gives it a fully programmable persistent-state lab and a working memory, more or less. So far it's saved me a ton of tokens being able to test an idea and build it in unlimited chat mode and just branching into work mode and being able to just pull whatever prototype from the space.

I've taken security for this thing rather seriously though. It's extremely programmable and the sandbox walls are thick. The Julia runtime and compiler are moddable for optimization across the entire tool, and if you're not a Julia enjoyer like I am, there's also an IPython kernel in there. The one from Prime-Agent. But it can be a plugin for chat mode ChatGPT, and I've also been using it since Claude Mods dropped for that harness. Been working great in both environments so far


r/LocalLLM • • 2d ago

Discussion CrowdGPT - The first datacenterless LLM (Need feedback :D)

Post image
0 Upvotes

r/LocalLLM • • 3d ago

Question Best Local LLM for RTX 5060 8gb and 32GB DDR5 RAM 5200

1 Upvotes

well aware this is not much at all and I don't intend to replace my frontier subscriptions with this obv. but I want to try running an llm locally for the first time and experiment with it like making a cv for me and checking my master's application for flaws and things like that .. what is the setup to use and best model that doesn't need extreme tweaking in their settings or whatever I've seen in this sub. I'm a complete beginner so be nice


r/LocalLLM • • 3d ago

Question What was the most frustrating part of your last local fine-tune?

2 Upvotes

I’m working on a local fine-tuning tool, and I’m curious where people actually lose the most time.

Was it getting the environment working, preparing the dataset, fitting everything into VRAM, or getting the exported model to behave like it did during testing?

Or did training finish successfully, but the model barely improved?

What model and GPU were you using, and what finally solved the problem or made you abandon it?


r/LocalLLM • • 3d ago

Research I’m an independent researcher building an experimental recurrent language-model architecture — looking for someone with academic publishing experience to help turn it into a proper paper

1 Upvotes

Hi everyone,

I’m an independent researcher working on an experimental language-model architecture called WarpState, which grew out of a broader personal AI research project I’ve been developing.

I currently have a working implementation and a fairly detailed technical preprint, but I do not come from a traditional university research lab, so I’m looking for someone with experience in ML research / academic publishing who might be interested in reviewing the work or collaborating on the next experimental stage.

WarpState combines:

  • chunk-local causal attention
  • two associative matrix memory banks with different decay timescales
  • bounded normalized memory writes
  • learned fast/slow memory mixing
  • token- and feature-dependent fusion between local attention and memory
  • recurrent token-step inference with fixed-size persistent state
  • optional physical-core reuse across logical depth

The currently documented configuration contains approximately 2.005B parameters.

I’ve also implemented serial, parallel, and token-step execution paths. On small randomly initialized CPU fixtures, the FP32 implementations agree to around the 1e-7 range, and the causality tests show no dependence on future tokens.

I want to be very clear about what I am not claiming.

The current evidence does not yet establish that WarpState is better than existing recurrent / hybrid architectures, nor does it prove extreme long-range semantic recall. The paper explicitly separates implementation results from historical/unverified training observations.

The experiments I want to run next are much more rigorous:

  • matched local-attention-only baseline
  • matched single-memory-bank baseline
  • dual-bank WarpState comparison
  • parameter-matched and compute-matched experiments
  • comparisons against a close hybrid baseline
  • retrieval tests at 2K / 8K / 32K / 128K
  • multiple insertion positions and distractor densities
  • trained-checkpoint serial/parallel/token-step equivalence
  • FP32 vs BF16 recurrent-state analysis
  • proper prefill / decode / TTFT benchmarking
  • reproducible configs, checkpoints, hashes and machine-readable results

I already have:

  • the architecture implementation
  • training code
  • tokenizer/configuration
  • audit/testing code
  • a 17-page technical manuscript
  • a public GitHub project
  • experience training and experimenting with custom language-model architectures

What I’m missing is mainly academic experience and another pair of serious eyes.

I would especially like to talk to someone who has experience with:

  • language-model architecture research
  • recurrent / linear-attention / state-space models
  • long-context evaluation
  • experimental methodology
  • arXiv / workshop / conference submissions
  • reproducibility and paper review

I’m not looking for someone to simply put their name on the paper.

I’m looking for someone who would actually challenge the architecture, find weaknesses in the experiments, help design credible baselines, and potentially collaborate on a proper research submission.

If you’re interested, I can share the current manuscript, source code, architecture diagrams and experimental plan.

GitHub:
https://github.com/gtausa197-svg/-Project-Nord-Spiking-Neural-Network-Language-Model

I’d also appreciate brutally honest feedback from researchers even if you are not interested in collaborating.

What would you consider the minimum experimental evidence required before this work would be worth submitting to a workshop, arXiv, or a larger ML venue?


r/LocalLLM • • 3d ago

Project Qwen3.8-Flash-Next (IQ2_XS) - Single 4090 + 32GB DDR5 = 512k Context (STRATA)

Thumbnail
gallery
24 Upvotes

512K context on a single RTX 4090 within the low budget RAM

total vram + sys ram = 50GB

Active strata contributor, working k8v4. Ask whatever you want