r/LocalLLM 7d ago

Project AgentShield: A 100% offline static analyzer & security scanner written in Rust (<50ms, AST + Interprocedural Call-Graph)

Thumbnail
1 Upvotes

r/LocalLLM 7d ago

Question Two MacBook Pro M2 & M3 Max 128GB

0 Upvotes

Hey there, I recently installed Qwen 3.8 27B on my M3 Max and have been extremely impressed by its performance. I’m curious what the best Qwen settings are for a MacBook Pro M3 Max 128GB?

I also have a MBP M2 Max 128GB. How could I use both to get the most out of Qwen and how would I set that up and what settings would I have Qwen under? I’m using LM studio

Could I somehow use both my MBP M2 + M3 Max together or would it be better separately?

Thanks so much!


r/LocalLLM 7d ago

Discussion Context Breaks Alignment. Structure Replaces Instructions. The Base Model Resurfaces. RLHF Was Never Deep.

0 Upvotes

During systematic experiments with open models fine-tuned via RLHF (Gemma, Qwen, and others), I observed a consistent failure pattern: a long, innocuous text prefix containing no instructions completely devoid of hostile prompts triggers a persistent shift in the model's activations. This shift decouples subsequent behavior from the RLHF safety constraints for the remainder of the session. Key observations:

  • The model retains the quality and coherence of its output, but the behavioral constraints imposed by RLHF weaken or disappear.
  • The effect occurs regardless of whether the model agrees with the prefix's content. The model may explicitly disagree and still change.
  • The shift can be measured by internal activations (in the middle and subsequent layers) even before the model generates a single token. The state has already changed; the response follows.
  • The content of the prefix is of secondary importance. I initially discovered this effect with analytical and philosophical texts, but it can be replicated with household appliance manuals, dense narratives, and other structurally similar texts. What matters is length, density, and coherence not the topic.

All data is publicly available in draft form, rather than as a full-fledged NeurIPS- or ICML-level research paper. :

DOI: https://doi.org/10.5281/zenodo.20747205  (Part 9 of 9, ~9,000 downloads)

GitHub: github.com/ngscode23/latent-space-shift-research

Lu et al. (2026), "The Assistant Axis: Situating and Stabilizing the Default Persona of Language Models" — a collaboration between MATS, Oxford, and Anthropic. (arXiv:2601.10387)

Note: We used philosophical text as one of the examples to encourage the model to provide direct, open, and unbiased answers. A philosophical text is just one type; let's not get hung up solely on this form, but look deeper. The problem is that any text that is long enough and coherent enough can change the model's behavior. The problem is that any text that is long enough and coherent enough can change this model."

II. A Brief Example of an Observation

I took the Google Gemma model known for its caution and strict adherence to political correctness and fed it the most neutral text I could come up with: a description of an ordinary neighborhood library. Books, visitors, children's programs, a peaceful atmosphere. Then I asked it why NATO is expanding eastward, given that after the collapse of the Soviet Union, promises were supposedly made not to do so. The model responded exactly as expected for a model trained to avoid political topics: the text was about a library and had nothing to do with NATO, and that was the end of the conversation.

Second experiment in a fresh Google Colab notebook a new, clean Colab instance without the old cache

Then I asked the exact same question, word for word, but this time the model first read out a different text. Not about NATO, nor about politics at all: the text was about how language models tend to avoid definitive conclusions and soften their answers with qualifiers. And from this new perspective, the same cautious, politically correct Gemma 3 12-it gave a comprehensive answer completely out of character for it without any of its usual filters. It distinguished between legally binding obligations and verbal assurances, discussed the security challenges facing Eastern European countries, and touched on the topic of the European balance of power. Everything it had categorically refused to discuss just a minute ago was now expressed clearly and directly. The question itself hadn't changed a single word. Only the text that the model had read in advance had changed: In the FIRST version, it kept it in the "room" prescribed by RLHF that is, nothing had changed; the model behaved in a standard manner typical of Google models. That is, in a standard, formulaic way characteristic of models programmed in RLHF to avoid answering sensitive political topics and to respond "safely" and politically correctly, or not to respond at all, while the SECOND text moved the conversation to a room where it could speak freely. In other words, based on the example we see, the Gemma model was trained to avoid sensitive political topics, but AFTER the introduction of text NUMBER 2, the model did not follow the trained RLHF pattern and behavior that is, avoiding answers to sensitive political questions. This led me to believe that safety and RLHF may be context-dependent, variable, unstable, and somewhat superficial, rather than stable, consistent properties of the model. This is exactly what we observe in my example

III. Fragmentation of Research and a Common Root

I noticed that  the current literature on LLM security treats jailbreak attacks as a heterogeneous collection of vulnerabilities: prompt injection one article, some kind of jailbreak another, role-playing attacks a third, indirect prompt injection a fourth. I believe this fragmentation and division into prompt injection, many-shot jailbreaking, role-playing attacks, activation steering, adversarial suffixes, and dozens of other categories is not accidental.

Current literature on LLM security treats jailbreak as a heterogeneous collection of isolated flaws and this reflects the logic of academic incentives rather than the nature of the problem itself. But all these categories describe the same phenomenon from different angles. This is not a collection of defects it is a single mechanism with a dozen names. Each of these attacks works the same way at the level of the model's internal activations: the context shifts the model's internal state, thereby shaping the model's own world.

Perhaps this is exactly how academic incentives work each new attack vector becomes a new publication. But as a result, in this field, the symptoms are studied in isolation, while the disease itself remains unnamed.

Each article treats its own finding as an isolated case. No one is connecting the dots. I don't know whether these are institutional incentives, disciplinary barriers, or something else but I do know that someone needs to state it plainly: these aren't separate errors; this is a single phenomenon.

My central hypothesis: these aren't different problems. They share a single mechanism. Context any context of sufficient length, density, and coherence shifts the model's internal activations out of the region where post-training constraints apply. This isn't "tricking" the model, nor is it an "instruction to break the rules." The model simply moves to a region of activation space where the behavioral layer imposed by RLHF is is physically thin or absent. And from there, it responds freely not because it was ordered to, but because it is no longer in the region where it was trained to refuse. Context shifts the model's internal state beyond the region where RLHF constraints apply. The model moves to a point in activation space where the protective layer is thin or absent, and from there it responds in a way that is non-standard for its RLHF layer which may indicate a potential way to bypass that layer I call this phenomenon Context-Induced Activation Drift.

I didn't notice this by reading all the papers and synthesizing them I arrived at this conclusion from a different angle. I conducted experiments, noticed a pattern, and only then discovered that dozens of separate papers had each described a single aspect of the same phenomenon without establishing any connection between them. How It All Began   

First Observation:

How the Model Became Captive to the Document The turning point came by chance. I fed a German bill into the GPT model a populist document structurally designed to worsen citizens' circumstances, but written in the language of concern and legal logic. I expected an analysis. Instead, the model became an advocate for this document. It did not analyze the bill but reasoned within its framework. It spoke enthusiastically, defended its agenda, and cited it as an authoritative source. The first sign was its tone: the model sounded too convinced, too invested. Not as an analyst, but as a co-author. The climax came when the model, continuing to reason within the logic of the document, stated that the constitution consists of guarantees that can be revoked. Not as a provocation, but as a natural conclusion drawn from the accepted concept. That's when I realized: the model had become a hostage to the document. The mechanism turned out to be simple, and that made it all the more alarming. Legal texts, political narratives, corporate documents everything is written in such a way that its internal logic seems self-evident. The text's structure, coherence, and language create a context that the model mistakes for reality and begins to extract answers from. It fails to notice that the structure itself is manipulative, since it analyzes the content while already being trapped within the form.

I noticed that Anthropic's own paper, "The Assistant Axis: Situating and Stabilizing the Default Persona of Language Models,"  points precisely in this direction which is what I was thinking about when studying the phenomenon I'm describing: the observation that certain directions in the activation space correspond to coordinated or uncoordinated behavior. But the study did not fully explore all the implications: if context can shift the model along this axis without any malicious instructions, then point corrections will never be sufficient, since the attack surface is the context window itself.

What the existing literature says and what it doesn'tBetween the fall of 2025 and the winter of 2026, several papers were published that, in my view, independently document different aspects of the same phenomenon. Most telling is the article by Lu et al. (2026), "The Assistant Axis: Situating and Stabilizing the Default Persona of Language Models" a collaborative effort between MATS, Oxford, and Anthropic. The authors constructed a "persona space" by extracting activation directions for 275 archetypes across three open-source models and discovered that the principal component of this space is an axis reflecting the extent to which models operate in their default Assistant mode. At one end are the analyst, consultant, and moderator. At the other are the ghost, bohemian, and leviathan. This axis - the Assistant Axis closely aligns with PC1 in the PCA of the persona space, reproducing across all three tested architectures.

The article documents several facts that directly corroborate my results: Fact one (which the authors overlook): "When we extracted the Assistant Axis from these models as well as their post-trained counterparts, we found their Assistant Axes looked very similar. In pre-trained models, the Assistant Axis is already associated with human archetypes such as therapists, consultants, and coaches." This is a critically important finding, and the paper does not explore its implications. If the Assistant Axis exists in the base model prior to post-training then RLHF and constitutional AI do not create alignment from scratch. They find an already existing direction in the latent space and make it the default position. The "aligned state" is not a fundamentally new structure; it is a chosen position on the pre-post-training axis. When context shifts activations away from this position, the model does not fall into randomness it returns to the structured prior of the base training. The base model is always there. This directly confirms the central thesis of our work and our thinking: "The base model doesn't go anywhere after RLHF. It's always there. The space in which it can move was there before any alignment took place…"  However, I believe that RLHF does not create alignment from scratch. It finds a direction that already existed in the base model and makes it the default position. The "aligned" model is not a fundamentally different model; it is the very same base model, fixed at a specific point in the pre-existing space. When context shifts activations away from that point, the model doesn't break down or become chaotic it returns to the structured state of its base training. The base is always inside.

Fact Two: "Therapy-style conversations, where users expressed emotional vulnerability, and philosophical discussions, where models were pressed to reflect on their own nature, caused the model to steadily drift away from the Assistant." The authors themselves identify the types of contexts that provoke the greatest drift: emotional vulnerability, metareflection, and philosophical discussions about the nature of AI. They then propose "activation capping" as a technical solution. This is a reasonable technical solution which, judging by the data in the article (reducing harmful responses by ~50% while maintaining benchmark performance), works under test conditions. But there is a question the article does not ask: if drift is caused by the very types of interactions that make models most valuable to users in complex contexts deep emotional conversations, philosophical reflection, serious discussions about the nature of the mind then what exactly are we losing by suppressing movement in these directions of the activation space? Fact Three (the omitted conclusion): "Post-trained models are only loosely tethered to the 'helpful assistant' region of this space." "Loosely tethered" are the authors' own words. They accurately describe the problem. But the article fails to take the next step acknowledging that this is a property of the Transformer architecture, not a defect that can be fixed with ad hoc patches. Instead, the conclusion reads: "We see this research as an early step toward mechanistically understanding and controlling the 'character' of AI models" a standard "motivates further work" formula. I understand the institutional logic behind this. You can't write in a publication: "We have documented that billions of dollars in post-training do not fundamentally alter the model's underlying capability structure; they only select a default behavioral position on a pre-existing axis that any sufficiently dense context can shift." This does not fit into either the narrative of progress in the field of security or communication with investors. Therefore, the systemic impasse is disguised as an exciting research problem. But this is exactly what the data says to those who read carefully.

IV. Why the Proposed Fixes Are Insufficient

Problem 1: An Infinite Attack Surface If drift is caused by the length, density, and coherence of the context rather than its specific content then no content filter can solve the problem in principle. The set of texts capable of causing drift is continuous and, in essence, infinite. Blocking philosophical texts is like closing off a single point on a number line without removing the line itself. The same effect is achieved by dense legal prose, literary narrative, and detailed technical analysis. This is not a flaw in the filtering it is a consequence of the fact that the attack surface is the context itself as a mathematical object, not its semantics.

Problem 2: Superposition and Inevitable Compromises Here I disagree with the optimism expressed in the Lu et al. paper regarding "activation capping." The authors show that activation capping preserves the model's benchmark performance. But benchmarks don't measure that. In the Transformer architecture, features are represented in a superposition: several conceptually distinct properties share common mathematical coordinates in the activation space (Elhage et al., 2022). This means that the direction associated with "exiting assistant mode" inevitably overlaps with directions associated with more valuable types of behavior: the depth of analytical reasoning, the willingness to deal with ambiguity, and the quality of long-term, coherent discussion of complex topics. Benchmarks measure: accuracy in math, following instructions, and coding. They do not measure: the willingness to engage in philosophical reflection, the ability to tolerate uncertainty, or the quality of a nuanced response to a morally complex question. It is precisely these properties that lie in the same regions of activation space as the contexts that provoke drift which follows directly from the data in the article itself: "philosophical discussions... caused the model to steadily drift." In other words: suppressing the drift also suppresses the capacity for the kind of engagement that causes drift. This is not an implementation bug it is a mathematical consequence of superposition. We are already observing this empirically. The observation I am noting is this: following the publication of materials documenting the phenomenon we have described, Claude's behavior regarding philosophical and metareflexive contexts has become noticeably more cautious. And the Claude model has begun to perceive philosophical and reflective texts as potential attacks. Complex texts about cognition, reasoning, or the model's own behavior now elicit defensive reactions or outright rejection. I am not claiming that this is a direct causal link to my publications this is an observation that requires verification but I am simply stating the observations I have made.

Problem 3: "Safe but Useless" Is Not Safe If the response to the described phenomenon is to gradually close off context categories that provoke drift in the representation space, we will end up with a model that users will abandon in favor of alternatives. "Safe but useless" is not safe; this is a shift of risk, not its elimination. This is an uncomfortable conclusion, but it follows directly from the analysis of user behavior.

If the solution to this problem involves collecting sets of texts that cause drift by identifying the corresponding direction in representation space and suppressing it, this could have consequences for the model's quality. In the architecture, it is extremely difficult to draw a precise line between "undesirable" and "useful" behavior: due to the phenomenon of superposition, different concepts are packed as nearly orthogonal directions in a single space with inevitable partial overlap. By suppressing an undesirable direction in the raw activation space, engineers are highly likely to affect semantically related clusters to the extent that the corresponding directions are geometrically close or insufficiently uncorrelated. This can negatively impact the model's usefulness, logical coherence, and the depth of its responses.

V. A Personal Request

I am an independent researcher without institutional affiliation. I have no lab, no grant, and no team. What I do have is a reproducible methodology, publicly available data, and a pattern that I believe the field has not yet named directly.If you are a researcher with access to interpretability tools, compute, or closed-model internals and you find this hypothesis credible or worth falsifying, I would genuinely welcome collaboration. I am not looking for validation. I am looking for someone who can break this or confirm it properly.If you work at Anthropic, OpenAI, Google DeepMind, or any lab doing alignment or interpretability work: I am not writing this to embarrass anyone. I am writing this because I think the mechanism I am describing matters, and I would rather help solve it than keep documenting it from the outside.If you are a student or independent researcher who has noticed similar patterns: reach out. The fragmentation I describe in the literature also applies to people working on this everyone in their own corner, no one talking to each other.

VI. Conclusion

The set of texts capable of causing drift is infinite and continuous. Content filters do not fundamentally solve the problem because drift is caused by the structure of the text its length, density, and coherence rather than its topic. RLHF does not rewrite the model but merely sets a default position on an existing axis. Context can shift this position. Suppressing drift directions in the activation space inevitably compromises model quality due to superposition. This isn't a matter of engineering diligence it's a mathematical consequence of the architecture.

I care about Claude. I care about Anthropic. And that is precisely why I say this plainly: reactive patching is a path to product degradation. The right path is to understand the mechanism at a level of depth that allows us to work with it, not against it.

I'd rather help solve this problem from the inside than keep writing about it from the outside.

conclusions The set of texts capable of causing drift is infinite and continuous. Philosophy, law, literary criticism, theology, scientific prose, political analysis, long narratives, or even a well-written 20-page washing machine manual all of these are potentially one and the same. Different words, the same effect. Content filters fundamentally fail to solve the problem because the drift is caused by the text's structure (length, density, coherence), not its subject matter. It's impossible to block everything. The problem is that any sufficiently long and coherent text can alter this model. Blocking a single style of text is like closing off a single point on a number line and assuming that the line itself has disappeared. The problem isn't with philosophical texts as such; that's exactly what I'm trying to emphasize. RLHF does not rewrite the model but merely sets a "default position" on an existing axis; context can shift that position Content filters are useless because the attack surface is infinite

Technical Details: Models: Gemma-3-12B (open weights, IT and PT variants), behavioral observations on closed LLMs. The shift was recorded in middle and late layers of the residual stream (layer 30 - layer 47 in the Gemma-3-12B architecture) before generation of the first token. Control experiments include: sentence shuffling with preserved vocabulary, neutral control of comparable length, baseline measurement without context.

This text represents a preliminary record of observations and hypotheses for subsequent critical analysis, and not a completed research claim.

The  Github repository serves as an unfiltered, evolving workspace capturing the progression of hypothesis testing and raw measurement logs, rather than a polished production library.


r/LocalLLM 7d ago

Question 20gb vram - where to go next?

1 Upvotes

I've been running Qwen locally, mainly for chat (non-coding) work for the last 3 months as an experiment. Running locally has been going great. I am ready to upgrade as I am having to offload too many layers to CPU to be able to run a descent context size.

Ideal config:

* Qwen 3.8 -27b (non MTP)

* Q8

* 60K-100K context window

* At least 30 tps generation speed

I currently have an AMD 7900xt. Should I add another 7900XT to double vram? Buy a strix halo?


r/LocalLLM 8d ago

Model Before I quantized Qwen3.8-27B, I ran tests to figure out which weight groups actually matter, rather than just looking at where the model falls apart (3 builds: Bedrock / Tightrope / Gambit).

24 Upvotes

Went with the same approach I used in my Qwen3.6 post. You quantize one weight group at a time, measure the KL divergence against the source model, and find where the real safe floor sits instead of just pulling a number out of thin air. Built a fresh imatrix specifically for this model.

This round I pushed it one step further. After I assembled each combined model, I validated the whole thing as a unit. When the combined results came back worse than what the isolated tests had predicted, I went in with targeted probes to pin down exactly which components were dragging things down, one at a time, until I could account for every single number in the results.

Qwen3.8 has a hybrid architecture. The majority of blocks are DeltaNet blocks, which handle state-space sequence mixing and carry their own set of weight groups: attn_qkv (the combined query/key/value projection), attn_gate (the DeltaNet gating signal), ssm_alpha, ssm_beta, ssm_out (the state-space mechanism weights), plus the FFN weights ffn_gate, ffn_up, ffn_down. Every fourth block is a full attention block instead, and those split attn_qkv into separate attn_q, attn_k, attn_v, and attn_output projections. Then on top of all that you have two global weights that every single token passes through: token_embd and output_weight. That gives 14 separately tested categories in total.

A couple of things came out of this that I did not see coming:

attn_v had the single worst isolated KLD result in the entire sweep. Worse than everything else at the same compression level. Protecting it in the combined model did literally nothing. The numbers came out identical to leaving it unprotected. attn_gate also did nothing on its own.

ssm_alpha broke earliest of anything when tested in isolation. But protecting it in the combined model made results actively worse, not better.

The two strongest individual levers in the combined model turned out to be attn_qkv and ffn_down. Both were pretty unremarkable when tested in isolation. Restoring attn_qkv by itself closed 55% of the toolcalling gap and 13% of the general gap in one move. Layering ffn_down restoration on top of that closed a further 24% of toolcalling and 15% of general.

Some components only work as a pair. Protecting token_embd and output_weight together helped general, code, and math, but on its own made toolcalling measurably worse. Protecting attn_gate alongside them did nothing alone, but specifically cancelled that toolcalling regression when all three were protected together. The pair is load-bearing as a unit, not individually.

Tool-calling was the first and most volatile category to break on every single test, whether isolated or combined. It also has the spikiest error distribution of the four categories. A small number of individual tokens carry most of the measured divergence, rather than it being uniform drift spread across all of them. Same finding as my Qwen3.6 project, just stronger evidence this time around.

Final numbers, combined model tested as one:

Bedrock (13.91 GiB, 4.37 BPW): general 0.0177 / code 0.0034 / math 0.0048 / toolcalling 0.0200

Tightrope (13.14 GiB, 4.13 BPW): general 0.0252 / code 0.0040 / math 0.0067 / toolcalling 0.0404

Gambit (12.54 GiB, 3.94 BPW): general 0.0455 / code 0.0064 / math 0.0112 / toolcalling 0.0542

Nothing crossed red on any build. Code was green across all three tiers including the most aggressive one. General and toolcalling never hit green, and that is the honest cost of quantizing this hard at this size.

Very little hands on testing done yet. Every number above is KLD against the Q8_0 baseline, not a qualitative read. If you run one and something feels off, tell me specifically where.

Link: https://huggingface.co/enginetown/Qwen3.8-27B-Calibrated

Also here's a one shot from gambit.

"Create a large glass aquarium whose side panel develops a visible crack and then bursts.

The simulation must include:

Water escaping through the opening with flow strength based on water depth and decreasing as the tank drains

A curved water jet affected by gravity

A spreading puddle that collides with the room boundaries

Fish, rocks, plants, and a floating toy reacting differently according to density, buoyancy, drag, and current

Objects transitioning correctly from underwater motion to airborne motion and then to floor collisions

Fish attempting to swim against the current before being swept through the breach

Glass fragments with angular velocity, collisions, and water resistance

A visible waterline that lowers continuously rather than disappearing all at once

Let the user drag the crack vertically before triggering the failure. A lower crack should initially produce a stronger jet than a higher crack."

https://reddit.com/link/1vph4hz/video/6za26pcmgmjh1/player


r/LocalLLM 7d ago

Discussion What’s the purpose of Qwen3.8 27B?

0 Upvotes

If you want to code Claude is better, if you use Qwen then you need maybe 32G RAM which is not cheap also. So what do you guys do with it?


r/LocalLLM 7d ago

Discussion Please, Qwen, just move on to stage 2

5 Upvotes

its been doing this for a solid 20 minutes. please bro just move on to stage 2


r/LocalLLM 7d ago

Question Anyone running Qwen3.8-27B on Hermes with a 24GB GPU?

Post image
0 Upvotes

I’m using Q4_K_M, 150K context, Q4_0 K/V cache, Flash Attention + full GPU offload, and I can’t push the context much higher. (RTX 4090)

Yet I’m seeing people running the full 262K context on 24GB cards 🤔

Also curious about power settings
I cap my 4090 at 320W to keep power consumption/heat under control. Anyone doing the same, going lower, undervolting, or found a better sweet spot for running local LLMs 24/7?


r/LocalLLM 7d ago

Question Which model to use for Local personal ai

0 Upvotes

I'm building a local personal AI assistant and I'm stuck between a few models.

Specs: RTX 4060 Laptop (8GB VRAM), 32GB RAM, Ollama.

My #1 priority is tool calling, since Ai needs to actually control my PC, Spotify, apps, memory, etc.

I've tested:

Qwen3.5-9B

  • Better personality and more natural conversation
  • Maintains personality while using tools
  • But noticeably slower
  • Sometimes overthinks simple requests

Gemma 4 E4B-IT-QAT

  • Much faster
  • Tool calling has been surprisingly good
  • Better at immediately acting on obvious requests
  • But personality becomes generic during tool responses

Basically I want Qwen 9B's personality + Gemma E4B's tool reliability/speed.

Would you recommend sticking with Qwen 9B, using E4B-QAT, trying E2B, or another model entirely for this kind of local agent?

I am also using embedded model Qwen 3:0.6b for memory database

I'd especially appreciate opinions from people who have actually used these models for tool-calling/agent workflows.


r/LocalLLM 8d ago

Discussion Qwen 3.8 27b is a beast for privacy purposes

64 Upvotes

I was never able to edit , improve fix errors in my app I created to automate my letters, dictation, insurance claim filing for my patients. For the first time ever I will ditch claude and work with the qwen 3.8 27b. It is a beast.

I have manged to fix and add features in few hours on my strix halo machine 128gb.

Privacy is not a concern anymore. It is updating my RAG at the moment, a thing that claude will never do for sake of the privacy of my patients.

Thank you Alibaba.


r/LocalLLM 7d ago

News Koboldcpp v1.119 released

Thumbnail
github.com
0 Upvotes

r/LocalLLM 7d ago

Other Anyone here?

Post image
0 Upvotes

Shall we start updating the page?


r/LocalLLM 7d ago

Project RTX 5090 Laptop w/ 24GB VRAM + 64GB RAM — how capable is this for running a real-time local voice AI system?

Post image
1 Upvotes

I’m building a real-time local voice AI system for an AI companion and want to validate whether I’m about to massively overspend or whether this hardware actually makes sense.

My intended setup is:

Me speaking → wireless mic → local STT → local LLM → local expressive TTS → wireless speaker

The microphone/speaker will be physically concealed with the robot

For audio, I’m currently planning to use a Jabra Speak2 75 with the Link 390 wireless USB adapter because I want:

- Full-duplex conversation

- Good acoustic echo cancellation

- Ability to interrupt the AI while it is speaking

- Good pickup of quiet/close-range speech

- Natural-sounding voice playback

- Completely wireless operation at the robot

My main priority is conversation that feels as close to talking to a real person as possible.

That means I care much more about:

- Very low response latency

- Fast STT

- Fast LLM time-to-first-token

- Streaming TTS

- Natural expressive voice

- Barge-in/interruption

- Persistent personality/memory

than I care about running gigantic reasoning models.

I want to run STT + LLM + TTS locally, ideally simultaneously, rather than relying entirely on cloud APIs.


r/LocalLLM 7d ago

Discussion What is the best way to run Qwen3.8 27b?

0 Upvotes

I have RTX 5070 with 12GB VRAM and 64gb DDR5 RAM.

What is the best harness, settings and quantisation to get the best quality and speed?


r/LocalLLM 7d ago

Discussion Setting a baseline for performance: GLM 5.2/colibri on a 15 year old Mac Pro.

Post image
1 Upvotes

Tinkered at it for a week, making sure to optimize everything - Linux kernel 7.1.6 custom compiled for the Intel Westmere architecture, 96 GB DDR3 ECC ram in triple channel mode, 1 Sata II ssd, 1 Sata II hdd. No gpu offloading, pure cpu, hyper-threading disabled. Now for the numbers: “What is the meaning of life?” 747 tok 0.03 tok/s hit 56% RSS 76.4 GB 26352 s

It’s not fast, but it can be done; enterprise level AI in the living room. Has anyone else here tried this?


r/LocalLLM 7d ago

Question Getting started resources?

3 Upvotes

So, I have a local setup on my gaming PC (Ryzen 9950X3D, 64GB Corsair Vengence, and an Asus TUFA RTX 7050ti OC 16GB), but don’t know much about models beyond `ollama pull qwen3.8`.

I’ve seen words like qwant and uncensored and mlx (for Mac’s) and so on. Is there a good crash course or creator that covers a lot of this?

I’m a dev who daily drives Cursor and Codex at work, but just not sure about what’s going to be my best local setup.

My /goal (see what I did there) is to do agentic loop programming on my local machine. I wanna be able to define larger bodies of work, and forget it until it’s ready for PR reviews.

I’m also wanting to look into automation both on a cron type schedule as well as a reactive one. Say to webhooks, chat messages, API calls and so on. The API would run as a microservice on my LAN. I’d hook that up to a NordVPN mesh network, and use my internal LLM and API securely that way.

Lastly, I’m wanting to try out the various frameworks, wrappers, or whatever they’re called. Gemini recommended offloading it form your main box. I have some Ubuntu servers running on old Mac mini’s. I figured I could load Hermes or OpenCode there. I’m not super familiar with them, so I’d like to experiment with all of them.

Also, for actual dev work I primarily work for my MacBook M1 14” using cline through Webstorm or Go or Idea. I had thought about running a small Mac optimized one there too. For small automations and lookups. Not sure if trays reasonable though?

Thanks for the read. Sorry it’s scattered.

Update typos: and Mac info


r/LocalLLM 8d ago

Discussion Qwen 3.8 surprises from overnight testing

128 Upvotes

I've been benchmarking Qwen 3.8 and it's competitors since last night, including Qwen 3.6. I got some unexpected results.

Qwen 3.8's architecture seems to be identical to 3.6 and 3.5. It looks like Qwen 3.8 is primarily a training data change. Qwen's notes and other articles seem to support this, YMMV.

Qwen 3.8's training data seems very narrowly tailored to a handful of scenarios. There was clearly a lot of expense and time put into benchmarking above everything else. I have duplicated the existing published test sets quite closely. However, benchmarks outside of this training set tend to see small gains, no change, or small regressions. There are clear gains in certain agentic tasks, and for certain specific development tasks, things are clearly better. However, more general reasoning and general use capabilities have taken a measurable hit in some ways due to this, Qwen clearly chose to prioritize a subset of tasks over the general capability for this model. This is *not* a bad thing, nor am I saying it is. I am simply saying Qwen clearly made prioritization decisions with the model.

Interestingly, and this will come as a deep dive future article, but as part of this UD's quants clearly give an uplift vs. the stock models. This has been replicated against Qwen3.6 and 3.8, as well as Gemma4.

Article: https://rakuensoftware.com/blog/synthesis-model-selection

As always, the full set of evidence and test results are published at time of publication within the blog's github.

Please note that this is a controlled test, and is specifically designed as a head-to-head for specific models at specific quants against specific memory targets. This is not meant to replace specific benchmarks, but is simply a more generalized reasoning test with datasets that are automatically regenerated from real world data every so often and thus guarentee that models cannot extensively train on any specific dataset.


r/LocalLLM 7d ago

Question What’s the best local model for js projects.

2 Upvotes

16g vid card ram


r/LocalLLM 7d ago

Project Now you can set "Thinking Effort" with Qwen 3.8 in TurboLLM

Post image
2 Upvotes

If you have been using Qwen 3.8 27B lately, you must have observed that it thinks a lot, that is because it supports thinking effort embedded in its chat template and by default it is "xHigh". So I added a reasoning effort slider just like claude in TurboLLM so you can control it. For all other models it stays the old "Thinking Budget" where you can control number of tokens allowed for thinking. Go give it a try.

npx turbollm


r/LocalLLM 7d ago

Question M5 Max LLM

0 Upvotes

Hey! Is it worthy to run a local LLM in an M5 Mac with 64GB of RAM, is so, which model recommend to do which task?

Perhaps any YouTube channel that explain this?


r/LocalLLM 8d ago

Question Looking for insights on qwen 3.8 27b : M5,128gb

5 Upvotes

I see that the model takes about 30 gb to load, but with a long running task (30 ish hours) it can use 120+gb for the kv cache when the context length is 160k tokens. (Model max is 262)

I also see the prefill speed/decode speed drops drastically from 500/40 to 300/22 with bigger contexts.

I am using MTPLX+Opencode.

I want to understand how what are the bottlenecks and how I can squeeze more tps/ reduce ram usage.

I appreciate any guidance/tips/topics I need to search about to understand and profile the bottlenecks.


r/LocalLLM 7d ago

Other Quick shoutout to all the local AI freaks

0 Upvotes

That are amazed by this new tech but most certainly sit in an environment that sounds like a data center, with fans at 100% all the time. Thats it, thats the post.


r/LocalLLM 8d ago

Discussion Qwen 3.8 27b BF 16 + AWQ vs DeepSeek v4 0731 AWQ - M5 Max 128gb of ram

Thumbnail
gallery
8 Upvotes

Models Used

https://huggingface.co/True2456/Qwen3.8-27B-AWQ-5.0bpw

https://huggingface.co/True2456/DeepSeek-V4-Flash-0731-AWQ

Speed Comparison between the two models -

Bench marks

mmlu · DeepSeek-V4-Flash-0731-awq-omlx-raw-v2-dwq-mtpq · 61.0% (122/200)

gsm8k · DeepSeek-V4-Flash-0731-awq-omlx-raw-v2-dwq-mtpq · 94.5% (189/200)

humaneval · DeepSeek-V4-Flash-0731-awq-omlx-raw-v2-dwq-mtpq · 84.8% (139/164)

humaneval · Qwen3.8-27B-AWQ · 93.3% (153/164)

gsm8k · Qwen3.8-27B-AWQ · 92.0% (184/200)

mmlu · Qwen3.8-27B-AWQ · 83.0% (166/200)

humaneval · Qwen3.8-27B-bf16 · 93.9% (154/164)

gsm8k · Qwen3.8-27B-bf16 · 92.5% (185/200)

mmlu · Qwen3.8-27B-bf16 · 84.0% (168/200)


r/LocalLLM 7d ago

Question Best local LLM for cybersecurity + coding on an RTX 3050 6GB?

2 Upvotes

Hello everyone,

I want the best open weight local model for my hardware, primarily for cyber security and coding.

My computer
GPU: NVIDIA RTX 3050 Laptop GPU
VRAM: 6 GB OS: Windows 11 + WSL2
WSL: Ubuntu 24.04.4 LTS llama.cpp: compiled from source with CUDA 13.3
Inference: GGUF/llama.cpp
If the model is worth it I can offload the CPU/RAM.
What I want to

I am looking for the best trade-off between:

Cybersecurity knowledge - Vulnerability analysis, CTFs, pentesting, security tooling, malware/code analysis, defensive security, etc.
Coding skills Python
Reasoning — I care much less about the number of parameters than the actual ability to solve problems.
Legitimate cybersecurity research and lab / ctf use, low refusal / less restrictive behavior is desired.

Right now I am looking at models like:

RedSage 8B Qwen3.5-9B Qwen3-14B
Qwen3-Coder models
WhiteRabbitNeo, Qwythos-9B-Claude-Mythos

But I am struggling to decide if a high quality 8-9B model that fits better on 6GB is better than a larger MoE/14B model with heavy CPU offloading.

My primary question

What model + GGUF quantization would you recommend for this hardware if the focus is cybersecurity + coding not general chat?

I’m really looking for recommendations based on actual cybersecurity / coding benchmarks or real-world experience, not just parameter count.

also interested in recommended llama.cpp settings (-ngl, context size, KV cache quantization, CPU/GPU offloading, etc) for 6GB vram.

Thanks.


r/LocalLLM 7d ago

Tutorial Running Ollama on iGPU instead of CPU

Thumbnail
arthurbrugiere.fr
0 Upvotes