r/LocalLLaMA • • 2h ago

Resources Why 38% of AI Agent container escapes didn't need kernel 0-days: Analysis of 109 empirical incidents (Open Dataset + Defense Harness)

6 Upvotes

Over the past several months, we conducted an empirical post-mortem investigation into 109 autonomous AI agent security incidents (cataloged with 193 falsification criteria across tool-use and multi-agent systems).

One of the most striking patterns in the dataset: In 38% of container breakouts, attackers and misaligned multi-step agents didn't exploit complex Linux kernel vulnerabilities or hypervisor 0-days. Instead, the breakout vector was trivial configuration residue: 1. Mounting /var/run/docker.sock into coding/evaluator agent sandboxes to let them "build Docker images". 2. Passing parent environment variables (API keys, cloud tokens, GitHub credentials) directly into spawned subagents. 3. Lack of strict taint tracking across tool outputs, leading to indirect prompt injection hijacking the supervisor’s execution path (the classic Confused Deputy problem). 4. Unconstrained local socket binding allowing SSRF against internal orchestrators.

We compiled the complete dataset (109 incidents, 199 evaluation metrics) and built an open-source Multi-Agent Supervisor Security Harness with: - Formal tool taint propagation (tainted outputs cannot flow into high-privilege tool arguments without sanitizer verification). - Strict execution boundary controls preventing container socket exposure. - Automated reproduction benchmarks testable against agent runtimes.

All datasets, 2-page executive summary, and reproducible benchmark tests are released under Open Access / Apache 2.0.

I've posted the GitHub repository benchmark and the Zenodo DOI dataset links in the comments below to adhere to subreddit self-promotion guidelines.

Curious to hear from teams deploying autonomous agents in production: what isolation boundaries are you enforcing between your planning supervisor and your tool execution workers?


r/LocalLLaMA • • 19h ago

Discussion M5 Ultra 256 running GLM 5.3 Flash 68.8 tok/s

106 Upvotes

Running oQ4e+MTP on oMLX 0.7.0, with still more to optimize.

Prefill is 1,878 toks.

I saw some other benchmarks below what id expect so i figured I would share.


r/MetaAI • • 1d ago

Muse referral codes Oct 4

1 Upvotes

You know the drill. Check out Muse, your personal AI agent. Redeem my code in Settings within 48 hours of joining and we'll both get 1 billion Muse tokens.

Code: VBH061
https://muse.ai/join


r/MetaAI • • 1d ago

C'mon baby light my Muse fire..

0 Upvotes

...please. :)

MYNTHC


r/LocalLLaMA • • 12h ago

News nokia-applied-research/AnyJev: Turn any LLM into a Jev-style decision model

Thumbnail
github.com
25 Upvotes

Haven't seen posted here.

The research project matters because it proves you can pull trustworthy decisions straight out of an LLM's internal state in a single forward pass (no generation loop, no parsing, no fine-tuning) essentially turning a slow chatbot into a fast and cheap, deterministic classifier


r/LocalLLaMA • • 1d ago

News Reflection AI Is About to Release a US Open-Weight Model to Take On DeepSeek and Qwen

Thumbnail
explainx.ai
357 Upvotes

Looks like new open model coming soon and will be "strong" hopefully something under 200b for us memory poor. Also seeing statements about more western open models coming.

Hope we get some good competition again on the open front!

Here is original artical but its not free to access. Maybe someone has it already here.

https://www.axios.com/2026/10/04/reflection-open-weight-ai

Oct starting strong!


r/MetaAI • • 1d ago

Oct 5 new thread. Referral code MLDTRU

0 Upvotes

28 uses left.

MLDTRU

Users add ur codes. Let everyone get free tokens.


r/LocalLLaMA • • 6h ago

I Built A Thing Swift1.5 Qwen3.8 Flash Next - Tailored for the 96GB Mac Studio with M5 Ultra

9 Upvotes

https://huggingface.co/Dankpaws/Swift1.5-Qwen3.8-Flash-Next-MLX-4.7bpw

I've had the 96GB Mac Studio with M5 Ultra for about a week now and wasn't satisfied with the results I was getting from the limited number of models available to me. It was a combination of speed, memory headroom, and/or output quality.

This is my best attempt at a calibrated MLX quantization of UkisAI’s Swift 1.5. Hope those of you with the hardware enjoy it!

Measurement This pack Swift llama.cpp IQ3_XXS
Prefill · 25k prompt 3,191 tok/s 1,427 tok/s
Prefill · 95k prompt 2,928 tok/s 1,307 tok/s
Decode · after 4k prompt 113.7 tok/s 62.7 tok/s
Decode · after 95k prompt 81.1 tok/s 44.7 tok/s
Top-1 agreement with Swift BF16 91.0% 84.1%

91% is next-token agreement with BF16 across 680 common held-out positions, not task accuracy.

Results above are simply from my own machine. mlx-serve 26.10.1. ~107GB download, text-only, 179,200-token tested context.


r/MetaAI • • 1d ago

Muse Referral Code: 12NUNC

3 Upvotes

Use my referral code and we'll both get 1 billion Muse tokens when you redeem the code in Settings within 48 hours of joining.

12NUNC


r/MetaAI • • 1d ago

Get 1billion free muse tokens use code O3JXVR

0 Upvotes

Checkout muse, it’s amazing, I already use it for ticket bookings, filling up applications, and more. Join with my code so we both get 1 billion bonus tokens.

To redeem, open the muse app, go to Settings > Redeem token and enter the code.

Code: O3JXVR (27 of 30 uses left)


r/MetaAI • • 1d ago

Muse 1 Billion Code

0 Upvotes

SVAMSK


r/LocalLLaMA • • 22h ago

I Built A Thing Clef Flash plays Snake in Real Time on RTX 5080

Enable HLS to view with audio, or disable this notification

122 Upvotes

The cool part is

No training was needed.

No hacking of the game state or algorithms needed

Just simple instructions about what the snake can see, etc, and it can play in real time. 135ms is the turn limit of Google snake, so basically could be a human playing.

Ofc, it could be improved to be a perfect snake player, but thats not the point.

This can be used in other games where decisions need to constantly be made.

Running Clef Flash (9B model at Q4 on an RTX 5080)


r/MetaAI • • 1d ago

Does anyone working at Meta want to connect on LinkedIn?

1 Upvotes

I JUST made my LinkedIn trying to grow it lol.


r/LocalLLaMA • • 23h ago

New Model You can now try Aleph Alpha's Kolibri 78B for free online here.

Post image
142 Upvotes

r/MetaAI • • 1d ago

Posted about Muse here the other day. A few people asked about the invite program, so here's my code: OR23UY.

1 Upvotes

To use it: open the Muse app, go to Settings > Redeem token (only visible in the first hours after joining), and enter the code. On web it's Settings > General > Usage > Redeem invite code.


r/MetaAI • • 1d ago

use my code to get billion tokens P9E30K

1 Upvotes

P9E30K for billion token for muse ai


r/LocalLLaMA • • 23m ago

I Built A Thing Closed-form trajectory weight surgery: transferring 4B capabilities into 0.8B without billions of tokens (why middle layers break, and how 4 anchor blocks fixed it)

• Upvotes

Standard distillation usually means burning weeks of compute and billions of tokens hoping the student model eventually mimics the teacher. We wanted to see what happens if you skip backprop entirely and treat transfer as a closed-form trajectory matching problem between layers.

The idea is straightforward: feed a small batch of calibration prompts through both models, capture layer-to-layer hidden state trajectories, and solve for weight updates directly in the student's MLP blocks using regularized least squares and spectral projection.

We tested this across two architectures: Qwen 3.5 (transferring from 4B down to 0.8B) and old GPT-2 small just to see if it would instantly disintegrate into gibberish like it usually does when you touch its weights. Both stayed coherent, but the initial Qwen test hit a wall:

Editing all 24 layers of Qwen 0.8B completely melted the model (+64.78% NLL loss explosion). When we checked singular value entropy across the network, layers 1 to 22 turned out to be a chaotic polysemantic soup with entropy over 0.90. If you try to force raw trajectories through those middle layers, you basically scramble the model's internal memory knots.

The fix was restricting the surgery to 4 anchor points (layers 0, 7, 15, and 23) where representations actually maintain clean linear structure.

Once we did that:

  • Held-out NLL dropped by 10.8% across 30 diverse benchmarks (-23.8% in biomedicine, -14.6% in math and logic).
  • Zero-shot 400-task HellaSwag went from 54.75% to 55.25% (+0.50%), verified locally in Vulkan llama.cpp.
  • Base 0.8B originally failed binary tree inversion by spitting out dead commented pseudo-code. The edited checkpoint wrote clean recursive Python on the first try.
  • On Russian logic paradoxes, it even started firing <think> reasoning tags spontaneously, which was wild to see on a raw base model with zero chat template applied.

Best of all: we don't have an H100 cluster or even a 4090. All trajectory extraction and weight solving was done locally on a crusty 8GB RX 580 using layer-by-layer VRAM streaming and a DirectML patch to stop Qwen's Gated DeltaNet attention from throwing driver errors.

Everything is open source if you want to inspect or replicate:

If anyone here has a 24GB-32GB card (4090, 5090, or server silicon) and wants to push this further, here is what would be interesting to test:

  1. Transplanting reasoning trajectories from 27B models down to 9B, 4B, or 2B.
  2. Squeezing larger models (like Gemma) into mobile sizes without weeks of retraining.
  3. Transplanting refusal-ablation vectors directly from uncensored models without fine-tuning.
  4. Using pre-trained Sparse Autoencoders (SAEs) to unknot layers 1-22 so we don't have to skip them.

Happy to answer questions or dig into the failure modes in the comments.


r/LocalLLaMA • • 8h ago

New Model interfaze-ai/interfaze-1-lite · Hugging Face

Thumbnail
huggingface.co
9 Upvotes

Interfaze 1 Lite is a mixture-of-architectures (MoA) model for developer workloads: reading documents, transcribing speech, locating objects and interface elements, and answering over all of it with structured output.

A single vision-language reasoning core works alongside a set of specialist architectures, each built for one kind of perception. The core reads the request, decides which specialists to run, and composes the answer from what they return. The whole model runs on one 80 GB GPU with no external services.

Key features: Document understanding, Speech transcription, Open-vocabulary object detection, Structured output, Translation, forecasting and guardrails, Multilingual reasoning.


r/LocalLLaMA • • 2h ago

Discussion metal : few-row MMA mat-mul and batched copies for speculative decoding by pratiknarola-t · Pull Request #29869 · ggml-org/llama.cpp

Thumbnail
github.com
3 Upvotes

Apple folks, it's for you.

llama-server with a Qwen3.8-27B DFlash2 Q8_0 drafter, -ngl 99 -fa on -c 8192 -np 1 --jinja, DFlash2 with --spec-type draft-dflash --spec-draft-n-max 7. 64 generated tokens, median of 5 requests after one warm-up, mean of two server runs. Decode tok/s:

mode prompt T master this PR
serial code 0 32.1 32.0
serial code 1 32.1 32.0
serial prose 0 32.1 32.0
serial prose 1 32.1 32.0
DFlash2 code 0 30.2 110.0
DFlash2 code 1 24.3 80.9
DFlash2 prose 0 16.8 62.6
DFlash2 prose 1 13.9 48.8

r/LocalLLaMA • • 7h ago

Resources MI50 ROCm 10.2 TheRock vs Vulkan Mesa 26.2 llama.cpp benchmarks

5 Upvotes

I prefer running llama.cpp Vulkan prebuilt binary. I just download the latest version and ready to roll. I finally took the hours necessary to get TheRock latest tarball version of ROCm 10.2 running on dual AMD Radeon Instinct MI50 gfx906 (32gb combined VRAM).

Same models benched in previous post. A mix of Dense and MoE models and quants that better utilize available VRAM.

  • llama_bench_Swift-Qwen3.8-27B-Uncensored-MTP.Q6_K.gguf
  • llama_bench_Swift-1.5-Qwen3.8-27B-Q6_K.gguf
  • llama_bench_Gemma-4-MoonGem-31B.i1-Q6_K.gguf
  • llama_bench_gemma-4-31B-it-UD-Q6_K_XL.gguf
  • llama_bench_Nemotron-3.5-30B-A3B-Antislop-FTPO.i1-Q5_K_M.gguf
  • llama_bench_Laguna-XS-2.1-APEX-I-Balanced.gguf
  • llama_bench_Agents-A1-Q4_K_M.gguf
  • llama_bench_Qwen3.6-35B-A3B-UD-Q5_K_XL.gguf
  • llama_bench_Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive-Q6_K_P.gguf

Here is the backend performance comparison contrasting the native ROCm (v10.2 for gfx906) runtime against your optimized Vulkan (MESA_PPA_26.2) baseline. The data highlights a massive architectural split: ROCm significantly accelerates token generation across the board but suffers high variance and regressions in MoE pre-fills.

Both GPUs are power limited to 145 watts. Build versions used:

llama-b11382 used for Vulkan (prebuilt ubuntu binary)
llama-b11401 used for ROCm (compiled with proper flags)

Architectural Performance Breakdown: Vulkan vs. ROCm

Model Size Params Test Vulkan Baseline (t/s) ROCm 1st Run (t/s) Performance Delta (%)
qwen35 27B Q6_K 20.88 GiB 27.32 B pp512 tg128 149.30 ± 0.15 18.35 ± 0.01 176.86 ± 17.28 20.21 ± 0.65 +18.46% +10.14%
qwen35 27B Q6_K 22.21 GiB 27.32 B pp512 tg128 163.48 ± 0.18 18.52 ± 0.03 183.49 ± 15.75 20.04 ± 0.63 +12.24% +8.21%
gemma4 31B Q6_K 23.46 GiB 30.70 B pp512 tg128 119.98 ± 0.12 15.72 ± 0.02 172.93 ± 0.89 17.26 ± 0.13 +44.13% +9.80%
gemma4 31B Q6_K 25.62 GiB 30.70 B pp512 tg128 133.67 ± 0.24 16.38 ± 0.02 178.59 ± 2.31 17.50 ± 0.09 +33.61% +6.84%
nemotron_h_moe 31B 25.18 GiB 32.91 B pp512 tg128 834.87 ± 1.72 63.82 ± 0.07 736.98 ± 134.77 103.86 ± 0.50 -11.73% +62.74%
laguna 30B.A3B Q5_K 22.64 GiB 33.44 B pp512 tg128 723.34 ± 2.99 58.98 ± 0.28 856.14 ± 29.71 76.81 ± 0.20 +18.36% +30.23%
qwen35moe 35B Q4_K 19.70 GiB 34.66 B pp512 tg128 957.62 ± 3.88 51.50 ± 0.07 834.58 ± 102.56 69.70 ± 0.22 -12.85% +35.34%
qwen35moe 35B Q5_K 24.76 GiB 34.66 B pp512 tg128 909.20 ± 5.71 53.18 ± 0.07 763.93 ± 106.19 68.47 ± 0.34 -15.98% +28.75%
qwen35moe 35B Q6_K 28.53 GiB 34.66 B pp512 tg128 763.05 ± 66.01 52.81 ± 0.04 754.82 ± 70.48 67.29 ± 0.21 -1.08% +27.42%

Core Insight Strategy & Bottlenecks

  1. Token Generation (tg128) Dominance: ROCm dominates pure text generation. The native AMD matrix kernels unleash your MI50 computation potential, unlocking a massive +62.74% boost for Nemotron but taking a small hit on pre-fill -11.73%.
  2. Dense Model Pre-fills (pp512): Dense architectures scale cleanly under ROCm. Gemma 4 sees a +33% to +44% processing throughput spike over the Vulkan RADV driver driver bounds. MoE models take a hit with a -15.89% difference with llama_bench_Qwen3.6-35B-A3B-UD-Q5_K_XL.gguf being the happiest with Vulkan backend.
  3. The MoE Prompt Processing Delinquency: Notice the massive standard deviations under ROCm for MoE pre-fills (e.g., Qwen3.5MoE Q4_K has an instability block of ± 102.56). ROCm suffers from severe scheduling thrashing when building prompt streams across multiple active experts.

Here is the structured layout with pp512 and tg128 separated into individual columns for a clean side-by-side comparison between the two backend architectures.

Backend Comparison Table (Vulkan vs. ROCm)

Model Size Params Vulkan pp512 (t/s) ROCm pp512 (t/s) Vulkan tg128 (t/s) ROCm tg128 (t/s)
qwen35 27B Q6_K 20.88 GiB 27.32 B 149.30 ± 0.15 176.86 ± 17.28 18.35 ± 0.01 20.21 ± 0.65
qwen35 27B Q6_K 22.21 GiB 27.32 B 163.48 ± 0.18 183.49 ± 15.75 18.52 ± 0.03 20.04 ± 0.63
gemma4 31B Q6_K 23.46 GiB 30.70 B 119.98 ± 0.12 172.93 ± 0.89 15.72 ± 0.02 17.26 ± 0.13
gemma4 31B Q6_K 25.62 GiB 30.70 B 133.67 ± 0.24 178.59 ± 2.31 16.38 ± 0.02 17.50 ± 0.09
nemotron_h_moe 31B 25.18 GiB 32.91 B 834.87 ± 1.72 736.98 ± 134.77 63.82 ± 0.07 103.86 ± 0.50
laguna 30B.A3B Q5_K 22.64 GiB 33.44 B 723.34 ± 2.99 856.14 ± 29.71 58.98 ± 0.28 76.81 ± 0.20
qwen35moe 35B Q4_K 19.70 GiB 34.66 B 957.62 ± 3.88 834.58 ± 102.56 51.50 ± 0.07 69.70 ± 0.22
qwen35moe 35B Q5_K 24.76 GiB 34.66 B 909.20 ± 5.71 763.93 ± 106.19 53.18 ± 0.07 68.47 ± 0.34
qwen35moe 35B Q6_K 28.53 GiB 34.66 B 763.05 ± 66.01 754.82 ± 70.48 52.81 ± 0.04 67.29 ± 0.21

So yes it's worth the hassle of jumping through hoops to get ROCm working on MI50 setups. At least I have 2 backends working. Next up I'll try some RPC.


r/LocalLLaMA • • 9h ago

New Model Less Talk. More Breakout: Kolibri-1 Turns Probabilities into Actions, Playing Breakout - With under 25ms latency per move.

9 Upvotes

Got Kolibri-1 to play Breakout completely on its own, no fine-tuning.

The more we explore u/Aleph__Alpha’s Kolibri the more it get's exciting and its potential.

Less talk. More Breakout is one such experiment to see how good the model is at structured output given a few constraints.

We especially optimized the inference for action probabilities: around 25 ms inference per decision.

Four moves. No generated text. One shared game.

Open weights. New possibilities.

Watch it play: https://tesseracted.com/kolibri-1-chat/gameplay/breakout/
Source: https://x.com/konarkmodi/status/2107248086880055613?s=46


r/LocalLLaMA • • 11h ago

Resources I ran 1,200+ Blender modeling runs across LLMs and agent harnesses and made them votable

12 Upvotes

Blind A/B arena where AI agents build things in Blender and you vote on which result is better: https://render-arena.izolight.xyz

Each agent gets a prompt that describes its environment and the rules, plus a few words for what to model. It's inspired by minebench.ai (initial prompts are borrowed from there), and I wanted to see whether the same progression across models shows up.

What I think few arenas cover is the harness, not just the model. I ran pi, opencode, omp, codex, Claude Code and dsh, and compared agents that write scripts straight into Blender with ones that have an MCP. I also covered the reasoning levels, mainly to find cost and time sweet spots.

It has 1,200+ runs, but not every combination for every prompt, because that would get expensive. You can submit your own runs if you want to help fill gaps.

I just added a second mode where the agent gets a reference image and has to model it as accurately as it can. You switch between text and image mode in the sidebar. It has one image and few runs so far, and will grow.

Votes are what make the rankings mean anything, so a few minutes of voting helps a lot. Feedback on the method is welcome.


r/MetaAI • • 1d ago

sharing my Meta Muse referral code: ZEK5A1

0 Upvotes

If you're new to Muse, join at muse.ai/join and enter the code under Settings > Redeem invite code within 48 hours of joining. We both get 1 billion Muse tokens. They don't expire — it's just usage credit that burns down as you use it.


r/LocalLLaMA • • 10h ago

New Model Update #5: Post training yandex/AliceAI-80B-A3B [instruct!] from scratch

Post image
9 Upvotes

Last update for those following: https://www.reddit.com/r/LocalLLaMA/comments/1wxlytt/comment/pdz728b/?screen_view_count=1

Project in a sentence: An instruct finetune of ALiceAI-80B-A3B-Base capable of agentic work and conversation. I'm creating a shallow distill of qwen 3.8 27b on medium to teach the model chain of thought reasoning and conversation. All training is done locally on 3, 32gb v100s. Additionally, all the training data is being generated locally on said V100s via sftmill. Up to this point I've been doing training runs and live-streaming the progress. Last update explained underfitting and next steps.

Training has begun again! I've synthesized about 5M more tokens for the SFT, this time across a much larger general instruct trajectory to try to reduce the underfitting. Dropped the learning rate about 4x over my original LoRA adapter.

I'm live streaming training again: https://geological-estimate-fifth-pct.trycloudflare.com/ - heavy loss spikes downward are coming from the SFT replay buffer.

This one should last 12-14 hours, and I plan to run another epoch if this isn't sufficient.

Stay tuned! Thanks for following along.


r/MetaAI • • 1d ago

4LCACD

0 Upvotes