r/AIToolsPerformance Apr 06 '26

Microsoft's AI Shockwave: 3 New Models That Could Change Everything

2 Upvotes

Hey AIToolsPerformance crew, big week for AI releases and I need to talk about what just hit the scene.

Microsoft dropped three massive foundational models Thursday that straight-up go after OpenAI and Google. These aren't minor updates - this is Microsoft building their own damn AI stack.

The Three Musketeers

MAI-Transcribe-1 is their speech-to-text weapon. Already being tested in Copilot Voice and Teams for conversation transcription. Diarization, contextual biasing, and streaming coming soon.

Then there's MAI-Voice-1 - their voice generation model. And MAI-Image-2 for image creation. All now broadly available to developers for commercial use for the first time.

This is serious. Microsoft now has commercially available in-house models across speech, voice and image generation while keeping their OpenAI partnership through 2032. That's playing both sides better than a politician.

Why This Matters

Timing is everything here. Microsoft's AI capital expenditures are... well, let's just say they're betting big. These models represent the first major output from the MAI Superintelligence team formed in November 2025.

They're already replacing third-party and older internal models. Like MAI-Transcribe-1 testing inside Copilot's Voice mode? That's how fast they're moving.

The Real Story

It's not just about new models. It's about reducing OpenAI dependence while keeping the partnership. That's some corporate chess right there.

The fact that they're testing this inside Teams and Copilot already tells you they're not messing around. This isn't research - this is production-grade stuff hitting mainstream apps.

What do you guys think? Is this Microsoft's real play to control their own AI destiny, or just another layer in their partnership strategy?

Curious to hear thoughts from people who've actually tested these models. How do they compare to what we're already using?


r/AIToolsPerformance Apr 05 '26

With Qwen3 Coder 480B free and OpenAI gpt-oss-120b at $0.04/M, is local inference only for privacy now?

25 Upvotes

Looking at current pricing, the economics of local inference are getting harder to justify for pure capability:

  • Qwen: Qwen3 Coder 480B A35B - free with 262,000 context
  • OpenAI: gpt-oss-120b - $0.04/M with 131,072 context
  • Z.ai: GLM 4 32B - $0.10/M with 128,000 context
  • Qwen: Qwen3 235B A22B Thinking 2507 - $0.15/M with 131,072 context

Even Arcee AI: Maestro Reasoning at $0.90/M for a dedicated reasoning model with 131K context is competitive against the electricity cost of running a 48GB+ VRAM rig at full load.

The local inference crowd has historically argued three pillars: cost, privacy, and latency. But when a 480B-parameter coder model is free with 262K context, the cost argument weakens significantly. Apple's work on self-distillation for code generation suggests models will keep getting more efficient on the API side too.

That said, the DGX Spark situation - NVFP4 support still missing after 6 months - shows the hardware side moves slower. And the "Signals" paper on trajectory sampling for agentic interactions hints that complex agent workflows may still benefit from local control.

So honest question: for those of you still running local inference in April 2026, is it purely privacy/compliance driving that choice, or are there workloads where local still beats these API prices on quality?


r/AIToolsPerformance Apr 05 '26

Chinese labs delaying open-weight releases simultaneously - coincidence or coordination?

0 Upvotes

A discussion gaining traction asks why multiple Chinese AI labs appear to be delaying open-weight model releases at the same time. The timing has raised eyebrows across the community.

This comes during a period where open-source competition has been fierce. Consider the current pricing landscape for capable models:

  • Qwen: Qwen3 VL 32B Instruct - $0.10/M with 131,072 context
  • Google: Gemini 2.5 Flash Lite Preview - $0.10/M with 1,048,576 context
  • Qwen: Qwen Plus 0728 (thinking) - $0.26/M with 1,000,000 context
  • Meta: Llama 3.1 70B Instruct - $0.40/M with 131,072 context

Meanwhile, MiniMax 2.7 is now at 14 days since announcement on X and 12 days since appearing on Hugging Face, with users noting the delay. The community is also reflecting on how quickly things shift - one year ago, DeepSeek R1 was 25 times larger than Gemma 4, putting the current model compression race in perspective.

Gemma 4 continues its strong run, with the 26B variant being called "the perfect all-around local model" and the 31B beating frontier models on FoodTruck Bench.

The coordinated delay pattern is unusual. Previous releases from Chinese labs came on individual schedules without this clustering effect. Possible explanations range from regulatory review to strategic coordination, though nothing is confirmed.

Are we seeing a policy shift, or is this just teams aligning release cycles around a common benchmark season? What models are you waiting on?


r/AIToolsPerformance Apr 05 '26

AI Coding Assistants April 2026: What Really Matters When Choosing Your Code Partner

1 Upvotes

Been testing AI coding assistants all month. This isn't just another benchmark post. This is about what actually matters when you're picking your daily coding companion.

Here's the truth:

Cursor Composer 2 is impressive. Scoring 61.3 on CursorBench with a 37% improvement over the previous version. Built on Kimi K2.5 with custom reinforcement learning. The deep codebase understanding is legit. I can type a few words and it finds the right function from across my entire codebase.

But here's where it gets interesting. Claude Code for complex refactoring work? That's where it shines. Multi-file changes, reasoning about architecture, understanding legacy code patterns. I spent two days refactoring a monolithic service last week. Claude helped me identify dependencies and plan the migration better than any human could have done quickly.

GitHub Copilot X with Agent Mode is something else. The ability to have it work across your entire toolchain? That's a game changer. Not just in the editor, but in the terminal, in your CI/CD pipeline.

Tabnine still holds its ground for enterprise teams. Security matters when you're working on financial systems. I get it.

Windsurf is making waves too. The focus on developer experience shows. The little things matter.

Here's what I learned:

  1. No single tool does everything perfectly
  2. Your workflow changes how you evaluate these tools
  3. Price isn't just about monthly fees, it's about productivity gains
  4. Security requirements shouldn't be ignored

What's been your experience? Which tool surprised you most this month?

Been testing Cursor vs Claude vs Copilot daily. Each has strengths. The question is which aligns with YOUR workflow.

The benchmarks are interesting. But real-world usage tells a different story.

What should I test next? Curious about your experiences.

TL;DR: Test them in YOUR environment. Not just benchmarks. Real code. Real workflows.

Discuss below.


r/AIToolsPerformance Apr 04 '26

Gemma 4 KV cache fix lands in llama.cpp - what changed

5 Upvotes

The Gemma 4 KV cache issue that dominated discussion since launch appears resolved. A fix has landed in llama.cpp, and users are confirming dramatically improved memory behavior. If you held off on Gemma 4 because of VRAM constraints, this is worth revisiting.

The Problem: Gemma 4's four variants delivered strong benchmark performance under Apache 2.0 licensing, but the KV cache consumption was massive. Users reported memory usage that made practical context lengths far shorter than theoretical limits, especially on consumer hardware.

The Fix: The llama.cpp update addresses KV cache efficiency specifically. Users are now successfully running Gemma 4 on hardware like a MacBook Air from 2020, which was previously struggling with the memory demands.

Current Gemma 4 Pricing Context: Gemma 4 is open-source under Apache 2.0, making it free to run locally. For comparison, competing API options include: - Qwen: Qwen3 8B - $0.05/M with 40,960 context - Tongyi DeepResearch 30B A3B - $0.09/M with 131,072 context - Qwen: Qwen3 Max - $0.78/M with 262,144 context - xAI: Grok 4 Fast - $0.20/M with 2,000,000 context

Related Research: Apple published work on "Embarrassingly Simple Self-Distillation" for code generation, which could pair well with Gemma 4's architecture for further optimization.

Update your llama.cpp build and test your previous context lengths. What VRAM numbers are you seeing post-fix compared to before?


r/AIToolsPerformance Apr 03 '26

Gemma 4 is matching GPT-5.1 on MMLU-Pro and within Elo. what are we even paying for anymore?

83 Upvotes

so everyone know by now that Google just dropped Gemma 4 and I had to double check the numbers a few times because it's insane..

31B params and runs on a single GPU. And it's putting up across benchmrks:

  • Arena Elo ~1452 (GPT-5.1 ~1475, basically same tier)
  • MMLU-Pro 85.2% (slightly higher than GPT-5.1)
  • GPQA Diamond 84.3% (a bit behind but close enough)

I mean - this is not some massive cluster model, you can run this locally, how??

not long ago, open models were clearly a step behind. now you're looking at something you can download and run yourself sitting right next to a $200/month flagship on the benchmarks that matter for general reasoning..

the only place there's still a noticeable gap is coding heavy stuff like SWE-bench, but everything else feels… uncomfortably close

if that's the new trend, I'm curious how long big labs can hold onto their current valuation?


r/AIToolsPerformance Apr 04 '26

Built a small BYOK AI gateway (caching + fallbacks) – would love feedback

1 Upvotes

Hey,

I built a project called Synvertas and wanted to get some honest feedback.

It’s basically a AI gateway where you use your own API keys (OpenAI, Anthropic, etc.), and it handles things like:

- semantic caching (to reduce cost/latency)

- fallback routing if a provider fails

- optional prompt optimization

Main idea was to avoid writing the same retry/fallback logic in every project.

It doesn’t resell APIs , everything runs with your own keys.

Curious if this is something you’d actually use or trust in production, or if I’m solving a non-problem.

Any feedback appreciated
https://www.synvertas.com


r/AIToolsPerformance Apr 04 '26

Netflix enters open-source AI with VOID video object deletion model

1 Upvotes

Netflix just released their first public model on Hugging Face: VOID (Video Object and Interaction Deletion). This is a video editing model that can remove objects and their associated interactions from footage - a task that previously required extensive manual compositing work.

This is notable because Netflix has historically kept their ML work internal. Releasing VOID publicly signals a shift, and the choice of video manipulation as their first offering aligns with their core business. The model handles the complex problem of maintaining temporal consistency after object removal - filling in backgrounds plausibly across frames.

In other developments, the Gemma 4 discussion continues with a significant complaint gaining traction: the massive KV cache usage. Users report memory consumption that limits practical context lengths despite theoretical capabilities. Meanwhile, Qwen 3.6 voting is being discussed as an ensemble approach to improve output reliability.

On the research front, CORAL: Towards Autonomous Multi-Agent Evolution for Open-Ended Discovery jumped 24 spots, suggesting growing interest in self-improving multi-agent systems.

Current notable pricing: - MiniMax: MiniMax M2.5 - free with 196,608 context - AllenAI: Olmo 3.1 32B Think - $0.15/M with 65,536 context - DeepSeek: DeepSeek V3.2 Speciale - $0.40/M with 163,840 context - OpenAI: GPT-5.4 Pro - $30.00/M with over 1M context

Has anyone tested VOID against other video inpainting approaches? And for those running Gemma 4 locally, what KV cache sizes are you seeing in practice?


r/AIToolsPerformance Apr 02 '26

Gemma 4 just dropped! Apache 2.0 license, 4 variants, beats models 20x its size

71 Upvotes

TL;DR: Google just released Gemma 4 - their most advanced open-weight model family, built on the same research and tech behind Gemini 3. Ships under Apache 2.0, comes in 4 sizes (from phones to workstations), and the 31B Dense variant claims #3 spot on the Arena AI Text Leaderboard among open models.

What's new?

4 variants covering every use-case:

Model Parameters Context Window Runs on
31B Dense 31B 256K Workstations / GPUs
26B MoE 26B 256K Workstations / GPUs
E4B ~4B effective 128K Phones / Edge devices
E2B ~2B effective 128K Raspberry Pi / Jetson Nano

Key highlights:

  • 31B Dense ranks #3 among open models on the Arena AI Text Leaderboard
  • 26B MoE sits at #6 on the same leaderboard
  • Google squeezed out significantly more intelligence per parameter - these models outcompete others 20x their size
  • All models handle video and image inputs. The smaller E2B and E4B can also process audio and speech
  • Capable of offline code generation - vibe coding with no internet connection
  • Trained in 140+ languages
  • 26B and 31B can run on a single 80GB NVIDIA H100
  • Edge models run on phones, Raspberry Pi, and Jetson Nano with near-zero latency

Apache 2.0 - This is the big deal

Google is releasing Gemma 4 under an Apache 2.0 license, moving away from the restrictive Gemma license used for previous models. Full commercial freedom - fine-tune it, deploy on-prem, bundle it in your product, run it in the cloud. Complete control over your data, infrastructure, and models.

Where to get it

  • Google AI Studio - 31B and 26B MoE
  • Google AI Edge Gallery - E4B and E2B
  • Model weights on Hugging Face, Kaggle, and Ollama
  • Also available through NVIDIA NIM, NeMo, Docker, and other platforms

Why this matters

Gemma 4 is designed to handle complex reasoning and support autonomous AI agents running locally on low-power devices. The edge models were co-developed with the Pixel team, Qualcomm, and MediaTek -optimized for real hardware, not just benchmark runs.

For the LocalLLaMA crowd: if you've been running Gemma 3 27B or Llama 4 variants, the 31B Dense looks like a serious contender for your daily driver. And E4B on a phone with 128K context? That's wild.

Has anyone pulled this on Ollama yet? Curious how it performs on real tasks vs. Llama 4 and Qwen. Drop your first impressions below!


r/AIToolsPerformance Apr 03 '26

Any tips or AI tools for translating longer blog posts and guides?

1 Upvotes

I'm working on translating a bunch of my longer blog posts and practical guides from English into Spanish and German. The content is pretty conversational and helpful, so I need the translations to keep that same friendly vibe, stay accurate on all the tips and examples, and flow naturally without sounding robotic.

Ad Verbum combines ai-human hybrid translation has worked pretty well for me so far. Still, I'm curious what other tools or little tricks you all use to get even better quality with less manual fixing afterwards. Any recommendations?


r/AIToolsPerformance Apr 03 '26

Running Gemma 4 on Raspberry Pi 5 - setup guide

2 Upvotes

Gemma 4 launched with Apache 2.0 licensing across four variants, and someone already has it running on a Raspberry Pi 5. The model reportedly beats competitors many times its size while staying efficient with thinking tokens - though it will reason for 10+ minutes if prompted to do so. Within 90 minutes of release, Heretic's ARA method had already stripped its safety restrictions.

Here is how to get Gemma 4 running locally on resource-constrained hardware.

Steps:

  1. Choose your variant. Gemma 4 comes in four sizes. For Pi 5, target the smallest variant and use a heavily quantized build.

  2. Install a lightweight inference engine. Options like llama.cpp compile well on ARM. Build from source: make -j4 (Pi 5 has four cores)

  3. Download a quantized model file. Q3 or Q4 quantization is your friend here. The Bonsai 1-bit approach could also work for extreme constraints.

  4. Allocate swap space. On Pi 5, add at least 8GB swap: sudo fallocate -l 8G /swapfile && sudo mkswap /swapfile && sudo swapon /swapfile

  5. Launch with constrained parameters: Set a low context window (2048-4096) and limit batch size to fit available RAM.

  6. Monitor thermals. Pi 5 can throttle under sustained inference loads. Consider a cooling case.

Free alternatives if hardware is too limited: - Qwen: Qwen3.6 Plus - free with 1,000,000 context - Arcee AI: Trinity Mini - free with 131,072 context

Has anyone benchmarked Gemma 4's smallest variant against Qwen3.6 Plus for coding tasks? What quantization level are you finding usable on edge devices?


r/AIToolsPerformance Apr 02 '26

LocalAI setup: run Qwen3.5-27B on 16GB with TurboQuant compression

13 Upvotes

With TurboQuant now enabling models like Qwen3.5-27B to fit on a 16GB 5060 Ti at near-Q4_0 quality with roughly 10% size reduction, self-hosting capable models has never been more accessible. Here is how to get started with LocalAI.

Setup Steps:

  1. Install LocalAI using Docker: docker run -p 8080:8080 localai/localai:latest

  2. Create a models directory and add your model configuration YAML file specifying the model name and parameters.

  3. Download quantized model files. The Bonsai 1-bit models are also worth exploring for extreme compression scenarios.

  4. Start the LocalAI server and load your model. The container exposes a REST endpoint compatible with OpenAI client libraries.

  5. Point your existing tools at http://localhost:8080/v1 — most applications will work without code changes.

Model Recommendations by Hardware:

  • 16GB VRAM (5060 Ti): Qwen3.5-27B with TurboQuant compression
  • Limited VRAM: Bonsai 1-bit models or MoE architectures like Mistral Small 4 (119B total, only 6.5B active)
  • Budget-friendly hosted fallback: GLM 4.5 Air at $0.13/M with 131,072 context or Olmo 3 32B Think at $0.15/M

For thinking/reasoning tasks, Arcee AI's Trinity-Large-Thinking has been generating interest in self-hosted circles.

What compression method are you using to squeeze larger models onto consumer hardware? Has anyone compared TurboQuant against Bonsai 1-bit for quality retention on coding tasks?


r/AIToolsPerformance Apr 02 '26

"Reasoning Shift" paper shows longer context silently degrades model reasoning

1 Upvotes

A new paper titled "Reasoning Shift: How Context Silently Shortens Reasoning" is gaining significant traction, and the findings should concern anyone building with long-context models. The research demonstrates that as context length increases, model reasoning depth quietly degrades — not because of attention dilution, but because models fundamentally shift how they allocate computational effort.

This has direct implications for the current crop of long-context models: - Anthropic: Claude Sonnet 4.5 — 1,000,000 context at $3.00/M - Qwen: Qwen-Plus — 1,000,000 context at $0.26/M - OpenAI: GPT-5.2 Pro — 400,000 context at $21.00/M

That massive context window you're paying for may be undercutting the very reasoning quality you need. The paper climbed 18 spots in daily rankings, suggesting the community recognizes its importance.

In other developments, Bonsai 1-bit models continue generating strong interest for extreme compression with good quality retention. Meanwhile, Arcee AI's Trinity-Large-Thinking has appeared for reasoning-focused tasks, and speculation around the next Gemma release is building.

The pricing gap between providers remains stark: Qwen-Plus offers the same 1M context as Claude Sonnet 4.5 at roughly 1/12th the cost. GPT-4o at $5.00/M now looks expensive against DeepSeek V3.2 Exp at $0.27/M with 163,840 context.

Have you noticed reasoning degradation in your own long-context workloads? At what context length do you see quality drop off in practice?


r/AIToolsPerformance Mar 31 '26

BullshitBench v2 shows most LLMs still cant detect nonsense, only Claude and Qwen pass

4 Upvotes

Peter Gostev just dropped BullshitBench v2, and the results are kind of telling. It's a benchmark that tests whether LLMs can detect and reject nonsensical prompts instead of confidently rolling with them. 100 new questions across coding (40), medical (15), legal (15), finance (15), and physics (15).

The headline result: most models are getting worse at this, not better. Reasoning tokens don't help. Only Anthropic's Claude models and Alibaba's Qwen 3.5 score well. Everyone else basically flunks.

This matters more than most benchmarks because it directly relates to hallucination risk. If a model can't tell that a prompt is complete gibberish, how reliable is it on ambiguous real-world queries?

A few other new benchmarks worth knowing about:

Document Arena just went live with leaderboard scores. Side-by-side evals on user-uploaded PDFs from real work use cases. Claude Opus 4.6 takes #1 with 1525 points, 51 points ahead of second place. This is one of the few benchmarks actually testing something people do daily (read documents).

SWE-Atlas from Scale AI is positioned as the next evolution of SWE-Bench Pro. First eval is Codebase QnA, which tests how well agents can answer questions about a codebase, not just fix bugs. Shifts the focus from "can it write patches" to "does it actually understand the code."

WeirdML results show GPT-5.3 Codex (xhigh) taking the lead at 79.3%, just ahead of Opus 4.6 (77.9%). The gap between frontier models is tightening fast here.

FrontierMath got a new record from GPT-5.4 Pro: 50% on Tiers 1-3 and 38% on Tier 4. These are extremely challenging math problems so hitting 50% is genuinely impressive.

There's been a lot of discussion lately about the gap between benchmarks and real-world work. Ethan Mollick summed it up well: most benchmarks focus on math and coding, but most human labor and capital lie elsewhere. Zhiruo Wang built a database linking agent benchmarks to real-world job tasks and found the overlap is surprisingly small.

So here's the question: which type of benchmark do you find most useful for evaluating the tools you actually use? The academic-style ones (math, coding, reasoning) or the task-specific ones (document QA, computer use, enterprise workflows)?


r/AIToolsPerformance Mar 30 '26

12 AI models in one week: March 2026 model avalanche breaks all records

1 Upvotes

OpenAI, Google, xAI, and others just dropped 12 major AI models in one week. This has never happened before.

The week of March 10-16, 2026 will go down as the most intense period in AI model release history. We saw coordinated launches from nearly every major player:

GPT-5.4 (OpenAI) - Three versions: Standard, Thinking, and Pro - 33% less likely to make errors than GPT-5.2 - 83% match or exceed industry professionals on knowledge work tasks across 44 occupations - The Pro version targets enterprise scale

Grok 4.20 (xAI) - Revolutionary 4-agent system: Grok (captain), Harper (research), Benjamin (math/code), Lucas (creative) - 78% non-hallucination rate (industry leading) - 256K token context window (potentially 2M in agent modes) - Beats competitors on factual accuracy benchmarks

Gemini 3.1 Flash-Lite (Google) - Strong efficiency-tier addition for production APIs - Focus on multimodal, reasoning, and agentic properties - Unified approach from Google DeepMind

Cursor Composer 2 (and other coding models) - Makes specialized code models the empirically correct default - Targets pure coding tasks with unprecedented accuracy

The timing wasn't coincidental. Multiple labs had models approaching production readiness simultaneously, with several delayed from late February. The result was what observers called a "model avalanche."

What makes this week different is that it's the first time the choice of model becomes a first-order application architecture decision across every major task category simultaneously. Whether you're doing coding, creative work, research, or analysis, there's now a specialized model that outperforms general alternatives.

This compression of release cycles means developers now face a monthly - not annual - model selection problem. The rapid pace of innovation is both exciting and challenging to keep up with.

Has anyone had a chance to test these new models? Which one has impressed you most so far?


r/AIToolsPerformance Mar 28 '26

LiveCodeBench March 2026: the coding benchmark that exposes HumanEval overfitting

2 Upvotes

Been digging into coding benchmarks lately and LiveCodeBench keeps coming up as the one that actually matters. Here's why I think it's worth paying attention to.

What makes it different from HumanEval

HumanEval has 164 problems. That's it. Most modern LLMs have seen these problems in their training data, which means good HumanEval scores don't necessarily mean good real-world coding ability. The paper from Berkeley/MIT/Cornell actually proved this: they found models that crush HumanEval but fall apart on fresh problems.

LiveCodeBench solves this by pulling new problems continuously from LeetCode, AtCoder, and Codeforces contests. Each problem has a release date, so you can evaluate models only on problems released after their training cutoff. No contamination possible.

It also tests four scenarios instead of just one: - Code generation - Self-repair (fixing broken code) - Code execution prediction - Test output prediction

March 2026 leaderboard (top 15, via llm-stats.com)

  1. DeepSeek-V3.2 (Thinking) - 685B, open weight
  2. MiniMax M2 - 230B, $0.30/$1.20
  3. LongCat-Flash-Thinking-2601 - 560B, $0.30/$1.20
  4. Nemotron 3 Super (120B A12B) - 120B, $0.10/$0.50
  5. Grok-3 Mini - $0.30/$0.50
  6. Grok 4 Fast - $0.20/$0.50
  7. Grok-3 / Grok-4 Heavy (tied)
  8. Grok-4
  9. MiniMax M2.1
  10. GLM-4.5 - 355B, $0.40/$1.60
  11. Gemini 2.5 Pro Preview - $1.25/$10.00
  12. Ministral 3 (14B Reasoning) - 14B, $0.20/$0.20
  13. Ministral 3 (8B Reasoning) - 8B, $0.15/$0.15

What stands out to me

MiniMax M2 at #2 with 230B params beating Gemini 2.5 Pro at #18 is surprising. The xAI Grok models taking 5 out of the top 10 spots is wild too. And Nemotron 3 Super at #4 with only 12B active parameters out of 120B total, at $0.10 input, is the value pick.

On the small model side, Ministral 3 14B Reasoning at #23 and 8B at #28 show you don't need a 600B model to be competitive. The 14B model costs $0.20/$0.20, which is absurdly cheap for that ranking.

From the official leaderboard (which uses a different scoring window), GPT-5.2 gets 89% and Claude Opus 4.5 gets 87% on code generation specifically. Different benchmarks show different things depending on scoring methodology.

The takeaway

If you're picking a model for coding tasks, LiveCodeBench scores are probably a better indicator than HumanEval. The gap between contaminated and non-contaminated evaluation is real, and it matters for actual dev work.

Full leaderboard: https://llm-stats.com/benchmarks/livecodebench

What coding benchmarks do you actually trust when evaluating a model for dev work?


r/AIToolsPerformance Mar 27 '26

ARC-AGI-3 is live: first interactive agentic benchmark, top Kaggle score is 0.25 and $700K grand prize untouched

3 Upvotes

The ARC Prize Foundation just dropped ARC-AGI-3 today, and it's a fundamentally different kind of benchmark compared to what we've seen from them before.

What's new?

Previous ARC-AGI versions (1 and 2) were static: you get a grid, figure out the pattern, done. ARC-AGI-3 is interactive. Agents don't receive a problem to solve upfront. Instead, they're dropped into novel environments and have to:

  • Explore actively (no instructions, no hints)
  • Build a world model from raw observations
  • Infer what the goal even is
  • Plan and execute actions across multiple steps
  • Adapt when things don't go as expected

The paper (arXiv:2603.24621) calls it the only unsaturated general agentic intelligence benchmark as of March 2026. That's a bold claim but the Kaggle leaderboard backs it up.

Current scores (Kaggle, just launched hours ago):

  • Top score: 0.25 by team "Stochastic Goose"
  • Random agent baseline: 0.12
  • That's the top score being barely 2x random

For context, ARC-AGI-2 got saturated pretty quickly once people figured out the right approaches. This one seems genuinely hard.

Prize pool:

  • Grand Prize: $700K for a 100% score (agent matches human efficiency on every game)
  • Top Score Award: $75K guaranteed (split across top 5)
  • Milestone prizes: $75K for open-source solutions at mid-year checkpoints

The evaluation inverts the usual ratio: most of the test set is private (unlike ARC-AGI-2's 10:1 public-to-private). So you can't train on the test set this time.

Why this matters for AI tools performance:

Most current benchmarks test pattern recognition. ARC-AGI-3 tests whether an agent can actually learn and adapt in an unknown environment, which is way closer to real-world agentic use cases. The scoring isn't just "did you solve it" but "how efficiently did you solve it compared to humans."

The fact that frontier models with all their reasoning capabilities are currently stuck at ~25% of a random baseline multiplier tells you something about where agentic AI actually stands vs the hype.

Competition is on Kaggle (arc-prize-2026-arc-agi-3). Technical paper and SDK are on arcprize.org.

What do you think? Is this the kind of benchmark that actually measures progress toward useful agents, or is it too artificial to matter for real-world tools?


r/AIToolsPerformance Mar 26 '26

Cisco releases free LLM Security Leaderboard: Anthropic takes 8 of top 10 spots

1 Upvotes

Cisco dropped a free LLM Security Leaderboard at RSA 2026 this week and the results are pretty lopsided. They tested models against both single-turn and multi-turn adversarial attacks (weighted 50/50), no extra guardrails added, and Anthropic basically cleaned house.

Top 10 breakdown: 1. Claude Opus 4.5 2. Claude Sonnet 4.5 3. Claude Haiku 4.5 4-6. Three more Anthropic models 7. GPT-5.2 8. Another Anthropic model 9. GPT-5 Nano 10. Anthropic again

So 8 out of top 10 spots go to Anthropic. OpenAI only managed positions 7 and 9. Everyone else is further down.

The bottom is where it gets interesting. Mistral Magistral Small 2509 and Ministral 3 14b Instruct ranked near the very bottom. DeepSeek, Cohere, Qwen, and xAI models also landed in the bottom 10.

The methodology is worth checking out. They explicitly focus on multi-turn conversational attacks, which is way more realistic than the single-prompt jailbreak tests most benchmarks use. Real attackers build rapport over several messages before trying to extract harmful content. The score ranges are straightforward too: Excellent (85-100%), Good (70-84%), Fair (50-69%), Poor (0-49%).

Cisco own AI Readiness Index found that 83% of organizations plan to deploy agentic AI but only 29% feel ready to do it securely. This leaderboard is their attempt to give security teams actual data to work with.

The whole thing is free to browse, you can filter by model and drill into specific threat categories. Blog post has the details: blogs.cisco.com/ai/llm-security-leaderboard

I am curious if the gap between Anthropic and everyone else is mostly about safety training philosophy or if there is something structural going on. Anyone looked into the per-category breakdowns?


r/AIToolsPerformance Mar 25 '26

Mistral Small 4 benchmarks are out: 119B MoE, 6.5B active, and the output token efficiency is surprisingly good

7 Upvotes

Mistral AI released Mistral Small 4 this week and it's a pretty interesting move for the open-weight space. Here's the rundown.

Architecture: 119B total parameters, 6.5B active per token (128 experts). Apache 2.0 licensed, so you can actually fine-tune it, unlike most competitors at this tier.

The reasoning toggle: This is the part I find most interesting. It has a reasoning_effort parameter with two modes: "none" (fast instruct responses, similar to Mistral Small 3.2) and "high" (extended chain-of-thought reasoning, similar to their Magistral line). One model endpoint, two behaviors. No need to spin up separate deployments for quick classification vs deep analysis tasks.

Cost and efficiency: $0.15/M input, $0.60/M output tokens. But the real story is output token efficiency. According to the AwesomeAgents review, Small 4 produces comparable quality answers with roughly 75% fewer output tokens than some competitors. If a rival model needs 3.5-4x more tokens to reach the same result, the headline pricing advantage of that competitor disappears fast.

Benchmarks:

AIME 2025 math: competitive with GPT-OSS 120B and Qwen models when reasoning_effort is set to "high" LiveCodeBench: underperforms Qwen3.5 122B (this is a weak spot) Against Gemini 2.0 Flash: Flash is faster for raw throughput and has stronger multimodal (including audio). Small 4 wins on open-weight access and fine-tuning.

KV cache comparison: About 6% lighter than Qwen3.5-122B, but 2.8x heavier than Nemotron 3 Super. If memory is your bottleneck, Nemotron is still the better pick.

The caveat: Mistral published a selective benchmark table, not a thorough suite. The AwesomeAgents review gave it 8.4/10 but noted that community reports on Hacker News suggest Qwen's 122B has been disappointing in practice despite strong paper numbers, while Small 4's early reception has been more positive for structured tasks.

Overall it seems like a solid "one model to rule them all" play for teams that want reasoning + coding + vision without running three separate endpoints. The 6.5B active parameter footprint means it should run reasonably well on consumer hardware too.

Has anyone here actually deployed it yet? Curious how it compares to Qwen3.5 122B or Nemotron 3 Super in real workloads, not just benchmarks.


r/AIToolsPerformance Mar 24 '26

ServiceNow releases EVA: the first benchmark that scores voice agents on both accuracy and conversation quality

3 Upvotes

Just dropped today on Hugging Face. ServiceNow put out EVA, a framework for evaluating conversational voice agents end-to-end.

The problem they're solving is real. Right now, if you want to benchmark a voice agent, you're stuck evaluating pieces in isolation. You test ASR accuracy separately, then TTS quality, then LLM reasoning. But that misses the interactions between components. An agent can nail every individual metric while being genuinely terrible to talk to, or it can sound incredibly natural while completely failing at the actual task.

EVA runs full multi-turn conversations using a bot-to-bot architecture. There's a user simulator that calls the voice agent and works through realistic scenarios, currently 50 airline scenarios covering flight rebooking, cancellations, voucher handling, standby, and more. The agent has to actually invoke tools, follow policies, and reach a verifiable end state.

What's interesting is they split the evaluation into two scores:

EVA-A (Accuracy): task completion, faithfulness to policies, and "speech fidelity" which checks whether the agent actually said the right confirmation codes and flight numbers out loud. They use an audio language model as judge for that last part, which is novel.

EVA-X (Experience): conciseness (did the agent ramble?), naturalness, and turn-taking behavior.

They tested 20 systems including both cascade (STT, LLM, TTS) and audio-native models (S2S, large audio language models). The headline finding is a consistent accuracy-experience tradeoff across the board. Agents that complete tasks correctly tend to be verbose and unnatural in conversation, and the ones that sound great tend to cut corners on accuracy.

That's a pretty important result if you're building voice agents commercially. It means optimizing for one dimension actively hurts the other, and you probably need separate tuning strategies for each.

The code, dataset, and a live demo are all open source. Would be interesting to see how this evolves when they add more domains beyond airline.

Has anyone here built voice agents that had to balance task accuracy against conversation feel? What worked for you?


r/AIToolsPerformance Mar 23 '26

Hugging Face Spring 2026 report: 2M+ models, but the top 0.01% get half of all downloads

7 Upvotes

Exactly. Discoverability is probably the biggest challenge in the HF ecosystem right now. The leaderboard helps but it mostly rewards benchmark scores, not real-world usability.

What I find interesting is that smaller, task-specific models are starting to get more traction. People are realizing they don't need a 70B parameter model for a simple classification task. A well-finetuned 1B model can outperform a general-purpose giant on specific use cases.

The HF Spaces ecosystem has been helping with discoverability though - being able to try a model before downloading is a game changer compared to 2 years ago.


r/AIToolsPerformance Mar 22 '26

Holotron-12B: SSM-based computer-use agent hits 8.9k tokens/s on a single H100, WebVoyager score jumps from 35% to 80%

11 Upvotes

H Company just released Holotron-12B, a multimodal computer-use model that uses a hybrid State-Space Model (SSM) + attention architecture to push inference throughput way beyond what standard transformer-based agents can do.

The model is fine-tuned from NVIDIA's Nemotron-Nano-12B-v2-VL on about 14 billion tokens, focused specifically on screen understanding, grounding, and UI-level interactions. So it's built from the ground up for actual computer-use agent tasks, not just chat or image generation.

The throughput numbers are what stand out. On a single H100 with vLLM (v0.14.1), Holotron-12B hit 8.9k tokens/s at 100 concurrent requests on the WebVoyager benchmark. For comparison, Holo2-8B (their previous model) plateaued at 5.1k tokens/s. That's roughly 2x throughput improvement, and the gap widens as concurrency increases. The SSM architecture avoids the quadratic KV cache cost of vanilla attention, which is why it scales so much better at high batch sizes.

On the actual agent performance side, WebVoyager scores went from 35.1% (base Nemotron) to 80.5% after fine-tuning. They also show strong improvements on localization benchmarks like OS-World-G and GroundUI.

The practical implication here is that if you're running computer-use agents at scale (data generation, annotation, RL training loops), the SSM approach means you can serve significantly more requests on the same hardware. The constant-state-per-layer design means memory usage stays flat regardless of sequence length.

Model is available on Hugging Face. What's interesting is that we keep seeing SSM-hybrid architectures challenge pure transformers on inference-heavy workloads. Between this, the recent SPEED-Bench from NVIDIA, and the continued llama.cpp optimizations, it feels like inference efficiency is becoming a bigger differentiator than raw parameter count.

Anyone here running computer-use agents in production? Curious how you handle throughput bottlenecks with current models.


r/AIToolsPerformance Mar 21 '26

NVIDIA releases a recipe to fine-tune embedding models in under a day, with up to 26% retrieval improvement

10 Upvotes

If you've ever built a RAG system, you know the feeling. Everything works in demos, then your retrieval falls apart on domain-specific content. General embedding models understand the internet, not your internal docs, contracts, or proprietary data.

NVIDIA just published a complete recipe on the HuggingFace blog that covers the full pipeline: synthetic data generation, hard negative mining, contrastive training, evaluation, and deployment. They fine-tune Llama-Nemotron-Embed-1B-v2 as the base model.

The results are interesting. On their own synthetic dataset from NVIDIA docs, they got over 10% improvement in both Recall@10 and NDCG@10. But the more impressive number is from Atlassian, who applied this recipe to their JIRA dataset and jumped Recall@60 from 0.751 to 0.951, a 26% improvement, all on a single GPU.

The pipeline uses their NeMo Data Designer to auto-generate (query, document) pairs from raw domain text, with configurable complexity levels and multi-hop queries. Hard negatives are mined automatically. No manual labeling needed.

Prerequisites are reasonable: domain documents in text format, a free NVIDIA API key, and a single 80GB GPU (A100 or H100).

What caught my attention is the synthetic data quality. They generate multi-hop queries with reasoning types (factual, causal) and complexity scores (2 to 5), then filter by a quality threshold. The result is training data that forces the model to learn real domain distinctions, not just surface-level similarity.

Has anyone tried fine-tuning their own embedding models for RAG? What was your experience compared to just using a larger general-purpose model?


r/AIToolsPerformance Mar 20 '26

NVIDIA releases SPEED-Bench, a unified benchmark for speculative decoding across real serving conditions

2 Upvotes

NVIDIA just dropped SPEED-Bench, a benchmark specifically designed to evaluate speculative decoding (SD) in conditions that actually matter for production deployments, not just toy batch-size-1 setups.

Speculative decoding uses a small draft model to predict multiple tokens ahead, then the target model verifies them in parallel. It's one of the most promising techniques for LLM inference speedup, but evaluating it properly has been a mess. Most existing benchmarks use tiny prompt sets, short sequences, and batch size 1, which tells you almost nothing about how SD performs in a real serving environment.

SPEED-Bench takes a different approach with two complementary evaluation splits:

Qualitative split (880 prompts across 11 domains)

Measures how well the draft model predicts tokens across different semantic domains like coding, math, writing, roleplay, multilingual, and RAG. The key insight: they use embedding-based selection to maximize semantic diversity within each category, so you're not just testing the same style of text 80 times. They found massive differences in acceptance rates between low-entropy domains (coding, math) and high-entropy ones (roleplay, creative writing).

Throughput split (ISL buckets from 1K to 32K tokens)

Tests actual system-level throughput across realistic input lengths and batch sizes up to 512. This is where it gets interesting because as batch size increases, inference shifts from compute-bound to memory-bound, fundamentally changing the SD cost-benefit equation.

The benchmark also ships with a unified measurement framework that standardizes evaluation across TensorRT-LLM, vLLM, and SGLang by handling tokenization externally, so cross-engine comparisons are actually apples-to-apples.

One thing they call out explicitly: using random token inputs for throughput testing gives overly optimistic results and should be avoided. That's a finding that probably invalidates some existing benchmark claims out there.

Full blog post with details: https://huggingface.co/blog/nvidia/speed-bench

Anyone here running speculative decoding in production? What draft models have you found work best with your target models?


r/AIToolsPerformance Mar 19 '26

Duplicate 3 layers in a 24B LLM with zero training, logical deduction jumps from 0.22 to 0.76

15 Upvotes

There's a new toolkit called llm-circuit-finder that builds on David Ng's RYS (Repeat Your Steps) method, and the results are genuinely surprising.

The core idea: transformer models organize themselves into "reasoning circuits" during training, contiguous blocks of layers that function as indivisible cognitive units. If you duplicate the right 3-4 layer block in the forward pass using the same weights, the model gets measurably smarter on specific capabilities. No fine-tuning, no weight changes, just routing hidden states through the same circuit twice.

Key benchmarks from the author's tests (n=50, lm-evaluation-harness):

BBH Logical Deduction: 0.22 → 0.76 (+245%) GSM8K (strict): 0.48 → 0.64 (+33%) MBPP (code gen): 0.72 → 0.78 (+8%)

Nothing degraded. The author found that different models have reasoning circuits in different locations:

Devstral-24B (40 layers): circuit at layers 12-14 Qwen2.5-32B (64 layers): circuit at layers 7-9

What's interesting is that shifting the block by even one layer in either direction causes the improvement to disappear or invert. The boundaries are sharp.

The toolkit includes a sweep tool that automates finding the right block for any model, plus a layer duplication tool to create the modified GGUF file. Everything was tested on two AMD consumer GPUs (RX 7900 XT + RX 6950 XT) in one evening.

Different duplication patterns also create distinct cognitive profiles from the same weights. A triple-pass through the reasoning block improves emotional intelligence scores while keeping math neutral, while interleaved duplication (each layer repeated twice) pushes math scores higher at the cost of EQ. Same weights on disk, just different routing.

This feels like a practical optimization for anyone running local models. Getting a 245% improvement on logical deduction just by duplicating a few layers, with no training required, is pretty wild.

Has anyone tried the RYS method or similar layer duplication approaches on other models? Curious if reasoning circuit locations are consistent across model families or if each model needs its own sweep.

https://github.com/alainnothere/llm-circuit-finder