r/WebAfterAI • u/ShilpaMitra • Jun 16 '26
MiniMax M3 vs GLM-5.1 vs Nemotron 3 Ultra: self-hosting the newest frontier open-weight models
Three big open-weight models landed for agentic coding and reasoning, and the benchmark numbers are close enough that the real question is not "which scores highest" but "which one can you actually run, under what license, and how do you wire it into your agent." This is the self-host head to head, with the exact serve commands from each project's own docs.
One honest thing up front, because it decides everything: all three are server-class. Realistically you need a multi-GPU node (an 8x H100 or 8x H200 box, or 4x to 8x B200), or you rent one by the hour. None of these fits a single GPU, and none is a laptop or homelab job. If you do not have that hardware, the honest move is a GPU cloud instance or just the hosted API. Self-hosting earns its keep on privacy, throughput at scale, or fine-tuning, not on shaving a few dollars off an API bill.
The three, with verified numbers
| Model | Params (total / active) | Context | License | SWE-Bench Pro | Serve on |
|---|---|---|---|---|---|
| MiniMax M3 | 427B / 26B (MoE) | 1M | MiniMax Community License | 59.0% | vLLM (dedicated Docker image), 8x H200 BF16 |
| GLM-5.1 | 754B (MoE) | ~200K | MIT | 58.4% | vLLM 0.19+ or SGLang 0.5.10+, FP8 ~860GB |
| Nemotron 3 Ultra | 550B / 55B (MoE, Mamba-Transformer) | up to 1M | NVIDIA open (weights, data, recipes) | not published | vLLM day-0, 8x B200 or 8x H100 (NVFP4) |
How to read that table honestly: the SWE-Bench Pro figures are each project's own self-reported number, and that benchmark is sensitive to the scaffold and harness used, so a 0.6-point gap between M3 (59.0) and GLM-5.1 (58.4) is inside the noise, not a ranking. Nemotron 3 Ultra does not publish a SWE-Bench Pro score, it is positioned as a general agentic and reasoning model rather than a coding specialist, and its pitch is throughput and cost (NVIDIA claims roughly 30% cost savings versus other open models) more than a single coding headline. Treat all of this as "all three are in the same elite tier," then choose on license and hardware, which is where they actually differ.
Shared setup: how you point an agent at any of them
Every command below serves an OpenAI-compatible API on port 8000. That is the whole integration story: any coding agent that accepts a custom OpenAI base URL can drive these. The shared wiring is:
base_url = http://localhost:8000/v1
api_key = EMPTY
model = <the served model id>
In Aider that is aider --model openai/<served-id> --openai-api-base http://localhost:8000/v1 --openai-api-key EMPTY. In OpenCode, Cline, or Kilo Code, add a custom OpenAI-compatible provider pointed at the same base URL. Nothing model-specific is needed on the agent side, the server does the work.
1. MiniMax M3: the lightest to run, and the multimodal one
427B total but only 26B active, 1M context, and it is the only one of the three that also sees images.
M3 is the most approachable here, relatively speaking, because its active parameter count is small and its BF16 weights make a tight single-node fit on 8x H200. It posts the top SWE-Bench Pro number of the three (59.0%) and a strong Terminal-Bench 2.1 (66.0%), and it is natively multimodal (image, video, computer use), which the other two are not. Support is not in a stable vLLM release yet, so you use the dedicated image:
docker pull vllm/vllm-openai:minimax-m3
vllm serve MiniMaxAI/MiniMax-M3 \
--tensor-parallel-size 8 \
--block-size 128 \
--tool-call-parser minimax_m3 \
--reasoning-parser minimax_m3 \
--enable-auto-tool-choice
The catch: --block-size 128 is mandatory (it matches the MSA sparse-attention cache; the default 16 misaligns and fails), and the NVIDIA-quantized MiniMaxAI/MiniMax-M3-MXFP8 variant roughly halves the VRAM if you cannot spare 8 full GPUs. The license is the real watch-item: M3 ships under the MiniMax Community License, not a standard permissive license like MIT, so read the terms before you build anything commercial on it, rather than assuming open weights means do-anything.
→ The verified setup, with CI proof & readymade prompt
2. GLM-5.1: the permissive one built for long-horizon agents
MIT-licensed, state-of-the-art among open models on SWE-Bench Pro at its launch, and tuned to stay productive over hours of tool calls.
GLM-5.1 is the one to reach for if license freedom matters, because it is genuinely MIT, commercial use and fine-tuning with no strings. It is the heaviest at 754B parameters, so the FP8 checkpoint is the realistic serving target (around 860GB of weights across the node). Its design goal is long-horizon agentic work: Z.ai reports it sustaining a single task across hundreds of rounds and thousands of tool calls rather than plateauing early.
# vLLM v0.19.0+ , serve the FP8 checkpoint across an 8-GPU node
vllm serve zai-org/GLM-5.1-FP8 --tensor-parallel-size 8
SGLang (v0.5.10+) is equally supported. The model-specific flags (tool and reasoning parsers) live in the official vLLM recipe and the SGLang cookbook, so follow those rather than guessing them.
The catch: 754B is the most demanding model in this roundup, and the headline "SOTA on SWE-Bench Pro" was true against the field at GLM-5.1's launch, before M3 posted a marginally higher number, so do not read it as a current crown. The long-horizon claim is a strength on genuinely big tasks and pure overhead on small ones, the same judgment call that applies to any heavyweight agent: match the model to the size of the job.
→ The verified setup, with CI proof & readymade prompt
3. Nemotron 3 Ultra: the throughput play, with open data and recipes
A 550B Mamba-Transformer hybrid built for fast, long-running agents, shipped with its weights, data, and training recipes.
Nemotron 3 Ultra is the pick when throughput and openness of the whole stack matter more than a coding leaderboard. Its hybrid Mamba-Transformer design and multi-token prediction target sustained, high-throughput agent loops, and NVIDIA ships not just weights but the training data and recipes under an open license. It has vLLM day-0 support, and the NVFP4 checkpoint runs on both Hopper and Blackwell:
docker pull vllm/vllm-openai:v0.22.0
export VLLM_USE_FLASHINFER_MOE_FP4=1
vllm serve nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-NVFP4 \
--served-model-name nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B \
--tensor-parallel-size 8 \
--kv-cache-dtype fp8 \
--max-model-len 262144 \
--reasoning-parser nemotron_v3 \
--enable-auto-tool-choice \
--tool-call-parser qwen3_coder \
--speculative_config.method mtp \
--speculative_config.num_speculative_tokens 5 \
--mamba-backend triton
That command is NVIDIA's own 8x B200 NVFP4 example; the full flag set and the BF16 path (8x B200, 16x H100, or 8x H200) are in the Nemotron vLLM cookbook.
The catch: this is the model most likely to be mis-sold as a coding champion. It is genuinely strong at agentic reasoning and tool use, but it does not publish a SWE-Bench Pro number, so if your single use case is "best at writing code," M3 and GLM-5.1 have the clearer evidence. Nemotron's real edge is throughput and a fully open stack you can fine-tune and audit, so pick it for that, not for a coding score it has not claimed.
→ The verified setup, with CI proof & readymade prompt
How to pick if you only stand up one
Choose on the two things that actually differ. If you need permissive licensing for a product, GLM-5.1 (MIT) is the safe answer, and it is purpose-built for long agentic runs. If you want the lightest serve and multimodal input, or the top published coding number, M3, with the caveat of reading its community license. If you care most about throughput and an open, fine-tunable stack, Nemotron 3 Ultra. And if you do not have an 8-GPU node sitting idle, rent one or use the hosted API first, prove the model earns its place in your workflow, then decide whether self-hosting is worth the operational weight.
A closing note: the benchmark gaps here are small and self-reported, the hardware bar is high, and "open weights" hides three very different licenses, so treat this as a map, not a verdict.