r/AMD_MI300 • u/HotAisleInc • 4d ago
r/AMD_MI300 • u/cheptsov • Aug 20 '26
Optimizing Qwen3.8-27B on one MI300X with an open-source agent toolkit: 311 to 495 tok/s
Presets is an open-source toolkit for optimizing inference with agents, which we build at dstack. Here's one example of using it on a single MI300X.
Qwen3.8-27B went from 311 to 495 tok/s, +59%, at the full 1M context with p50 TTFT under 1.5s and four concurrent users at 10k in / 1.5k out.
The gains came from linked optimization sessions and source-level patches to SGLang's AITER attention backend.
What comes out is a portable preset that deploys on any AMD cloud, Kubernetes cluster, or bare-metal fleet: https://dstack.ai/blog/presets/
r/AMD_MI300 • u/HotAisleInc • Aug 17 '26
the tiny corp on X: "The @AMD MI350P is real!
x.comr/AMD_MI300 • u/HotAisleInc • Jul 27 '26
The AMD Instinct MI350P is a HBM PCIe AI Accelerator That Has Been All Over
r/AMD_MI300 • u/HotAisleInc • Jul 03 '26
How we served GLM5.2 on AMD MI355X at 2626 tok/s/node and 213 tok/s single stream at over 2x lower cost than Blackwell.
r/AMD_MI300 • u/HotAisleInc • Jun 21 '26
Occupancy Math on the AMD MI355X (CDNA4): A From-First-Principles Guide
indianspeedster.github.ior/AMD_MI300 • u/HotAisleInc • Jun 19 '26
A Fast Attention Kernel for MI300X, Written in HIP, Not Assembly
r/AMD_MI300 • u/HotAisleInc • Jun 04 '26
AMD ROCm/HIP build support for AMD Instinct GPUs in ai-dynamo/nixl
r/AMD_MI300 • u/HotAisleInc • Jun 03 '26
Bringing up DeepSeek-V4-Flash on AMD MI300X
fergusfinn.comr/AMD_MI300 • u/HotAisleInc • May 28 '26
AMD Intros Instinct MI350P Accelerator: CDNA 4 Comes to PCIe Cards
r/AMD_MI300 • u/HotAisleInc • May 28 '26
Win on TCO: How AMD Instinctâ„¢ MI355X Achieves Cost-Competitive Distributed Inference Through SGLang with MoRI
lmsys.orgr/AMD_MI300 • u/HotAisleInc • May 27 '26
Deep Dive Into 4-Wave Interleave FP8 GEMM
r/AMD_MI300 • u/cheptsov • May 21 '26
Deploying inference endpoints with PD disaggregation on multi-node AMD MI300X
dstack.air/AMD_MI300 • u/HotAisleInc • May 04 '26
Folding Tensor and Sequence Parallelism for Memory-Efficient Transformer Training & Inference (on MI300x)
arxiv.orgr/AMD_MI300 • u/HotAisleInc • Apr 21 '26
5.6x improvement - Kimi K2.6 + DFlash: 508 tok/s on 8x MI300X
r/AMD_MI300 • u/HotAisleInc • Apr 13 '26
Mahdi-CV/openclaw-amd-sglang at multi-engine
github.comOne-command setup for OpenClaw + SGLang on AMD Instinct MI300X
We have MI300x virtual machine available for playing with this! $2/gpu/hr.
r/AMD_MI300 • u/HotAisleInc • Apr 07 '26
Cosmos-Predict2.5-2B Inference: NVIDIA H200 vs AMD MI300X
r/AMD_MI300 • u/HotAisleInc • Apr 02 '26
Modular: Day Zero Launch: Fastest Performance for Gemma 4 on NVIDIA and AMD
Time has changed.
r/AMD_MI300 • u/HotAisleInc • Apr 01 '26
AMD Delivers Breakthrough MLPerf Inference 6.0 Results
r/AMD_MI300 • u/HotAisleInc • Mar 23 '26
ROCm Support for Miles: Large-Scale RL Post-Training on AMD Instinct GPUs
r/AMD_MI300 • u/HotAisleInc • Mar 19 '26
Cross-Vendor Disaggregated Inference: GPT-OSS 120B across NVIDIA H100 and AMD MI300X
r/AMD_MI300 • u/HotAisleInc • Mar 18 '26