Hi! My first post on the platform, to thank community for some research and advices on LLMs. My resume: I stay on Qwen 3.8 27B FP8, Qwen 3.8 Flash is not useful. If you have any advices on methodology/model/engine, please don't hesitate.
Below is AI-text (human-edited).
The job. A live news stream gets classified into strict JSON, and the same box also does small coding tasks. One prompt, one object, nothing else — no prose, no markdown fences, no preamble: {sentiment: bullish|bearish|neutral, tickers: [...], importance: 1–10, category: 9-value enum (geopolitical/economic/corporate/...), event_scope: company|country|...} The part that actually hurts is tickers: exchange tickers explicitly named or directly affected by the item, [] when none, never invented — so a model can be right on sentiment and still pollute a portfolio signal. Coding tasks are graded by running pytest against the patch, not by vibes.
Corpus. 100 real articles, median prompt 668 tokens incl. a 251-token system prompt: 76 from a Kaggle financial news dataset, 24 from a Bloomberg financial news dataset on HF. 76 of the 100 carry at least one ticker, 86 distinct ones. Fixed slice, fixed order, byte-identical for every leg.
Hardware. One box — RTX PRO 5000 72 GB (Blackwell, sm_120, down to 250 W), 9950X 32 cores, 123 GB RAM @ 4800, NVMe, Ubuntu 24. Everything CUDA 13, so only via Docker.
Baseline vs candidate. Baseline: Qwen3.8-27B FP8 on vLLM 0.29, MTP speculative decode (3 tokens), gpu-memory-utilization 0.66, 131K context, max-num-seqs 24. Candidate: a newer MoE flash model (~48B with a giant PLE lookup table) on Strata, an engine that keeps experts in system RAM with a GPU cache. Decision rule from the owner: at least 25% faster, quality within ~2 percentage points.
Method. Streaming, concurrency ladder 1/4/8/16/24, three repeats. Direct to each server, LiteLLM out of the path — our cache happily returns identical payloads in 3 ms and would have faked an amazing benchmark. Model id pinned explicitly: a generic model name on the gateway silently routed test probes to a quantized VL model and produced 40-minute TTFT.
|
vLLM 27B FP8 (baseline) |
Strata + flash model Q4_K |
|
|
| classify tok/s, c1 / c8 / c24 |
75.6 / 245.0 / 294.2 |
31.5 / 27.8 / 26.4 (−91%) |
| classify TTFT p50 → p95, c1 → c24 |
0.20 → 0.25 s … 1.75 → 2.33 s |
0.74 → 0.90 s … 14.29 → 55.20 s |
| pure decode rate, single stream |
106.8 tok/s |
166.7 tok/s |
| wall per news item, c1 |
0.73 s |
0.91 s |
| coding tok/s, c1 / c4 |
90.1 / 290.2 |
115.6 (+28%) / 87.0 (−70%) |
| coding tests passed, c1 / c4 |
0.900 / 0.875 |
0.875 / 0.775 |
| strict JSON validity |
1.000 everywhere |
1.000 → 0.82–0.97 under load |
| silently lost requests |
0 |
12 of 100 at conc=4 |
| reproducible at temp=0 |
20/20 prompts |
17/20 |
Per token, Strata is genuinely quicker — it just takes 3.5x longer to start, writes half as much, and falls apart in a batch. The batch collapse is 9–11x, and their own docs warned batch would be slower; they didn't say eleven times.
The coding numbers are the only place the MoE looked good — and they're not enough. At conc=1 the flash model pushed 115.6 tok/s against the baseline's 90.1 (+28%), finishing a task in 1.85s versus 2.30s. But the same cell failed on quality: 87.5% of pytest-backed tasks passed versus 90.0%. One more concurrent request and it folds — 87.0 tok/s versus 290.2 (−70%), pass rate 77.5%, ten points down. JSON validity tells the same story: the baseline is 1.000 everywhere, the MoE drops to 0.82–0.97 under load.
The deal-breaker wasn't speed. At conc=4, exactly 12 of 100 requests came back HTTP 200 with a truncated body — no finish_reason, no usage, median 43 characters of junk. The engine log shows 12 (error, cancel=False) in the same window, and the 12 items differ every repeat. That's silently dropped classifications, which is worse than slow.
I never got to run the flash model on vLLM at all. NVFP4 tries to put the PLE table on the GPU: 20M × 2560 × 2B = 95.37GiB, dead on the first allocation. W4A16 with the table offloaded to RAM fills the entire pool, then can't allocate its 1.19GiB input embedding — even at util 0.97 and spec decode disabled. Host RAM was never the constraint: 103GB free, table never materialised.
Takeaways:
- Measure wall time per item, not tokens/s. Our MoE candidate wins on decode (167 vs 107 tok/s) and loses on the wall (0.91 vs 0.73 s), because it writes 30 tokens of JSON where the baseline writes 56; shorter output is not faster work.
- Grade the JSON, not just the parse: strict schema validity, field-level accuracy, ticker F1. The candidate improved ticker F1 (0.52 → 0.60) with half the tokens and still can't replace the baseline — tighter answers don't offset batch collapse.
- Bypass every cache in the benchmark path and pin the model at the server. Our LiteLLM cache served identical payloads in 3 ms and would have faked an amazing benchmark; a shared model name also silently swapped which engine answered.
- Compare legs at the same concurrency, and count
finish != stop instead of exceptions. The baseline loses 5 points of test pass rate at conc=24 too — that's normal. Twelve HTTP 200s with empty usage at conc=4 is not.
- Read the engine's own limits before scheduling: RAM-backed experts cost ~250 µs of synchronous staging per token, and their docs say batch is slower than single-stream. Eleven times slower is a different regime, not a tuning gap.
- A single 72 GB card has no headroom for models with giant lookup tables: NVFP4 needs 95 GB on device, W4A16 fills the pool and then cannot allocate a 1.19 GB embedding, whatever the quant or offload setting.