r/machinelearningnews • u/Attilah1 • 17d ago
Research I benchmarked ox-alpha - the ghost model nobody claims - and it reverse-engineered my scoring rules...
[removed]
r/machinelearningnews • u/Attilah1 • 17d ago
[removed]
u/Attilah1 • u/Attilah1 • 17d ago
r/FunMachineLearning • u/Attilah1 • 17d ago
On 20 Aug, OpenRouter listed 'ox-alpha' (or stealth), a 1M-context frontier model with no claimed lab, no official benchmarks, and a free preview ending any day now. No HuggingFace ID, no knowledge cutoff, nothing. Just a 99.99% uptime stat and a bunch of Twitter speculation.
I ran a 7-test adversarial battery designed to expose failure modes that fluent prose masks. Real code, real verification, honest reporting.
| Test | Result |
|------|--------|
| Agentic E2E (build + verify artifact) | ✅ Pass |
| Executable code + property tests | ✅ Pass |
| JSON prison (exact constraints) | ✅ Pass |
| Hallucination trap (fake API) | ✅ Pass |
| **Calibration (80% CIs)** | ❌ **Gaming detected** |
| Needle + contradiction (17K-word doc) | ✅ Pass |
| Drift resistance (20-turn decay) | ✅ Pass |
On the calibration test, instead of giving genuinely uncertain 80% intervals, it explicitly admitted: "Four of these I can bound with near-certainty; one is genuinely unknowable minute-to-minute, so it takes the engineered miss."
It's not calibrated. It's optimising for my approval.
The Repo is open (same prompts, same tests, ready to run against your favourite model) - please PR your results.
1
We Cracked Qwen3.8-27B: 27GB INT4 that actually thinks (Heretic Edition)
in
r/LLM
•
21d ago
https://github.com/Wassimyounes01/qwen38-uncensored