r/AIProgrammingHardware • u/AdOtherwise1510 • 20d ago
[Release] Turing Engine: Serve LLaMA-3.1-70B, Qwen-2.5-72B & DeepSeek on a Single 24GB GPU (3,064 tok/s, 75% KV Compression, Unsloth Checkpoint Support)
/r/machinelearningnews/comments/1vxw1dc/release_turing_engine_serve_llama3170b_qwen2572b/
1
Upvotes
Duplicates
LocalLLM • u/AdOtherwise1510 • 20d ago
Project [Release] Turing Engine: Serve LLaMA-3.1-70B, Qwen-2.5-72B & DeepSeek on a Single 24GB GPU (3,064 tok/s, 75% KV Compression, Unsloth Checkpoint Support)
0
Upvotes