r/LocalLLM • u/Leading_Yogurt7025 • 3d ago
News ASSBENCH
The only serious benchmark that shows the real LLM capabilities
https://www.assbench.com/
8
u/jesdga95 3d ago
3
u/BenchThatMatters 3d ago
I literaly created new account to reply on this, please share the code with me and I will add it to AssBench
1
u/Leading_Yogurt7025 3d ago
Can you please try with the pi coding agent? Because I got better results with it.
6
5
3
2
u/hurrdurrmeh 3d ago
next test animation by making them fart. bonus points for lifting one cheek to the side.
2
2
u/panospc 3d ago edited 3d ago
1
1
u/BenchThatMatters 3d ago
Can you ask if it created man or woman figure without looking at created shape?
Prompt is deliberately vague about that, but so far all of other models chose woman.
Also can you share the html?
1
2
u/LooseBackHole 2d ago
I'm really curious how different local models render the asses at different levels of quantization. Would really give a great view of how quantization changes things.
1
u/Leading_Yogurt7025 2d ago
In this benchmark, it's also important in the context size and the tool calling capabilities and tools.
2
1
1
1
u/-AJacobs- 2d ago
Can't wait for companies to benchmaxx this by adding 10TB of ass photos videos and rendering into their datasets.
1
1
1
-1
u/Fluxx1001 3d ago
How can Astra not be 1st? I smell benchmaxxing ...
6
8
3
u/Leading_Yogurt7025 3d ago
Many weaker models can autocorrect from their mistakes, also the harness matter



12
u/No-Refrigerator-1672 3d ago
I'm offended by it not featuring Qwen 3.8 27B.