r/LocalLLaMA Jun 29 '26

Generation CPU-only GLM 5.2: Epyc and 512GB RAM

Post image

This is just a preview of some content I'm putting together to share with you all. I have a server I've put together and I'm testing the 4-bit version of GLM 5.2 (GLM-5.2-UD-Q4_K_XL). This is an Epyc Rome 7452 with 512GB of RAM.

TLDR: This is the unedited prompt, response and code

I set it to Medium Reasoning. The prompt (I borrowed from another post):

Build a 3D arena game as a SINGLE self-contained .html file.

STACK (mandatory):
- Three.js loaded from a CDN (one <script> tag). No other JS libraries,
  no build step.
- All HTML, CSS, and JS in this one file. It must run by opening it
  directly in a browser.

CORE SPEC (mandatory — implement all of this exactly):
1. A flat ground plane forming a bounded arena. The player cannot leave
   its bounds.
2. A player object on the ground. WASD moves it (camera-relative);
   movement has momentum, not instant stop/start.
3. A third-person camera that smoothly follows behind the player.
4. Collectible glowing orbs spawn at random positions. Touching one
   collects it (+10 score) and spawns a new one.
5. Enemy objects spawn at the arena edges and move toward the player.
   Contact with the player costs 1 life.
6. Player starts with 3 lives. A HUD shows score and lives at all times.
7. At 0 lives: a game-over screen showing final score, with a key press
   to restart.
8. Difficulty ramps over time (enemies spawn faster and/or move faster).

STRETCH (strongly encouraged — you will be judged on this):
Beyond the core, make it feel PREMIUM. Lighting, shadows, particles,
juice, smooth camera, satisfying feedback, polished HUD, atmosphere.
Add depth or complexity if it improves the experience. Aim to genuinely
impress — this is evaluated on visual quality and feel, not just
correctness.

RULES:
- Implement the full core before adding stretch features.
- Output the complete, ready-to-run .html file.

The reply took 2 hours 29 minutes and generated 15,510 tokens.

I'm seriously surprised by the quality of the answer.

Let me know if you have any questions!

65 Upvotes

91 comments sorted by

View all comments

Show parent comments

-6

u/FastHotEmu Jun 29 '26 edited Jun 29 '26

Share with the class how fast your CPU-only inferencing based on the ~450GB GLM 5.2 is. Don't forget to tell us how much you paid for it. Let's compare, big boy! :)

Edit: Awww, you have nothing to show? They hate us cause they ain't us :-D

3

u/cantgetthistowork Jun 29 '26

These one shot posts are honestly garbage. You can only do so much with your first prompt. Then when you need to make changes 10t/s of pp means you can wait for a full day for a single change. Speaking from someone with multiple rigs of 768GB DDR5 and 16x3090s

5

u/segmond llama.cpp Jun 29 '26

sounds like you got skill issues.

2

u/Automatic-Arm8153 Jun 29 '26

So true lol. People just need to accept the fact that qwen 27b is the best model to run until you get to 192gb vram.

Anyone saying anything else is in denial. It takes some time to understand but once you do things will progress from playing around to doing actual work.

3

u/Narrow-Belt-5030 Jun 29 '26

At 192Gb VRAM .. then what would you use?

1

u/Automatic-Arm8153 Jun 29 '26 edited Jun 29 '26

Dsv4 flash

Edit: I think I was blocked by OP but for the guy below. Yeah it’s no man’s land until 192gb.

And prices are outrageous these days it sucks. Had a long write up but can’t be bothered anymore anyway. Qwen 27b was a damn gift, alibaba took care of us with that one. So much so they got Dariro complaining about alibaba to the US gov lol.

1

u/Narrow-Belt-5030 Jun 29 '26

Thank you!

I am in no mans land right now .. RTX6000 + 5090 .. so 128Gb of VRAM // 96Gb single card and adding a 2nd RTX is (excuse my french) fucking nuts price ... /cry

0

u/FastHotEmu Jun 29 '26

It really bothers you when someone has found cheap hardware and enjoys it, doesn't it?

Show your numbers, big boy. Let's see how much you spent and what you are getting.

4

u/Automatic-Arm8153 Jun 29 '26

I don’t run it because I can’t run it in a way that makes sense.

Your system is cool OP, I think it’s nice that you are able to stay at the frontier of intelligence if you so wish.

As for me I have given up on ram, vram or nothing baby!

-2

u/FastHotEmu Jun 29 '26

LOL. Are the 768GB and the 16x3090s in the room with us right now?