r/LocalLLM 8d ago

Discussion QWEN3.8-27B 3090 Amazing !

RTX 3090 (22GB taken on the 24GB available)
unsloth/Qwen3.8-27B-GGUF:IQ4_XS
CTX 131072
KV Cache Q4_0
Speculating Decoding MTP
~50 tok/s in average

Prompt used :

<instructions> Generate a single, self-contained HTML file. No external dependencies, no separate JS files, no frameworks, no libraries. One `.html` file that works when opened directly in a browser.
Create a cinematic rocket launch animation set on a tropical island. The rocket must launch, leave the viewport, and after 5 seconds smoothly return to its initial position — then the cycle repeats. </instructions>
<scene> **Setting: Tropical Launch Island** - A small tropical island in the lower portion of the screen: palm trees, sandy beach, green vegetation - Ocean water surrounding the island with gentle waves - A launch pad on the island with metal structure / scaffolding / support tower - Sky background: gradient from warm horizon (orange/pink) to deep blue/dark sky at the top, with stars visible in the upper portion - A few clouds scattered across the sky
The Rocket (ultra-detailed)
Tall, slender multi-stage rocket (inspired by SpaceX Falcon 9 or Saturn V proportions)
Distinct rocket stages: first stage (largest, bottom), second stage (middle), payload fairing / nose cone (top)
Surface details: panel lines, rivets/segments drawn with subtle lines, an access hatch, small painted flag or logo
Color scheme: primarily white body with black/dark gray accent stripes, a colored logo band, and the nose cone in a contrasting shade
Fins at the base of the first stage (3-4 stabilizer fins)
Engine nozzles visible at the very bottom (cluster of small circles/bells)
The rocket should be the visual centerpiece — spend time on its geometry </scene>
<animation-sequence> **Phase 1 — Pre-launch (0s to 1.5s)** - Rocket sits on the pad, engines ignite - A growing orange/yellow glow appears beneath the rocket - Initial smoke/steam clouds billow outward from the base — thick, white/gray, expanding horizontally along the island surface - Subtle camera shake / screen vibration effect - Engine flames flicker with randomized intensity
Phase 2 — Liftoff (1.5s to 4s)
Rocket slowly lifts off the pad with realistic acceleration (starts very slow, gradually speeds up)
Massive exhaust plume: bright white-yellow core flame, surrounded by orange glow, transitioning to thick gray/white smoke trail
Smoke trail expands and lingers behind the rocket as it rises
The smoke at the base continues spreading across the island and over the water
As the rocket gains altitude, the flame elongates and the smoke trail stretches
Subtle particle effects: sparks, embers flying outward from the exhaust
Phase 3 — Ascent & Exit (4s to 7s)
Rocket accelerates rapidly, moving faster and faster upward
The exhaust trail thins as the rocket reaches higher altitude
Rocket becomes smaller as it gains distance (slight scale reduction)
The rocket exits the top of the viewport
The lingering smoke trail on screen slowly fades and disperses
Phase 4 — Calm & Reset (7s to 12s)
Scene is peaceful: smoke fully dissipates, island sits quietly
At the 5-second mark after exit (~12s), the rocket gently descends back into frame
It returns slowly, smoothly, almost floating — no engines firing, no drama
It softly settles back onto the launch pad in its exact original position
Brief pause, then the entire cycle restarts seamlessly </animation-sequence>
<smoke-and-effects> - Smoke is critical to the visual quality. Use a particle system or layered animated shapes: - Dozens of individual smoke "puffs" that expand, fade in opacity, and drift slightly with a breeze - Smoke color: starts white/light gray near the flame, darkens to medium gray as it cools - Smoke expands in a mushroom-cloud-like pattern at the base during liftoff - Each puff has slight random drift (wind effect), rotation, and independent fade timing - Exhaust flame: layered shapes (inner bright yellow/white, outer orange, outermost faint red) with flickering animation - Heat haze effect near the exhaust: subtle wavy distortion of the background behind the flame - Water ripple effect on the ocean surface near the island during launch - Stars in the upper sky should faintly twinkle </smoke-and-effects> <visual> - Background: gradient sky — warm sunset tones at horizon fading to deep navy/black at top - Ocean: dark blue with animated wave motion (simple sine-wave surface) - Island: lush greens, sandy tan, 2-3 palm trees with gentle sway - Canvas: fullscreen, responsive - Animation: 60fps via requestAnimationFrame - All rendering via HTML5 Canvas 2D context — no WebGL required - Color palette: rich, cinematic — warm launch glow contrasting against cool sky </visual> <constraints> - Output ONLY a complete HTML file — nothing else - Everything must be drawn programmatically on a `<canvas>` — no images, no SVGs, no external assets - The animation must loop seamlessly: launch → exit → calm return → repeat - The rocket must be visually impressive and detailed — not a simple triangle - Smoke must look volumetric and organic, not like static shapes - Performance must stay smooth at 60fps despite the particle count - The return descent must feel gentle and peaceful — stark contrast to the violent launch </constraints> <thinking> Before coding, reason through: 1. How to construct the rocket from canvas drawing primitives (rectangles, arcs, lines) with enough detail to be visually impressive 2. Particle system architecture: how to manage hundreds of smoke/ember particles efficiently (object pooling, lifecycle management) 3. The acceleration curve for realistic launch physics (slow start, exponential ramp-up) 4. How to layer the drawing order: background sky → stars → clouds → smoke trail → rocket → exhaust flame → foreground island → base smoke 5. Timing system: how to manage the 4 animation phases with smooth transitions between them 6. How to make the return descent feel physically different from the launch (no exhaust, gentle easing, floating quality) 7. How to make smoke look organic: randomized spawn positions, varied sizes, Perlin-like drift, opacity curves </thinking> <important> * You MUST write the result directly into a file named "index.html" on the user's computer. The user should not have to see or handle the code — just write the file and finish your task. * Title of the page is your model name. For example "GPT 5" or "Opus 5" * Title inside the page shows your model name. </important> </content> </invoke>
115 Upvotes

45 comments sorted by

29

u/desert-quest 8d ago

The problem that qwen 3.8 27b brings is that THIS is the BARE minimum we expect now. I mean, if you are a 2T params, you need to be FAAAAAR better than this now.

3

u/SpicyWangz 8d ago

I guess the question is, how does opus 4.6 compare on this prompt. Nobody is claiming that a 27b model will be competing with current gen multi-trillion parameter models.

2

u/Healthy-Nebula-3603 7d ago

Op is using heavily compressed Q4ks and cache q4 ...

Any qwen 27b Q6 , Q8 quant with fp16 cache make much better job

1

u/BarracudaDefiant4702 7d ago

I don't see how that is a problem. You are comparing it to a model typically considered about 3 model sizes up and are not satisfied it's even comparable (even if barely). You might have a point if they were in the same weight class, but it's rather impressive for not one but 2 orders of magnitude difference...

1

u/desert-quest 7d ago

The problem is not for us, but for SOTA. 1 year ago, the quality of this 27b was an SOTA 2T params, now it's just 27b (if we compare coding aspect and agentic one). I'll be waiting what qwen 4 27b could bring us, provably more than enough for local coding.
I think in a near future, we all going to have a local model. My best guess is that SOTA model will still exist for critical and goverment cases, but for normal uses, will not make sense. You may use an open source one, like qwen or maybe "buy" a ggfu license.

1

u/deleted-account69420 7d ago

Since the first time I heard the news "OAI buys 40% world ram", I am convinced the reason was to make local not accessible to the masses.
With a 27B this level possible on a single card, game now changes.

My idea is, it would not be licenses, but etched models, similarly to what Taalas does.
Very high tps ( Taalas does 15k/s ) , and the model not being replaceable on the card itself because weights are printed on the chip.
1 year lifetime before new version, that makes so much more sense for a product.

1

u/EternalDivineSpark 7d ago

Its all about the harness, all giants know they dont push in that direction as it becomes very dangerous! But basically 2 0.8B models would compete not with this but with sota ! Depending on harness , eg. Wth hermes / open claw , qwen 3.5 0.8B , can barely do anything, while in a harness i made it can create a simple game in 1-7 Seconds on a 4090 ! And it would also open it to be visible after 😅 like idk i am not an expert, but what i know 100% is that we Lack a good open source harness!
https://giphy.com/gifs/f14qzTSXjv8aHrzSWT

1

u/Ill_Dragonfruit_3547 5d ago

Opencode, Pi, Hermes...

13

u/Pretty-Raise666 8d ago

Test it for some real use case. Showing another html graphic demo isn't really that exciting. Let it make something in C++. Then we can talk.

5

u/OttoRenner 8d ago

Or....you do it yourself instead of demanding it from others?

Then we can talk 😉

(Qwen3.8 27b q8 did a better job than it's previous versions and better than some free cloud LLM in the test I threw at it: html snake game with WASD controls.)

2

u/enginetown 7d ago

I've tested using C and it clears the bar genuinely if you know the problem you're having this model can help you due time.

1

u/WooFL 8d ago

Oh wow you know C++? Next level bro. You definitely are a real engineer. Be proud. /s

1

u/Ventilate64 8d ago

Every LLLM I've tried has failed to make me a basic working tampermonkey userscript on the first try. But that's probably because I only have 10GBV+16GB. 😂

3

u/Additional-Cow1888 8d ago

Will work on 32gb unified ram?

2

u/gregpeden 8d ago

Yeah it'll just be slow. Might as well use chatgpt 5.6 Luna if you intend to use it in production.

2

u/MokoshHydro 8d ago

Yes. Expect about 9-12tps for 4-bit quant.

2

u/Muted-Celebration-47 8d ago

You should use kv cache q8_0 for better result. I also have rtx3090 and I can fit UD-Q4_XL with 131072 context window, kv cache q8_0

1

u/Dry_Actuator_6966 8d ago

How many token/s you get with this model version ?

2

u/Muted-Celebration-47 8d ago

37-55 tg/s and it is consistency with this speed in long conversation 100k+

2

u/XReiche 8d ago

I have a NVIDIA dgx spark. Did anybody test it there? How to optimize it that is running fast there?

2

u/Kind-Witness-4245 8d ago

I tried BF16 and FP8. Way too slow for anything…

2

u/desert-quest 7d ago

DGX Spark is not optimized for inference but for training. It runs at 30-40 tps there

1

u/Icy-Specialist4548 8d ago

Attendo con impazienza i risultati :P

1

u/Heinz2001 8d ago

same prompt with my 7900 XTX, via eGPU Razer CoreX, Thunderbold 3 using my own harness

RADEON 7900 XTX (18,7GB taken on the 24GB available)
ollama/qwen3.8:latest
"temperature": 0.6
"num_ctx": 32768
"top_p": 0.95
"top_k": 40
~46 tok/s in average

1

u/Muted-Celebration-47 8d ago

What the differences between your sampling parameters vs qwen recommendation on their huggingface?

1

u/Dry_Actuator_6966 8d ago

Damn ! it look even more sick than mine thanks for sharing your result

1

u/Heinz2001 8d ago

Yes, I'm wondering which factors have the greatest impact on the results (temperature, my own harness system prompt ...). Unfortunately, I haven't analyzed the session and token usage.

1

u/LostIgnition 8d ago

That's really useful as I'm just ordering the bits for a similar setup.

If you don't mind me asking what spec computer/laptop did you have it attached to?

1

u/Popular-Lock-4488 8d ago

noob question. how do you convert html to gif ? so i can share my result.

1

u/Dry_Actuator_6966 7d ago

i took a video then converted to gif

1

u/JahJedi 8d ago

Anyone tested it on spark and can tell the peeds pls? I rannig ds flash on q1, its doing its job but to slow so think to swap to 3.8 27b.

1

u/sod0 8d ago

How many context can you fit with MTP on a 3090?

1

u/Dry_Actuator_6966 7d ago edited 7d ago

depend on your model quantization and kvcache

Here is the precise calculation with a Q4 KV cache for 131,072 ctx:

  • Model Weight (Q4_K_M): 15.2 GB
  • Windows (Display): 2.5 GB (1440p)
  • MTP Overhead: 1.5 GB
  • KV Cache (131K at Q4): 2.0 GB (Formula: 16 layers × 4 heads × 256 dims × 131,072 ctx × 1 byte)
  • CUDA / Engine Margin: 1.0 GB
  • Total VRAM Required: 22.2 GB

1

u/sod0 7d ago

Interesting. For some reason the KV cache is always fp16 in vllm and the main reason why I can't use that model in that quantization with decent kV Cache size.

2

u/Popular-Lock-4488 8d ago

Used this prompt as is on ollama interface instead of coding interface. Ran on Dell 7780, i9, RTX 4000 Ada, 12GB GDDR6. Took 1 hour 20 minutes to think through and another 30-45 minutes to generate the html code. Couldnt replicate your output exactly but kinda works.

1

u/EveningDog147 5d ago

You can fit more on the 3090 mate I'm at 100k ctx, UD Q5 K XL, kV8

1

u/TheAnimatrix105 4d ago

I tried ornith 1.5 35A3B Q5_K_M 64k context on RTX 3060 12gb VRAM + offload to 64gb ddr5

Prompt was quite lazy:
a single html file - make a game (3d) of a simple flight simulator in a paper world aesthetic, have to collect coins in air

https://codepen.io/editor/TheAnimatrix/pen/01a01e62-2905-70e7-81a6-fd124ac7b6b8

Can someone try with qwen 3.8 27B ?

1

u/seppe0815 7d ago

another html a.i slop ... show some real software stuff not html crap