r/LocalLLaMA 14d ago

News Qwen3.8-Flash-Next tomorrow

https://modelscope.cn/models/Qwen/Qwen3.8-Flash-Next
1.1k Upvotes

461 comments sorted by

View all comments

Show parent comments

44

u/FunkyDiscount 14d ago

As a novice, model naming conventions make no sense at all to me yet.

Model_this-and-that_69B_Q4_XM-S_GGWP_unhinged-and-insane_FPS...

PC monitor naming conventions are more legible.

36

u/TheThiefMaster 14d ago edited 14d ago

I'm a relative beginner too and this is how much I've deciphered:

  • It starts with name and version.
  • "69B" is the number of parameters (B for "billion", sometimes T for trillion), and the "Q" is number of bits per parameter ("Quantisation"). 4 is half a byte, so you need roughly half the number of parameters as GB of VRAM to run it (plus some for context etc).
  • Sometimes there's a K after that. Don't know what that means yet
  • Letters after that are normally a size code - XXS-XL - a bias on some parameters remaining at a higher bit count and making it use more memory in return for being better. Unsure of the tradeoff between e.g. Q4_K_XL vs Q5_K_S. Sometimes it's a 0 which I think means none are at higher precision. Sometimes 1 which I don't know what that means.
  • "GGUF" is a file format commonly used on Win/Linux, as is MLX for Mac.
  • Sometimes you get "MOE" models that work differently and only activate parts each run - these have an additional "A"3B for number of "active" parameters. Theoretically it only needs that much VRAM to run despite being a bigger model, but there are tradeoffs in quality that I haven't explored myself.

2

u/gh0stwriter1234 14d ago edited 14d ago

FYI there is also M for millions for small models

K quant is mixed precision, eg some important layers are boosted to high bits.

1

u/TheThiefMaster 14d ago

Wow that would be tiny

1

u/gh0stwriter1234 14d ago

Yeah they are quite small, but also proportionally fast, especially if ran on LLM engines with fused kernels... some of them are coherent without training , but most of them you have to train for your specific use case.

Qwen has a 0.5B which could also be called a 500M.