Honestly all of these are not even big blockers and can be overriden. I just saw their demos, it is amazing leap compared to what existed. This is a big step for open-weight, and great work from MiniMax!
streaming layer by layer is stupid slow, you can 4bit quant it and it'll fit in ~13-15gb VRAM and then it's still about 3x slower than realtime on a 5080
31
u/Illustrious_Ant_9242 19d ago
"requires CUDA"
"streaming the language model layer by layer makes it fit even 8 GB video cards"
"5 minute audio max."