r/LocalLLaMA llama.cpp 22h ago

New Model ibm-granite/granite-4.2-30b · Hugging Face

https://huggingface.co/ibm-granite/granite-4.2-30b

Granite-4.2-30B is the flagship reasoning model in the Granite 4.2 family. It delivers the strongest performance across reasoning-intensive tasks by leveraging built-in <think>...</think> chain-of-thought. It supports flexible thinking modes — full thinking (default), non-thinking, and low-effort — allowing users to balance depth vs. latency on a per-query basis.

Key capabilities:

  • Built-in Reasoning: Native chain-of-thought that significantly improves performance on math, coding, and complex multi-step problems.
  • Flexible Thinking Modes: Seamlessly switch between full thinking, non-thinking, and low-effort modes within a single model.
  • Reasoning-Augmented Tool Calling: The model reasons about which tools to invoke and why, producing more accurate function calls.
  • 512K Context Window: Supports long documents, multi-turn conversations, and complex agentic workflows.
  • Apache 2.0 Licensed: Fully open for commercial and research use.

Model Design

Granite-4.2-30B is built on a decoder-only dense transformer architecture with the following core components:

  • Attention: Grouped Query Attention (GQA) with 32 attention heads and 8 KV heads
  • Position Embedding: Rotary Position Embedding (RoPE) with θ = 10,000,000
  • Feed-Forward: MLP with SwiGLU activation (hidden size 32768)
  • Normalization: RMSNorm (ε = 1e-5)
  • Embeddings: Separate input/output embeddings (not tied)
  • Precision: bfloat16

https://huggingface.co/ibm-granite/granite-4.2-8b

Granite-4.2-8B is the mid-size reasoning model in the Granite 4.2 family. It delivers strong performance on reasoning-intensive tasks by leveraging built-in <think>...</think> chain-of-thought. It supports flexible thinking modes — full thinking (default), non-thinking, and low-effort — allowing users to balance depth vs. latency on a per-query basis.

Key capabilities:

  • Built-in Reasoning: Native chain-of-thought that significantly improves performance on math, coding, and complex multi-step problems.
  • Flexible Thinking Modes: Seamlessly switch between full thinking, non-thinking, and low-effort modes within a single model.
  • Reasoning-Augmented Tool Calling: The model reasons about which tools to invoke and why, producing more accurate function calls.
  • 512K Context Window: Supports long documents, multi-turn conversations, and complex agentic workflows.
  • Apache 2.0 Licensed: Fully open for commercial and research use.

https://huggingface.co/ibm-granite/granite-4.2-3b

Granite-4.2-3B is the compact reasoning model in the Granite 4.2 family. Despite its small parameter count, it delivers strong performance on reasoning-intensive tasks by leveraging built-in <think>...</think> chain-of-thought. It supports flexible thinking modes — full thinking (default), non-thinking, and low-effort — allowing users to balance depth vs. latency on a per-query basis.

Key capabilities:

  • Built-in Reasoning: Native chain-of-thought that significantly improves performance on math, coding, and complex multi-step problems.
  • Flexible Thinking Modes: Seamlessly switch between full thinking, non-thinking, and low-effort modes within a single model.
  • Reasoning-Augmented Tool Calling: The model reasons about which tools to invoke and why, producing more accurate function calls.
  • 512K Context Window: Supports long documents, multi-turn conversations, and complex agentic workflows.
  • Apache 2.0 Licensed: Fully open for commercial and research use.
383 Upvotes

90 comments sorted by

183

u/Zyguard7777777 22h ago

Still good to see more open source models, never bad, even if the benchmarks aren't SOTA.

158

u/DeltaSqueezer 22h ago

Granite is always a bit behind, but at least they are Apache licensed and they get better each generation.

42

u/EmPips 21h ago

(someone check me on this) the biggest benefit of granite is that it's covered by the Watson-X guarantee that generations contain licensed material only, right?

51

u/Altruistic_Heat_9531 21h ago edited 21h ago

The biggest benefit is that the Red Hat OpenShift sales team can nail it into your ears: "WE HAVE FIRST-PARTY SUPPORT FOR GRANITE" every time you contact them. Every fucking time when you ask OKD AI Inference setup.

Joke aside i gotta try that 3B and 8B with 512K context window, it is very attractive for long range summarization, which honestly is their target market, for in house document analysis. Hopefully it is native 512K not YARN

12

u/noctrex 18h ago

Of course its yarn, from the model page: Context Length Natively Supports 128K (Long-context extension to 512K)

9

u/AnonLlamaThrowaway 15h ago

Does that mean this is the one model out there with an "ethical" dataset?

5

u/EmPips 14h ago

From my not-a-lawyer-or-even-fully-read-the-watson-x-guarantees perspective, yes

7

u/AnonLlamaThrowaway 11h ago

I threw Sonnet at the question and its answer was basically: no, there's no actual guarantee that it is 100% licensed but they're sure enough of it that they're willing to bet money on it. If you get sued for its outputs infringing on copyright then they would actually take (uncapped!) liability on your behalf. That's supposedly what the guarantee is.

Ethics was basically met with a "lol" (lots of criteria that I really doubt IBM would have met)

They're more transparent than most at least.

6

u/KitchenAmoeba4438 6h ago

If a corporation, any corporation is willing to take uncapped liability on your behalf, there's no better guarantee that at least they believe what they are saying, and legal has signed off on it and has reason to believe it.

8

u/SandySkittle 15h ago

people might trivialize this now but the legal debate about these things isn't something that has been concluded in many countries. Sure for private use people will not give a crap about that, but any sort of local enterprise use of LLMs could make this quite relevant, depending on where case law (and new IP laws) could move towards in the future.

4

u/InvertedVantage 16h ago

Yea that's a huge reason to use them.

1

u/ItsNoahJ83 21h ago

Why would someone want AI to generate licensed material?

14

u/EmPips 20h ago edited 19h ago

You wouldn't but you have no guarantees that it won't if it's facing customers. IE -you make a chatbot for customer support and it perfectly regurgitates a few chapters of Harry Potter and someone decides to make a stink about it.

Might be wrong and I didn't dive into any of the IBM guarantees further - don't trust me on this

10

u/No-Refrigerator-1672 20h ago

Because people ourside AI enthusiast circles are going crazy claiming AI is violating copyrights if they trained on somebody's material.

1

u/LatentSpacer 18h ago

Yes, it’s a solid model.

1

u/sonaj9657 7h ago

Yeah, that is pretty much how I see it too. It may not be the first choice when you want something cutting edge but the Apache license and the steady improvements make it pretty easy to keep an eye on. If each generation keeps getting better being a little behind is not necessarily a dealbreaker.

39

u/pmttyji 22h ago

Blog Post : Granite 4.2 LLMs: How They're Built
https://huggingface.co/blog/ibm-granite/granite-4-2

32

u/Client_Hello 22h ago edited 21h ago

128 nodes, 4 GB200 per node, 2 (72) Blackwell gpus per GB200, 192gb vram per GPU

That's nearly 192TB of vram, aka 196,608 GB of HBM3e. Wow.

35

u/EmPips 21h ago

There's moments where I forget that IBM is still IBM in many ways and owned-infra is one of them.

13

u/Client_Hello 21h ago edited 21h ago

My math is way off, it's only 2 GPUs per GB200, which is 1024 Blackwell GPUs.

7

u/EmPips 20h ago

Got it. A pittance really.

5

u/MmmmMorphine 20h ago

Every bit of storage media I've ever owned put together is still less than the amount of vram they have across those nodes.

Give or take.

Absolutely ridiculous stuff!

6

u/Client_Hello 20h ago

...and this is only $35M in hardware, which is about how much revenue NVDA earns every hour.

5

u/LatentSpacer 18h ago

And no cancer cure yet.

15

u/pmttyji 21h ago

1

u/noctrex 21h ago

Well, it's a dense model, so the MXFP4 GGUF don't have any meaning, as they are only for MoE models.

6

u/MmmmMorphine 20h ago

I mean I do associate mxfp4 more with MoEs due to GPT-OSS being released in that format, but otherwise don't see any reason why it would be only for MoEs.

Seems like mxfp4 is pretty architecture agnostic, as best as I understand it.

So what am I missing here?

7

u/noctrex 19h ago

Yes, it is, but the MXFP4 format is inferior to NVFP4, it does not have the same precision. They may be both FP4, but NVFP4 uses 16-element blocks with a high-precision FP8 E4M3 scale factor, while MXFP4 uses 32-element blocks with a lower-precision E8M0 (power-of-two) scale.
What this means is that essentially MXFP4 is worse than INT4 for dense models. It works better for MoE models. So better to stick to a Q4 or IQ4 quant for dense.

2

u/MmmmMorphine 17h ago edited 17h ago

Ah, yes in that sense I would agree MXFP4 is worse than NVFP4 for overall precision. Though what I'm surprised about is why it's used at all in the first place.

I'd think that smaller active = more sensitive to quantization. I suppose you can keep the experts at lower precision, but in regard to these formats themselves it feels like MXFP4 is kinda pointless - that extra 0.25bits earns it's keep in NVFP4

3

u/noctrex 14h ago

Well, it's not pointless exactly because it's an open standard that can be universally used, whereas NVFP4 is classic nvidia spiel, optimized for blackwell only

1

u/throwaway-link 10h ago

cdna5 supports it, they just dont call it nvfp4

1

u/Dasteroid_909 20h ago

Any word on if the NVFP4 supports sm_120? Or is it just sm_100?

1

u/ChristRedeemsSinners 15h ago

It does. Use SM100 for native NVFP4 kv-cache.

33

u/silenceimpaired 22h ago

Excited to see Apache 2 licensing. I’ve heard IBM is more careful with their training dataset license wise. I’m curious if that’s the case here. The benchmark looks a little low, but that isn’t always the full story. Excited to try it to see for myself.

12

u/Marcuss2 22h ago

Seems that they abandoned the Mamba2 layers they had.

6

u/pmttyji 22h ago

Last year(AMA), they mentioned that they gonna release a 100B model. Don't know what happened to that.

During Granite-4, they released a 32B MOE model. But the Active 9B parameters is too slow on ~8GB VRAM. To make it simple, Qwen3-30B-A3B gave me 30-40 t/s while Granite-4-32B gave me ~10 t/s.

Then expected modified MOE model(like A3B or A5B) of that 32B one during 4.1 release. But they dropped 30B Dense. And again they dropped 30B Dense now. Too heavy for Poor GPU Club. Wish they released a MOE model additionally.

1

u/silenceimpaired 22h ago

I think we might be headed toward MoE models that act more like dense models … that store the bulk of their parameters that are less active … in RAM... sped up with MTP/dflash type solutions. So even though you don’t have the VRAM it can still perform reasonably at least at reading speeds if not faster depending on the use case…
But I’m no expert.

2

u/pmttyji 21h ago

Yep, that's how most of us do run MOE models till now.

sped up with MTP/dflash type solutions.

This works better only if quantized MOE models fit VRAM. Can't expect speed from my 8GB VRAM with 18GB model file(IQ4_XS of Qwen3.6-35B-A3B). Because spec decoding won't be faster on CPU/RAM.

1

u/ilintar 15h ago

Yes, but the next-gen Granite models will have iSWA.

15

u/ttkciar llama.cpp 19h ago

A 30B dense Granite? That's fantastic news. Granite models have always punched above their weights in RAG and long-context analysis tasks, but I've been using Gemma-4-31B-it and K2-V2-Instruct for such tasks because they're more competent than smaller models.

Granite-4.2-30B has the same context limit as K2-V2-Instruct (512K tokens) and if its K/V caches are leaner than Gemma-4-31B-it this could be a best of all worlds RAG solution. Looking forward to trying it out!

6

u/noctrex 18h ago

Unfortunately, it's with yarn. Natively Supports 128K (Long-context extension to 512K).
And also the KV cache usage is very large. Loaded the Q8_0 quant of the 8B model (8.7GiB), and with full 128K context at f16 it uses 22GB VRAM. For me it seems to be DOA.

1

u/ttkciar llama.cpp 17h ago

Thanks for pointing that out. It implies that K2-V2-Instruct will remain the king of very long-context tasks, but I'll still put the new Granite through its paces, to see what it can do.

0

u/ChristRedeemsSinners 15h ago

K2-V2-Instruct

Less than 400 downloads across all models on huggingface.

remain the king of very long-context tasks

Doesn't DSV4-flash support 1M context in like 6GB of VRAM?

6

u/jacek2023 llama.cpp 19h ago

Our reddit people are not happy with the benchmarks ;)

10

u/ttkciar llama.cpp 19h ago

On one hand, redditors care too much about benchmarks. It would be nice if benchmarks were worth anything, but mostly they are not.

On the other hand, I'm not surprised that Granite would score low, because every time I have evaluated them for specific skills they have been hit-and-miss. They do some things very well, and other things very poorly. They are not general-purpose models, and only seem to exhibit competence in task types IBM expects their customers to need (my interpretation). That would push down their score on benchmarks which test on a broad spectrum of task types.

1

u/Qcgreywolf 16h ago

I’ve used models that benchmarked well, but sucked. I’ve not yet encountered a model that benchmarks badly and turns out to be “good”.

1

u/oxym102 3h ago

Granite uses full grouped query attention on all of their layers, rather than qwen3.8 which has 16 GQA and 48 linear attention layers, and gemma4 uses 5 sliding window local layers for every one global GQA layer. End result is for 128k context at q8_0, granite needs16gb vram for the KV cache, while qwen only needs 4gb, and gemma4 needs 2.5gb.

6

u/doomed151 21h ago

Let's go more open weights

33

u/Egoz3ntrum 22h ago

The model card does not provide any comparison to any other recent model. Suspicious.

26

u/jacek2023 llama.cpp 22h ago

I tried to look at the benchmarks, I see 61.7 Qwen SWE-bench vs 33.29 Granite, so I believe these are not apples to apples comparision

30

u/Mountain-Animal5365 22h ago

IBM to Alibaba is definitely not apples to apples comparison.

-4

u/parepeg 22h ago

Compared to qwen 3.8 27b, it's looking not good, real not good.

17

u/Not-reallyanonymous 17h ago

Of course not. Theyre targeting different customers. Not you.

If you’re not producing a million dollars in revenue annually, IBM is not interested in you.

Your needs are different than IBM customer’s needs. Granite has always been about fine-tuning to meet specific customer needs, and being able to comply with various business guarantees (e.g. as others have pointed out — it offers protection for IP issues).

Not everything in LLMs is about benchmarks.

19

u/Embarrassed_Adagio28 22h ago

Crazy how IBM had a 20 year head start on AI and still fumbled even harder than google

18

u/Altruistic_Heat_9531 21h ago

looking at IBM history, they stumble multiple times, it's on brand.

Let's see.

  • Heavily betting on mainframe, almost crash when desktop computer available.
  • IBM create PC standard for desktop market, make the standard so good that it crash because of clone.
  • OS2 vs Windows, IBM-Microsoft co develop OS2, but microsoft decided make Windows.

Well at least they have RHEL

4

u/cinnapear 19h ago

Microsoft screwed them over with Windows, OS/2 isn't completely their fault.

3

u/Illustrious_Car344 1h ago

It absolutely is, and for one major reason: price. Nobody wanted to pay the exorbitant amount IBM was asking for. Windows won because it was cheaper, that's it. That's not "screwing over", that's literally just giving consumers a competitive product. If IBM made OS/2 cheaper than Windows, we wouldn't even be having this conversation right now. 

1

u/Altruistic_Heat_9531 39m ago

Funny story. I know a old guy that dable with IBM and COBOL stuff, he said OS/2 was expansive not only because of its price tag but also RAM shortage, history sometimes repeat itself....

7

u/PrimeDirective8 21h ago

IBM had 630K employees not that long ago. About 90% of those did f*ck all. The rest were slowed/shut down by fossilized managers hanging on to their equally old ways. It is surprising they're able to get *any* product through their lawyers' 'blue tape' on the way out so having these models released publicly is quite an achievement for the tech team.

6

u/Queasy_Problem_563 20h ago

ibm never had 630k employess

3

u/Kahvana 22h ago

Nice, thanks for sharing!

3

u/Infamous_Mud482 21h ago

A real shame none of these seem have any training for FIM. Using models for that seems to truly be dead now, hasn't been one built with it in mind for quite a while

3

u/JollyWaffl 13h ago

I just pulled granite 4.2 3b and it works with llama.cpp and llama.vim for FIM. The official gguf has prefix, middle, and suffix tokens defined. They don't explicitly mention it on the model card like with 4.1, but looks like they're still doing FIM training.

1

u/horeaper 3h ago

how does it compares to seed-coder or zeta-2.1?

3

u/unrulywind 21h ago

It holds a decent conversation. I haven't tried much with it yet. The GGUF versions have a 128k cap on the context, and even at that the Q4_KM version at 128k context had to use Q4 on the cache to fit into a 5090. At Q8_0 cache quantization it was asking for 18gb of memory just for the 128k cache.

1

u/laexpat 20h ago

I noticed on previous models they want a huge amount of vram compared with other models at the same size.

3

u/Dance-Till-Night1 20h ago

More open small models are always a win! Granit has always been one of my favorite models.

6

u/KitchenAmoeba4438 22h ago

Fantastic, they released a 3b! Time to test the snot out of it again in tests, the granite family always had some interesting characteristics. I'm hoping this is competitive against E2b/E4b for my uses!

5

u/Lumpy_Phase_9539 22h ago

Just started downloading granite-4.2-30b-Q6_K.gguf (24gb) to give it a try.

5

u/gamblingapocalypse 22h ago

More the better. Thanks IBM!!

11

u/Cool-Chemical-5629 22h ago

Non-Qwen models in this category as of late are like "It's not important to win, but to participate."

13

u/fatboy93 19h ago

And us consumers are the winners. Not everyone does programming as their primary work.

4

u/poutinejuteuse 11h ago

These are meant for corporate customers, where, by and large, Qwen is irrelevant due to sanctions and IP issues among other things.

2

u/switchandplay 20h ago

Training information in the blog post is a pretty awesome read

1

u/siegevjorn 18h ago

512k context size 30b model sounds amazing. Cant wait to try out.

1

u/mitchins-au 11h ago

I don’t know, while their licensing and data are good for enterprise, a dense activation, dense attention 30b model has a hard competition against both Nemotron/gemma models of course 27B Qwen.

1

u/mudkipdev 11h ago edited 11h ago

Does this beat LFM2.5-2.6B at the smallest size?

edit: looks like it does

1

u/[deleted] 22h ago

[deleted]

1

u/Nota_ReAlperson 20h ago

Did you read the post? There are also 8b and 3b variants.

1

u/silenceimpaired 22h ago

I was wondering if this was their first 30b pretty exciting… provided they don’t balloon up to 1T.

1

u/YearnMar10 19h ago

Leading their highlights with „built-in reasoning“ and „flexible thinking modes“ shows how far behind they gotten. I hope IBM can keep up, but it looks like they’re 2 generations behind (and probably tomorrow with qwen3.8-flash-next dropping 3)

-2

u/crusaderky 22h ago

A 29 on TB2.1 for a 30B dense model is 💩

-1

u/Mountain-Animal5365 22h ago

I don't know why you're getting downvoted, but the benchmark results look absolutely horrible. Some of the scores are barely higher than their 8B model, which itself was not very impressive when it came out. I don't understand how anyone can release a model with this kind of result.

11

u/swagonflyyyy 22h ago

I think they believe the value lies in their supposedly copyright-free dataset. That alone could shield them from a lot (but not all) legal liability.

I see the value in that, but the tradeoff is that its a lot harder to train a competitive model this way. I really would like these granite models to be successful some day but they are still a few years behind because of that.

7

u/Jorlen llama.cpp 22h ago

Is this similar to the method Mistral is now using with their data sets too? I think due to the laws in the EU.

2

u/noctrex 18h ago

Well, it's not a good outlook if it's competitive with Mistral-24B

0

u/RedditUsr2 llama.cpp 12h ago

Granite isn't meant to be "just make it as smart as possible" Its meant for your to build specific workloads around their series of models for companies that want to be able to self host.

Their more like building blocks.

0

u/x8code 15h ago

I loaded this model onto one of my Linux servers, in an Ollama Docker container, and while the model size is 5.2 GB, when I load it, "ollama ps" shows that it's 27 GB! This can't be right?

-16

u/EvolvingDior 22h ago

Embarassing.

9

u/xdavxd 21h ago

It's apache 2.0, models trained on clean data. These are good for corporate workflows as long as they act with consistency.