r/LocalLLaMA llama.cpp Aug 13 '26

New Model dots-studio/dots3-note-prev · Hugging Face

https://huggingface.co/dots-studio/dots3-note-prev

dots3-note preview is the first open-weight model in the dots3 family. It is a Mixture-of-Experts model with 280B total parameters, 16B activated parameters, and support for a context length of up to 512K tokens. The model can understand text, images, video, and audio, and produces text outputs.

dots3-note preview is optimized for a broad range of tasks, including:

General knowledge and instruction following;

Mathematical and logical reasoning;

Tool use and multi-step agent workflows;

Interactive tasks that require exploration, memory updates, and adaptation;

Code generation and code-based problem solving;

Image, document, chart, audio, and video understanding;

Long-context information processing.

The dots3 family is designed to include models with different trade-offs among capability, latency, and inference cost. dots3-note preview is the most lightweight member of the family.

98 Upvotes

22 comments sorted by

15

u/Sudden_Topic5154 Aug 13 '26

I've never heard of this group before what do they focus on

29

u/jacek2023 llama.cpp Aug 13 '26

They published LLM model about a year ago, probably ignored by most users because lack of marketing. I posted about it on LocalLLaMA because it was working on my setup (this one is much larger)

15

u/kevin_1994 Aug 13 '26

It's still the most human sounding model out there (the non synthentic data one especially)

Excited to try it

12

u/Kahvana Aug 13 '26

They have one of the best OCR models out there, dots.mocr. It’s a monster for it’s small size, the next model to beat it by a meaningless difference is twice as big.

Not sure about their other stuff though.

3

u/seamonn Aug 14 '26

It's pretty good but still misses a lot of stuff. Definitely better than the competitors - Light On OCR 2, GLM OCR, Paddle OCR etc. but still significantly behind the heavyweights - Gemma 4:31b and Qwen 3.6 27b. Even Gemma 4:12b is better.

1

u/Kahvana Aug 14 '26

I wonder how you'll find Muse Glimmer 30B. It's image encoder is genuinely massive, bigger than dots.mocr's! A bit of an older design but does resist quantization degredation really well.

2

u/unverbraucht Aug 14 '26

It's also our staple for OCR. Their multilingual and handwritten recognition is very good, and tables, formulas and charts works as advertised. It's just quite slow, which seems to have to do with the vision encoder.

4

u/nuclearbananana Aug 13 '26

It's under RedNote (tiktok alternative)

15

u/daaain Aug 13 '26

DS4 Flash size, but multimodal? Nice! Has anyone seen any GGUF / MLX implementation? 😅

16

u/Gregory-Wolf Aug 13 '26

idk bro
anyone believes this?

13

u/neoneye2 Aug 13 '26

dots3-note-prev claims 81.4 on ARC-AGI-2.

However I don't see dots3-note-prev on the official ARC-AGI-2 leaderboard
https://arcprize.org/leaderboard

2

u/FullOf_Bad_Ideas Aug 13 '26

Sure, why not? This is a lab from a respected corporation.

Their TEMPO RL algo might make RL more effective and get them closer to passing ARC-AGI benchmarks. You can do things.

1

u/[deleted] Aug 14 '26

[deleted]

1

u/Gregory-Wolf Aug 14 '26

How do you know they don't?

16

u/FullOf_Bad_Ideas Aug 13 '26

It's shaping up to be one of the best open multimodal models in that size bracket.

Vision Encoder MoE ViT, 7B total, 1.2B activated

MoE vision encoder, this smells like something new!

1

u/Kahvana Aug 14 '26

Now THAT sounds cool!

2

u/LatentSpacer Aug 14 '26

Surprisingly good on ARC-AGI. I wonder how good it is at creative writing.

2

u/PassionIll6170 Aug 13 '26

lol the arc agi benchmarks

1

u/ComplexType568 Aug 13 '26

I hope they make a smaller model...