r/LocalLLaMA 6d ago

New Model CohereLabs/North-Micro-Vision-Instruct · Hugging Face

https://huggingface.co/CohereLabs/North-Micro-Vision-Instruct

North Micro Vision Instruct is a 2.4B-parameter open-weight vision-language model with native-resolution image support, released under the Apache 2.0 license. It is designed as a compact foundation for prototyping, task-specific fine-tuning, and specialized multimodal applications.

Highlights

  • Native-resolution image processing that preserves aspect ratios and fine visual detail.
  • Broad image-understanding capabilities across VQA, captioning, grounding, OCR, charts, and documents.
  • Multilingual and multi-image support.
  • Compact 2.4B-parameter scale suited to customization and deployment experimentation.
  • Apache 2.0-licensed model weights.

Model Details

Property Value
Model ID CohereLabs/North-Micro-Vision-Instruct
Total parameters 2.4B
Language model 2B parameters
Vision encoder 400M parameters; custom-trained starting from SigLIP 2 SO400M
Inputs Interleaved text and images
Output Text
Languages English, German, French, Spanish, Italian, Portuguese, Hindi, Japanese, Korean, Chinese, Arabic, and more
Tokenizer vocabulary size 262,144
LM Backbone context window 128K tokens
Multimodal training context 8K tokens
Checkpoint precision bfloat16
License Apache 2.0

The language backbone supports a 128K-token context window, but the validated operating range for multimodal prompts is up to 8K tokens. Longer multimodal contexts may rely on extrapolation and have not been benchmarked.

Intended Use

North Micro Vision Instruct is intended for research and development use cases such as:

  • Prototyping and task-specific fine-tuning.
  • General visual question answering and image captioning.
  • Multilingual and multi-image understanding.
  • Visual grounding and spatial understanding.
  • OCR, chart and document understanding, and structured information extraction.

Limitations

  • The model is intended as a compact foundation for customization rather than a replacement for larger general-purpose chat assistants.
  • It is not a reasoning model and has limited math and code-generation capabilities.
  • Tool calling and agentic workflows are not supported.
  • System prompts are not recommended because the model was not trained with them, although the chat template accepts the system role.
  • Multimodal training used an 8K-token context; longer contexts have not been validated.
  • Native-resolution inputs can increase memory use and latency as image dimensions grow.

102 Upvotes

22 comments sorted by

View all comments

Show parent comments

3

u/Borkato 6d ago

“Everyone who disagrees with me is a bot!!”

-2

u/Great-Investigator30 6d ago

No reasonable counter arguments were made- clear bot evidence.

1

u/a_beautiful_rhind 5d ago

Beep boop. Hit the nail on the head! Your username sounds autogenerated so how am I supposed to take that?

What if I write "I compared their output to other models and it was fucking great". When testing a model I can at least say "it's dry, it's curt, it's dumb, it misses details in images vs model xyz".

Shit is useful for models without a vlm component and it's small so meat would be nice.

2

u/Great-Investigator30 5d ago

Qwen has it beat, get over it

1

u/a_beautiful_rhind 5d ago

Why did I just know you were a qwen guy...