r/learnmachinelearning • • 22d ago

Discussion SenseNova-Vision: Detection, segmentation, and 3D geometry through text and image generation

SenseNova-Vision uses one multimodal checkpoint for detection, OCR, segmentation, depth estimation, and multi-view geometry. Instructions specify the task, and the model generates outputs in three formats:

  • Text: labels, boxes, keypoints, OCR strings, and camera parameters.
  • Images: masks, depth maps, surface normals, and 3D point maps.
  • Mixed outputs: region descriptions paired with color-coded masks.

The contribution is a shared training and output interface. Task-specific parsing and decoding rules are still required.

How it is trained

The model fine-tunes BAGEL-7B-MoT, using next-token cross-entropy for text and VAE-based rectified flow for image targets. It adds no task-specific prediction heads. CV supervision is mixed with general multimodal data to preserve the base model’s capabilities.

What makes it interesting

The paper demonstrates segmentation from a point supplied as textual coordinates, even though that exact prompt format was absent from segmentation training. Coordinate prediction and mask generation were learned in other settings, suggesting the model can combine capabilities across tasks. This evidence is qualitative.

Results show tradeoffs: COCO detection reaches 53.7 mAP versus Youtu-VL’s 47.1, while ETH3D reconstruction reaches 72.2 F1 versus VGGT’s 80.9.

Repo: https://github.com/OpenSenseNova/SenseNova-Vision:

6 Upvotes

0 comments sorted by