Artificial intelligence

CohereLabs Launches Open-Weight North Micro Vision Model for Native-Resolution Image Processing

CohereLabs has launched North-Micro-Vision-Instruct, an open-weight vision-language model with 2.4 billion parameters that supports processing images at their native resolution under an Apache 2.0 license. The model targets document and chart understanding, OCR, and custom fine-tuning applications, and also includes community support for MLX-VLM and an AutoModel recipe for deployment on NVIDIA graphics processing units.

2026-08-14
4 min read
14 views
فريق تحرير certi.news
CohereLabs Launches Open-Weight North Micro Vision Model for Native-Resolution Image Processing

CohereLabs announced the launch of North-Micro-Vision-Instruct, an open-weight vision-language model with 2.4 billion parameters that supports native-resolution image input under an Apache 2.0 license. The company says the model is its smallest vision-language model to date and was designed as a compact foundation for specialized multimodal applications.

The model and its weights are available on Hugging Face in the CohereLabs/North-Micro-Vision-Instruct repository. General support for vLLM is still being prepared, while CohereLabs used an internal version of vLLM during its evaluations.

Focus on Documents and Native-Resolution Images

North Micro Vision preserves an image’s aspect ratio and fine details instead of first reducing every image to a small square. This enables it to handle small text, document layouts, tables, charts, screenshots, and forms. The training also covers multiple languages and visual domains, including documents, charts, and natural images.

These capabilities target developers who need a model that can be fine-tuned for specific visual domains or run under local and edge constraints. CohereLabs indicates that models of this size, with a suitable inference package and quantization techniques, may support experiments beyond server deployment to include laptops and devices categorized as edge or mobile.

Three-Component Architecture

The model consists of a native-resolution vision encoder, a projection module, and a compact language model. It combines a company-trained vision encoder with 400 million parameters and the internal North Micro LLM language model with 2 billion parameters.

The language model follows the Command A+ architecture, alternating between three sliding-window attention layers that use rotary positional embeddings and one global attention layer without positional embeddings. The vision encoder uses 2D RoPE and learned one-dimensional positional embeddings to preserve spatial structure in native-resolution images. The projection module injects representations from multiple layers of the vision encoder into corresponding early layers of the language model.

Multistage Training

The model was trained in four stages. The process began by initializing the vision encoder and projection module, followed by increasing the resolution in two stages while training the encoder, projection module, and language model together. The model then underwent instructional fine-tuning, and training concluded with a simplified version of Mixed Preference Optimization to improve safety, alignment, and response quality.

Training of the vision encoder progressed from a resolution of 384×384 pixels using 10 million examples, to 1024×1024 pixels using 13 million examples, and then to a native maximum resolution of 1654×2339 pixels using 10 million examples. This maximum corresponds to an A4 page at 200 dots per inch, allowing the model to process a single document page while preserving its aspect ratio.

During instructional fine-tuning, 50 million multimodal samples were used, distributed primarily among native OCR, charts and tables, localization and counting, OCR question answering, and general image understanding. The data also included image captioning, knowledge, text-only, mathematics, and science. The process relied on publicly available datasets and the company’s internal multilingual document collection.

Evaluation Results and Available Integrations

CohereLabs evaluated the model on tasks covering general, multilingual, and multi-image understanding; chart, document, and OCR understanding; science and technology; localization and counting; hallucination resistance; and textual capabilities. According to the published results, the model scored 0.921 on DocVQA, 0.808 on ChartQA, 0.792 on OCRBench, and 0.725 on CountBench, while its HallusionBench performance reached 0.615.

The company provides MLX-VLM-compatible weights with a community contribution from Prince Canuma and Neywa. In partnership with NVIDIA, it also ships an AutoModel recipe that enables developers to fine-tune and deploy the model on NVIDIA graphics processing units, along with support for custom fine-tuning through the community’s Axolotl framework.

News source
Hugging Face Blog
Open original source ↗
ف
Author

فريق تحرير certi.news

In the same category

You may also like

View all news