Follow the latest coverage, related explainers and connected technology stories.
JetBrains has launched the open Mellum2.1 model under the Apache 2.0 license, focusing on running coding agents and sub-agents locally or on users’ own infrastructure. The release is based on reinforcement-learning training in real software environments, with major improvements in handling code repositories and inference speed.
A CNCF article argues that operating enterprise AI infrastructure is not about deploying a single model or cluster, but about managing a shared fleet of GPU units for multiple teams while balancing utilization, isolation, and cost. The article presents technical layers covering provisioning, allocation, scheduling, networking, storage, monitoring, and billing.
Perplexity launched the Portable Computer platform to run its agents locally on users’ devices, starting with Nvidia DGX Spark and Linux computers equipped with RTX cards containing at least 24GB of memory. By default, the system keeps models, files, and work on the device, with the option to request user approval before escalating a specific step to a more powerful cloud model.
CohereLabs has launched North-Micro-Vision-Instruct, an open-weight vision-language model with 2.4 billion parameters that supports processing images at their native resolution under an Apache 2.0 license. The model targets document and chart understanding, OCR, and custom fine-tuning applications, and also includes community support for MLX-VLM and an AutoModel recipe for deployment on NVIDIA graphics processing units.
VIDRAFT’s AX-Ray project presents a model-evaluation method that goes beyond performance scores after identifying and repeatedly reproducing defects it described as causal leakage in the Zyphra/Zamba2-1.2B and nvidia/Nemotron-H-8B-Base-8K models. The framework places these cases among defects that prevent deployment, distinguishing them from serving-path or API issues.
Liquid AI announced the availability of the LFM2.5-2.6B and LFM2.5-2.6B-Base models on Hugging Face, with support for tool calling and multi-step workflows locally. The model targets running fast agents on computers and phones, while using less than 2.5 gigabytes of memory.