Tools

vLLM

Follow the latest coverage, related explainers and connected technology stories.

Latest coverage

CN
JetBrains Launches Mellum2.1 as a Fast Open Model for Coding Agents

JetBrains Launches Mellum2.1 as a Fast Open Model for Coding Agents

JetBrains has launched the open Mellum2.1 model under the Apache 2.0 license, focusing on running coding agents and sub-agents locally or on users’ own infrastructure. The release is based on reinforcement-learning training in real software environments, with major improvements in handling code repositories and inference speed.

CN
Why Do “AI Factories” Need More Than One Kubernetes Cluster?

Why Do “AI Factories” Need More Than One Kubernetes Cluster?

A CNCF article argues that operating enterprise AI infrastructure is not about deploying a single model or cluster, but about managing a shared fleet of GPU units for multiple teams while balancing utilization, isolation, and cost. The article presents technical layers covering provisioning, allocation, scheduling, networking, storage, monitoring, and billing.

CN
Perplexity and Nvidia Launch Portable Computer to Run AI Agents Locally

Perplexity and Nvidia Launch Portable Computer to Run AI Agents Locally

Perplexity launched the Portable Computer platform to run its agents locally on users’ devices, starting with Nvidia DGX Spark and Linux computers equipped with RTX cards containing at least 24GB of memory. By default, the system keeps models, files, and work on the device, with the option to request user approval before escalating a specific step to a more powerful cloud model.

CN
CohereLabs Launches Open-Weight North Micro Vision Model for Native-Resolution Image Processing

CohereLabs Launches Open-Weight North Micro Vision Model for Native-Resolution Image Processing

CohereLabs has launched North-Micro-Vision-Instruct, an open-weight vision-language model with 2.4 billion parameters that supports processing images at their native resolution under an Apache 2.0 license. The model targets document and chart understanding, OCR, and custom fine-tuning applications, and also includes community support for MLX-VLM and an AutoModel recipe for deployment on NVIDIA graphics processing units.

CN
AX-Ray Reveals Causal Leakage Defects in Two Public AI Models

AX-Ray Reveals Causal Leakage Defects in Two Public AI Models

VIDRAFT’s AX-Ray project presents a model-evaluation method that goes beyond performance scores after identifying and repeatedly reproducing defects it described as causal leakage in the Zyphra/Zamba2-1.2B and nvidia/Nemotron-H-8B-Base-8K models. The framework places these cases among defects that prevent deployment, distinguishing them from serving-path or API issues.

CN
Liquid AI Releases LFM2.5-2.6B Model to Run Agents Locally on Devices

Liquid AI Releases LFM2.5-2.6B Model to Run Agents Locally on Devices

Liquid AI announced the availability of the LFM2.5-2.6B and LFM2.5-2.6B-Base models on Hugging Face, with support for tool calling and multi-step workflows locally. The model targets running fast agents on computers and phones, while using less than 2.5 gigabytes of memory.