AI News · September 9, 2026

vLLM v0.29.0 defaults Model Runner V2, DeepSeek cuts prices

57,684 stars today74 new projects+3,433 archify

vLLM v0.29.0 ships with Model Runner V2 enabled by default, marking the completion of a major rollout across all supported models. The release includes 594 commits from 277 contributors, 91 of them new.

The ecosystem added 57,684 GitHub stars today across 10,000+ tracked open-source AI projects. 74 new projects entered the index. This week 2,258 new projects and 100 new Hugging Face spaces were added. The top gainer was "archify" with +3,433 stars.

DeepSeek cut input prices on multiple models while other providers raised theirs. The largest reduction was DeepSeek V3.1 dropping 55% to $0.25 per million tokens.

Releases of the day

vLLM v0.29.0 makes Model Runner V2 the default inference engine for all models, completing a transition that began in earlier releases. The change should improve performance and memory usage across the board. OpenCode v1.18.30 added the Astra system prompt for GPT-6 models and fixed Bedrock DeepSeek model ID resolution. Cline desktop-v0.0.24 fixed a live chat stream bug that doubled text and dropped messages mid-turn. Milvus v3.0.1 is a patch release with a changelog pending. Pydantic AI v2.42.0 now rejects invalid DeferredToolResults.approvals values. Instructor v1.17.0 includes previously planned fixes and changes cache key format. LoopX v1.0.2 introduced single-owner Todo authority and automatic recovery.

Pricing moves

DeepSeek led price cuts: V3.1 input down 55% to $0.25, V4 Flash down 11% to $0.08, V4 Pro down 1% to $0.95. Other providers moved upward: Kimi K2.7 Code up 8% to $0.71, Qwen3.5 397B A17B up 41% to $0.55, GLM 4.6 up 28% to $0.55, Qwen3 14B up 92% to $0.23, and DeepSeek V3 0324 up 16% to $0.29.

▼ 11%DeepSeek V4 Flash 0423DeepSeek · input $0.09 → $0.08 per 1M tokens
▼ 1%DeepSeek V4 Pro 0423DeepSeek · input $0.96 → $0.95 per 1M tokens
▲ 8%Kimi K2.7 CodeMoonshot AI · input $0.66 → $0.71 per 1M tokens
▼ 55%DeepSeek V3.1DeepSeek · input $0.55 → $0.25 per 1M tokens
▲ 16%DeepSeek V3 0324DeepSeek · input $0.25 → $0.29 per 1M tokens
▲ 41%Qwen3.5 397B A17BAlibaba · input $0.39 → $0.55 per 1M tokens
▲ 28%GLM 4.6Z.AI · input $0.43 → $0.55 per 1M tokens
▲ 92%Qwen3 14BAlibaba · input $0.12 → $0.23 per 1M tokens

Research pick

NeoHorse-1 presents a family of agent-native models designed for recursive self-improvement. The method uses a routing harness to let an AI system observe its own capabilities and convert that evidence into the next round of learning. The paper has 163 upvotes on Hugging Face.

The bigger picture

Today's facts show two converging trends: inference infrastructure matures with vLLM's default switch, while model pricing becomes more volatile as providers adjust to shifting demand. The NeoHorse-1 paper points toward a future where models improve themselves through structured iteration.

Source: olud.ai tracking of 10,000+ open-source AI projects, 300+ models and live provider pricing. Figures are measured, not estimated. All releases · Live pricing · Latest in AI

← All AI news