vLLM v0.29.0 defaults Model Runner V2, DeepSeek cuts prices
57,684 stars today74 new projects+3,433 archify
By olud.ai editorial · built from our own tracking data, published daily
vLLM v0.29.0 ships with Model Runner V2 enabled by default, marking the completion of a major rollout across all supported models. The release includes 594 commits from 277 contributors, 91 of them new.
The ecosystem added 57,684 GitHub stars today across 10,000+ tracked open-source AI projects. 74 new projects entered the index. This week 2,258 new projects and 100 new Hugging Face spaces were added. The top gainer was "archify" with +3,433 stars.
DeepSeek cut input prices on multiple models while other providers raised theirs. The largest reduction was DeepSeek V3.1 dropping 55% to $0.25 per million tokens.
Releases of the day
vLLM v0.29.0 makes Model Runner V2 the default inference engine for all models, completing a transition that began in earlier releases. The change should improve performance and memory usage across the board. OpenCode v1.18.30 added the Astra system prompt for GPT-6 models and fixed Bedrock DeepSeek model ID resolution. Cline desktop-v0.0.24 fixed a live chat stream bug that doubled text and dropped messages mid-turn. Milvus v3.0.1 is a patch release with a changelog pending. Pydantic AI v2.42.0 now rejects invalid DeferredToolResults.approvals values. Instructor v1.17.0 includes previously planned fixes and changes cache key format. LoopX v1.0.2 introduced single-owner Todo authority and automatic recovery.
DeepSeek led price cuts: V3.1 input down 55% to $0.25, V4 Flash down 11% to $0.08, V4 Pro down 1% to $0.95. Other providers moved upward: Kimi K2.7 Code up 8% to $0.71, Qwen3.5 397B A17B up 41% to $0.55, GLM 4.6 up 28% to $0.55, Qwen3 14B up 92% to $0.23, and DeepSeek V3 0324 up 16% to $0.29.
NeoHorse-1 presents a family of agent-native models designed for recursive self-improvement. The method uses a routing harness to let an AI system observe its own capabilities and convert that evidence into the next round of learning. The paper has 163 upvotes on Hugging Face.
Today's facts show two converging trends: inference infrastructure matures with vLLM's default switch, while model pricing becomes more volatile as providers adjust to shifting demand. The NeoHorse-1 paper points toward a future where models improve themselves through structured iteration.
Source: olud.ai tracking of 10,000+ open-source AI projects, 300+ models and live provider pricing. Figures are measured, not estimated. All releases · Live pricing · Latest in AI