AI News · August 16, 2026

DeepSeek and GLM slash input prices, llama.cpp ships LoRA bounds fix

238,868 stars today133 new projects+6,339 speech-to-speech

DeepSeek V4 Flash 0423 input pricing dropped 57% to $0.06 per 1M tokens, the largest cut among five price reductions today. GLM 5.2 from Z.AI fell 61% to $0.46 per 1M tokens, while GLM 5.1 dropped 31% to $0.97. Kimi K2.6 from Moonshot AI edged down 4% to $0.54, and Qwen Plus 0728 thinking from Alibaba fell 35% to $0.26.

The open-source ecosystem gained 238,868 GitHub stars across the 10,000+ tracked projects today, with 133 new projects entering the index. The week totals 2,131 new projects and 100 new Hugging Face spaces. The "speech-to-speech" category leads GitHub with +6,339 stars today.

Tencent's Hy3 preview moved against the trend, with input price up 200% to $0.18 per 1M tokens. The price cuts concentrate in the low-cost tier, where DeepSeek now sits at $0.06, undercutting most rivals.

Releases of the day

LiteLLM v1.97.0 ships signed Docker images using cosign, with every release verified by a key introduced in commit 0112e53. Users can now confirm image integrity before deployment. llama.cpp b10451 adds a check that LoRA tensor data stays within file bounds, closing a potential out-of-bounds read in src/llama-adapter.cpp. Lobe Chat v2.2.14 is a minor release auto-published from PR #18373, with changes detailed in the pull request description. ms-swift v4.5.2 fixes an incorrect size configuration for Qwen3.8-35B-A3B, which had been mistakenly copied from the Qwen3.5/3.6 series during development.

Pricing moves

Five models saw input price reductions today, led by GLM 5.2 at 61% and DeepSeek V4 Flash at 57%. The cuts bring DeepSeek to $0.06 per 1M tokens, the lowest among the affected models. Tencent's Hy3 preview rose 200% to $0.18, a notable outlier in an otherwise downward pricing day.

▼ 57%DeepSeek V4 Flash 0423DeepSeek · input $0.14 → $0.06 per 1M tokens
▼ 61%GLM 5.2Z.AI · input $1.19 → $0.46 per 1M tokens
▼ 4%Kimi K2.6Moonshot AI · input $0.56 → $0.54 per 1M tokens
▼ 31%GLM 5.1Z.AI · input $1.4 → $0.97 per 1M tokens
▼ 35%Qwen Plus 0728 (thinking)Alibaba · input $0.4 → $0.26 per 1M tokens
▲ 200%Hy3 previewTencent · input $0.06 → $0.18 per 1M tokens

New models

No new models were announced in today's release facts. The price movements apply to existing models DeepSeek V4 Flash 0423, GLM 5.2, GLM 5.1, Kimi K2.6, Qwen Plus 0728 thinking, and Hy3 preview.

Emerging projects

No new projects were detected beyond the 133 that entered the index today. The "speech-to-speech" category gained the most stars, but no specific project names were provided.

Research pick

No research papers were included in today's facts.

The bigger picture

The pricing moves show aggressive competition at the low end of the input token market, with DeepSeek and GLM leading cuts. The steady stream of minor releases across Lobe Chat, LiteLLM, ms-swift, and llama.cpp indicates continued maintenance momentum. The speech-to-speech star surge suggests growing interest in voice interfaces, though no specific project details are available.

Source: olud.ai tracking of 10,000+ open-source AI projects, 300+ models and live provider pricing. Figures are measured, not estimated. All releases · Live pricing · Latest in AI

← All AI news