AI News · September 18, 2026

DeepSeek V4 Flash price down, new KV cache compression paper

126 new projects

DeepSeek V4 Flash input price dropped 12% to $0.07 per 1M tokens, while DeepSeek V3 input price rose 23% to $0.32 per 1M tokens. NVIDIA raised Nemotron 3 Ultra by 5% to $0.63 and Nemotron 3 Nano by 20% to $0.06. Meta cut Muse Glimmer 30B by 14% to $0.3.

A new research paper from DeepSeek-AI on KV cache compression for DeepSeek-V4.1-Flash gained 36 upvotes on Hugging Face. The work targets input-heavy agent workloads, aiming to reduce prefill cost and large KV cache overhead.

This week 2,162 new projects entered the index and 100 new Hugging Face spaces were added.

Releases of the day

n8n shipped 2.39.8 with bug fixes for data-encryption keys and sub-execution cleanup. Cline released desktop-v0.0.32, fixing a launch failure that could prevent the app from reaching the workspace picker. Promptfoo 0.123.1 added bedrock auth refresh across HTTP adapters and support for Gemini 3.8 and Vertex Live. Pydantic AI v2.45.0 introduced TypeSafeModel for TypeSafe's Jev and passed through xhigh effort. Phoenix v20.14.0 added the Harbor benchmark for MCP and CLI tools and updated built-in model token prices.

Pricing moves

DeepSeek V4 Flash input price fell 12% to $0.07 per 1M tokens. DeepSeek V3 input price rose 23% to $0.32. NVIDIA raised Nemotron 3 Ultra by 5% to $0.63 and Nemotron 3 Nano by 20% to $0.06. Meta lowered Muse Glimmer 30B by 14% to $0.3. Users of these models face adjusted costs.

▼ 12%DeepSeek V4 Flash 0423DeepSeek · input $0.08 → $0.07 per 1M tokens
▲ 23%DeepSeek V3DeepSeek · input $0.26 → $0.32 per 1M tokens
▲ 5%Nemotron 3 UltraNVIDIA · input $0.6 → $0.63 per 1M tokens
▲ 20%Nemotron 3 Nano 30B A3BNVIDIA · input $0.05 → $0.06 per 1M tokens
▼ 14%Muse Glimmer 30BMeta · input $0.35 → $0.3 per 1M tokens

Research pick

DeepSeek-AI published "DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression" on Hugging Face. The paper addresses the rising cost of long-context computation for input-heavy agent workloads. It proposes methods to reduce prefill expense and shrink the KV cache, which has become a bottleneck as context lengths grow.

The bigger picture

Today's price moves show divergence across providers, with some models getting cheaper and others more expensive. The research focus on KV cache compression signals that efficiency in long-context handling is a key battleground. Together, these facts point to a market where cost and performance pressure are driving both pricing adjustments and technical innovation.

Source: olud.ai tracking of 10,000+ open-source AI projects, 300+ models and live provider pricing. Figures are measured, not estimated. All releases · Live pricing · Latest in AI

← All AI news