1M Context Models Weekly Benchmark - Sep 5 2026: 9 Models Compared at 500K-1M Tokens
September 5 2026 weekly benchmark of all 1M context AI models in production: DeepSeek V4-Pro, GLM-5.3, Kimi K3, Qwen3.8-Max, mimo-v2-pro, Gemini 3.7-Flash, GPT-5.6 Sol, MiniMax M3, Claude Sonnet 5. Tests retrieval accuracy, multi-hop reasoning, agent task reliability, and price-per-million-token at full context. This week: NVIDIA acquires Hugging Face - expect NIM API pricing changes for HF-hosted models within Q4 2026.
Frequently Asked Questions
What changed for 1M context models in the week of Sep 5 2026?
Three major events: (1) NVIDIA acquires Hugging Face for $12.93B - expect NIM pricing changes for HF-hosted models in Q4; (2) OpenAI announces GPT-6 Astra preview with native 1M+ context and multi-agent; (3) GitHub HydraFusion multi-agent framework enters preview. None of the existing 9 1M context models changed pricing this week; all stayed flat from Aug 27 baselines.
Which 1M context model is the cheapest right now (Sep 5 2026)?
DeepSeek V4 Flash at $0.14/$0.28 per 1M tokens remains the cheapest. GLM-5.3-Flash at 50% off extended through Sep 30 ($0.075/$0.25) is even cheaper but only has 128K context. For true 1M context, DeepSeek V4-Pro at $0.435/$0.87 (permanent 75% discount) is the value pick. Most expensive: Claude Sonnet 5 at $3/$15 per 1M.
Does NVIDIA acquiring Hugging Face affect API pricing for DeepSeek / GLM / Qwen?
Short-term: no. Hugging Face inference API pricing for DeepSeek V4 / GLM-5.3 / Qwen3.8 stays unchanged. NVIDIA NIM free tier also remains. Long-term (Q4 2026+): expect bundled pricing where NIM + HF Inference API get combined discounts for enterprise customers. TokenRhythm and OpenRouter (which route to HF endpoints) will likely renegotiate their aggregator rates.
Want more AI pricing comparisons?
Check our price changelog for daily updates on AI model pricing, new releases, and deprecations.