The number that matters is $0.14: the price per million input tokens of DeepSeek's V4-Flash-0731, released July 31 on Hugging Face under an ungated MIT license, with output at $0.28 per million, roughly a third of what the V4-Pro tier charges. The base is a 284-billion-parameter mixture-of-experts with 13 billion activated per token and a 1-million-token context window; the 304-billion figure on the model card includes the attached DSpark speculative-decoding draft module. The official line is that architecture and size are unchanged from the preview and that the gains come from re-post-training alone.

The independent measurement, and it is deliberately small: Simon Willison ran his standard pelican test through OpenRouter and got a disappointing pelican at default reasoning effort and, in his words, something much better with reasoning set to high. Artificial Analysis ranks the model ahead of MiniMax's M3, a 428-billion model, and frames it as possibly the best value-per-intelligence available. DeepSeek's own benchmark table is bigger and reads like marketing; the useful warning attached to it is MarkTechPost's, that agent scores are harness-sensitive, so independent runs may diverge. In this newsroom's shorthand: a lot is claimed, a little is measured, and what is measured is favorable. Running it yourself remains cheap at these prices.

MiniMax's H3, launched the same day, is the other half of the week. It is an omni-modal video model that reads text, images, video and audio as one unified context, and its real differentiator is native stereo sound: audio is a first-class input, so you can hand it a reference clip and ask the character in an image to sing it, and it emits stereo audio with the video instead of relying on a separate dubbing stage. Output is 2K at 4 to 15 seconds, integer durations only. Under the hood: a captioning scheme that distills roughly 100K tokens of source material to 4K, a new VAE it credits with a 4x gain in effective sequence length, and an in-context regeneration step instead of a bolt-on super-resolution module.

The caveats are the useful part. H3 today is reachable only through MiniMax's API and the Hailuo app; the open weights are promised 'in the coming days' and, as of publication, do not exist, which is exactly the kind of promise this newsroom prices at zero until it ships. The price claims are MiniMax's own (less than a third of mainstream models at 2K); third-party trackers put pay-as-you-go at about $0.13 per second, reported rather than confirmed. And the South China Morning Post, citing Artificial Analysis, has H3 leading in video editing while trailing Google's Gemini Omni Flash in text-to-video and sitting behind Seedance 2.0 in image-to-video. Two drops, one pattern: the interesting action in models this week is happening at the edges, open weights and video, not at the closed frontier.