DeepSeek spent two years being the floor under everyone's prices. On Monday it stops.
The model landed first, quietly. At around 3am Beijing time the API docs flipped: the fingerprint behind deepseek-v4-pro became the 0813 build, 111 days after the preview, with no blog post and no changelog. Architecture untouched — 1.6 trillion total parameters, 49 billion activated, 1M context, 384K max output. Everything that moved, moved in post-training.
And it moved a lot. DeepSWE, DeepSeek's own software-engineering benchmark, went from 12.8 to 62.7. Cybergym 52.7 to 83.3. Terminal Bench 72.1 to 87.9. Those are all long-horizon code, terminal and tool-calling tasks — exactly where the preview fell over. On pure knowledge reasoning it still trails: HLE 42.7 without tools against Claude Opus 4.8's 49.8 and Fable 5's 53.3.
Then came the invoice. DeepSeek is introducing peak and off-peak billing, off-peak at half the peak rate, new prices live at 16:00 UTC on 16 August. Reuters read the statement as rates ranging from 50% to 1,100% above current prices depending on model, token type and hour. A developer's tabulation on Hacker News put V4-Flash peak output at $1.32 per million against $0.28 today.
Look at when peak is. 01:00–04:00 and 06:00–10:00 UTC — the Chinese working day, morning and afternoon. Cache reads take the worst multiple, and cache reads are what agent loops burn. That is not margin expansion; that is load-shedding by price, from a lab that started hiring electrical and HVAC engineers for its own data centres two weeks ago.
Why now is in the leaked minutes. ChinaTalk has been working through notes from a four-hour meeting between Liang Wenfeng — the quant fund manager who founded DeepSeek and is now, on paper, richer than Dario Amodei or Sam Altman — and his investors. Enterprise revenue this year in the hundreds of millions of dollars. Net profit soon. Then an IPO.
A company priced for a listing cannot subsidise the planet's inference. The February 2025 disclosure that R1's API earned $562,027 a day at a 545% margin was a flex; it is now a floor to defend. Developers on V2EX spent the day joking that the rheostat needs a new setting and complaining that models were shipped and pulled, a mess. The cheap era was a customer-acquisition budget, and it has been spent.
What Beijing is signalling
Washington published 25 pages instead. The White House Office of Trade and Manufacturing Policy released "The Great Transshipment Scam", alleging Chinese exporters have laundered origin through more than 40 third countries since 2018 via relabelling, re-invoicing and minor assembly. It estimates the annual value at $40bn to $303bn and, assuming $75bn, claims 450,000 displaced American jobs.
The targets are the middlemen. Korea, Japan, Taiwan, the EU, India, Mexico and Canada are filed as "diversified scale leaders"; Vietnam, Malaysia, Thailand, Turkey, Brazil and Indonesia as deeply China-integrated ones. Nothing in today's mainland feed touches it — the Chinese press is on Zhu Rongji and V4-Pro. Xi arrives in Washington next month with this on the table and the investment board still unbuilt.
Nuclear is the AI answer nobody can sanction. CCTV Finance puts 57 reactors under construction and 135 GW operating-plus-approved, the largest build programme in the world. Eight units were approved in late July — the first batch of the 15th Five-Year Plan. Nuclear is under 2% of installed capacity but nearly 5% of generation, and above 15% in Guangdong, Fujian, Liaoning and Hainan.
Set that against yesterday's number: 800 billion kWh of compute electricity demand by 2030. Reactors approved now come online in the 2030s. Beijing is answering a question about tokens with twenty-year concrete.
And state TV is now bragging about exploit agents. A Fudan University team's Whitzard agent scored 91.2% on CyberGym, second globally and first among universities, ahead of Anthropic's security-focused frontier model and teams from Microsoft and Google, across 1,507 real vulnerabilities in 188 large open-source projects. One base model, path reachability recast as reasoning constraints, total inference cost under 5,000 yuan (~$700). Framework and benchmark open-sourced, and CCTV News carried it.
The AI race
Tencent's free cash flow went negative, on purpose. Q2 revenue 204.79bn yuan (~$28.8bn), up 11%; capex 52.78bn yuan (~$7.4bn), up 176% year on year and 65% on the quarter; R&D up 35%. The market hated it — shares closed down 4.46% the next morning.
The interesting claim is what the money buys. 36Kr's read is that Tencent is hoarding high-spec memory as much as accelerators, with Samsung, SK Hynix and Micron effectively sold out and expansion running one to three years behind demand. Hunyuan's next generation is meant to be bigger. Racing the clock on DRAM is a stranger bet than racing it on GPUs, and a more honest one.
DeepSeek shipped Harness the same day. A developer preview, one npm install, everything a plugin, built on an existing open-source plugin framework. What US developers on Hacker News fixed on was the raw reasoning trace being visible where American labs obfuscate theirs.
The leaked shortlist is the real strategy. Of 17 accepted ecosystem projects, about 70% are near-zero-star personal repos — sandboxes with credential injection at the gateway, cost-aware model routing, persistent "failure notebooks", multi-agent schedulers. Famous high-star projects were rejected; medical, legal and financial verticals are under 4%. DeepSeek is buying the execution layer and leaving applications to everyone else.
Meanwhile, the least glamorous work of the week. The FlagOS community got Alibaba's 2.4-trillion-parameter Qwen3.8 running day-zero on nine different accelerator families — Nvidia plus Ascend, Moore Threads, Kunlunxin, Hygon, MetaX, Enflame and others — largely by adding INT8 quantisation paths, because more than 80% of China's installed AI capacity cannot do FP8. A frontier model reaching hardware bought two generations ago is worth more than a leaderboard position.
The vibe
The Unitree lottery is the story of the week in China, and it isn't about robots. The humanoid maker's STAR Market allotment rate finished at 0.01809759%, a record low, with more than 9.78 million accounts applying for 19,414 winning numbers of 500 shares each. One Chengdu winner told reporters he expects about 200,000 yuan (~$28,000) and will pay cash for a Xiaomi car.
The detail that lands is what he won't do. He won't post it to WeChat Moments — afraid colleagues will resent him and relatives will ask for loans. Also in the book: Liang Wenfeng's two funds took the largest private-fund allocation by account count, 198 products for about 38.94m yuan (~$5.5m).
Someone in Hangzhou built a flying skateboard. Kufei launched a standing single-person eVTOL — no seat, no cabin, a flattened black platform with caged rotors, about 11 minutes aloft with a 70kg pilot — plus a seated model doing 100 km/h for 40 minutes, on a dual-redundant in-house flight controller with three aviation-grade IMUs per channel. It flew live at the Spring Festival Gala's Harbin stage at minus 30.
The route to market is American. Roughly 100 paid deposits overseas, aimed at FAA Part 103 ultralights, which need no airworthiness certificate and no pilot licence — plus a package from PICC and CPIC covering aircraft, occupant and third parties. A Chinese consumer aircraft finds its first legal home in a 1982 US rule.
Read differently
One release, two entirely different stories. Chinese tech press led on benchmarks and explicitly reported the price as unchanged at 3 yuan in, 6 yuan out per million (~$0.42/$0.85) — QbitAI called it a rehearsal of the 2025 Spring Festival. English-language developer forums skipped the benchmarks and went straight to the hike, doing substitution arithmetic against OpenAI's Luna price cut and noting DeepSeek is still roughly twice as slow.
Only the Chinese write-ups tested it properly. GeekPark found thinking mode consuming 13,394 of 16,384 output tokens on a pathfinding task and truncating the HTML mid-file; with thinking off, 715 clean lines in 54 seconds. The chain of thought itself has degenerated into clipped telegraphese to save tokens. Great for agents, probably fatal for prose.
Threads we are pulling
- Called it, and fast. Yesterday's note on a registered Harness account and data-centre job ads; the developer preview shipped within the day.
- Still missing: Alibaba's promised 27B dense companion to Qwen3.8. The multi-chip port covers only the 2.4T text checkpoint.
- Lin Junyang's new lab is confirmed as an angel round at a $2bn valuation, co-led by Gaorong and HSG with Tencent and a Shanghai future-industries fund.
- Zhu Rongji: tributes, including a personal one from Malaysia's Anwar Ibrahim, but no funeral grade yet. Who eulogises is still the thing to watch.
- Insta360 outgrew its own category: revenue approaching 10bn yuan against a global panoramic market Leiphone puts under 7bn. Which explains the pivot.