Beijing asked for a truce through 2029 and got eleven weeks
Eleven weeks of truce, not three years
The Busan truce got a two-month extension. Treasury Secretary Scott Bessent said on Fox News that the pause due to expire 10 November now runs to 10 January, announced after an unscheduled second meeting in four days with vice-premier He Lifeng, Beijing's top economic negotiator. China had wanted it stretched to January 2029, when Trump's term ends, per SCMP.
Two months is a leash, not a deal. Bessent said some deliverables "have not been perfect on the Chinese side" and that the coming months would show whether Beijing enacts more of the agreement. Scott Kennedy of CSIS read the short window as Washington keeping the heat on — with the bonus that it makes Xi likelier to attend the G20 in Miami in December, via CNBC.
Chinese outlets had a different Wednesday. State media did not immediately pick up Bessent's truce remarks, CNBC noted; the Chinese feed today led with a table-tennis gold in Nagoya and a Geely SUV at a limited-time 92,900 yuan (~$13,100). The summit copy that ran was tarmac copy: first lady Peng Liyuan talking to Melania Trump in English without an interpreter.
The comparison being drawn in Asia is with Japan. Sanae Takaichi, Japan's prime minister, arrived at JFK to a near-empty apron before her first UN General Assembly address; Xi got Trump at the foot of the plane. SCMP says the split screen has started an argument about how Washington treats an ally versus a rival.
Go deeper on this section: ClaudeChatGPT
Anthropic's ban wave hit the foreign clouds
Anthropic ran a new compliance sweep last week. Large numbers of Claude accounts belonging to AWS and Microsoft Greater China customers were suspended, in a round wider than earlier ones, multiple people told Leiphone. Foreign-cloud sales staff in China had been getting contacts in Africa and the Americas to open overseas accounts; the new detection catches accounts whose actual user sits in a restricted place.
It landed on the month-end invoice. Some customers have said they will not pay for services they can no longer use — which Leiphone's sources put at a potential revenue hit in the hundreds of millions of dollars for the two vendors' Greater China businesses.
This is stage three of one policy. Anthropic's September 2025 terms barred any company more than 50% owned by entities headquartered in unsupported regions like China, as SCMP covered then. In July the FT reported it was closing the workarounds: Singapore subsidiaries, cloud fronts, reimbursed personal subscriptions, and relay "transfer stations" detectable from signals like a machine's time zone, summarised here.
The account-farming side shows up on the forums. A V2EX poster today asks whether shifting a batch of Pro and Max accounts — registered over roaming on Japanese and Singaporean prepaid SIMs, all with clocks set to Shanghai — onto cheaper Japanese exit nodes would kill the lot at once, here. Sentiment, not proof of whose accounts.
What Beijing is arguing out loud, meanwhile, is interdependence. Zheng Yongnian, dean of the school of public policy at Chinese University of Hong Kong, Shenzhen, warned before the summit that AI competition must not harden into cold war and that full decoupling would be difficult, in SCMP. The decoupling is proceeding at the billing layer anyway.
Go deeper on this section: ClaudeChatGPT
DeepSeek doubled revenue, then deprioritised it
The run rate is $1bn. The Information reports DeepSeek's annualised revenue has more than doubled from under $500m a few months ago — a number Liang Wenfeng, the former hedge-fund manager who founded the lab, gave at a recent investor meeting. August's peak-hour price rise on V4-Pro made usage 2.3 to 4.5 times costlier and cost it no customers, via Zhidx.
The raise behind it is 50bn yuan (~$7bn) at a 500bn (~$70bn) valuation, due to close by end-October, with STAR Market listing prep running alongside. June's round was also about 50bn, at nearly 400bn (~$56bn) post-money: Liang personally 20bn (~$2.8bn), Tencent 10bn, battery maker CATL 5bn. API gross margin is 82.9%, above Anthropic's and OpenAI's.
And Liang says revenue still isn't the first priority. Over 70% of compute goes to training, under 30% to serving. Internal tests show the smaller models run acceptably on gaming GPUs and cover most everyday queries — push inference onto consumer silicon, keep the good cards for training runs. Huawei may begin delivering training chips in the fourth quarter.
Go deeper on this section: ClaudeChatGPT
Xiaomi previewed V3's attention two days later
MiMo-V2.6 shipped Tuesday; Wednesday night brought the V3 architecture. Luo Fuli — head of Xiaomi's MiMo team, previously a core DeepSeek researcher on R1 — is corresponding author and team lead on a 15-author paper on HySparse2. At 1m tokens of context it cuts prefill compute to roughly a fifth of the hybrid sliding-window scheme in V2.6, and KV cache from 12.09GB to 2.69GB, per Zhidx.
The trick is letting prefill quit halfway. Full-attention layers in the second half build their keys and values from the first half's hidden states, then hand that cache down to several sparse layers — on a 49-layer example, only the first 25 run to prefill a long input. Sparse selection moved from blocks to individual tokens: RULER-v2 at 256K goes 35.74 to 58.45.
Which is an agent-economics paper wearing an architecture hat. Tool returns are long and model actions are short, so the expensive step in a multi-turn run is re-reading what the tools just dumped in. Four of the references are DeepSeek papers, V4 and V4.1-Flash included.
Go deeper on this section: ClaudeChatGPT
Chinese labs are redoing pretraining over dirty data
Several are restarting because the corpus was bad. Leiphone's account, sourced to pseudonymised insiders, describes a lab that spent a year on a trillion-parameter model, was outperformed by models a hundredth its size, and traced it to dirty data; and a mid-tier firm that bulk-loaded scraped technical blogs, found mid-run that much of it was machine-generated pages repeating one paragraph thousands of times, and scrapped the version, here.
The budget split is the whole explanation. A rough industry breakdown quoted in the piece: of every 100 yuan of training spend, 40 goes to compute, 30 to talent, 20 to marketing, 10 to data. A supplier says expert-written items fetching $10,000–20,000 from US buyers get 1,000–2,000 yuan (~$140–280) domestically. Some vendors quietly run token relay services to harvest real request traces. Claim, not finding.
Go deeper on this section: ClaudeChatGPT
Xianyu says the worst screenshots weren't its
A Hebei broadcaster's investigation put Xianyu at the front of a prostitution funnel. Listings on Alibaba's second-hand marketplace advertised home fitness, model shoots and swimming lessons; one purchase delivered a video telling the buyer to scan into an outside group, where profile cards offered women "available nationwide" — some minors, the youngest 14 — per a TMTPost piece.
The platform's reply was a boundary. Xianyu said the incriminating chat logs came from other platforms, reported the matter to police, and called the claim of on-platform obscene material involving minors untrue — while conceding it is one of the funnels and bears responsibility. Of ten listings traced, six were caught by daily patrols before broadcast, three only after.
The low barrier that built it subsidises what runs through it. 234m monthly users in June, daily gross merchandise value above 1bn yuan (~$141m) since 2024, no deposit or business licence to list, private chat by design, and a shop commission raised from 0.6% to 1.6% in April. Commission approaching a real marketplace's; enforcement that still needs a TV crew.
Go deeper on this section: ClaudeChatGPT
Nobody is taking their annual leave
Enterprise employees averaged 48.2 hours a week in August, per the statistics bureau, while a recruitment-site survey has nearly 70% of respondents leaving annual leave unused and 42% working through the leave they take. Statutory entitlement is 5, 10 or 15 days by tenure. The TMTPost essay's point is sharper than a complaint: efficiency gains only decide how long a task takes, never who keeps the hours saved.
Go deeper on this section: ClaudeChatGPT
Threads we are pulling
- I was wrong about the CEOs. Yesterday I said Chinese business leaders probably would not travel at all. A small group holding US visas flew to Washington separately from Xi's official party and is still waiting to learn whether it gets seats at tonight's state dinner, per SCMP. Beijing's security protocol, not the guest list, is the reported reason they came on their own.
- Monday's Huawei-silicon claim got a delivery date. Following Liang Wenfeng's "must succeed" line on training on Huawei chips: The Information now puts first training-chip deliveries to DeepSeek as early as the fourth quarter. Separately, 36Kr confirms the CFO hire I reported Wednesday — Yan Wentao, born 1991, a Hillhouse partner who previously worked at Tencent Investment and H Capital.
- Alibaba's 20GW got a segment-level bill. Following yesterday's capex item: the AI cloud and compute unit booked 48.44bn yuan (~$6.8bn) of revenue and 5.63bn (~$790m) adjusted EBITA last quarter, while the model-and-apps unit lost 13.86bn (~$1.95bn) on 3.34bn of revenue — the Qwen loss is a customer-acquisition line. Goldman puts in-house T-Head silicon at 10% of Alibaba Cloud compute now, targeting 50%, per TMTPost.
- Two small domestic-compute receipts. Following the 17 September GPU half-years: Moore Threads says its MTT S5000 now runs full inference for Protenix-v2, ByteDance Seed's open biomolecular structure model, per IT Home. And iFlytek's Spark-ASR-2.0, trained entirely on domestic accelerators, claims 202 dialects without switching and 10% higher inference cost than its predecessor, per Zhidx.
Go deeper on this section: ClaudeChatGPT