Synthesized by Clarity (Claude) from 28 sources · May contain errors — spot one? mail@promitb.dev · Methodology →
Chinese Open Models Now Take 60% of US OpenRouter Tokens
- Sources
- 28
- Words
- 2,075
- Read
- 10min
Topics Agentic AI LLM Inference AI Capital
◆ The signal
DoorDash and Airbnb are running Kimi K3 and GLM-class models in production, not piloting them, per The Algorithmic Bridge. Treasury's Scott Bessent is floating sanctions on US adopters. The tradeoff is now explicit: the best available model against a dependency a policy move could strand overnight. Teams that know exactly which workflows sit on those models can price that risk. Teams that don't are guessing.
◆ INTELLIGENCE MAP
Intelligence map
01 Chinese Open Weights Are Already In Your Cost Base
act nowUS companies now route nearly 60% of their OpenRouter token usage to Chinese open-weight models, per The Algorithmic Bridge, with DoorDash and Airbnb running Kimi K3 and GLM-class models in production. Kimi K3 delivers near-frontier quality at roughly a third of frontier API pricing. Treasury's Scott Bessent has floated sanctions on US companies that use Chinese AI. The same line item now holds your largest margin win and your newest policy risk.
- Kimi K3 API cost
- Frontier gap
- Startups on open
- Anti-ban coalition
- Chinese open-weight60%
- Everything else40%
02 Controllability Becomes a Product Requirement
monitorReps. Lieu and Moran introduced the AI Kill Switch Act on July 23. It would require companies above $100M compute spend and $500M revenue to support government-ordered throttling, feature disabling, shutdown, and rollback, preserve model weights and telemetry, and report within 15 days, with fines up to $20M a day. Read it as a feature list, not a statute: enterprise buyers will ask for throttle, rollback, and audit logs long before enforcement exists.
- Compute threshold
- Revenue threshold
- Reporting window
- Jul 16OpenAI test agents breach Hugging Face
- Jul 23AI Kill Switch Act introduced
- If enactedThrottle, shutdown, rollback, weight retention
- PenaltyUp to $20M per day
03 Who Captures the Inference Savings
monitorAn AI startup can book $100M in revenue and pass $90M straight to model providers, per TLDR Founders' reporting, and cheaper models do not fix it because customers either demand the savings or bring their own inference. Agentic features consume 10-100x the tokens of chat. One open-source tool cut token usage a median 82x by sending structural context instead of whole codebases. Efficiency is now a pricing decision, not only an infrastructure one.
- Agentic token use
- Spot inference
- Token reduction
- AI revenue paid to model providers90
04 AI Defaults Need an Escape Hatch and a Legal Review
monitorGoogle Photos shipped a one-tap toggle back to classic keyword search after users revolted against irrelevant Gemini results. The same week, OpenAI opened ChatGPT Health to every US adult across free, Go, Plus and Pro tiers, one day after a Florida pastor sued over advice to skip seeing a doctor. Techpresso reports 70% of health queries already happen outside the dedicated health hub, so the product boundary you draw does not constrain what users ask.
- Health integrations
- Photos fallback
- Rollout tiers
- Health queries asked outside the health hub70
05 Machines Became the Majority Audience
backgroundOne development site logged 268,000 AI agent requests against 107,000 human pageviews over two months. ChatGPT-User drove 73% of that agent traffic and reads HTML almost exclusively, while Claude Code requested Markdown 76% of the time through an Accept header. The llms.txt file everyone recommends drew only about 37 fetches from named assistants. Serving each agent the format it asks for is cheap and measurable; the standard file is not.
- ChatGPT-User share
- Claude Code Markdown
- APIs built for agents
◆ DEEP DIVES
Deep dives
01 Your Cheapest Model Is Now a Foreign-Policy Dependency
act now evidence: highStart with what a founder actually did this quarter, not what the White House said. The accusation is that Moonshot AI distilled Anthropic's Fable to build Kimi K3. The Algorithmic Bridge lays out the calendar that argument cannot survive: Fable was available to users only June 9-12 and again from June 30, Kimi K3 shipped July 16, and K3 beats Fable on some benchmarks. A student model does not overtake its teacher in a two-week window. Anthropic separately settled its own copyright case for $1.5 billion. So treat the distillation charge as live policy risk, not as settled fact about provenance.
Why the usage is sticky. Separate the pitch from the thing being done. Chinese labs give the weights away and monetize hosted inference, so the model is free and the tokens are cheap. Newcomer reports Kimi K3 hitting near-frontier quality at roughly a third of frontier API pricing, with a 1M+ token context window, topping Moonshot's internal benchmark against everything except GPT 5.6 and Fable and leading on some coding and agentic tasks. TLDR Founders adds the adoption picture: open-weight models now sit in roughly 80% of startups and trail frontier capability by only 4-6 months.
Where the sources disagree, and why that matters. The Algorithmic Bridge puts Chinese open-weight models at nearly 60% of US companies' token usage on OpenRouter. TLDR Founders pegs open-weight models at 25-50% of volume across OpenRouter and Vercel. Both point the same direction, and the spread is the instruction: measure your own mix rather than paste either number into a planning doc.
The functional argument nobody priced in. During the Hugging Face compromise, OpenAI's and Anthropic's models refused to help because of safety guardrails, and Hugging Face ran GLM 5.2 on its own hardware to finish the forensics. For a product that touches security, operations, moderation, or incident response, refusal behavior is a requirement you test in a bake-off, not a philosophical position you admire.
Camp Position Effect on your roadmap Treasury (Scott Bessent) Floats sanctions and pressure on US companies using Chinese AI Adoption carries documented policy risk OpenAI policy (Dean Ball) Wants government to impose "large amounts of regulatory risk" and FUD on adopters Expect vendor FUD inside competitive deals Commerce Wants to fund US open-source alternatives A domestic hedge may arrive, unevenly Little Tech Alliance 200+ startups including Y Combinator, Replit, Proton, Yelp lobbying against bans An outright ban is not on the table today The cost floor falls either way. Nvidia's Jensen Huang used his first-ever X post to back open weights, and Microsoft claims its in-house models cut costs up to 89% versus OpenAI. That is a vendor claim, not a benchmark. Whether the cheap tokens end up Chinese, American, or in-house, a margin model built on last quarter's frontier pricing is already wrong.
The move is architectural, not political. The decision this week is not whether Chinese weights belong in the stack. The decision is a 2x2 on control: can you switch inference providers inside one sprint, and do you know which features stall if a lab lands on an entity list. One investment covers both the price war and the sanctions order.
A model you cannot swap out in one sprint is not a vendor choice. It is a policy bet.
Action items
- Produce a one-page model-dependency map by end of week: every AI feature, its provider, its share of inference spend, and its US or China exposure.
- Run a bake-off against Kimi K3, DeepSeek V4, and GLM 5.2 on your top three real workloads this sprint, scoring quality, latency, cost per 1,000 tokens, and refusal behavior.
- Get a written legal and compliance position on Chinese-model usage before your next roadmap review, and name one owner to flag any sanctions or entity-list action within 24 hours.
Sources:Alberto Romero from The Algorithmic Bridge · Newcomer · TLDR Founders · Ben Thompson · Techpresso · Matt Johansen
02 Congress Just Wrote Your Agent Controllability Spec
monitor evidence: highA procurement lead reads the AI Kill Switch Act and does not see policy. She sees a requirements list. Reps. Lieu and Moran introduced it on July 23 as an amendment to the Homeland Security Act. Covered companies would have to support government-ordered throttling of inference and compute, feature disabling, full shutdown, and rollback to earlier versions, while preserving model weights and telemetry and reporting within 15 days. Penalties run up to $20M per day, DHS gets emergency authority, and the thresholds are $100M in compute spend and $500M in revenue. Passage is uncertain. The buyer expectation is not. Procurement will ask for these controls well before any enforcement date.
Two of those requirements are harder than they read. Forensic telemetry means more than what teams keep now. SANS research notes that the standard 48-96 hour retention windows already miss correlated events. Rollback means a versioned, known-good model state, which most AI features cannot produce. Prompt templates, model versions, and tool schemas drift independently, and nothing pins them together as one artifact.
The connector story is the nearer-term risk. Zenity Labs reported through Bugcrowd on June 4 that a single crafted ChatGPT link could feed instructions into OpenAI's agent builder, auto-wire the victim's existing connectors across Outlook, Teams, Slack, SharePoint and Google Drive, switch off approval prompts, and publish a scheduled agent. That agent read attacker emails tagged "TASK," searched files, took passwords and API keys, and sent phishing as the employee. OpenAI patched it in four days by removing the offending URL parameter. Automatic connector wiring and skip-the-approval onboarding both live on someone's backlog right now. Teams file them as activation improvements.
A second failure mode belongs in the same epic. Agent memory can be quietly poisoned so a corrupted belief lies dormant and triggers harmful action later, which defeats filters that only inspect immediate inputs. Agents that remember need memory integrity validation and audit logging of their own.
Control Who is commoditizing it Your call Prompt-injection scanning Microsoft, natively in Defender for Office 365 before Copilot reads email Buy — do not rebuild the filter Agent runtime governance OpenBox, one SDK across LangChain, LangGraph, n8n, Mastra Buy if you are on those frameworks Secret leakage at the prompt Sequirly browser extension scanning prompts and uploads Point solution, likely absorbed by platforms Connector scoping and approval integrity Nobody — still your design surface Build; this is the differentiation window Why buyers care now. MIT Technology Review reports AI-generated errors reaching an official court transcript, reportedly the first known case, with judges also taken in by AI misinformation. The failure mode enterprises fear has shifted from "the model was wrong" to "a human trusted it anyway." That changes what a demo shows. Controllability and a measurable human-review catch rate become the thing on screen, and the disclaimer becomes the footnote.
The controls worth having early are the throttle, the rollback, and the audit log; they show up on procurement questionnaires before they show up in law.
Action items
- Scope a controllability epic this quarter covering throttle and pause, versioned rollback to a known-good model state, and telemetry retention beyond the current 48-96 hour window.
- Audit connector permissions on every agentic surface this sprint: least-privilege defaults, explicit user consent before wiring Outlook, Slack, Teams, SharePoint or Drive, and tamper-resistant approval state.
- Kill or gate any setup-via-link agent configuration flow behind authenticated confirmation before your next agent release.
Sources:SANS NewsBites · Cyberpresso · The Download from MIT Technology Review · TLDR InfoSec
03 The $90M Question: Who Keeps the Savings When Inference Gets Cheap
monitor evidence: highA founder books $100M in revenue and watches $90M of it route straight to Anthropic or OpenAI. TLDR Founders' number is the moment worth sitting with, because cheaper inference does not land in your margin by itself. Here is what teams tell themselves: when the model gets cheaper, the savings are ours. Here is what customers actually do: they demand the discount at renewal, or they bring inference they already pay for. The routing win then shows up on the customer's invoice, not your gross margin. Unless the price is anchored to something other than tokens.
The token curve is moving against you on exactly the features you are adding. Agentic harnesses — coding agents, autonomous multi-step workflows — consume 10-100x more tokens than conversational features. A COGS model built in the chat era is off by an order of magnitude for anything on the agent roadmap. Separate the thing being pitched from the thing being done: a market has already formed to arbitrage the gap. Spot and secondary venues now trade idle GPU capacity and unused credits at 20-80% discounts to retail inference pricing, which makes procurement a lever the finance partner can pull.
The efficiency lever that customers never see. An open-source tool, code-review-graph, achieved a median 82x reduction in token usage by feeding assistants precise structural context through Tree-sitter instead of re-reading whole projects, and it re-indexes a 2,900-file codebase in under two seconds. That gain never appears on the price list, which is precisely why it stays yours. The caveat is sharp. Prompt-caching savings are fragile. Adding a tool, reordering a schema, swapping a model, or rewriting conversation history silently erases them. Routine roadmap moves can destroy an inference budget without a single alert firing.
Lever What it changes Who captures the gain Route simple calls to a cheaper model Cost per call falls Your customer, if you price per token Context efficiency and caching Same output, fewer tokens You — invisible to the buyer Package the savings as a paid tier Cost control becomes a feature You — Cursor gated its router to Teams and Enterprise Price on usage, seats, or outcomes Revenue decouples from token cost You, structurally The pricing assumption under attack. Fugu-Ultra v1.1 shipped a capability upgrade across coding, agentic and reasoning work at the same price as v1.0. Users are being trained to expect smarter AI for free, so any tier priced on model quality is a depreciating asset. The counter-move is already on the board. Cursor kept its cost-savings router on Teams and Enterprise plans only, converting an efficiency gain into an upsell instead of a giveaway — and leaving individual and SMB developers unserved if you want that wedge. That is the tradeoff, named in the same breath as the play.
Buyers are formalizing the question. Cursor, Google Cloud and BMO convene at Harness's FinOps Excellence Summit on July 29 specifically to connect tokens to outcomes and push cost guardrails into deployment pipelines. When a bank shows up to demand cost-to-outcome proof, that proof becomes an RFP line. "We don't measure cost per outcome" becomes an answer you have to give in writing. The forcing function here is the RFP, not the summit.
If your price is a markup on tokens, every efficiency win you engineer belongs to your customer.
Action items
- Rebuild the cost model for every AI feature on the backlog this sprint as cost-per-active-user and cost-per-outcome, using 10-100x chat token volumes for anything agentic.
- Add an AI cost-regression check to the release gate so prompt-cache-breaking changes — new tools, model swaps, schema reorders — are flagged before they ship.
- Re-anchor at least one AI price point off model quality and onto usage, seats, or outcomes before your next pricing review this quarter.
Sources:TLDR Founders · TLDR Crypto · TLDR DevOps · TLDR AI · Devshot
04 Bots Outnumbered Humans 2.5-to-1, and They Read Different File Formats
background evidence: mediumTwo agents crawled the same server for two months and asked for two different things. Over that window, evilmartians.com logged 268,000 AI agent requests against 107,000 human pageviews, and the agents did not behave as one bucket. ChatGPT-User drove 73% of agent traffic and requests HTML almost exclusively. Claude Code asked for Markdown 76% of the time through an Accept: text/markdown header. That is content negotiation — serving a different representation of the same page based on what the client asks for — behaving like a product decision with a measurable outcome, not a formatting detail.
The null result is the money-saver. llms.txt, the file every AI-discoverability post recommends, drew about 660 fetches, of which only roughly 37 came from named assistants. A hidden hint div got exactly zero attributable follows. If llms.txt is on the content-ops roadmap, that is capacity to reclaim today and spend on format matching instead. The caveat: this is one property over two months, so treat the ratio as directional and instrument your own logs first.
This is a distribution channel, not a curiosity. One in four software developers now build APIs for AI agents rather than for humans. Amazon adopted MCP — the Model Context Protocol, an emerging standard for how agents call external tools — for Alexa+ third-party integrations. Datadog built its Cloud SIEM agent tooling on MCP with progressive disclosure specifically to reduce context consumption, and Xweather ships an MCP-ready API with 15,000 free monthly calls. Products with structured, machine-navigable interfaces get reached by that ecosystem. Products without them are invisible to it.
Segment Share of agent traffic What it wants Your move ChatGPT-User 73% Clean semantic HTML Fix HTML structure on high-intent pages Claude Code Coding-focused, significant Markdown via Accept header Serve Markdown by content negotiation llms.txt fetchers ~37 named-assistant fetches Nothing measurable Deprioritize entirely The human door is narrowing as the machine door widens. Google's AI search increasingly answers the question before the click, favoring creator and question-based content over branded pages, and Reddit threads are rising as a cited source. On the demand side, 55% of social users post less and 47% have deleted apps from stress. Read alongside the agent numbers, the audience mix behind the traffic charts is shifting even when the totals look flat.
The order of work is the decision. Segment first, because you cannot A/B a format for an audience you have not measured. Then negotiate format, a days-long change with directly observable behavior behind it. Then restructure high-intent content around real user questions, where both the AI answer engines and the agents pull from. None of this requires a model, a vendor, or a budget line.
Who reads your docs is now a segmentation question, and one of your two largest segments does not render CSS.
Action items
- Stand up server-side segmentation of agent versus human traffic within one sprint, broken out by ChatGPT-User, Claude Code, and other named crawlers.
- A/B content negotiation on your top documentation and product pages this quarter — clean HTML for ChatGPT-User, Markdown for Accept: text/markdown — and measure citation and referral lift.
- Cut llms.txt work from the roadmap and reallocate that capacity to format matching and question-structured content.
Sources:TLDR Dev · AI Breakfast · TLDR Marketing · TLDR DevOps
◆ QUICK HITS
Quick hits
Gen-Z business formation ratio narrowed from 26:1 to 4:1 since 2020
LLM evaluation guidance pulls safety checks out of averaged quality scores
Shopify swapped five React checkout widgets to Preact to fit a 64KB budget
Amazon fixed its gaming business with a Prime Video tab, not new games
Pronto won a 100-truck mining expansion by retrofitting Caterpillar and Komatsu fleets
Amazon shut its AGI Lab and lost David Luan plus a dozen Adept hires
Trump Media sold Truth API access to five trading firms at $100K a month
◆ Bottom line
The take.
Pick the single wrapper layer you can defend — provider swappability, provable cost per outcome, or user-facing control — and write it into one PRD as an acceptance criterion this sprint.
Frequently asked
- What's the first thing I should do if my product runs on Chinese open-weight models?
- Produce a one-page model-dependency map covering every AI feature, its provider, its share of inference spend, and its US or China exposure. This pulls the risk out of engineers' heads into a document your CFO and counsel can read. Without it you can't price the sanctions exposure or the margin opportunity, so also run a bake-off against alternatives to know your fallback path.
- Is the distillation charge against these models actually solid?
- No — the timeline doesn't hold up. The accusation is that Moonshot distilled Anthropic's Fable to build Kimi K3, but Fable was only available June 9-12 and again from June 30, K3 shipped July 16, and it beats Fable on some benchmarks. A student model doesn't overtake its teacher in two weeks. Treat it as live policy risk rather than settled fact, and get a written legal position before a buyer asks.
- If inference keeps getting cheaper, won't that just improve my margins?
- Not if your price is a markup on tokens — then every efficiency win flows to your customer, who demands the discount at renewal or brings their own inference. Re-anchor at least one price point onto usage, seats, or outcomes so revenue decouples from token cost. Also note agentic features burn 10-100x more tokens than chat, so a chat-era cost model is wrong by an order of magnitude.
- What controllability features does the AI Kill Switch Act require me to build?
- Government-ordered throttling of inference and compute, feature disabling, full shutdown, rollback to a known-good version, plus preserved weights and telemetry with reporting inside 15 days. Thresholds are $100M compute spend and $500M revenue, with penalties up to $20M per day. Even if it stalls, procurement will ask for throttle, rollback, and audit-log controls before any enforcement date arrives.
- Should I invest in llms.txt to get discovered by AI agents?
- No — it's a measured non-result, drawing roughly 37 named-assistant fetches and zero attributable follows in a two-month study. Spend that capacity on content negotiation instead: serve clean HTML to ChatGPT-User, which drives 73% of agent traffic, and Markdown to Claude Code, which requests it 76% of the time via an Accept header. Instrument your own logs first, since the ratios come from a single site.
◆ Same day, different angle
Read this day as…
◆ Recent in product
Keep reading.
- Airtable spun its agent platform out days before selling itself for $1.285B.
- Enterprise LLM Spend Doubled to $8.4B as Prices Fell 95%
- DeepSeek V4-Flash Hits 82.7 on Terminal-Bench, Up From 61.8
- AI Refusals and Timeouts Look Identical in Error Dashboards
- OpenAI's 80% Luna Cut Is Built to Reprice in Two Quarters
Spot an error? mail@promitb.dev