Synthesized by Clarity (Claude) from 35 sources · May contain errors — spot one? mail@promitb.dev · Methodology →
Big 3 AI Labs' 8x Token Premium Is a Price Level, Not a Moat
- Sources
- 35
- Words
- 2,108
- Read
- 11min
Topics AI Capital LLM Inference AI Regulation
◆ The signal
The premium everyone has been treating as a moat is a price level, which is a less comfortable thing to own. Cursor's router matched frontier output quality at roughly 60% lower cost, and OpenRouter is now fielding multibillion-dollar takeover interest, which suggests someone with a balance sheet agrees that routing is where the margin went. This is probably too early to call, but anything in the book whose margin or revenue passes through frontier per-token pricing deserves re-underwriting now rather than after the next price cut.
◆ INTELLIGENCE MAP
Intelligence map
01 The Model-Layer Rent Is Being Routed Away
act nowVercel's AI Gateway data, surfaced by Exponential View, shows OpenAI, Anthropic and Google capturing 90% of spend while serving 52% of tokens — roughly 8x the revenue per token everyone else earns. Cursor's router reproduces frontier-perceived quality at about 60% lower cost, and The Information reports OpenRouter fielding multibillion-dollar takeover interest. Every position whose gross margin runs through frontier API pricing needs a sensitivity table before its next mark.
- Big 3 token share
- Big 3 spend share
- Router cost delta
- Big 3 token share52%
- Big 3 spend share90%8x rev/token
02 AI App-Layer Multiples Inverted
monitora16z disclosed marks that invert the usual pattern: Harvey sits at $11B on hundreds of millions of ARR (~25–45x), while Decagon is $4.5B on an eight-figure base after 18 months (~50–150x), having tripled in under six months. The same piece describes founders discounting steeply — and sometimes paying customers to adopt — for logo revenue that 'barely registers.' At 50–150x, a 30% error in revenue quality is a catastrophic pricing error in your comp tables.
- Harvey
- Decagon
- Hebbia penetration
03 Nvidia Moves From Supplier to Royalty Holder
monitorThe Information reports Nvidia will take a cut of some customers' cloud revenues — converting a hardware sale into a perpetual claim on downstream revenue. In the same cycle, The Information Briefing notes Nvidia fell 5% on a $500B SK partnership that was substantively a re-run of early-June announcements and rests on letters of intent only. The reaction function flipped: the market now charges for AI capex promises instead of paying for them.
- Headline partnership
- Binding status
- Korea AI cloud
- Sept 2025Nvidia-OpenAI 'landmark' LOI signed, abandoned within months
- Early June 2026Ohio campus financing first reported
- Current period$500B SK headline re-run; Nvidia falls 5%
04 General Scaling Erases Two Claimed Moats
backgroundImport AI reports Claude Opus 4.7 completing a MirrorCode task in 14 hours for $251 of inference — work Epoch AI and METR estimate at 2–17 human weeks. Separately, Anthropic's Project Fetch had Opus 4.7 finish quadruped robot tasks autonomously in 9 minutes 35 seconds against a 181-minute human-plus-AI record, with Anthropic stating this was not the result of any robotics-specific effort. Two diligence templates break: code as a moat, and proprietary teleoperation hours as a moat.
- Unsolved targets
- Sunday Robotics
- Robot task speedup
05 Security Capital Rotates From Prevention to Access Architecture
monitorVerizon's 2026 DBIR ties ransomware to 48% of breaches, and CSO reporting has all four leading perimeter vendors — Palo Alto, Fortinet, Citrix and Check Point — under active campaigns, with a CVSS 9.3 Check Point SmartConsole flaw granting unauthenticated admin. CyberScoop reports Senator Wyden asking CISA for a binding two-year deadline to eliminate internet-facing legacy federal VPNs. Appliance-renewal durability and mandate-dependent compliance ARR both deserve a haircut this quarter.
- Perimeter vendors hit
- Proposed VPN deadline
◆ DEEP DIVES
Deep dives
01 The Routing Layer Just Got Its Price Discovery Event
act now evidence: highWhat a router actually owns
A router owns no model. It owns the model-selection decision, which is where the spread between token volume and token spend gets captured, and along the way it accumulates something no single lab can replicate: a per-model, per-task performance table. Devshot's reporting puts numbers on how valuable that table is. Matching edit format to model swings agent success from 66% (DeepSeek, unified diff) to 94% (Doubao, JSON Patch). Reviewer pairing is asymmetric in the same direction. Claude checking Codex lifts pass rates from 71.6% to 89.7%, while Codex checking Claude drops accuracy from 91.4% to 82.8%. Those are not tuning footnotes. That is the routing table, and routing tables compound with usage.
Two independent cost proofs, one bid
The substitution here is measured rather than theoretical, which in this market is rarer than it sounds. Cursor's Auto Intelligence mode produces output users judge as good as the premium model at roughly 60% lower cost, per Exponential View's read of Vercel gateway data. Cursor separately cut a browser-building experiment from $10,000 to $1,300, about 87%, by routing routine work to cheap execution models behind a frontier planner. And Lenny's Newsletter reports a blind seven-model practitioner benchmark in which Claude Sonnet 5 scored 77 against Opus 5's 78, which is a vendor's own mid-tier model within one point of its flagship.
Then the price discovery. The Information reports OpenRouter fielding multibillion-dollar takeover interest. Strategic acquirers are paying for demand aggregation across models, cross-model performance data, and the switching friction that appears once a router sits in the critical path of production inference. The market spent two years dismissing that as a moat argument.
Where the sources genuinely disagree
The deflation story is not clean, and the disagreement is the useful part. Anthropic shipped Opus 5 at unchanged pricing of $5 per million input and $25 per million output tokens while claiming efficiency gains, passed none of them through as a price cut, and opened a paid latency tier at roughly 2.5x speed for 2x the rate. That is price discrimination, not a price war. ChinAI reports Moonshot going the other way entirely: Kimi K3 launched at $2.30 per million blended tokens, a 3.5x output-price increase over its own predecessor, thirteen times DeepSeek V4 Pro's $0.18.
If frontier prices stop falling, the automatic COGS tailwind penciled into your AI application models disappears — and the only remaining path to software-grade gross margin is architectural.
Two readings survive, and this is probably wrong, but the first is the one to underwrite. Either capability parity is converting into pricing power at the model layer, in which case app-layer margin expansion has to be engineered rather than waited for. Or Moonshot's increase is a capacity constraint wearing a strategy costume, since it suspended subscriptions within 48 hours of launch because demand exceeded compute. Allocation looks identical under both. Durable margin sits with whoever decides which model runs, not whoever trained it.
The reflex to avoid
The cheap version of this trade funds another gateway. The expensive version keeps crediting 'we use the best model' as defensibility while the routing question goes unasked in diligence. Founders who have already tested output parity at 60% traffic diversion are structurally cheaper to scale, and the re-architecture money not spent on them later is money available for the next position. Founders who have not have just told you something about their rigor.
Action items
- Commission a routing-adjusted gross-margin case on every inference-exposed position this month, modelling a premium collapse from 8x to 3x revenue per token and 50%+ of volume diverted to cheap execution models.
- Add one gating diligence question to every AI application memo: what is gross margin if 60% of traffic routes to non-frontier models, and has output parity been tested with human scoring?
- Map five routing, model-selection and eval-observability teams for first meetings this quarter, prioritising those holding proprietary per-model performance data over integration surface.
Sources:Azeem Azhar, Exponential View · Devshot · The Information · TLDR Marketing · AI Breakfast · ChinAI Newsletter
02 Harvey at $11B, Decagon at $4.5B: The Multiple Went Backwards
monitor evidence: mediumRead the marks as a statement of what is being bought
a16z published a go-to-market framework, which is fine, but the investable content is the valuation data stapled to it. Harvey: $11B on hundreds of millions of ARR, call it 25–45x. Decagon: $4.5B on an eight-figure base built in 18 months, call it 50–150x, after tripling in under six months on the back of 100+ new enterprise customers in 2025. The company with less revenue carries the richer multiple. What is being priced is not the revenue base. It is logo velocity as a proxy for future net revenue retention.
That is precisely the bet a16z's own failure list warns about, which is a pleasant sort of honesty. 'Dying of indigestion', waking up with 200 customers and 50 underwater, is the specific failure mode of a velocity multiple. The non-linearity is what matters for recovery scenarios: 50 unhappy customers is churn, 500 is a reputation problem in a market where references travel faster than renewals.
The revenue-quality problem is described from inside the category
The uncomfortable disclosure sits mid-post, where such things usually sit. Founders are burning through rounds courting first Fortune 100 customers, offering steep discounts and in some cases paying customers to adopt, for revenue that 'barely registers.' Roughly 500 marquee accounts are being pitched by effectively every AI startup, and those accounts extract concessions because they can see the seller is desperate. If that is happening at scale, a meaningful slice of the ARR in the pipeline is discounted, pilot-stage or non-recurring revenue wearing an ARR label.
At 50–150x revenue, a 30% error in revenue quality eats the entire return.
The saturation risk on the other playbook
The lighthouse cohort has the mirror problem. Hebbia is already past 40% of the largest asset managers by AUM, including KKR and BlackRock. Deep penetration of a concentrated, status-legible market is exactly what makes those logos travel, and it also means over half the addressable cohort is already consumed. The counter-thesis is that those references open the mid-market at near-zero cost, and that is genuinely possible. A terminal-growth assumption on a lighthouse asset still requires an explicit down-market or international plan with unit economics attached, not a TAM slide.
The same tell shows up one layer over. The Information reports Mercor's growth concentrated in the largest AI labs, growth funded by a handful of counterparties who could each build the capability in-house. That is a services business wearing SaaS clothing. It deserves a customer-concentration haircut and an insourcing scenario, not a software multiple.
Where the competition is thinnest
The framework's most useful output for sourcing is geographic rather than strategic. Against those 500 contested accounts sit roughly 50,000 companies outside normal venture networks, or rather the more interesting version of that number: the manufacturers and distributors in Ohio and Texas nobody bothers to call. Michigan is on the list as well, and no associate is flying there this quarter. That opacity is the mispricing, and most sourcing funnels structurally cannot see it. Caveat worth holding: every positive case study in the piece is a16z-affiliated, and the operating metrics cited are vendor-supplied and unaudited.
Action items
- Require ACV distribution, discount-to-list, services-versus-recurring split, top-five customer concentration and pilot-to-paid conversion in every AI app-layer data room going forward.
- Re-underwrite every late-stage lighthouse asset in the book this quarter against an explicit cohort-saturation case, requiring a costed down-market or international expansion plan before crediting terminal growth.
Sources:a16z · The Information
03 $251 of Inference and a 9-Minute Robot: Two Diligence Templates Just Broke
background evidence: mediumStart with the eight targets nobody finished
Import AI reports Claude Opus 4.7 completing a MirrorCode target in 14 hours for $251 of inference, against Epoch AI and METR estimates of 2–17 human weeks, or roughly $11.5k–$98k of fully-loaded senior engineering labour. The benchmark asks for reimplementation of real software — Apple's 61,000-line pkl configuration language among them — from command-line access alone, no source code, no web.
Everyone will quote the arbitrage. The screen, or rather the more interesting version of the screen, sits in the residual. Eight of 25 targets were never solved to 100%, and the failures cluster diagnostically: ruff (Python linter), giac_subset (mathematics), mailauth (email authentication). What they share is dense implicit specification — correctness defined by thousands of undocumented conventions and standards interactions that black-box observation will not surrender.
Which converts neatly into a grading rule for app-layer software. High specification density (tax and regulatory compliance, clinical protocol, payments and settlement rails, security primitives, standards-heavy interoperability) keeps its moat, because correctness is unobservable from outside. Low density does not: CRUD workflow tools, thin orchestration, single-function utilities. If a model can probe the API, it can rebuild the product for the price of a lunch.
The robotics moat moved, and two independent sources say so
Anthropic's Project Fetch Phase Two is the second break. Opus 4.1 could not perform the quadruped tasks at all in August 2025; by May 2026 Opus 4.7 completed all but one autonomously in 9 minutes 35 seconds against a 181-minute human-plus-AI record. Anthropic's own framing is the load-bearing part: 'this progress is not the result of a concerted effort to improve the robotics capabilities of our models.' Sunday Robotics reports the same mechanism from the startup side — scale pretraining, then hill-climb with minimal in-house data — at 99.1% success across 778 folds and nine garment types, with a family beta committed for Fall 2026.
Two independent confirmations of one causal claim is a thesis, and the thesis contradicts the most common robotics moat slide in venture: proprietary teleoperation hours at volume.
What this changes in the diligence template
Retire 'teleoperation data hours owned' in favour of base-model access, post-training loop velocity, and a real-world recovery-data flywheel, then re-score the robotics deals already passed on the old criterion. The opportunity cost is the point: capital committed to buying teleoperation volume is capital not committed to the loop. And note the overhang, which sits on every robotics cap table whether or not anyone discloses it — frontier labs are now latent robotics platform players with zero robotics capex.
Two caveats belong in the memo, and this thesis is probably wrong in at least one of them. Demo tasks are not commercial reliability, and MirrorCode is flagged by its own authors as 'perhaps a little too easy' with 17 of 25 targets achieving perfect runs at release. Sunday concedes the residual gap is long-tail failures only visible after repeated real-world runs. The Fall beta will show them or it will not.
Action items
- Grade every app-layer deal in the active pipeline high, medium or low on specification density before the next investment committee, using the unsolved benchmark cluster as the reference for what remains hard.
- Rewrite the robotics diligence template this quarter to score base-model access and post-training loop velocity, and re-score deals previously passed on teleoperation-data grounds.
Sources:Jack Clark from Import AI · Lenny's Newsletter
04 Nvidia Extends Its Chokepoint Into Its Customers' P&L
monitor evidence: mediumThe structural item nobody indexed
The Information reports that Nvidia will take a cut of some customers' cloud revenues, which sounds like a footnote and is actually the business changing underneath the footnote: from selling an asset that depreciates over roughly six years to holding a royalty on the downstream revenue that asset generates. If it standardises, every GPU-cloud model built on buy-once-monetise-for-six-years is wrong at the gross-margin line, and terminal value resets across neoclouds, colocation platforms and GPU-collateralised lending.
Nobody has published the contract language. That is the diligence gap worth closing with portfolio CFOs directly, rather than inferring it from a headline that was written to be read quickly.
The reaction function flipped in the same cycle
The Information Briefing has the tape. Nvidia fell 5% on Monday on a $500 billion SK partnership headline and reporting about financing OpenAI's Ohio campus, and neither item was substantively new — the SK announcement largely reprised two early-June announcements, the Ohio financing was reported in early June. The one genuinely new fact, the Korea AI cloud doubling from 1GW to 2GW, was bullish and arrived inside a bearish tape.
Six months ago this headline adds five percent. In this cycle it cost five percent. The public comp set that late-stage AI infrastructure marks lean on has already repriced; the marks have not.
Price the instrument, not the number
Nvidia and SK have signed letters of intent only, and Nvidia declined to say which party contributes how much of the $500 billion, which is the interesting part rather than the number. The base rate is not theoretical: Nvidia and OpenAI signed an LOI in September 2025 for a 'landmark strategic partnership' and abandoned it within months. Any pipeline narrative resting on a non-binding agreement with Nvidia, a hyperscaler or a sovereign entity is carrying an option priced as a commitment.
The confirming cash number
Turing Post supplies the constraint that makes the rest legible. Alphabet posted its first-ever negative quarterly free cash flow of -$5.9B, even with Google Cloud revenue up 82% to $24.8B, against $195–205B of 2026 capex. AMD, meanwhile, is investing up to $5B in Anthropic as Anthropic commits to up to 2GW of MI450, and landed Anthropic and Azure in the same week. Vendors financing their own customers is a late-cycle demand-inflation marker; two hyperscaler wins for the second supplier is a genuine break in the single-accelerator narrative.
The honest counter-read: vendor financing may simply be cheap insurance on genuinely scarce supply, and AMD's two wins may stay at two. The claim that holds either way is narrower — cash conversion gets priced before any of this resolves.
Action items
- Pull actual contract language from every compute-exposed portfolio CFO within 30 days and stress-test margins at a 5–15% revenue-share overlay.
- Run an LOI audit across the portfolio and pipeline this quarter, converting every non-binding agreement with a chip vendor, hyperscaler or sovereign into a probability-weighted line item.
Sources:The Information · The Information Briefing · 🔳 Turing Post · Techpresso
◆ QUICK HITS
Quick hits
Hyperliquid's real-world-asset perps hit $26B, 54% of platform volume
Waymo is exiting Uber in Austin and Atlanta, effective January 2028
Blackstone takes Jersey Mike's public Thursday at a potential ~$8B initial cap
Global Laser Enrichment is about to run its first commercial-scale test
LG bans residential-proxy SDKs after 42% of webOS store apps were found routing traffic
Thinking Machines Lab has lost four of its six cofounders to Meta and OpenAI
◆ Bottom line
The take.
Stop underwriting model access and start underwriting who holds the model-selection decision — then re-price revenue quality at the application layer before your next mark.
Frequently asked
- What should I do about portfolio companies whose margins run through frontier per-token pricing?
- Re-underwrite them now, before the next price cut rather than after. Model a routing-adjusted gross margin where the frontier premium collapses from 8x to roughly 3x revenue per token and half or more of volume diverts to cheaper execution models. Public comps and vendor pricing have already moved, but private marks and board decks still assume frontier pricing and automatic cost deflation.
- If routing is where the margin went, what actually makes a router defensible?
- The defensible asset is the proprietary per-model, per-task performance table, not the orchestration code. That data captures which model wins on which task — agent success swings from 66% to 94% depending on edit-format-to-model matching, and reviewer pairings are asymmetric, with Claude checking Codex lifting pass rates to 89.7% while the reverse lowers them. These tables compound with usage and no single lab can replicate them.
- Why does Decagon carry a richer multiple than Harvey despite booking less revenue?
- The market is pricing logo velocity as a proxy for future net revenue retention, not the current revenue base. Harvey sits near 25–45x ARR ($11B on hundreds of millions), while Decagon runs 50–150x ($4.5B on an eight-figure base). The exposure is revenue quality: at those multiples a 30% error in what counts as recurring revenue can erase the entire return, and discounted pilot or paid-to-adopt revenue often hides under an ARR label.
- What separates AI app-layer companies that keep their moat from those that don't?
- Specification density — how much of correctness depends on undocumented conventions and standards interactions that can't be observed from outside. High-density domains like tax and regulatory compliance, clinical protocols, payments rails and security primitives stay defensible because a model can't infer correctness by probing an API. Low-density products such as CRUD tools and thin orchestration can be rebuilt cheaply once a model probes the API.
- Why does Nvidia taking a cut of customer cloud revenue matter to infrastructure valuations?
- It converts a depreciating hardware sale into a perpetual royalty on the downstream revenue that hardware generates, which resets terminal value across neoclouds, colocation platforms and GPU-collateralized lending. If the terms standardize, every GPU-cloud model built on buy-once-monetize-for-six-years is understated at the gross-margin line. No contract language is public yet, so confirm terms directly with portfolio CFOs and stress-test a 5–15% revenue-share overlay.
◆ Same day, different angle
Read this day as…
◆ Recent in investor
Keep reading.
- Airtable cleared at 2.7x ARR in an all-cash sale, 88% below its 2021 mark.
- Palantir's $2.1B Cash Still Doesn't Earn Software Economics
- SpaceX Trades 20% Below IPO Price at 51x Forward Revenue
- UEFA Killed FIFA's $4.2B Carve-Out in 4 Days With No Equity
- Situational Awareness Sold $10B to Citadel Despite 439% Gain
Spot an error? mail@promitb.dev