Model mechanics, equations, and evidence
This document is the implementation-level companion to the in-model manual. It describes how the current version converts demand, powered infrastructure, serving costs, modeled task value, human use, and distribution into the results shown on screen.
This is a scenario model, not a forecast. Its inputs combine measured benchmarks, external proxies, projections, and explicit modeling judgments. A source may support the direction or scale of an assumption without measuring the exact default.
1 · Computation sequence
The model runs at eight steps per year from 2026 through 2032. Each interface update follows the same order:
- Compound raw output-token demand from the Q4 2025 proxy.
- Apply token efficiency, serving-tier mix, capability drift, local serving, active-parameter weights, and context cost to obtain datacenter compute-unit demand.
- Simulate powered infrastructure by installation cohort and derive available fleet capacity.
- Calculate utilization, proportional rationing, prices, revenue, capital spending, and operating costs.
- Value delivered tasks on a task ledger whose class anchors are frozen at the 2026 reference ladder.
- Split modeled creation between engagement-produced and autonomy-produced flow.
- Apply a distribution regime and close its realization and affluence feedback to a fixed point.
A failed realization solve is rejected and shown as an error rather than being allowed into the charts.
2 · Demand and compute units
Raw demand and efficiency
Demand begins at a proxy of about 1 billion output tokens per second in Q4 2025. The demand-level slider spans approximately 0.3–3.3× around that proxy. Annual multipliers compound and interpolate geometrically within each year. The default 2026–2031 path is 8×, 6.8×, 5.8×, 5×, 4×, and 3.4×.
Token efficiency reaches 3× by 2030 by default. It divides the raw token requirement while multiplying modeled task value by the same factor. Induced demand is not automatic; represent Jevons-style recapture with a hotter demand curve.
Reference compute unit
One compute unit is one reference output token per second at 49 billion active parameters, 8K input context, and the reference serving setup. For serving tier k:
With q = input context / 8K and sparse share sh:
The 0.70 and 1.41 exponents are calibration choices. Sparse adoption changes serving cost and capacity pressure, not task value.
Serving classes, drift, and local execution
With capability drift off, tier shares interpolate from editable 2026 values to editable 2030 values. With drift on, the 2026 mix becomes the state and the 2030 mix is computed. The drift curve moves eligible work toward cheaper serving classes. A persistent floor is replenished in its original class, while the knowledge ratio slows movement for knowledge-intensive work.
Local-serving shares interpolate separately. Local tokens leave datacenter demand and provider revenue but remain in the task-value and mediation ledgers. Local hardware supply, cost, and power are outside the model. Mastermind is a separate 400B-active class introduced at the chosen date and carved proportionally from the legacy task mix.
3 · The watts supply ledger
State and calibration
The conserved supply state is powered watts grouped by installation step. Each group retains its installation date, watts, and capability per watt. Fleet capacity is the sum of capability-weighted watts. GPU-equivalents, gross build, capital spending, operating costs, retirement, and low-capability watt share are derived from that state.
Two inputs calibrate new-unit tokens per watt on an unconstrained 3.4× reference path:
- 2026 powered footprint: 35 GW by default.
- Mid-2027 capacity pin:
3.8 × (2M GPU-equivalents × 10K output tok/s).
The approximately 10K tok/s/GPU term is benchmark-measured. The footprint, 2M fleet scale, and 3.8× aggregation are projections. Capability per watt on newly installed units improves at 2.0×/year by default.
Build-out, power availability, and retirement
The build-out curve is attempted gross GW/year. Delivered deployment is limited by current firm-power availability, banked reserve, and recycled retired watts. Unused availability retains 80% per year. Named capacity-growth presets are translated back into gross GW/year through the live ledger and therefore include replacement builds.
Infrastructure leaves the frontier fleet at the decommission age, six years by default. Its interconnection returns to the availability pool. Under steady growth, retirement changes gross build, footprint, and capex more than capacity. Under deceleration or binding power, older low-capability watts can accumulate and retirement timing matters more. Secondary-market hardware diffusion is not represented.
4 · Rationing, prices, and cash
Utilization is compute-unit demand divided by fleet supply:
When R > 1, served datacenter volume is multiplied by 1/R. The model applies this reduction proportionally and does not separately allocate queues, priority users, speed tiers, or down-tier routing.
Serving-tier price per million output tokens is:
This is a cost-anchored scenario price, not a provider quote or an estimate of observed marginal cost. Commercial and uncovered personal traffic use the price ladder. Covered personal traffic produces subscription revenue from modeled subscribers and monthly price. Capex prices gross new capacity, including replacements; opex is charged per powered watt.
5 · Creation and capture
The model deliberately keeps two ledgers:
- Capture is modeled cash at actual scenario prices: API revenue, subscription revenue, payments, and retained surplus.
- Creation is modeled uplift from delivered task value.
Each task class receives a value anchor frozen at the 2026 reference ladder:
Creation then applies a separate context-value term (input context / 8K)γ, token efficiency, the 2026 value anchor, complementarity, delivered volume, and the regime realization factor. The default context-value elasticity is γ = 0.20.
Capability drift changes the serving class, and therefore cost and capacity pressure, without changing the task-class anchor. Scarcity premia transfer cash between users and owners; rationing can still destroy modeled creation by reducing delivered volume. The exchange-valued twin prices delivered datacenter flow at the current serving-class reference ladder and is a diagnostic, not measured consumer surplus.
6 · Mediation and realization
Total flow for mediation includes served datacenter tokens and local tokens. Engaged adults are the smaller of implied users and the modeled addressable population, except that Basic compute can extend access to the full adult population.
Flow inside the envelope is engagement-produced. Flow beyond it is autonomy-produced. The model applies an autonomy realization coefficient δ to the latter. The default begins at 0.60 in 2026 and converges linearly to 1 in 2034.
The addressable population begins at an illustrative 250 million adults, grows at 2.5%/year baseline threshold crossing multiplied by an uplift feedback, and is capped at 5.5 billion. Because realized uplift changes effective output, which changes addressable population, which changes realized uplift, each regime and build-out counterfactual closes this loop independently. Convergence tolerance is 1e−4 within 16 iterations.
7 · Distribution regimes
The distribution layer begins with 240 quantiles from a stylized lognormal baseline: median $7.5K, dispersion σ=1.21, top-10 threshold about $35K, top-1 threshold about $125K, and bottom-half mean about $3.5K. This is an approximate WIR 2022-era calibration and materially understates current WIR 2026 thresholds.
| Regime | Mechanism represented |
|---|---|
| BAU | 5% capture, 55% pass-through, and a $50K access threshold. |
| Enclosure | Lower capture and pass-through, concentrating access and retained value. |
| Dividend | Routes 15% of capture as an equal cash payment. |
| Basic compute | Spends 15% in kind and grants universal modeled access. |
| Broad ownership | Routes half of the capital stream as a universal dividend. |
The optional human-capital layer adds latent ceilings, development, access-gated pipeline compression, task tenure, motivation, and firm-side integration. These are sensitivity parameters, not estimated causal effects. The regime frontier therefore asks what follows if the chosen allocation and capability-development mechanisms hold.
8 · Solver and counterfactuals
Solve: no crunch iteratively chooses annual attempted deployment to target utilization of 1 on the actual ledger. The 2026 build is treated as committed; 2027 is the first responsive year and targets balance in 2028. Retirement, reserves, power availability, and the 0.1–2,900 GW/year bounds remain active. A shortfall forced by the anchor or power curve is reported rather than erased.
Ramp limits constrain only the solver. They cap implied capacity growth on a log-linear schedule from 3.6× in 2026 to the selected 2031 endpoint. Manual drags and named build-out presets are not constrained by this toggle.
Every named regime and build-out card receives its own realization solve. A custom dragged path is not automatically identified with a named card.
9 · Chart semantics
- Rationing strip and verdicts: quarterly utilization; crunch timing is interpolated between simulation steps.
- GDP effect by build-out path: independent build-out counterfactuals under the active regime.
- The race and utilization: raw-token demand, compute-unit demand, supply, and utilization.
- The watts ledger: powered footprint by installation age; growing older bands indicate deceleration.
- Fleet demand by tier: compute-unit demand shares by serving class.
- Uplift by production mode: engagement-billed, engagement-local, and autonomy-produced modeled value.
- The tranches: population shares in fixed resource bands plus a non-additive capability-dividend diagnostic.
- The regime frontier: independent 2029→2032 arrows comparing bottom-half resources and realized uplift.
- Price ladder and cash flows: scenario prices, revenue, capex, opex, and exchange-valued flow.
10 · Evidence and calibration
- Measured benchmark SemiAnalysis InferenceX reports roughly 10.2–11.1K output tok/s/GPU for the stated GB300-NVL72 workload. The model rounds to 10K.
- External proxy Epoch AI estimates about 3.3× annual growth in AI compute stock since 2022 and tracks sales rather than deployments. The model uses a rounded 3.4× reference path.
- External proxy The IEA reports that AI-factory capacity more than tripled over 18 months. This supports rapid-growth direction, not the exact 35 GW default.
- External proxy Google I/O 2026 reports 3.2 quadrillion monthly tokens and 7× year-over-year growth. The default demand path begins at 8× as a scenario assumption.
- Legacy calibration Distribution targets approximate World Inequality Report 2022 thresholds after approximate currency relabeling. WIR 2026 reports materially higher current thresholds.
Named power, build-out, hardware, and policy presets are constructed counterfactuals, not announced schedules or forecasts. The 2.9 TW/year build ceiling is a numerical display limit; reaching it indicates an infeasible path.
11 · Known limits
- No labor market: displacement, wages, employment, and labor-supply responses are absent.
- No local-hardware supply or cost: local execution does not compete with the datacenter fleet for components or power.
- Legacy income baseline: the distribution is not calibrated to current WIR 2026 levels.
- Exogenous demand and regimes: prices, budgets, politics, concentration, and capability do not select their own path.
- Judgment-heavy value layer: value anchors, context elasticity, complementarity, addressable population, δ, and human-capital parameters are not jointly identified from an empirical system.
- Uniform rationing within tiers: users and tasks are not prioritized inside a serving class.
- Basic compute cost feedback omitted: additional universal access does not add fleet demand.
- PPP-style resource bands: technology can change the meaning of the consumption basket.
- A $1/adult-year numerical floor keeps log charts finite; visible binding marks a distributionally infeasible scenario.
- One aggregate autonomy coefficient applies across reference-valued streams.
12 · Technical glossary
- Compute unit: one reference output tok/s at 49B active parameters, 8K input context, and the reference serving setup.
- GPU-equivalent: ledger hardware capacity before the serving-stack multiplier, divided by 10K tok/s.
- Footprint: powered watts in the modeled frontier fleet.
- Binding ratio: attempted divided by delivered deployment in a year, capped at 3× for scarcity-cost calculations.
- Context penalty: blended serving-cost multiplier relative to 8K context.
- Mediation share: fraction of total flow inside the modeled human absorption envelope.
- Oversight ratio: total steered tokens per engaged adult per month.
- δ: value assigned to an autonomy-produced token relative to an engagement-produced token.
- Task value anchor: per-class value frozen at the 2026 reference ladder.
- Exchange-valued flow: delivered datacenter flow valued at the current serving-class reference ladder.
- Creation vs capture: modeled uplift from delivered task value versus modeled cash at actual prices.
13 · Version history
v2.8.1 gives every regime and build-out counterfactual its own realization fixed point, rejects non-convergence, and adds executable regression checks. v2.8 freezes task value in task space, makes drift a serving-assignment map, separates the 2026 value anchor from complementarity, and adds exchange-valued flow. v2.7.1 separates context-length value from the sparse/dense cost blend. v2.7 closes affluence on realized creation. v2.6 adds the human-capital ledger. v2.5 makes δ a trajectory. Earlier v2 releases introduced the watts ledger, Mastermind class, mediation, distribution regimes, the creation/capture split, and solver ramp limits.