Seven short reference sheets on how AI usage actually works — who burns the most, why the build-out is so massive, whether self-hosting is greener, what you're billed on, why it's all getting cheaper anyway, whether the whole thing's a bubble, and what it costs to build in the first place. Same engine underneath; very different views depending on where you stand.
A person typing into a chatbot moves tokens at human speed; a company running APIs and agents moves them at machine speed, around the clock. Plotted on a log scale — each gridline 1,000× the last — the gap between an individual and a single enterprise runs to roughly a million-fold. The upside for you: that machine-scale demand funds the infrastructure and price drops behind your cheap personal plan.
Every token runs on a real chip drawing real power — and demand is racing so far ahead of supply that it's driving the biggest infrastructure build-out in a generation. The capex, the grid, and why power (not silicon) is now the thing being built. The upside: that scarcity is fueling record efficiency gains, and smart pricing keeps AI flowing to everyone.
A measured look at one always-on home AI box — built from its own usage logs. The honest answer turns on how busy the box is: it gets greener the more you use it, and it already wins on embodied carbon, water, and privacy. The same query runs 2.6–43 Wh depending on how you count idle. Includes both attribution models, the break-even math, and the cheap wins that close the gap.
Two billing universes running across every major provider. One sells a flat-rate bucket of usage with rolling caps; the other meters every token you burn. Side-by-side consumer plans and API rates from the premium labs down to the cheap open-weight disruptors — plus how image inputs are counted and why generating an image is billed nothing like reading one.
The rest of this collection looks at the hard limits — soaring demand, finite power, carefully shared compute. This is the upbeat counterpart: AI is getting cheaper and more efficient faster than almost any technology in history. Inference costs fall ~10× a year, energy per prompt is collapsing, and small models keep eating big-model jobs. Plus the honest catch — why efficiency gains get partly eaten by demand.
The most-asked question — and the one with the least honest answers, because it mixes up two different things: is the technology real? (almost certainly yes) and is the spending a financial bubble? (genuinely contested). The bull and bear cases weighed with real 2026 numbers — capex vs. revenue, GPU depreciation, circular financing — plus what the dot-com, Cisco, and railway busts actually teach us.
Every other sheet is about the running meter — what each query costs. But before a model answers anything, someone spends hundreds of millions teaching it. Training is AI's other cost, with the opposite shape: enormous and one-time. The twist — spread across a model's life it's a fraction of a cent per query, yet inference still wins the lifetime total. Why "AI is wildly expensive" and "AI costs ~nothing" are both true.
Building this collection with an AI assistant wasn't free either. Here's the rough footprint of the whole back-and-forth that produced it — estimated the same honest way as the sheets above, and small enough to be encouraging.
Estimate — A genuine back-of-envelope, not a meter reading. The token count is approximate; energy / carbon / water are derived from published per-token cloud-inference figures (plausible range ~150–650 Wh) and a ~330 gCO₂/kWh grid, using the same 1.08 mL-per-Wh ratio as the cited cloud benchmark. Real numbers swing with the model, data center, and grid. The hopeful part — prompt caching (sheet 04) and the efficiency trend (sheet 05) mean the same work keeps getting cheaper and greener — this is already a small fraction of what it would have cost a couple of years ago.
About — A small collection of standalone reference sheets on AI token economics. Plain HTML/CSS with one shared stylesheet for navigation — no build step, no frameworks. Sources are cited at the foot of each individual sheet (Anthropic docs, Google I/O 2026, IEA, Gartner, Menlo Ventures, and others).