Enterprise Field Notes · Issue #9

The SpaceX IPO Is a Bet on Your AI Bill

Intelligence has never been cheaper. So why is your AI bill going up? The fix is not a cheaper model. It is knowing when to stop reaching for the biggest one. And it turns out to be the same move that decides what all of this costs the rest of us.

By Ben Pickett · June 18, 2026

New here? Subscribe to Enterprise Field Notes, one new issue every week.

The river that runs the data centers is the one a lot of us want still running for the next generation.
The river that runs the data centers is the one a lot of us want still running for the next generation.

By Ben Pickett

Last week, SpaceX went public in the biggest IPO in history. Seventy-five billion dollars raised, a valuation that blew past two trillion within days. A rocket company. So why could I, someone who writes about enterprise AI for a living, not look away?

But one number stuck with me. The same week of that IPO, Google agreed to pay SpaceX roughly 920 million dollars a month for access to about 110,000 graphics chips, because even Google, sitting on more money and more data centers than almost anyone alive, cannot get its hands on computing power fast enough to keep up with demand for its own AI. A rocket company is suddenly one of the places the biggest names in tech turn just to find compute. The biggest IPO in history is, in part, a bet on how desperate the world has become for AI horsepower.

Everyone is watching the stock. I want to talk about that desperation instead, because it is the same scarcity you pay into every time you reach for the biggest AI model when a smaller one would have done the job. It lands on your desk as a bill that keeps climbing.

Now the strange part. While the giants fight over hardware, the price of the intelligence itself is falling off a cliff. Epoch AI, which tracks this closely, puts the cost of GPT-4-level answers at roughly ten times cheaper every year. About 20 dollars per million tokens in late 2022, closer to 40 cents today. The exact slope depends on the task, and nothing I am about to say hinges on the precise number.

So intelligence is getting dramatically cheaper. And yet almost every company I talk to is watching its AI bill go up, not down.

That is the paradox I want to unpack. The reason your bill is climbing is not the price of the model. It is how many times you reach for the most expensive one. And the move that pulls the bill back down turns out to be the same move that makes you a good steward of something I care about a great deal. By the end you will have one question to ask your own team this week.

The trap is not the price. It is the habit.

The instinct everyone has is to shop. Prices are falling, so wait for the next cheaper model, or hunt for the cheapest tokens on the market. Wrong lever.

The real driver is simpler. We stopped asking AI single questions and started handing it whole jobs. An agent does not make one model call to finish a task. It makes ten or twenty. Now multiply that by a habit almost everyone has fallen into without noticing: defaulting to the biggest, most powerful model for every step, even the trivial ones. Classifying an email. Pulling a date out of a document. Deciding whether a message is a question or a complaint.

Cheaper tokens, times far more calls, times the most expensive model on every call. That does not add up to a smaller bill. The price of intelligence went down and the spend on it went up. This was never a pricing problem, which is why you cannot shop your way out of it. It is an architecture problem, and you cannot shop your way out of an architecture problem.

Economists even have a name for the trap. Jevons paradox. When a resource gets cheaper to use, we use so much more of it that the total bill climbs instead of falling. Cheaper steam engines burned more coal, not less. Cheaper tokens spend more money, not less.

Efficiency per call is not the same as efficiency on the invoice.

I learned this years before the chatbots showed up

Before Swa, I spent most of my career in engineering, a lot of it leading large-scale data and reliability platforms at Nike. And I learned the same lesson over and over, long before anyone was typing into a chat box.

You do not put your most powerful, most expensive infrastructure behind every job. The heavy, costly setup is for the work that genuinely needs it. Everywhere else, you use something lighter and cheaper, and you automate the parts people used to do by hand. The teams that got this right did not just spend less. They moved faster, because they were not dragging the heavy machinery into work that never required it, and they freed their best people to sit on the decisions that actually mattered.

That is the whole game with AI today, just wearing a new costume. Most of what you ask AI to do is routine. Classifying an email. Pulling a date out of a document. Deciding whether a message is a question or a complaint. None of that needs the most powerful model on the market. It needs a small, fast, capable one. You reserve the frontier model for the calls that genuinely earn it, and you keep a person on the decisions that carry real weight.

The industry has a name for this. Right-sizing, or model routing. One 2025 study estimated that choosing the right-sized model for each task could cut the energy behind AI inference by roughly a quarter, with the efficient models giving up under four percent of their quality. On mature, well-understood tasks the savings ran far higher. It is an early, single-study figure and depends heavily on the job, but the direction matches what anyone who has tuned this in production already sees.

Picture the support inbox for almost any business. A hundred messages come in. The lazy build sends every one of them, at every step, to the most powerful model on the market. Read it, classify it, draft a reply, check the reply. Four expensive calls times a hundred messages, and most of those messages are someone asking where their order is.

Now right-size it. A small, fast model reads all hundred and sorts them. The seventy that are routine, where is my order, how do I reset my password, change my address, it handles start to finish. The twenty-five that are genuinely tricky escalate to the frontier model. The five that carry real weight, a furious customer, a refund dispute, a question with legal teeth, go to a person who is now free to handle them well. Same hundred messages handled. A fraction of the bill, and arguably better service, because your best people are spending their time where it counts.

The catch, and anyone who has shipped this will tell you, is that the sorting is the hard part. You earn that split by tuning it on real traffic and setting honest confidence thresholds, not by hoping a cheap model guesses right. But that is engineering you do up front and revisit as traffic shifts. Paying premium prices to answer where is my order is a tax you pay forever.

In our own systems, over a recent ninety-day stretch across our production traffic, right-sizing the model to the task cut our compute cost by roughly half compared with routing everything to premium models. On money alone, that is the whole case.

Why this one is personal

But there is a second reason I care about this, and it is not about money at all.

My family came west on the Oregon Trail six generations back and homesteaded in Douglas County, in southern Oregon. My great-great-great-grandfather, James Swifty Pickett, brought his son James Riley there as a boy. We put down roots and never left. I have spent my whole life in the Pacific Northwest. I also spent years as mayor of Jefferson, Oregon, where I dealt with water rights more than I ever expected to. You learn quickly that water is not infinite, and that everyone downstream is counting on the choices you make upstream.

So none of this is abstract for me. A lot of the data centers training and running these models sit right along the Columbia River, powered in large part by the hydro and the Gorge wind that light up my corner of the country. The intelligence we are all racing to deploy runs on my river.

The bigger bill

Right-sizing is more than a budgeting trick. Every one of those model calls does something beyond cost you money. It pulls electricity off a real grid, water from a real river, and minerals out of a contested supply chain.

The power is the headline. The International Energy Agency projects that the electricity data centers draw will roughly double by 2030, growing more than four times faster than overall electricity demand, with AI driving almost half of that growth in data center demand.

The water is closer to home than most people realize. A peer-reviewed study estimated that training one older model, GPT-3, in US data centers could evaporate around 700,000 liters, roughly 185,000 gallons, of clean freshwater just for cooling. Run twenty to fifty everyday questions through a chatbot and you have quietly used about a bottle’s worth.

And the materials are finite and concentrated. Roughly 95 percent of the gallium that goes into this hardware is refined in a single country, which has already shown it will use that as leverage.

And before someone says the answer is to just put the data centers in space, that race is already on. Google has a feasibility project. SpaceX has filed to launch up to a million compute satellites. A startup up the road from me in the Seattle area has already trained a model and run Gemini in orbit. Free solar, no water for cooling. It is a genuinely exciting idea, and I hope some version of it works. It also costs several times what a data center on the ground does today, there are nowhere near enough rockets to do it at scale, and the serious versions are years away. The moonshot does not make the resource bill disappear. It moves it off the planet, after burning an enormous amount of resources to get there.

Which brings me back to the one lever you actually have this quarter. Right-sizing. The same move that pulls your AI bill down is also one of the most practical levers you actually control over what AI takes from the grid, the water, and the ground. You do not have to choose between running a tight operation and being a decent steward. The efficient architecture is both. This is not about saving the planet. It is about not being wasteful with what we all share.

What this comes down to

Cheaper tokens will not save you. Right-sizing will.

And the same decision that controls your bill eases what AI draws from the grid, the water, and the supply chain. Good operator and good steward, the same move.

One question to ask this week

Here is the gift, and it costs nothing.

Pull up your most-used AI workflow this week. Walk it step by step and ask one question. Which of these steps actually needs the most powerful model, and which ones are we sending there out of pure habit? Route the habits down to a smaller, efficient model. Keep your people on the judgment calls. That single pass will usually find more savings than any pricing negotiation ever will.

That is it. The question is free. What you do with the answer is the whole job.

I would like all of this to still be here, and running well, for the people who come after us. The way to get there is not a grand gesture, and it is not a rocket. It is a thousand small, smart decisions about when we actually need the biggest hammer. Start with one this week.

Header artwork generated with Swa.


References

  1. Google–SpaceX compute deal. Google will pay SpaceX ~$920M/month for ~110,000 Nvidia GPUs, Oct 2026–June 2029, as “bridge capacity” for Gemini demand. TechCrunch, June 2026.
  2. Inference price decline. Epoch AI, “LLM inference prices have fallen rapidly but unequally across tasks.” GPT-4-level performance fell from ~$20 to ~$0.40 per million tokens, roughly 10x/yr. Epoch AI.
  3. Agentic token use. Gartner (March 2026): agentic AI uses 5–30x more tokens per task than a chatbot. Gartner.
  4. Right-sizing / model selection. “Small is Sufficient: Reducing the World AI Energy Consumption Through Model Selection,” arXiv:2510.01889 (Oct 2025): model selection could cut AI inference energy ~27.8%; efficient models avg ~66% more efficient at under 4% quality loss; up to ~70–93% on mature tasks. An early, single-source, idealized scenario. arXiv.
  5. Data center electricity. IEA, Energy and AI (April 2025): data center electricity projected to roughly double to ~945 TWh by 2030, growing >4x faster than total electricity demand; AI “accelerated servers” drive almost half the net increase. IEA.
  6. Training water. Li, Yang, Islam & Ren, “Making AI Less Thirsty,” Communications of the ACM (2025), DOI 10.1145/3724499: training GPT-3 in US data centers could directly evaporate ~700,000 liters of freshwater (modeled estimate); ~20–50 queries ≈ a half-liter bottle. CACM.
  7. Gallium concentration. IEA, Energy and AI: ~95% of gallium refining is in one country; China’s July 2023 gallium/germanium export controls. IEA.
  8. Data centers in space. Google “Project Suncatcher” feasibility study (Nov 2025); SpaceX FCC filing for up to 1M orbital data-center satellites; Starcloud (Seattle-area) ran a model and Gemini in orbit. Orbital compute costs several times ground today; years from scale. MIT Technology Review.
  9. Columbia River data centers. Google (The Dalles) and Amazon (Boardman/Hermiston/Umatilla) draw on BPA Columbia River hydropower; Columbia Gorge wind feeds the same grid. Stanford, And the West.

About the author

I'm Ben. I write Enterprise Field Notes, and by day I'm COO at Swa, after years running reliability, data protection, and database operations at Nike. The lesson that keeps proving itself: anything you cannot run without, and cannot walk away from, is a risk you have not priced yet. What is yours?

Header artwork generated with Swa.

Read more of Ben's Enterprise Field Notes at benpickett.com.