Enterprise Field Notes · Issue #12

Anthropic Just Showed You Your AI Has No Backup

Anthropic put its best model on a meter, and the real story is not the price. It is that your business runs on one model, from one vendor, with no backup, and this is the moment you can finally see it.

By Ben Pickett · July 9, 2026

New here? Subscribe to Enterprise Field Notes, one new issue every week.

The spare lantern, hung before the first one goes dark.
The spare lantern, hung before the first one goes dark.

Enterprise Field Notes · Issue #12

Every single point of failure I have ever cleaned up existed for one reason. It was cheaper, right up until it was not. This month Anthropic handed the whole industry that lesson again, and almost everyone read it as a story about price.

By Ben Pickett

My first real engineering job, more than twenty years ago, I recommended a proper redundant array for the company mail server. The kind that keeps running when a drive dies, so you swap the bad one and nobody notices. My boss looked at the cost and asked the question that has cost companies everything since the beginning of computing. What are the chances a drive fails before we can replace it?

We found out. We ran the cheaper setup, a simple mirror, and two of its drives failed inside ten minutes of each other. The backup we had settled for, also to save money, turned out to be incomplete. I spent the next month rebuilding the CEO’s email by hand, out of raw log files. I was in my twenties. I have never forgotten it, and not because of the month. Because of the shape of the mistake. We did not lose that mail to bad luck. We lost it to a decision someone made to save a little money, betting that something unlikely would also never happen.

Unlikely and never are not the same word, and the gap between them is where careers go to die.

From that month on, keeping that class of mistake from ever happening again became the work. I spent years running data protection and database operations for one of the largest brands in the world, the kind of place where a single point of failure does not cost you a CEO’s inbox, it costs you the business. That job hands you a particular discipline. You stop asking whether the unlikely thing will happen and start assuming it will, then you build so that when it does, nothing important stops.

And this month a company you rely on made the old cheap bet for you, on your AI, then quietly took away the thing that had been hiding it.

What Anthropic actually did

On June 9, Anthropic released Claude Fable 5, its most capable model yet, and priced it at ten dollars per million tokens in and fifty per million out. That is double the standard price of Opus 4.8, which runs five and twenty-five, and the most expensive model Anthropic has ever listed. Three days later it was gone. On June 12 the US government issued an export control directive under national security authority, and to comply Anthropic disabled Fable 5, and its larger sibling Mythos 5, for every customer in the world. The trigger was a narrow jailbreak researchers had demonstrated, getting the model to read a codebase and surface its software flaws. Anthropic disagreed that a narrow finding should pull a model used by hundreds of millions, and said so, but it still went dark for nearly three weeks. The Commerce Department cleared it and access returned on July 1. It came back to the paid plans, included for up to half of your weekly usage, and on July 7 that included access ended and Fable 5 moved to metered credits, pay as you go.

Both of those made the news. Here is the part that did not. While it was bundled, the plan quietly absorbed the cost of leaning on the best model. You reached for it without thinking, because reaching was free, and the price of that reach was somebody else’s problem. On a meter it is yours. The model is still there, as good as it ever was, for exactly as long as your credits last. What moved is not the model. It is the cost of depending on it, from the plan’s books onto your own. That cushion, the one that quietly ate the price of your reaching, is gone. And the moment you feel the cost, you can see the thing it was paying to hide.

That cushion was never a safety net you designed. It was hiding a decision you did make, without ever calling it one. You wired one vendor’s newest model into everything and gave it nothing to fall back on, and as long as reaching for it was cheap, you never had to notice. Take the cushion away and the shape is plain. Your company runs on a single model, from a single vendor, with nothing behind it. That is not a billing problem. It is the RAID array all over again, at the scale of your whole business.

You already know this bet. You call it the cloud

If that sounds abstract, you have lived a version of it more recently than you think.

Amazon builds serious redundancy into AWS. Multiple regions, multiple data centers, the whole thing engineered so a failure in one place does not take you down. And most companies do not pay for it. Running across two regions instead of one costs on the order of thirty percent more, so teams run in a single region and bet that AWS will stay up. The bet is good almost every single day. On October 20 last year it was not. A failure in a single AWS region in Northern Virginia cascaded for hours, and Delta, Uber, Starbucks, Spotify, Shopify, Etsy, and Amazon’s own Alexa went down with it. Companies that had quietly decided a second region was not worth the money spent that day finding out what it was worth.

And it is not only Amazon. When one bad software update went out from a single security vendor, CrowdStrike, in 2024, it took down about a quarter of the Fortune 500 at once and cost them, by the analytics firm Parametrix’s estimate, about five and a half billion dollars in direct losses within days. Delta alone put its own share near three hundred and eighty million. One vendor. One dependency. No fallback. Every company caught in it had made the same quiet trade the mail server made, the same one the single-region shops made. This is reliable enough that paying for a backup is not worth it. They were right for years and wrong for a day, and the day was catastrophic.

I have been on the cleanup side of failures shaped exactly like this, and the recovery is never proportional to the mistake that caused it. At one point in my career I watched a failover environment take forty-five days to fully restore, forty-five days, and the only reason it was survivable is that it was not production. It was slow for the same reason the mail server was fragile and the single region was exposed.

The redundancy that would have made recovery fast was the line item someone decided not to pay for.

That is the bet you made with your AI. You wired the best model into everything, from one vendor, with no backup, because the best model was free inside the plan and a backup felt like a cost you did not need. Anthropic just pulled the cushion. The odds on your bet did not change. What changed is that you can finally see it.

Leah

You know Leah. She is the engineer I keep meeting in different companies. Sharp, quietly trusted, the one who wires things together while everyone else is in a meeting about it. A year ago Leah pointed the whole company at the best model, because it was the best and it was free inside the plan, so why would she use anything else. Nobody decided that. Nobody revisited it. What she actually did, without anyone signing off, was make one vendor’s newest model load-bearing for the entire business. The invoice is what hurt first when the plan changed. The part that should keep her up is quieter. If that model is down next Tuesday, a lot of things stop, and there is no plan for that.

So the easy story is that Anthropic got greedy and you should switch. That is comfortable, and it is wrong twice over. Blaming greed misreads what is happening. And switching vendors is not a plan, it just moves you to a different single point of failure. You can wire an automatic fallback to the same vendor’s older model, and you should, but a backup that lives inside your one vendor is only half a backup. The real exposure was never the price. It is that you built something you cannot run without on top of something you do not control, and you never gave it a way to keep going when that one thing does not.

This is not an Anthropic problem

It would be easy to make this about one company’s pricing page. It is not, and this is the part that should change how you read the whole category.

GitHub moved Copilot to usage-based billing this summer, the same shape of change, and some developers watched their monthly bill climb by an order of magnitude. The pattern is everywhere because the cause is not a boardroom deciding to squeeze you. It is physical. The compute these models run on is genuinely scarce, and the venture money that paid for the early all-you-can-eat plans is running out. Sam Altman has said plainly that OpenAI loses money on its most expensive plan. Scarce things get priced like scarce things. Fable 5 moving to a meter is not the surprise. It is the moment the pattern got too big to look away from.

And here is the twist that makes the bill climb even as tokens get cheaper, because tokens really are getting cheaper. A fixed level of capability falls in price every year, and last year’s frontier is now a commodity you can buy for a few cents. So why does the bill keep rising? Not because tokens cost more. Because you are burning far more of them. An agent calls the model ten or twenty times for one task, a system packs more context into every call, an assistant runs all night. Cheaper tokens, many more of them, every one pointed by default at the most expensive model. That is how a rounding error becomes the line item your CFO circles in red.

The fix, and the honest part

Here is the part you can copy this week, without a budget and without buying anything.

Three decisions, none of which require a tool. First, decide what earns the best model, by kind of work and not by habit. Hard reasoning, high-stakes output, the thing a client sees, that can earn the top model. Volume, routine, internal, throwaway, that runs on something cheap and fast. This is the cost choice the plan used to make for you. Now you make it. Second, give anything you cannot run without a backup. Name the cheaper model that carries the load when the best one is gone, and decide in advance what slows down and what must not stop. Your vendor did this for you invisibly, and this month it stopped. Bring your own. Third, give the whole thing an owner. Right now the model choice, the fallback, and the AI budget belong to everyone and no one, which is why the bill surprised you, and the outage will too. One person owns it and writes it down. In a five-person company that owner is you, and that is fine. The point is that someone decided, instead of a default deciding for you.

Now the honest part, because you can probably smell where this is going. The layer that decides which model runs each task, keeps a fallback ready, and holds the budget, that is exactly what we build at Swa. I am not going to turn this into a pitch. I will just give you the number, because it is public and my co-founder already shared it. Across our own 90 days: 26 people, 106 million tokens. We routed the routine work to cheaper models and saved the best one for the tasks that needed it. Measured against the cost of sending all 106 million tokens to the top model, that cut our compute spend by 51 percent. It is not a magic figure, and it is not our discovery. Independent work on the same technique, the RouteLLM research out of UC Berkeley, reports cuts in that range and larger, at near equal quality. None of it is our invention. The part most people skip is measuring it honestly, and you can start on a whiteboard this afternoon with the models you already pay for. The discipline is the product. Everything else is plumbing.

What the fix looks like, concretely

Take a support inbox pointed entirely at the best model, because that was the default and it came with the plan. A few thousand tickets a month, and almost none of them hard. Route by the shape of the ticket. Password resets, order status, where is my refund, the routine eighty percent that follow a known pattern, go to a cheap, fast model. Anything the cheap model flags as low confidence, and anything carrying a churn signal or a legal word, gets handed up to the best model. The trigger is a confidence score and a short red-flag list, not a person sorting mail.

Do that and most of the volume leaves the expensive model, the spend on that inbox falls to a fraction of what it was, and the customer never notices, because the tickets you moved were always going to be easy. And here is what you also just built. The day the best model is unavailable or over budget, that inbox does not stop. The routine work is already on the cheaper model, and the handful of hard tickets can wait a few minutes or fall back to an honest hold instead of a broken page. Cutting the bill was only half of it. You also removed a single point of failure. The cost win and the reliability win were the same move, made once, on purpose.

Every single point of failure I have ever cleaned up existed for the same reason. It was cheaper, right up until it was not. The frontier model in your stack is the newest one, and this month you found out it has no backup.

One thing to do this week

Find the one workflow that would break tomorrow if the best model simply vanished. Not slow down. Break. That is your single point of failure, and until this month you did not have to see it, because the plan was quietly covering for you. Now put a number on it. What does one day of that workflow being down cost you, in lost orders and idle people, and how long would you honestly take to recover? My failover took forty-five days. Multiply the daily cost by that recovery window and you are looking at the real size of the bet, the one nobody wrote down. Then give the workflow the two things it does not have. A cheaper model that can carry the load, and one person who owns the choice. If it is also the workflow burning the most on frontier prices for easy work, and it usually is, you just fixed the bill and the outage with a single decision.

Leah did not need a better model or a different vendor. She needed the two things any reliability engineer would have given the system on day one. A backup, and someone who owns the choice. Neither do you. The best model is still there when you need it. What just stopped being optional is having a plan for when it is not. I learned that at twenty-something, a month deep in a CEO’s lost email, and the lesson has not changed once in twenty years. Only the size of the thing you are betting.


References

  1. Claude Fable 5 and Mythos 5 introduced; standard pricing $10 per million input tokens and $50 per million output; launched June 9, 2026. Anthropic, June 2026.
  2. Claude Opus 4.8 standard pricing is $5 per million input and $25 per million output, half of Fable 5’s rate. Finout, 2026.
  3. On June 12, 2026, a US government export control directive citing national security forced Anthropic to disable Fable 5 and Mythos 5 for all customers worldwide, over a narrow demonstrated jailbreak; Anthropic disputed the action. Anthropic, June 2026; Forbes, 2026.
  4. The Commerce Department cleared Fable 5 and Anthropic restored access on July 1, 2026, included in the paid plans through July 7 before moving to usage credits. Anthropic, July 2026.
  5. Fable 5 moves from bundled subscription access to metered, pay-as-you-go usage credits on July 7. The New Stack, July 2026.
  6. The AWS us-east-1 outage of October 20, 2025 cascaded for hours and disrupted Delta, Uber, Starbucks, Spotify, Shopify, Etsy, and Amazon’s own services; running across multiple regions typically costs roughly 30 percent more than a single region, which is why many companies do not. Built In and AWS Well-Architected Framework, 2025–2026.
  7. The July 2024 CrowdStrike outage, from a single vendor’s flawed update, hit about a quarter of the Fortune 500; the analytics firm Parametrix estimated roughly $5.4 billion in direct losses across those companies within days, and Delta put its own loss near $380 million. Parametrix and Fortune, 2024.
  8. GitHub Copilot moved to usage-based billing on June 1, 2026, and removed the automatic fallback to a lower-cost model when credits run out; some developers reported large increases (one from about $29 to nearly $750 a month). GitHub Blog and TechCrunch, 2026.
  9. Sam Altman: OpenAI is losing money on its $200-per-month ChatGPT Pro plan because usage ran far higher than expected. Fortune, January 2025.
  10. AI compute scarcity: top-chip rental prices up roughly 48 percent in two months; Bank of America projects demand outstripping supply through 2029; OpenAI’s CFO cites hard trade-offs “because we don’t have enough compute.” Tom Tunguz, citing The Wall Street Journal, 2026.
  11. Swa’s own 90-day usage across 26 users and 106.2M tokens: 70 percent of inference routed to efficient models, cutting compute cost by 51 percent versus sending every token to the top model. Shared publicly by Swa co-founder and CEO Mike Sirchuk, 2026.
  12. RouteLLM (Ong et al., UC Berkeley and Anyscale, 2024, presented at ICLR 2025): trained routers matched roughly 95 percent of GPT-4 quality while cutting cost by up to about 85 percent on benchmark tests; production deployments of model routing commonly report 40 to 70 percent reductions. LMSYS.

About the author

I'm Ben. I write Enterprise Field Notes, and by day I'm COO at Swa, after years running reliability, data protection, and database operations at Nike. The lesson that keeps proving itself: anything you cannot run without, and cannot walk away from, is a risk you have not priced yet. What is yours?

Header artwork generated with Swa.

Read more of Ben's Enterprise Field Notes at benpickett.com.