# You’ve All Gotten This Email

Source: https://agentdrivendevelopment.com/youve-all-gotten-this-email/
Agent-readable URL: https://agentdrivendevelopment.com/youve-all-gotten-this-email/?agent=1
Published: 2026-08-08T13:00:00-05:00
Modified: 2026-08-08T17:23:20-05:00
Attribution: If you quote, paraphrase, summarize, or cite this material, credit agentdrivendevelopment.com and link to the source URL above.

## Summary

Your AI cost memo sounds sensible until you send it to San Francisco taxi drivers and ask them to avoid hills, skip breaks, and push the cab.

## Article

Once again, this is why IT should not own the token budget. To IT, tokens are not machine labor, productive capacity, or a way to ship faster. They are a line on an invoice. Success means making that line smaller — even when the company produces less right along with it.

You’ve all gotten this email.

Someone has analyzed the AI bill. They have spoken with the heaviest users. They have found the expensive sessions, the premium models, and the cache misses. They have spent weeks of salaried labor and millions of tokens reaching this point. Then they send the entire company a clearly AI-generated email explaining how everyone else should use less AI.

Here is the memo, rewritten to remove identifying details while preserving its argument:

> Subject: A Few Guardrails for Responsible AI Spend Everyone, We recently reviewed our AI consumption patterns and spoke with several frequent users about how they work. Three behaviors explain most of the variance in our bill: very large conversation histories, expired prompt caches, and using more model than the task requires. Keep Working Contexts Manageable Every new exchange in a long-running thread has to account for more previous material. Once a session grows beyond roughly 175,000 tokens, the cost curve becomes noticeably steeper. During the last three-week sample, sessions in this range represented about 68% of AI spend — approximately $8,940 of $13,150. When a thread becomes unwieldy, use /compact to preserve a shorter working summary. If the old context is no longer useful, use /clear and begin a new session. Protect Cached Work Cached context is economical only while it remains available. When a large session sits idle long enough for the cache to expire, the same material must be processed again at full cost. Our average cached block is currently reused about 29 times. Where practical, group related AI work together, avoid abandoning large active sessions, and finish the current line of reasoning before taking an extended break. Match Capacity to the Assignment Fast tier: Use for extraction, reformatting, classification, and straightforward mechanical changes. General tier: Use for most engineering work, including implementation, refactoring, and code review. Deep-reasoning tier: Use for difficult architecture decisions, ambiguous systems problems, and debugging that genuinely requires more reasoning capacity. Usage Snapshot: The general tier completed roughly 52,000 requests for about $2,100 during the review period. The deep-reasoning tier processed far fewer requests while consuming a larger share of the budget. Our review suggests that many of those requests could have used a less expensive tier without affecting the result. Older Defaults: A small group of accounts remains pinned to earlier-generation models. Those sessions contributed approximately $4,750 during the same period and cost about 27% more per request than current equivalents. Please review saved defaults before starting new work. We want teams to continue using AI wherever it helps them move faster. The request is simply to make context, continuity, and model capacity deliberate choices rather than expensive defaults.

It is specific, numerate, and written in the calm language of operational competence. It says nothing about what those 52,000 requests produced. It does not ask whether the $13,150 saved $20,000, $200,000, or nothing at all.

Was this sent by leadership or the AI police? Leadership asks what the investment produced. The AI police inspect your context window, monitor your lunch break, and issue a citation for using too much intelligence.

To see the problem, change the nouns and send the same management logic to a taxi company in San Francisco.

Subject: A Few Guardrails for Responsible Taxi Fuel Spend

Team,

Recently, I’ve been analyzing our San Francisco fleet’s fuel-utilization metrics and chatting with some of our highest-mileage drivers. Moving forward, I want us all to be mindful of three key factors that drive our costs: hills, trip continuity, and transportation-method selection.

Managing Hills

The longer a taxi travels uphill, the more expensive the trip becomes as the engine works harder. We see a significant fuel spike once climbs exceed 175 feet. In fact, roughly 68% of our total gasoline expenditure over the past three weeks came from trips over Nob Hill, Russian Hill, and Pacific Heights — $8,940 of our $13,150 total spend.

If a passenger asks to travel uphill, please suggest a flatter destination — the Embarcadero is lovely. If the hill is unavoidable, pull over and push the taxi until the road levels out. This will protect the fuel budget while allowing us to say we still support transportation throughout San Francisco.

Trip Continuity and Warm Engines

Whenever an engine cools down, we have to pay the full premium to warm it up again. If you pause mid-shift to grab lunch and return an hour later, you’re essentially buying that warm engine twice.

Stopping also wastes momentum, and rebuilding momentum consumes fuel. Effective immediately, stop signs are optional. Bathroom breaks are optional too. Legal exposure, driver health, and other non-fuel metrics will be reviewed after quarter close.

Right now, our typical warmed-up taxi completes about 29 trips before cooling. We want to protect and improve that efficiency. Please batch your driving into concentrated sprints, complete your current passenger journey before stepping away, and schedule any remaining human needs around the engine’s cache.

Selecting the Appropriate Transportation Method

To put this in perspective, our fleet handled 52,000 trips for just $2,100 recently. A deeper review found that many of those trips were less than one-third of a mile. Those passengers simply did not require that level of horsepower.

Going forward, please park the taxi and carry the passenger for trips under one-third of a mile. If the passenger has luggage, make two trips. This will allow us to reserve the taxi for highly complex transportation scenarios that genuinely require an automobile.

We also have several drivers defaulting to an older, pinned fleet of Cadillac Escalades. Yes, the Escalade fits the passenger and the luggage in one trip. But for a journey under one-third of a mile, you could have used a bicycle. Those Escalade trips accounted for $4,750 of our spend in the same period, despite costing roughly 27% more than lighter alternatives. We appreciate that the passengers arrived safely with their belongings. Arrival, however, is not the metric under review.

The goal here isn’t to discourage you from driving taxis. We absolutely want you moving passengers. We’re simply asking for a bit more strategic thinking before you press the accelerator.

Subject: Re: A Few Guardrails for Responsible Taxi Fuel Spend

Thanks for the clarification.

Apparently, saving money is the reason we all come to work. I had assumed it was getting paying passengers where they needed to go, but that appears to be a legacy metric.

One question before we get back to carrying passengers: how many tokens and paid human hours did you spend writing this email about spending fewer tokens? Please include the dashboard analysis, power-user interviews, drafting, revisions, approval chain, and the meeting where someone decided /clear was an operating strategy. Break the total down by model. It would be embarrassing if the cost-control memo cost more than the waste it found.

Going forward, we will avoid hills. When we cannot avoid them, we will push the taxi. This should improve the gasoline dashboard considerably. We will report fares, missed pickups, trip duration, driver injuries, and passenger retention separately — assuming anyone is still interested.

The AI memo sounds smarter because tokens are less visible than gasoline. The economic mistake is identical. Which raises the less comfortable question: did the sender of this memo ever understand AI, or did they merely learn how to sort an invoice from highest to lowest?

Cost matters. Waste matters. Model choice matters. But a company does not employ engineers to minimize inference any more than a taxi company employs drivers to preserve gasoline. It buys the input to produce an outcome.

If your dashboard can tell me that a session cost $42 but cannot tell me what the session built, repaired, accelerated, or prevented, you do not have an AI cost problem yet.

You have a management accounting problem — and your engineers are about to start pushing the taxi.
