7 min read
A friend of mine is a senior leader inside a large organization. We were talking about the company’s AI policy when he told me the monthly AI inference budget for a senior engineer.
It was worth less than one Starbucks drink every other working day. Not quite the whole coffee, either.
That was the AI Coffee Allowance.
He could probably buy his whole team a premium coffee every morning, put it on his corporate credit card, and nobody would care. Coffee may accelerate a developer. AI inference may accelerate the same developer. One expense fits neatly inside the culture budget. The other is apparently a bridge too far for a company describing AI as strategic.
I do not know who approved the number. I assume it was the same person responsible for the toilet paper in the corporate bathroom. Procurement can prove it is on the roll and costs eleven percent less per sheet, but after using it you still have one question: is that even toilet paper?
My friend had also been encouraged to coach anyone approaching the cost of one decent cup of coffee a day in inference. Those engineers needed to reduce their spending. If they did not, disciplinary action could follow.
This is a real story. I am leaving out his name and the company because neither one is the point. A six-dollar model run now receives more management attention than six weeks of delivery delay.
I asked if he was worried. He was not. Despite the silliness, he was still going to receive his maximum bonus. The organization was hitting the measures used to calculate his compensation.
The policy made no sense, and the incentive system was telling him the company was performing exactly as designed.
Since we are friends, this is where the conversation became sarcastic.
“If the organization cannot afford it, turn it all off,” I told him. “Turn off all the AI.”
That started as a joke. Then I imagined the next all-hands.
The CEO opens with the first slide: WE’RE ALL IN ON AI. AI will reinvent the company and define the next era of work. Every leader should model the behavior. There may even be applause.
Click.
The next slide says: BUT DON’T SPEND A DOLLAR.
In smaller type: AI inference has exceeded the AI Coffee Allowance. Please return to building the future manually.
Then the CEO hands the meeting to the head of the AI Center of Excellence, who explains that adoption remains below target and the culture needs to change. The Q&A ends early because there is no time for questions.
That all-hands is fictional. The contradiction is not.
Your Center of Excellence Is Becoming a Center of Normalcy
Policies outrank speeches. Engineers know which message carries consequences.
Most readers also read: If you cannot afford the tokens, can you afford to build it?
They use weaker models and spend human time repairing the result. They stop an agent before the work is finished. Some will move to personal accounts, converting a cost-control exercise into an unlogged security problem. The careful ones do less and wait.
The dashboard improves.
This is how an AI Center of Excellence becomes a Center of Normalcy (CoN), perhaps the first transformation office named honestly. Its job is no longer to discover better ways of working. Its job is to prevent anyone from behaving differently enough to create an uncomfortable number.
The CoN can eliminate a six-dollar variance, preserve a six-week delay, and report that the rollout remains on track. That is the operating model.
A Six-Dollar Model Run Needs Four Minutes
Take a senior engineer with a fully loaded annual cost of $180,000. Divide that across 2,080 working hours and the company spends about $87 an hour for the engineer’s capacity.
A six-dollar inference run needs to return roughly four minutes to break even on labor capacity.
The coaching conversation costs more than the alleged problem. A manager who spends fifteen minutes explaining why an engineer used seven dollars instead of six has already consumed several months of one-dollar savings. Add the usage report, exception request, and disciplinary paperwork and you have corporate performance art with a spreadsheet attached.
The model dollar has a meter. The human hour hides in payroll. One receives governance while the other disappears into the quarter.
I am not arguing for unlimited inference. Govern model routing, data boundaries, risk, and accepted outcomes. The AI Coffee Allowance does none of that. It is anxiety with a receipt.
Turn It Off Without Cheating
If leadership believes the expense is too high, stop debating it. Turn off the AI used to build software for 90 days.
Keep the commitments, people, quality bar, service levels, and security requirements fixed. Do not replace the capability with a vendor team, personal subscriptions, reduced testing, or delivery dates quietly moved into the next quarter.
Record the previous 90 days, then measure accepted business outcomes, lead time, aging work, defects, incident recovery, senior review effort, external capacity, customer impact, and total cost to production.
The experiment will not be academically clean. It only needs to be more useful than staring at an invoice and deciding the line feels large.
There are three results worth discussing.
If Performance Improves, Leave AI Off
Your AI program may be producing more software than customers need, increasing review load, degrading quality, or accelerating weak portfolio decisions. Removing it may reduce work in progress and stop the company from manufacturing features nobody wanted.
If accepted outcomes improve and total delivery cost falls, leave AI off. Cancel the tools. Then ask why faster software production was making the business worse.
AI did not choose the destination. Leadership did.
If Performance Declines, Price the Lost Capacity
If cycle time rises, accepted output falls, incidents take longer to resolve, or external capacity returns, price the difference.
Do not compare the AI invoice with zero. Compare the AI-assisted production system with its replacement. If eliminating $100,000 of inference removes $500,000 of delivery capacity or sends the work to a consulting partner carrying a seven-figure statement of work, the original expense was not high. It was badly attributed.
Restore the capacity where the economics support it. Give teams outcome-based envelopes. Use cheaper models for routine work and better models when finishing safely matters more than making the usage dashboard look tidy.
If Nothing Changes, Software Was Never the Constraint
This is the result an executive team should fear.
The tool may be useless. It may also be improving two days of coding inside a 28-day system where work waits for architecture, security, environments, product clarification, release coordination, and a meeting to schedule the next meeting. Making coding twice as fast does not rescue the other 26 days.
Paul David’s history of factory electrification describes factories that replaced a central steam engine with one electric motor while keeping the belts, shafts, layout, and management system designed around steam. The larger gains arrived when factories reorganized around distributed power.
AI gets trapped the same way. Add it to one step, preserve every queue around that step, and business performance stays flat.
Someone will probably claim this proves the company produces bespoke, handcrafted software like the finest handmade furniture. People pay more for a hand-cut walnut table because they can see the craft, value the material, and want the builder’s signature underneath it. Nobody pays a premium for an inventory workflow because a senior engineer lovingly typed every character while waiting for Architecture Review Board approval.
You have not pioneered handcrafted software. You have ordinary software that costs more and arrives later. The customer does not frame the release notes and pass them down to their children.
(Editor’s note: Yes, I remember the software craftsmanship movement. Art and software were apparently always destined to meet here. I am probably not scoring many points with the craftsmanship crowd, but if artisanal code now commands an artisanal price, I would like to see the price list.)
My Advice to Him Was Not to Fix the Company
The 90-day shutdown is my advice to the executive team. It was not my advice to my friend.
My formal advice to him was simpler: do not bring this argument to leadership.
The organization has already explained what it values. It created the AI Coffee Allowance, attached disciplinary language to it, and still awarded him the maximum bonus. This is not a misunderstanding waiting for a better slide. The system is working for the people who own it.
He should follow the policy, do his job, and avoid becoming the unpaid transformation office for a company that does not want to transform. He should not move company code or data into unapproved tools. He should use approved access, personal projects, open source, and non-company work to keep learning how to build with AI.
Then he should use that skill to get a better job.
This does not sound like an organization he can fix from his seat. It may not be an organization he can influence at all. The rational move is to collect the bonus, build the capability the market will value, and leave before the Center of Normalcy adds AI learning to the list of reimbursable beverage violations.
The executive team can keep protecting the system that made better software production irrelevant. My friend does not have to keep building his career inside it.
Companion
