Sang's Blog

Tokenmaxxing and the price of a token

The term “tokenmaxxing” entered the tech vocabulary sometime in 2025, born from a peculiar practice: companies publishing leaderboards of how many AI tokens each employee consumed, tying performance reviews to token burn, and sometimes spending more on API calls than on salaries. The reaction was predictable. The tech press called it Tulip Mania. Executives defended it as necessary to force adoption. Engineers called it the stupidest metric ever devised. Both sides had evidence for their position, and neither was entirely wrong. But both were debating the surface of a phenomenon that was never really about AI adoption at all.

The token is a strange unit of account. It is not a measure of output — no one has ever shipped a feature per token, or closed a deal per token, or fixed a bug per token. It is a measure of processing, of thinking, of the raw computation that goes into producing an answer. A token is to cognition what a watt is to a machine: a unit of energy consumed, not of work done. And the tokenmaxxing phenomenon, at its core, was the first large-scale experiment in treating cognitive energy as a line item on a balance sheet. Not the energy of human workers — that has always been abstract, priced by salary, hard to meter granularly. But the energy of machine cognition, which can be metered down to the individual token and billed to the cent.

This is a genuinely new development in the history of knowledge work. Every previous tool of the knowledge worker — the book, the compiler, the spreadsheet, the IDE, the programming language, the operating system, the cloud instance — was priced as a fixed cost or a subscription. You bought the book once. You paid for the IDE license annually. You provisioned a server and paid for uptime, not for each instruction it executed. The cost of using the tool was decoupled from the cost of thinking with it. You could spend an hour in an IDE producing nothing useful and the tool would not charge you extra. You could iterate a hundred times on a design and the cloud instance would cost the same as if you had iterated three times. The tools were infrastructure for thought, and like any infrastructure, their cost was sunk, not variable.

AI tokens broke this arrangement for the first time. Every request to a language model is individually metered. Every prompt, every refinement, every correction, every loop in which the model retries and fails and retries again — each one adds to the bill. The cost of thinking with the tool is now a variable cost that scales with the quantity of thinking done. And this is the deeper truth that tokenmaxxing exposed, even if nobody involved in the debate articulated it: the unit of account for AI is input, not output. You pay for how much the machine thinks, not for what it produces. And when you build an incentive system around an input metric, you get exactly what tokenmaxxing produced — people maximizing the input because they have been told, implicitly or explicitly, that input is what matters.

The mistake in the public debate is to ask whether tokenmaxxing was strategic or stupid. The strategic camp points out that it forced employees to learn AI tools, that familiarity is a precondition for adoption, that the spend was an investment in organizational learning. The stupid camp points out that burning money on tokens without measuring output is obviously idiotic, that any manager who thinks token consumption correlates with productivity has never managed anything, that the whole episode was FOMO dressed up as strategy. Both camps are making valid observations about the same phenomenon, and both camps are missing the structural condition that made the phenomenon inevitable.

The condition is the incompatibility between how knowledge work has historically been valued and how AI is priced. Knowledge work is valued by output: the feature that ships, the deal that closes, the problem that gets solved, the decision that turns out to be correct. Output is difficult to measure, context-dependent, and often delayed — a decision made today may prove right or wrong only months later. But it is the only meaningful unit of value in knowledge work, because knowledge work is defined by the production of non-standard outcomes. If the outcome were standard, the work could be automated. AI, by contrast, is priced by input. The API bill is a function of tokens consumed, not of features shipped or bugs fixed or revenue generated. A team that ships a major feature in one API call pays less than a team that spends a thousand calls iterating on a trivial script that never gets deployed. The pricing model rewards volume of machine thinking, not quality of human outcome.

This incompatibility is not a bug in the pricing model. It is a feature of the technology. Language models are stochastic — they produce different answers to the same question at different temperatures, with different contexts, with different seeds. The reliability of the output depends on the quantity of computation invested. Spend one token on a complex problem and the answer is likely wrong. Spend a thousand tokens iterating and refining, and the probability of correctness rises. The relationship between input and output is real. But it is probabilistic, not linear. You cannot say “this feature cost fifty thousand tokens to build” the way you can say “this house cost fifty thousand bricks to construct.” The tokens do not map to units of value. They map to units of search, of exploration, of trial and error. Paying for tokens is paying for the right to search a space of possibilities. The value emerges from the search, but the cost is in the search itself.

The perverse outcome that tokenmaxxing produced — employees pumping tokens through the system regardless of whether the output was useful — is not a failure of management or a sign of moral weakness in the workforce. It is the predictable result of an incentive system that rewards a measurable input while the corresponding output remains unmeasurable in the same accounting framework. The employee who burns tokens on endless loops is not gaming the system. They are responding to the system exactly as the system was designed. The leaderboard measures what it can measure. What it cannot measure — the quality of the work, the soundness of the decisions, the durability of the code — disappears from the incentive structure. And what disappears from the incentive structure stops happening at scale.

This pattern is not new. It has appeared in every domain where a proxy metric replaced a true metric. In healthcare, billing by procedure produces more procedures, not healthier patients. In education, teaching to the test produces higher scores, not better learning. In advertising, paying by click produces more clicks, not more conversions. The pattern is so well documented that it has a name — Goodhart’s Law, Campbell’s Law, the cobra effect, depending on which tradition you cite. The mechanism is always the same: when a measure becomes a target, it ceases to be a good measure. Tokenmaxxing is not a novel failure of management. It is the same failure that has been discovered and rediscovered in every domain where input and output can be decoupled. The only novelty is the unit of account. The token is new. The pattern is as old as measurement itself.

The significance of tokenmaxxing is not that some companies wasted money on AI. The significance is that the waste exposed a structural incompatibility that will not go away when the leaderboards are removed. The leaderboards were a crude and stupid instrument, but the underlying condition — input-based AI pricing meeting output-based knowledge work — is permanent. Every company that adopts AI will eventually confront the same tension: how do you pay for a tool whose cost scales with use when the value of the use is not proportional to the cost? The answer cannot be “don’t use AI” — the technology is too useful, and the competitive pressure is too strong. The answer cannot be “measure output” — output in knowledge work is famously hard to measure, which is why we pay salaries instead of per-feature bounties. And the answer cannot be “trust employees to use the tool responsibly” — trust is not a scalable coordination mechanism in large organizations, which is why metrics exist in the first place.

The real answer, if there is one, is that the pricing model for AI will have to change. It will have to move from input-based pricing toward something that better aligns with the value the tool produces. This could take many forms — subscription pricing that decouples cost from usage, outcome-based pricing where the API provider shares the risk, agentic pricing where a fixed fee covers a bounded task regardless of the tokens consumed. The industry is already moving in this direction with features like max spend caps and project budgets. But these are patches on a fundamentally input-based model, not a reconceptualization of what is being paid for. The token is a convenient unit for the provider because it maps directly to compute cost. It is a terrible unit for the customer because it maps to nothing that the customer values.

The tokenmaxxing episode will be remembered as a farce — a brief period when supposedly rational companies measured performance by how much they paid for machine thinking. But the farce is not the story. The story is what the farce reveals: we have built a tool that can think, and we are paying for it by the thought, and we have no idea how to reconcile that payment model with a world where human thinking is not metered, where knowledge work is valued by what it produces, not by how many times the worker iterated. The token is the first unit of account for machine cognition. It will not be the last. But the way we resolve the tension between input-based pricing and output-based value will determine not just how AI is used in organizations, but how knowledge work itself is structured in a world where thinking can be metered. The leaderboards are gone. The problem remains.

← Prev Post Next Post →