A minimalist cover on a near-black canvas: a solid cream semicircle on the left dissolves rightward into a grid of small squares that thin out and scatter — some solid, some outlined — like capacity breaking into billable units. A fine ruler scale runs beneath, labelled 'Cognition, metered' in spaced white capitals, with a small 'Fig. 01' mark top-left.

When cognition becomes metered

A firm can now rent increments of thinking.

Not metaphorically. When an organisation pays for tokens, it is buying bounded units of cognitive work — drafting, review, translation, retrieval, pattern-recognition — metered and priced the way we already price compute and storage. Whether machines “think” is a question for another discipline. What matters commercially is that something which behaves like thinking can be purchased by the unit, on demand, with no notice period, no onboarding and no exit interview.

That sounds like a pricing detail. It isn’t. It changes what the cost structure of a firm is made of.

Three pools of capacity

For most of the modern corporate era, a firm’s productive capacity came in two forms, and finance built its vocabulary around them. People: headcount, utilisation, attrition. Software and machines: licences, assets, depreciation. Hire, buy, amortise. That was the toolkit.

There is now a third pool. If people are embodied judgement and relationships, and software is frozen logic, then metered cognition sits between the two: configurable, rented per unit, endlessly elastic — and, in ways I’ll come back to, not actually yours.

Each pool has a different character of cost. Human capacity is lumpy and path-dependent; you hire in whole people, they take months to become useful, and their value compounds in ways that never appear on an invoice. Software is capex-shaped: expensive to build or buy, cheap to run, brittle when the world changes. Metered cognition is pure variable opex, expensed as consumed.

Notice, in passing, what that last point does to the accounts. A firm can materially expand its thinking capacity and leave no trace on the balance sheet. No asset, no depreciation schedule, no headcount line — just a subscription and usage charge dissolving into operating costs. The capacity is real; the accounting is nearly silent. We have been here before with cloud computing, but cloud replaced servers. This replaces, or at least reshapes, work.

What a unit of thinking actually costs

Take something mundane: reviewing an inbound NDA.

Done by a paralegal on £45k — call it £60k fully loaded — that’s roughly £32 an hour. A standard NDA takes perhaps 45 minutes: £24 of labour, plus whatever the two-day queue in legal’s inbox costs the deal it’s holding up.

Done by a frontier model, the same review consumes maybe 25,000 tokens. Call it 40 pence.

A sixty-fold saving, apparently. And almost entirely beside the point — because 40p is the cost of the tokens, not the cost of the outcome. The model’s output is probabilistic, so someone senior still spot-checks it: ten minutes of a solicitor’s time, £15 or more. Add the escalation path for the reviews it gets wrong, and the one-off cost of designing the checking process in the first place, and the realistic cost per trustworthy review lands somewhere around £8–15. Still much cheaper, and dramatically faster. But the number that matters was never on the token invoice.

This is the first correction most organisations need to make. The economic object is not the token; it is the outcome a sequence of tokens contributes to — a contract reviewed, a forecast updated, a risk flagged. The question “what do tokens cost?” is procurement. The question “what does it cost to reach a satisfactory outcome when part of the cognition is rented?” is strategy. The token line is noise. Cost-to-outcome is the signal.

Scale stops behaving

Under the old model, more thinking meant more people, and people came with friction: hiring, training, culture, management span. Those frictions were annoying, but they also acted as a natural brake. Nobody accidentally hired forty analysts.

Metered cognition has almost no friction, which means the brake is gone. It can be dialled up overnight — and it will be, because cheap inputs obey Jevons’ logic: when something becomes dramatically cheaper, consumption expands to swallow the saving. Suddenly there is a machine-written post-mortem on every ticket, a summary of every meeting, a report nobody asked for and nobody reads. The token bill grows while the stock of decisions actually improved by it does not.

So the binding constraint shifts from availability to design. Not “can we afford this capacity?” but “where in the system does it belong, and which outcomes justify any spend at all?” That is a harder question, and it cannot be answered by a rate card.

Capacity you don’t control

Here is the part of the CFO brief that gets the least airtime and deserves the most.

When capacity is people, you have levers: retention, contracts, culture. When it’s a perpetual software licence, version 11 keeps behaving like version 11 for as long as you run it. Metered cognition offers neither comfort. The vendor can reprice it. They can deprecate the model your workflow was tuned to — this already happens on cycles of months, not years. They can change its behaviour with an update, mid-quarter, without asking. The characteristics of your capacity can shift without your consent, and there is no version you can pin and own.

Layer the probabilistic output on top and the right mental model becomes obvious: this is a sole-supplier dependency for a critical input, and it should be governed like one. Second-sourcing, abstraction layers so a model swap is a configuration change rather than a rebuild, fallback paths for the processes that cannot tolerate an outage. Supply-chain discipline, applied to thinking itself. Very few firms have made that connection yet.

Where finance ends up

One more consequence, and it is the quietest. Once cognition can be bought, it can be benchmarked. Internal work that looked “good enough” for years starts to look slow, noisy or inconsistent when set against a metered alternative. The comparison will not always favour the machine — often it shouldn’t — but it will increasingly be made, and finance will be the function holding the numbers when it is.

Which means a budgeting decision about token spend is never just a budgeting decision. It is, implicitly, a judgement about which problems are still worth solving primarily through people, and which are better treated as flows of information through a hybrid of code and rented intelligence. The firms that get this right won’t talk much about tokens. They’ll talk about the shape of work.

Once cognition can be bought in units, the cost structure of the firm is partly a philosophical choice. You are deciding what you believe human minds are for.