← Back to the series

Part 5

Why we do not charge per user or per token

Why the way a vendor charges decides how an organization behaves after the purchase, instead of being one more item in the commercial negotiation.

Published on September 07, 20267 min read

CostLicensingAdoption

Every AI vendor presents pricing the same way in the first meeting: you start small and pay as you grow. It is a reasonable offer, and early on it usually comes out cheap. The problem is not the amount charged, it is the behavior the pricing model induces once the contract is signed.

The previous piece in this series dealt with compliance, and this one deals with what comes immediately after. Once the lawfulness of the processing is settled, what still decides whether the project scales is the way it is billed. Processing that is lawful and rationed by seat stays a pilot.

The difference between the price and the incentive

Per-user pricing makes sense for software one person operates alone. With AI, the effect is different. The return shows up when the analyst, the legal team and the operations team all query the same body of material on the same day. Charging per seat makes the organization ration the very tool that only pays for itself at volume, and the pilot with a few dozen licenses never becomes an operation.

Per-token or consumption pricing has a different kind of problem. First, the obvious one. Nobody can predict at the start of a project how many tokens an operation will consume, and the budget turns into an estimate reconciled after the fact. Second, the more serious one. A badly designed flow that reprocesses the same document three times increases the vendor's revenue. Nobody has to act in bad faith for that to bend the product out of shape. It is enough that no one has a reason to fix it. Third, the one least visible on a spreadsheet. When every question carries an apparent cost, people ask only the questions whose payoff they already know, and the organization stops finding out where the tool is useful. The saving that comes from that restraint shows up on the invoice. What it cost in learning shows up nowhere.

The useful question is not "what does it cost per token?". It is "what happens to that bill when it works and everyone uses it, from the shop floor to the C-suite?". The Aptabit license is a flat fee, with no charge per user, per seat, per token or per API call. For public sector organizations there is also a perpetual licensing option, under which the right of use does not expire, and the piece that closes this series covers both what it settles and what it leaves open. That does not make the tool cheap. It makes the bill known before you start.

Where this choice costs the client something

Predictable cost is not low cost, and a piece about pricing that showed only the favorable column would be sales material. Three situations make this arithmetic less comfortable, and they are worth putting to any vendor before signing, including to us.

The incentive on our side. With a flat license, the vendor's revenue does not grow when your usage grows, so nothing in the vendor's own numbers pushes it to make each query cheaper to run. The infrastructure bill left over by that complacency is yours. The question to ask is who pays for the inefficiency in a flow, not whether the inefficiency exists.

Contracted capacity. The license is flat within a capacity defined in the contract, and the infrastructure that sustains that capacity is on the client's account. An organization that grows faster than it projected comes back to the table to revisit that capacity, and the term of the contract is not what decides it. A capacity ceiling and a contract term are separate clauses, and a license with no annual renewal still carries the sizing that was contracted for. Infrastructure has to be sized before it is used, and predictable does not mean small. Capacity cost and consumption cost cross at some point that depends on your usage. If that usage is occasional, it is worth modeling both bills before signing.

Cloud keys. The platform runs local models and also accepts your own keys with external providers. In that case per-token billing comes back, now directly with the provider. The predictable bill applies to the license and does not apply to the inference, and nothing on our side makes it predictable again.

The questions that reveal the real cost

None of these questions depends on who is sitting across the table. Asked of any vendor, they separate one proposal from another:

  1. What happens to my invoice if usage doubles in the second year, and what happens if it drops by half?
  2. What is charged beyond the license, and who pays for the infrastructure?
  3. Does the vendor make more money when I use the platform more?
  4. Is there a capacity threshold beyond which the price changes, and is it written down?
  5. To extend the tool to the whole company, do I need a new budget approval?

The second is the one that stings most on our side, because the answer includes an infrastructure bill that stays with you.

Why this bill is a product decision

There is no invoice of ours that grows with your consumption, and the reason is architectural before it is commercial. Charging by consumption requires a meter somewhere on the path of execution. When the platform runs on your infrastructure and the model runs locally, that path is entirely yours, and execution never passes through us. A flat license, in that design, is not a commercial concession. It is what is left when the execution is yours. The question to ask any vendor is what its product measures about your usage and where that measurement goes.

That shifts work onto the client, who now operates the capacity it used to rent by the call. The other half of the arithmetic is what happens when the contract ends, and that is the subject of a later piece in this series.

Still have a question the series doesn't answer?

Send it to our team. The questions that keep coming back become pieces in this series.

Talk to Aptabit