AI · TypeScript, Node.js, PostgreSQL

What did that AI generation cost?

The problem

Nobody on the team could answer a simple question: what did that AI-generated document cost us? Generations ran across Anthropic, OpenAI, Google and Perplexity models, and model names change, get new suffixes, and get retired.

The decision

I built the boring thing: a price table. It covers 56 model families, each mapped to US dollars per million input and output tokens. Every generation stores its token usage, and the cost is worked out when someone reads it.

Three rules in that small file mattered more than the big features.

An unknown model gets no price. The screen shows the token counts and nothing else. The rule is written into the code as a comment: wrong rupees are worse than no rupees. A guess with a currency symbol in front of it is still a guess.

Retired model names never leave the table. A document generated in June must still price correctly in August, even if its model was retired in July. A cost is a historical record, not a live lookup.

Matching is by longest prefix, not exact name. A preview build of a model should price as its family, not fall through to nothing. The whole algorithm is one sort and one loop:

// Illustrative prices, US dollars per million tokens.
const PRICES: Record<string, { input: number; output: number }> = {
  "gemini-2.5-flash": { input: 0.3, output: 2.5 },
  "gemini-2.5-flash-lite": { input: 0.1, output: 0.4 },
};

const KEYS = Object.keys(PRICES).sort((a, b) => b.length - a.length);

export function priceFor(model: string) {
  const key = KEYS.find((k) => model.startsWith(k));
  return key ? PRICES[key] : null; // unknown model: no price, never a guess
}

priceFor("gemini-2.5-flash-lite-preview") finds the lite price, because the longer key is tried first.

What failed

The worst bug was not in pricing at all. Reasoning models count their thinking tokens against the output limit. With a tight limit, the model spent its whole budget thinking and returned an empty string. An empty string parsed as an empty object, and that failed silently further down.

The fix was giving thinking models explicit headroom on top of the limit, so the visible answer always has budget left.

The result

The platform processes hundreds of AI-scored sales calls a day, and anyone can see what any generation cost, in rupees, for each version of a document. The lesson I keep relearning: the highest-leverage AI engineering is often not the prompt. It is the accounting.

All case studies