Skip to main content

Token Breakdown

One chat turn can involve several LLM invokes: the model asks for a tool, reads the result, asks for another, and finally answers. Each of those invokes resends the whole conversation so far, so a turn’s token count grows in a way the transaction history cannot explain on its own. This endpoint answers “which tokens produced this charge” for one credit ledger row, one row per invoke.
All API requests require a valid JWT token in the Authorization: Bearer <token> header. The API Gateway decodes the JWT and forwards auth context (user-id, organization-id, user-email, x-platform-role, x-org-role) as headers to downstream services.
Superadmin only. Non-superadmin requests receive 403 Forbidden. The endpoint is read-only and writes no audit record.

It reports tokens, not credits

The response carries no per-step credit or USD figure, and that is a deliberate product decision rather than an omission. Credit economics and token economics are driven by separate logic — credits are not derived from tokens — so splitting a turn’s single authoritative debit across the invokes that produced it can only ever be an estimate. An earlier build did split it proportionally by token count; that under-weights output tokens, which cost three to five times more than input tokens, and about a quarter of multi-step turns failed to reconcile against the ledger. The charge itself is already stated on the ledger row this breakdown was opened from. What this endpoint adds is the token decomposition behind it.

Get token breakdown

string
required
The organization_credit_ledger row id — the id of an entry returned by the credit ledger endpoint.
string
Optional organization scope. Omit it. The owning organization is resolved from the ledger entry itself, so a superadmin who has switched organizations still gets the right rows; supplying a mismatched hint short-circuits that resolution and returns an empty list. The parameter exists for direct API callers that already know the owning organization.

Response

boolean
array
One object per LLM invoke, ordered by iteration ascending.
Response

Reading inputTokensComputed against inputTokensReported

Computed counts use the OpenAI tokenizer (cl100k/o200k) for every model. On a Claude or Gemini turn, part of any gap between the two numbers is tokenizer mismatch rather than a real accounting discrepancy. Account for that before reading a gap as signal.
A large gap that is not explained by the tokenizer usually means a component is reaching the model that the breakdown does not attribute — a prompt section added outside the counted path, for example.

Data availability

The underlying table is forward-only with no backfill: historical prompts cannot be reconstructed, so a ledger row from before the feature shipped returns an empty list. Rows are written best-effort after the turn is billed — a failure to record the breakdown never fails a chat turn.

Errors