> ## Documentation Index
> Fetch the complete documentation index at: https://docs.brainstormer.io/llms.txt
> Use this file to discover all available pages before exploring further.

# Token Breakdown

> Per-iteration token decomposition behind a single credit ledger transaction.

# Token Breakdown

One chat turn can involve several LLM invokes: the model asks for a tool, reads
the result, asks for another, and finally answers. Each of those invokes resends
the whole conversation so far, so a turn's token count grows in a way the
transaction history cannot explain on its own.

This endpoint answers **"which tokens produced this charge"** for one credit
ledger row, one row per invoke.

<Note>
  All API requests require a valid JWT token in the `Authorization: Bearer <token>` header. The API Gateway decodes the JWT and forwards auth context (`user-id`, `organization-id`, `user-email`, `x-platform-role`, `x-org-role`) as headers to downstream services.
</Note>

<Warning>
  **Superadmin only.** Non-superadmin requests receive `403 Forbidden`. The
  endpoint is read-only and writes no audit record.
</Warning>

## It reports tokens, not credits

The response carries **no per-step credit or USD figure**, and that is a
deliberate product decision rather than an omission.

Credit economics and token economics are driven by separate logic — credits are
not derived from tokens — so splitting a turn's single authoritative debit
across the invokes that produced it can only ever be an estimate. An earlier
build did split it proportionally by token count; that under-weights output
tokens, which cost three to five times more than input tokens, and about a
quarter of multi-step turns failed to reconcile against the ledger.

The charge itself is already stated on the ledger row this breakdown was opened
from. What this endpoint adds is the token decomposition behind it.

***

## Get token breakdown

<ParamField method="GET" path="/api/auth/billing/ledger/{creditLedgerId}/token-breakdown" />

<ParamField path="creditLedgerId" type="string" required>
  The `organization_credit_ledger` row id — the `id` of an entry returned by the
  credit ledger endpoint.
</ParamField>

<ParamField query="organizationId" type="string">
  Optional organization scope. **Omit it.** The owning organization is resolved
  from the ledger entry itself, so a superadmin who has switched organizations
  still gets the right rows; supplying a mismatched hint short-circuits that
  resolution and returns an empty list. The parameter exists for direct API
  callers that already know the owning organization.
</ParamField>

### Response

<ResponseField name="success" type="boolean" />

<ResponseField name="data" type="array">
  One object per LLM invoke, ordered by `iteration` ascending.

  <Expandable title="Breakdown row">
    <ResponseField name="id" type="string" />

    <ResponseField name="conversationId" type="string" />

    <ResponseField name="messageId" type="string">
      The assistant message this turn produced.
    </ResponseField>

    <ResponseField name="iteration" type="number">
      1-based invoke index within the turn.
    </ResponseField>

    <ResponseField name="botId" type="string" />

    <ResponseField name="organizationId" type="string" />

    <ResponseField name="model" type="string" />

    <ResponseField name="systemPromptTokens" type="number | null">
      First iteration only; `null` on later steps, whose input is the previous
      step's input plus what was appended.
    </ResponseField>

    <ResponseField name="toolSchemasTokens" type="number | null">
      First iteration only. Total cost of the tool definitions sent to the model.
    </ResponseField>

    <ResponseField name="toolSchemasByTool" type="object | null">
      First iteration only. `{ [toolName]: tokens }` — which tool's schema costs
      what.
    </ResponseField>

    <ResponseField name="kbContextTokens" type="number | null">
      First iteration only. Retrieved knowledge-base context, already subtracted
      out of `systemPromptTokens` (it is injected into the system prompt).
    </ResponseField>

    <ResponseField name="kbChunksCount" type="number | null" />

    <ResponseField name="historyTokens" type="number | null">
      First iteration only.
    </ResponseField>

    <ResponseField name="historyTurnsCount" type="number | null" />

    <ResponseField name="userMessageTokens" type="number | null">
      First iteration only.
    </ResponseField>

    <ResponseField name="attachmentTokens" type="number | null">
      First iteration only. Images are counted at a flat approximation — real
      per-image cost varies by model and resolution.
    </ResponseField>

    <ResponseField name="inputTokensComputed" type="number | null">
      Our own count of the input, summed from the components above. First
      iteration only.
    </ResponseField>

    <ResponseField name="inputTokensReported" type="number">
      The provider's own input count for this invoke.
    </ResponseField>

    <ResponseField name="outputTokens" type="number" />

    <ResponseField name="cacheReadTokens" type="number | null" />

    <ResponseField name="cacheWriteTokens" type="number | null" />

    <ResponseField name="reasoningTokens" type="number | null" />

    <ResponseField name="totalTokens" type="number" />

    <ResponseField name="responseType" type="string">
      `tool_call` or `text`.
    </ResponseField>

    <ResponseField name="invokeReason" type="string">
      `initial`, `tool_round`, `round_cap` (tools withheld at the round cap) or
      `empty_retry` (the model returned nothing and was re-prompted).
    </ResponseField>

    <ResponseField name="toolCallsJsonb" type="array | null">
      `{ id?, name, args }` — what the model asked for on this step.
    </ResponseField>

    <ResponseField name="toolOutputMessagesJsonb" type="array | null">
      `{ toolCallId, name, tokens }` — the tool results appended after it.
    </ResponseField>

    <ResponseField name="appendedAssistantTokens" type="number | null">
      Steps 2+. The preceding assistant tool-call message added to the context.
    </ResponseField>

    <ResponseField name="appendedToolMessageTokens" type="number | null">
      Steps 2+. The tool results added to the context. Usually the dominant term
      in a turn that grew unexpectedly.
    </ResponseField>

    <ResponseField name="externalCostEventId" type="string | null">
      The `external_cost_events` row for the turn's provider charge. Note this is
      **not** a credit ledger id.
    </ResponseField>

    <ResponseField name="createdAt" type="string" />
  </Expandable>
</ResponseField>

```json Response theme={null}
{
  "success": true,
  "data": [
    {
      "id": "6f1c…",
      "conversationId": "c2a1…",
      "messageId": "9b77…",
      "iteration": 1,
      "botId": "b01f…",
      "organizationId": "0d3e…",
      "model": "openai/gpt-4o",
      "systemPromptTokens": 1200,
      "toolSchemasTokens": 800,
      "toolSchemasByTool": { "get_products": 800 },
      "kbContextTokens": 2500,
      "kbChunksCount": 3,
      "historyTokens": 1500,
      "historyTurnsCount": 4,
      "userMessageTokens": 30,
      "attachmentTokens": 0,
      "inputTokensComputed": 6030,
      "inputTokensReported": 6030,
      "outputTokens": 150,
      "totalTokens": 6180,
      "responseType": "tool_call",
      "invokeReason": "initial",
      "toolCallsJsonb": [{ "name": "get_products", "args": { "category": "shoes" } }],
      "externalCostEventId": "3a90…",
      "createdAt": "2026-09-29T10:00:00.000Z"
    },
    {
      "id": "7d24…",
      "iteration": 2,
      "systemPromptTokens": null,
      "inputTokensComputed": null,
      "inputTokensReported": 7500,
      "outputTokens": 400,
      "totalTokens": 7900,
      "responseType": "text",
      "invokeReason": "tool_round",
      "appendedAssistantTokens": 150,
      "appendedToolMessageTokens": 1200,
      "externalCostEventId": "3a90…",
      "createdAt": "2026-09-29T10:00:02.000Z"
    }
  ]
}
```

## Reading `inputTokensComputed` against `inputTokensReported`

<Warning>
  Computed counts use the OpenAI tokenizer (cl100k/o200k) for **every** model.
  On a Claude or Gemini turn, part of any gap between the two numbers is
  tokenizer mismatch rather than a real accounting discrepancy. Account for that
  before reading a gap as signal.
</Warning>

A large gap that is *not* explained by the tokenizer usually means a component
is reaching the model that the breakdown does not attribute — a prompt section
added outside the counted path, for example.

## Data availability

The underlying table is **forward-only with no backfill**: historical prompts
cannot be reconstructed, so a ledger row from before the feature shipped returns
an empty list. Rows are written best-effort after the turn is billed — a failure
to record the breakdown never fails a chat turn.

## Errors

| Status | Meaning |
| - | - |
| `400` | Missing `Authorization` header, or no organization could be resolved |
| `401` | Missing, malformed, or expired token |
| `403` | Caller is not a platform superadmin |
| `500` | Lookup failed |


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.