Tokenomics
Tokenomics connects application token usage to a practical improvement: why it happened, what to change, and what that might save.
Open Tokenomics in the sidebar or /tokenomics. Platform admins also see a Token efficiency | Platform recommendations switch under the page title; see Platform recommendations. Export evidence downloads the findings for the current filters as CSV.
Choose what to analyze
The toolbar sets the scope of the whole page:
- Application — All accessible applications, one application, or Unattributed traffic (requests without an application identity).
- Period — 7d, 14d, 30d or 90d. The UTC date range is shown at the end of the toolbar.
- Environment — All environments,
DEVELOPMENT,STAGINGorPRODUCTION.
When you open Tokenomics from FinOps Analytics with a model or agent filter, a From FinOps chip shows the carried model = or agent = filter. Clear it with the chip's close button. If the selection exceeds the analysis limit, a notice explains that results describe the latest requests only; select an application or a shorter period for a complete view.
The summary at the top shows requests analyzed, applications scored, diagnostic coverage and the modeled savings opportunity for the period.
See where input tokens go
Where input tokens go follows input tokens through three columns: 1. Applications → 2. Context component → 3. Finding. Up to five applications are shown, ordered by measured input; when there are more, the top four are shown and the rest are grouped as N other applications. Hover an application to follow its tokens.
Context components are system instructions, chat history, tool results and other / current request. Each input-side finding draws from one component:
- Repeated system prompts and uncached reusable input: system instructions.
- Long chat history: chat history.
- Large tool results: tool results.
- RAG over-retrieval, duplicate RAG content and low context relevance: other / current request.
Larger findings claim a component's tokens first, and a token is never counted twice. Tokens without a finding flow to No finding. Only input with a gateway breakdown is included; when part of the input has none, the subtitle shows the percentage that does. Existing requests are not reconstructed.
Read the score with its coverage
The Application scores tab lists each application's efficiency score, requests, average input / output tokens per request, observed cost and opportunity. Rows marked Own activity include only your own requests. The application selected in the toolbar, otherwise the one with the largest opportunity, is shown in detail below the table. Select another row to switch the detail without changing the filters.
The detail shows input composition, a daily Tokens per request trend (use Exact values in the table view for daily requests, tokens and USD per 1,000 requests) and Evidence coverage: application identity, diagnostics recorded, context breakdown and scored signal weight. Signal details lists each check's status.
The score measures observed token efficiency, not answer quality or employee productivity. A check without enough assessable requests is labelled Not measured. A score requires ten requests, enough measured signals and at least 20% evidence coverage. Below 80% coverage, it is explicitly partial. How the Token Efficiency Score works explains the formula and weights.
Application managers (ADMIN users and admins of the application's team or organization) can open Efficiency targets below the detail for the selected application and tune the advisory targets and candidate model to the workload. Targets do not change live provider limits or application behavior.
The fourteen checks cover oversized prompts, repeated instructions, chat history, RAG over-retrieval, duplicate retrieval, model fit, excessive output, retries, agent loops, uncached reusable input, repeated embeddings, reasoning fit, large tool results and low context relevance.
Use the action queue
Actions & evidence lists findings across the selected applications, largest modeled savings first. Every finding includes request evidence, a change to try and a quality acceptance check. Application managers can track actions as Open, Planned, Experiment, Resolved or Dismissed, with a note; other users see that only application managers can track the action. Starting an experiment captures a baseline; subsequent traffic appears in a before/after comparison. Changes in workload or pricing can affect that comparison, so it does not certify realized savings.
The savings slider is an explicit scenario. The aggregate takes the largest priced opportunity per request to avoid counting overlapping fixes twice. Estimates are gross and exclude engineering, evaluation, cache writes/storage and quality effects. Unknown pricing stays unknown. Fixed subscription charges remain separate from metered API savings.
Add the missing evidence
After deployment, the gateway automatically captures text component estimates, keyed fingerprints and configured price snapshots without retaining prompt text. For RAG and task requirements, attach an optional zeallm_tokenomics object to the gateway request:
{
"workload": "rag",
"complexity": "low",
"expectedOutputTokens": 300,
"retryAttempt": 0,
"ragChunks": 12,
"ragContextTokens": 6000,
"ragDuplicateTokens": 1200,
"cacheableTokens": 2000
}
The gateway strips this object before forwarding. Use actual counts for your request. Unknown fields are rejected. After evaluating an answer, POST /v1/tokenomics/observations with the same virtual key and the response's x-zeallm-request-id to attach ragChunksUsed or relevantContextTokens, plus an evaluation method (citations, relevance_eval, ablation, manual). Counts are bounded by the originally reported context. A fresh request may need a short retry while its log is persisted.
Citation and relevance metrics do not reveal internal model influence. Validate context pruning with grounded-answer or ablation tests. Multimodal and server-managed conversation state can lack enough visible information for a text breakdown.
See the in-app Developer guide tab for the complete request/observation example. The repository's docs/tokenomics.md contains the scoring formula, exact limits, deployment order and test commands.