Claude Code's own /usage command tells you what a session cost in total, split by model and by token type. It does not show you what was inside each request. To see the actual system prompt, every tool schema, each MCP server's contribution, the message history and the cache reads and writes, and to put a price on each of those parts, you have to look at the traffic between Claude Code and the API. The simplest way to do that on your own machine is a local proxy that records each request, redacts secrets, and prices it afterwards. That is what our open-source cost-xray does. The rest of this article covers both routes: the built-in commands that answer "how much?", and request-level capture that answers "on what?".
how to see exactly what Claude Code sends to the API and what each part costs?
Put a recording proxy between Claude Code and the model API, then price what it captured. The built-in tools stop at totals, so the request itself is the only place the per-part answer exists.
Every Claude Code turn is one or more Messages API calls. Each call carries the same large prefix (the system prompt, the definitions of every tool the agent may use, the instructions from any connected MCP servers, your CLAUDE.md) followed by the conversation so far and the latest tool results. Anthropic's tool use pricing is explicit that tool names, descriptions and schemas in the tools parameter are billed as input tokens, as are tool_use and tool_result blocks. So the question "what does each part cost?" is really "how many tokens does each block contribute to each call, and at which rate was that block billed?"
What cost-xray captures
What it does: records what Claude Code and Codex send to the model API and attributes the tokens and the cost of each call to the source that caused them.
What it needs: macOS or Linux. No API key, no account, no change to how you launch the agent.
The call:
curl -fsSL https://raw.githubusercontent.com/tigerless-labs/cost-xray/master/install.sh | bashOpen a new terminal, then run the agent the way you always do:
claudeWhen you want to look, run cx from any directory to open the live terminal UI.
What comes back: a drill-down that goes agent, project, session, category, MCP server, tool, per-turn call and finally the output. At each level you see context-window occupancy by source and cost split into fresh input, cache read, cache write and output. That is where an MCP server whose tools were injected into every call but never used shows up, or a tool result that was re-sent for thirty turns after it stopped mattering.
What it costs: nothing to run; it is MIT-licensed. The capture adds a local hop on 127.0.0.1, and pricing happens offline in a separate step, off the live path.
How accurate the split is
The per-call total is exact, because it is calibrated to the usage figures the provider returns with each response. The split between sources inside one request is an estimate: cost-xray sizes each block with a tokenizer, applies targeted corrections to blocks that are usually mis-sized (thinking blocks, tool schemas), then scales everything so the parts sum to the provider's reported total. Treat "this MCP server is about a fifth of your input" as a reliable direction, not an invoice line.
We build cost-xray, so weigh our description accordingly; the source is on GitHub, and the longer walkthrough of why we wrote it is in our launch post on seeing what Claude Code and Codex actually send.
What to do with what you find
Once you can see the parts, the fixes are ordinary and mostly free. Anthropic's cost guidance for Claude Code lists the same levers the capture points at:
A server you never call: disable it in
/mcp, or prefer a CLI such asghthat adds no per-tool listing.A long CLAUDE.md: move workflow-specific instructions into skills, which load only when invoked. Our autoharness plugin turns repeated sessions into skills for exactly this reason.
History that keeps growing:
/clearbetween unrelated tasks, and/compactwith instructions when you need continuity.Cache misses after breaks: the first message after the cache lifetime expires reprocesses the full context, so a cold cache is a cost event you can see in the capture.
claude code check usage?
Run /usage inside Claude Code. On an API key it shows the current session's cost; on a subscription it shows your plan usage bars and what is eating them.
According to Anthropic's documentation, the Session block at the top of /usage lists total cost, API and wall-clock duration, code changes, and usage by model broken into input, output, cache read and cache write tokens. Claude Code computes that dollar figure locally from token counts at list price, so it is an estimate; the Claude Console usage page remains the authoritative bill for API users.
Recent versions add two things worth knowing:
| What you want to know | Where /usage shows it | Who sees it |
|---|---|---|
| Session cost and token types | Session block | API-key users (subscribers see it but it is not their bill) |
| Cache health | Prompt cache (main) line: share of input from cache, misses, warm or cold | Everyone, main conversation only |
| What counts against plan limits | Attribution by skill, subagent, plugin and MCP server, as percentages | Pro, Max, Team, Enterprise |
| Behaviours driving usage | Flags for long context or cache misses above 10% of recent usage | Pro, Max, Team, Enterprise |
| Scheduled task cost | Loops rows, heaviest first | Pro, Max, Team, Enterprise |
Two companion commands fill in the rest. /context shows what is occupying the context window right now, which is the fastest way to spot a bloated prefix. /insights writes an HTML report about how you work across recent sessions rather than about tokens.
The gap is granularity. /usage attributes plan usage to an MCP server as a share of recent requests; it does not show you the schema text that server injected, how many tokens that text was on call 47, or whether it was billed as a cache write or a cache read. If the totals look wrong and /context does not explain why, that is the point at which request-level capture earns its place.
claude code token usage monitor?
For one developer, a local capture tool or the status line is enough; for a team, turn on Claude Code's built-in OpenTelemetry export and send the metrics to a dashboard you already run.
Claude Code's monitoring documentation describes the switch:
export CLAUDE_CODE_ENABLE_TELEMETRY=1
export OTEL_METRICS_EXPORTER=otlp
export OTEL_EXPORTER_OTLP_PROTOCOL=grpc
export OTEL_EXPORTER_OTLP_ENDPOINT=http://localhost:4317Two metrics carry the money:
claude_code., with atoken. usage typeattribute ofinput,output,cacheReadorcacheCreationclaude_code., in USD, incremented after each API requestcost. usage
Both carry model and query_source (main, subagent or auxiliary), plus agent, skill and plugin names where they apply. Prompt text is not logged unless you set OTEL_LOG_USER_PROMPTS=1.
How the monitors compare:
| Dimension | /usage | Status line | OpenTelemetry | cost-xray |
|---|---|---|---|---|
| Granularity | Session and plan totals | Session totals, live | Per request, per model and source | Per block inside each request |
| Shows request content | No | No | No (prompts redacted by default) | Yes, redacted, stored locally |
| Team roll-up | No | No | Yes, into your observability stack | No, one machine |
| Setup | None | A script | Env vars and a collector | One install command |
| Billing shape reported | List-price estimate | List-price estimate | List-price estimate unless contracted rates are set | Per-call total calibrated to provider usage |
The honest summary: OpenTelemetry tells a team who spent what and on which model; a capture tool tells one developer why a particular call was expensive. Many teams will want both.
claude code statusline show usage?
Run /statusline and describe what you want, for example "show model, context percentage and session cost", and Claude Code writes the script and the settings entry for you.
The status line documentation explains the mechanism: Claude Code pipes a JSON object describing the session to your script on stdin, and prints whatever the script prints in a row above the footer. The fields most people want are cost. for the running session cost and context_window. for how full the window is. A prompt cache object carries the same cache numbers /usage reports.
A minimal hand-written version looks like this, using jq:
#!/bin/bash
input=$(cat)
model=$(echo "$input" | jq -r '.model.display_name')
cost=$(echo "$input" | jq -r '.cost.total_cost_usd')
ctx=$(echo "$input" | jq -r '.context_window.used_percentage')
printf '%s | ctx %s%% | $%.2f\n' "$model" "$ctx" "$cost"The status line is the right tool for noticing: the moment context jumps from 30% to 70% after one tool call, you know that call pulled in something large. It is the wrong tool for explaining, because it only sees totals. Pairing it with a capture tool gives you the alarm and the evidence.
A subscriber note: on Pro and Max the session cost figure is a list-price estimate of what the tokens would have cost on the API, not a charge. It is still useful, because it tracks the same token volume that drains your usage windows.
how to see my codex usage?
In the Codex CLI, run /status for the current session's model, context use and plan limits; for what each Codex request contains and costs, cost-xray captures Codex the same way it captures Claude Code.
OpenAI's Codex command reference describes /status as showing the status of the current session, which is where context usage and your ChatGPT plan's rate-limit windows appear. If you run Codex on an API key instead of a plan, the bill lives in your OpenAI platform account's usage view.
Request-level capture works differently for Codex than for Claude Code. Claude Code supports a base-URL override, so cost-xray runs as a reverse proxy with no certificate at all. Codex needs a forward proxy, so cost-xray creates a local certificate authority that is trusted only by the codex command, not by your browser or the rest of the system. To choose which agents are captured, pick them when the installer asks, or set the choice ahead of time:
COST_XRAY_AGENTS=codexThe variable accepts claude, codex or all. After install, launch codex as usual and open cx; Codex sessions appear under their own agent at the top of the drill-down, so you can compare what the two agents send for similar work in the same project.
is cost-xray safe, does it send my API traffic anywhere?
No. The proxy binds to 127.0.0.1, sends no telemetry, and your traffic continues only to the model API it was already going to. Captured requests stay on your disk, with secrets removed before they are written.
The specifics, from the project's README and our cost-xray product page:
Where it listens: localhost only. Nothing on your network can reach it.
What leaves your machine: the same requests Claude Code or Codex would have sent directly, to the same provider. There is no cost-xray server.
What is stored: redacted requests and responses under
~/.. Authorization headers, API keys, cookies and secret-looking body fields are redacted before anything touches disk. Long sessions resend their history every turn, so each unique block is stored once with a small per-turn delta.cost-xray/sessions/ Certificates: none for Claude Code. For Codex, a local CA scoped to the
codexcommand only.If it breaks: if the proxy is down, the agent runs direct and uncaptured. Capture never blocks your work.
How to undo it:
cx stoppauses capture,cx uninstallremoves the services and shell wrappers, and deleting~/.removes every captured session.cost-xray/
- Run cx status to see which services are up, which ports are live and how many sessions were captured.
- Run cx stop before a session you do not want recorded; the agent then talks to the API directly.
- Run cx start to resume capture, or cx restart after editing ~/.cost-xray/env.
- Delete ~/.cost-xray/sessions/ when you want the stored captures gone.
One caution that applies to any capture tool: redaction removes credentials, not your code. The stored sessions contain file contents and tool output your agent read, because that is what was sent. If you work on regulated data, treat that directory like any other local copy of the repository. Teams handling health data may want a dedicated boundary check such as our phi-boundary-gate, which flags PHI candidates in provider requests; it reports candidates and is not a compliance certification.
When is a proxy the wrong way to see what Claude Code costs?
When you only need the number, when you need numbers for a whole team, or when policy forbids local copies of prompts. In each of those cases an official route wins.
You only want to know if you are near a limit.
/usageand a status line answer that with zero install. A capture tool is overkill for a glance.You manage spend for an organisation. The Claude Console usage page, the Claude Code Analytics API, or the org spend report on Team and Enterprise plans are the sources your finance team will trust, and OpenTelemetry streams per-user metrics into your own stack. cost-xray sees one machine; it does not roll up a team.
You need the bill, not an estimate of the split. The per-call totals in a capture tool match provider usage, but the only invoice is the one in your Console or cloud billing account. If your organisation pays contracted rates, Claude Code's
modelPricingmanaged setting aligns its own figures with your contract.You use Bedrock, Google Cloud or Foundry through a gateway. An LLM gateway such as LiteLLM already sits in the path and tracks spend per key; adding a second proxy on each laptop adds little.
Your security policy bans storing prompts locally. Redaction covers secrets, not source code. OpenTelemetry with prompt logging off is the safer choice there.
You are on Windows. cost-xray supports macOS and Linux only.
If what you are really after is fewer tokens rather than better visibility, context hygiene often does more than measurement. Keeping memory in files the agent reads on demand, as our agent-memory runtime does, and turning repeated instructions into skills both shrink the prefix every call carries. We wrote about why the harness around the model matters so much in the same model scoring 42 or 78 depending on its harness.
Conclusion
Claude Code's built-in tools answer "how much did this cost?": /usage for session and plan totals, /context for what is filling the window, a status line for a live figure, and OpenTelemetry for a team. They do not answer "what was in the request, and which part cost what?". For that you need the request itself, captured locally and priced after the fact.
A workable rule: if your totals look normal, stay with the official views. If a session cost more than the work justifies, or plan usage drains faster than you can explain, capture a few sessions and look at the prefix: tool schemas and MCP instructions that ride along on every call, history that is never cleared, and cold-cache rebuilds are the usual culprits, and each has a free fix.
A free next step you can take right now: open your current Claude Code session, run /context, and note the largest item that is not your own conversation. If it surprises you, the Tigerless Labs projects include the tools we use to chase that down, and more write-ups live on our insights page.
FAQ
how much does claude cost?
It depends on how you pay. Pro, Max, Team and Enterprise plans are flat subscriptions with usage windows, while API and cloud-provider access is billed per token, with separate rates for fresh input, cache writes, cache reads and output. Prices change, so read them on Anthropic's pricing page and the claude.com plans page rather than from a blog post. For a sense of scale, Anthropic's own Claude Code cost guide reports enterprise deployments averaging in the low tens of dollars per developer per active day.
claude token pricing?
Claude API tokens are priced per million, per model, in four buckets: base input, cache writes, cache reads and output. Relative to the base input rate, a 5-minute cache write costs 1.25x, a 1-hour cache write costs 2x, and a cache read costs a small fraction (0.1x on most models), according to Anthropic's pricing documentation. Output is the most expensive bucket per token, and extended thinking is billed as output. That is why a coding agent's bill depends so heavily on how well its cache holds.
claude usage dashboard?
There are three official ones, depending on how you pay. API users see spend on the Claude Console usage page and per-member Claude Code figures on the Console's Claude Code dashboard. Team and Enterprise admins see a spend report in org analytics on claude.ai. Individual Pro and Max subscribers see plan usage bars in /usage inside Claude Code and under Settings > Usage on claude.ai. None of them shows what was inside each request; for that you need a request-level tool such as cost-xray.
claude ai token pricing risk?
The main risk is not the rate card, it is volume you cannot see. A coding agent resends its whole conversation on every turn, so a long session, a cold cache after a break, or a large tool schema that rides along on every call can multiply token use without any change in price. The second risk is staleness: any dollar figure you write into a budget or a blog post goes out of date when Anthropic reprices, so link to the official pricing page and track the shape of your spend (fresh vs cached vs output) instead.

