Claude API Operations
Building applications and agents on Anthropic's API: the Messages API, tool use, prompt caching, structured outputs, batches, thinking/effort, and the Claude Agent SDK. For developers writing apps against the API — not for using Claude Code itself.
API surfaces move fast. Model IDs, parameters, and betas in this skill were
verified against platform.claude.com (2026-08). When in doubt — especially for
"latest model" or pricing questions — verify with WebFetch against
https://platform.claude.com/docs/en/about-claude/models/overview.md or query
the Models API (client.models.list()).
Current Models (verified 2026-08)
| Model | ID (exact, no date suffix) | Context | Max Output | Input $/MTok | Output $/MTok |
|---|---|---|---|---|---|
| Claude Fable 5 | claude-fable-5 | 1M | 128K | $10.00 | $50.00 |
| Claude Opus 5 | claude-opus-5 | 1M | 128K | $5.00 | $25.00 |
| Claude Sonnet 5 | claude-sonnet-5 | 1M | 128K | $2.00 | $10.00 |
| Claude Haiku 4.5 | claude-haiku-4-5 | 200K | 64K | $1.00 | $5.00 |
Use these alias IDs verbatim. Never append date suffixes (claude-sonnet-5-20260630
is wrong → 404). Haiku 4.5 is the one current model with a dated snapshot id
(claude-haiku-4-5-20251001) behind its alias; from the 4.6 generation on, the dateless
id is the pinned snapshot.
Legacy (still available, no longer current): claude-opus-4-8, claude-opus-4-7,
claude-opus-4-6, claude-opus-4-5, claude-sonnet-4-6, claude-sonnet-4-5. Migrating
off one: https://platform.claude.com/docs/en/models/opus-5/migration-guide.md (or run
/claude-api migrate in Claude Code). Live capability lookup:
client.models.retrieve("claude-opus-5") → .max_input_tokens, .max_tokens,
.capabilities dict.
Model Selection Decision Tree
What is the workload?
│
├─ Hardest problems, long-horizon agents, deep research, ceiling intelligence
│ └─ claude-fable-5 (premium ceiling) or claude-opus-5 (default flagship)
│
├─ Agentic coding, tool-heavy workflows, production assistants
│ └─ claude-opus-5 (quality) or claude-sonnet-5 (speed/cost balance)
│
├─ High-volume production: summarization, RAG answers, extraction
│ └─ claude-sonnet-5
│
├─ Classification, routing, simple Q&A, latency-critical
│ └─ claude-haiku-4-5
│
└─ Subagents inside a larger system
└─ One tier below the orchestrator (Opus loop → Sonnet/Haiku workers)
Tiering rule: route by task difficulty, not by uniform default. An Opus orchestrator dispatching Haiku classifiers is routinely 5-10x cheaper than Opus-everywhere with no quality loss on the simple legs.
Which Surface? (API vs Agent SDK vs Batches)
| Need | Use | Why | |---|---|---| | One request → one response (classify, summarize, extract, Q&A) | Messages API | Simplest; full control | | Multi-step pipeline, your code controls the logic | Messages API + tool use | You own the loop | | Custom agent with your own tools, your infra | Messages API + tool use (manual loop or SDK tool runner) | Max flexibility | | Agent that reads/edits files, runs commands, searches — without building tools | Claude Agent SDK | Claude Code's tools + agent loop as a library | | CI/CD automation, coding agents, production agent apps | Claude Agent SDK | Built-in tools, hooks, sessions, MCP | | Large non-urgent workloads (eval runs, backfills, bulk extraction) | Batches API | 50% discount, ≤24h turnaround | | Hosted agent, Anthropic runs loop + sandbox | Managed Agents (beta) | No infra; see official docs |
Rule of thumb: start at the simplest tier. Reach for an agent only when the task is genuinely open-ended (multi-step, hard to fully specify, errors recoverable, value justifies cost).
Messages API Quick Start
Everything goes through POST /v1/messages. Headers: x-api-key,
anthropic-version: 2023-06-01, content-type: application/json.
# pip install anthropic
import anthropic
client = anthropic.Anthropic() # reads ANTHROPIC_API_KEY
response = client.messages.create(
model="claude-opus-5",
max_tokens=16000,
system="You are a concise technical assistant.",
messages=[{"role": "user", "content": "Explain CRDTs in one paragraph."}],
)
for block in response.content: # content is a list of typed blocks
if block.type == "text": # always check .type before .text
print(block.text)
print(response.stop_reason, response.usage.input_tokens, response.usage.output_tokens)
// npm install @anthropic-ai/sdk
import Anthropic from "@anthropic-ai/sdk";
const client = new Anthropic();
const response = await client.messages.create({
model: "claude-opus-5",
max_tokens: 16000,
messages: [{ role: "user", content: "Explain CRDTs in one paragraph." }],
});
for (const block of response.content) {
if (block.type === "text") console.log(block.text); // narrow the union first
}
Streaming (default to it for long outputs — non-streaming above ~16K
max_tokens risks SDK HTTP timeouts):
with client.messages.stream(model="claude-opus-5", max_tokens=64000,
messages=[{"role": "user", "content": "Write a long report"}]) as stream:
for text in stream.text_stream:
print(text, end="", flush=True)
final = stream.get_final_message() # full Message after streaming
Full params, response shape, stop reasons, errors, retries, rate limits: references/messages-api.md
Thinking & Effort (quick reference)
- Adaptive thinking is ON BY DEFAULT on Fable 5 / Opus 5 / Sonnet 5 — send no
thinkingfield and you still get (and pay for) thinking. On the legacy 4.6–4.8 models it stays off until you setthinking: {"type": "adaptive"}. - Manual budgets are gone.
{"type": "enabled", "budget_tokens": N}returns a 400 on Opus 4.7 and every later model (Opus 5, Sonnet 5, Fable 5 included); deprecated on Opus 4.6 / Sonnet 4.6. Control depth witheffort, not tokens. - Turning thinking off: Sonnet 5 accepts
{"type": "disabled"}. Opus 5 accepts it only at efforthighor below — pairing it withxhigh/maxis a 400. Fable 5 rejects it outright; thinking there is unconditional, so budget for it. - Effort (GA):
output_config: {"effort": "low" | "medium" | "high" | "xhigh" | "max"}— nested inoutput_config, not top-level. Defaulthigh(identical to omitting it).xhigh: Fable 5, Opus 5, Opus 4.8/4.7, Sonnet 5.max: those plus Opus 4.6 and Sonnet 4.6. Haiku 4.5 does not supporteffortat all. - Sampling params removed on Opus 4.7 and later (so Opus 5, Sonnet 5, Fable 5):
temperature,top_p,top_kall return 400 — and the Python SDK v1.0+ doesn't define them, so passing them raisesTypeError. Steer with prompting + effort. - Forced tool_choice is fine with adaptive thinking. The auto/none-only
restriction applies to manual extended thinking (
{"type": "enabled"}) only; adaptive mode — including the models where it's on by default — accepts{"type": "any"}and{"type": "tool", ...}. - Thinking text is omitted by default on Fable 5 / Opus 5 / Sonnet 5 / Opus 4.8 /
4.7 — opt in with
thinking: {"type": "adaptive", "display": "summarized"}if you surface reasoning to users. Either way the blocks are billed, and must be echoed back unmodified (emptythinkingfield included) in a tool-use loop, or the next request 400s.
Details and gotchas: references/structured-outputs.md (thinking interplay) and references/messages-api.md.
Tool Use (quick reference)
tools = [{
"name": "get_weather",
"description": "Get current weather. Call when the user asks about weather conditions.",
"input_schema": {
"type": "object",
"properties": {"location": {"type": "string", "description": "City, e.g. Paris"}},
"required": ["location"],
},
}]
response = client.messages.create(model="claude-opus-5", max_tokens=16000,
tools=tools, messages=messages)
if response.stop_reason == "tool_use":
... # execute, send tool_result back, loop
tool_choice: {"type": "auto"} (default) | {"type": "any"} | {"type": "tool", "name": "..."} | {"type": "none"}. Add
"disable_parallel_tool_use": true to force at most one call per response.
The agentic loop, parallel tool results, pause_turn, is_error, server-side
tools, and SDK tool runners: references/tool-use.md
Cost Optimization Checklist
Work top-down; each item is independent:
- [ ] Right-size the model. Haiku for classification/routing, Sonnet for volume work, Opus/Fable for the hard 10%. Largest single lever.
- [ ] Prompt caching on stable prefixes (system prompt, tool defs, big docs):
cache_control: {"type": "ephemeral"}. Reads cost ~0.1x; up to 90% savings. Verify withusage.cache_read_input_tokens > 0— zero means a silent invalidator (timestamp in system prompt, unsorted JSON, varying tools). - [ ] Batches API for anything that can wait ≤24h: flat 50% off all tokens, stacks with caching.
- [ ] Cap output: set
max_tokensto what you need (256 for classification); stream + generous cap for long generation. - [ ] Tune effort down where quality allows:
mediumis often the sweet spot;lowfor subagents and simple tasks. - [ ] Count before sending:
client.messages.count_tokens(...)(never tiktoken — it's OpenAI's tokenizer and undercounts Claude by 15-20%). - [ ] Keep prefixes stable: order requests
tools→system→messages, volatile content last; don't swap tool sets or models mid-conversation.
Mechanics, breakpoints, TTLs, batch lifecycle, tiering math: references/caching-and-cost.md
Context Engineering
Prompt engineering asks what to write in the prompt. Context engineering asks what earns a place in the window on this call — including everything that lands there without you typing it: tool definitions, tool results, retrieved documents, prior turns, thinking blocks. It is iterative (every inference) where prompt engineering is discrete (written once). Target: the smallest set of high-signal tokens that gets the outcome.
The budget is real because attention degrades with length (context rot — n² pairwise relationships), not just because tokens cost money. A 1M window is a capacity, not a target.
The three tiers
Every candidate fact lives in exactly one place. Choosing deliberately is most of the job.
| Tier | Where | Cost | Use when |
|---|---|---|---|
| 1 — In context | tools / system / messages, every call | Paid every turn (≈0.1× cached) | It steers most turns |
| 2 — On disk, read on demand | A file the agent can read; only the path stays in context | Paid only when read | The agent can tell from a name that it needs this |
| 3 — Retrieved | Index / search tool behind a query | Paid only on a hit, plus a relevance gamble | The corpus is too large to enumerate |
When a prompt is too big, demote before you delete — a path is ~10 tokens; the file it names may be 10,000.
This repo already runs on the tier-1/tier-2 split. A skill's description is
always resident (tier 1, so it must carry the routing signal); SKILL.md loads on a
match; references/*.md load only when cited and needed. "Description is the
trigger", "body under 500 lines", "one concept per reference", "every reference must
be cited" are context-engineering rules wearing authoring clothes.
Cache-aware prompt architecture
Requests render tools → system → messages, and the cache is a prefix match.
So static prefix first, volatile content last — put the cache_control breakpoint
at the end of the stable part and let per-request content fall after it.
Reordering a prompt destroys the cache silently: no error, just a different prefix
hash, cache_read_input_tokens: 0, and a 1.25–2× bill where you expected 0.1×. The
usage block is the only symptom, which is why asserting cache_read_input_tokens > 0
in staging is a real test.
The compaction decision
Under modern prompt caching, keeping the full history has been measured to beat summarisation on cost, latency AND recall at the same time. A 2026 production-tutor evaluation (660 turns, 11 configurations) put keep-everything at 92–100% fact recall, $0.11/turn and 17 s TTFT, against 38–58% recall, $0.24/turn and 21 s for its clear-plus-summarise preset. Summarising rewrites the cached prefix and forfeits the 0.1× discount — the cheap move is usually to append. (That study ran on a non-Claude model; what transfers is the mechanism, and Claude's flat 0.1× cache read makes it stronger, not weaker. Full caveats in references/compaction.md.)
So: compact only as a deliberate response to a named constraint.
| Constraint | Diagnose | Try first |
|---|---|---|
| Context ceiling — it will not fit | Projected tokens > window | Cap tool output → payloads to files → server-side clearing |
| Cost ceiling — the bill is unacceptable | Compare against cached cost, not uncached | Verify the cache is hitting → tier down → cap tool output |
| Latency target — TTFT too slow at depth | Confirm growth is in the prefix | Cap tool output → lower effort → stream |
Capping tool output at the tool boundary is the underrated lever: it shrinks context without rewriting the cached prefix (the same study measured −38% cost/turn with no recall loss). Clearing and summarising both break the cache; they are what people reach for first and should reach for last.
First-party clearing is context_management (beta context-management-2025-06-27):
clear_tool_uses_20250919 and clear_thinking_20251015, applied server-side. Always
set clear_at_least — it stops a trigger paying a full cache re-write to save a
handful of tokens. Pair with the memory tool so durable conclusions are written out
before raw material is cleared.
Agentic specifics
- Tool results are the growth term, not the system prompt. Design tools to return decisions, not dumps.
- Summarise vs write-to-file: needed later in full → write to a file, return the path. Only the conclusion matters → summarise at the tool boundary (free of cache cost, unlike rewriting history after the fact).
- Sub-agents are context isolation, not just parallelism: 80K tokens of exploration are billed once inside the child and discarded; the parent sees a ~1–2K-token distillation. Costs: cold cache in the child, a lossy hand-off. Skip it when the subtask needs most of the parent's context to make sense.
Full doctrine — tiers, progressive disclosure, instrumentation:
references/context-engineering.md.
Compaction economics, context_management parameters, memory tool:
references/compaction.md.
For Claude Code's own context surface see the claude-code-ops skill; for
prompts re-sent on a cadence, loop-ops; for cross-provider fan-out, fleetflow.
Claude Agent SDK (quick reference)
# pip install claude-agent-sdk (Python >= 3.10)
import asyncio
from claude_agent_sdk import query, ClaudeAgentOptions
async def main():
async for message in query(
prompt="Find and fix the bug in auth.py",
options=ClaudeAgentOptions(allowed_tools=["Read", "Edit", "Bash"]),
):
if hasattr(message, "result"):
print(message.result)
asyncio.run(main())
// npm install @anthropic-ai/claude-agent-sdk
import { query } from "@anthropic-ai/claude-agent-sdk";
for await (const message of query({
prompt: "Find and fix the bug in auth.ts",
options: { allowedTools: ["Read", "Edit", "Bash"] },
})) {
if ("result" in message) console.log(message.result);
}
Built-in tools (Read/Write/Edit/Bash/Glob/Grep/WebSearch/WebFetch/...), hooks
(PreToolUse, PostToolUse, ...), subagents, MCP servers, sessions
(resume/fork), permission modes, and the SDK-vs-raw-API decision:
references/agent-sdk.md
Common Pitfalls
| Pitfall | Symptom | Fix |
|---|---|---|
| Date-suffixed or guessed model ID | 404 not_found_error | Use exact alias IDs from the table above |
| budget_tokens on Opus 4.7+ (incl. Opus 5 / Sonnet 5 / Fable 5) | 400 | thinking: {"type": "adaptive"} + effort |
| Assuming thinking is opt-in on Fable 5 / Opus 5 / Sonnet 5 | Unexpected thinking tokens billed | Adaptive thinking is on by default there; Fable 5 can't be disabled at all |
| thinking: {"type": "disabled"} at xhigh/max on Opus 5 | 400 | Drop effort to high or below, or leave thinking on |
| temperature/top_p/top_k on Opus 4.7+ | 400 (or TypeError on Python SDK v1.0+) | Remove; steer via prompt + effort |
| effort on Haiku 4.5 | 400 | Haiku 4.5 doesn't support the parameter |
| Rebuilding assistant turns in a tool loop (dropping empty thinking blocks) | 400 "thinking blocks cannot be modified" | Echo the content list back exactly as received |
| Assistant-turn prefill on Opus 4.7+ models | 400 | output_config.format or system-prompt instruction |
| Cache marker on <minimum prefix | Silent no-cache (cache_creation_input_tokens: 0) | Min 512-4096 tokens depending on model (see caching ref) |
| Not handling stop_reason: "tool_use" | Agent "stops" after first tool call | Loop: execute tools, append tool_result, re-request |
| Missing tool_result for a tool_use id | 400 on follow-up | One tool_result per tool_use block, ids matching |
| Non-streaming with max_tokens > ~16K | SDK timeout / ValueError | Stream + get_final_message() / finalMessage() |
| output_format top-level param | Deprecated | output_config: {"format": {...}} |
| tiktoken for Claude token counts | 15-20%+ undercount | messages.count_tokens endpoint |
| String-matching error messages | Fragile retries | Typed exceptions: anthropic.RateLimitError etc. |
| Raw string-matching tool input | Breaks on escaping changes | Always json.loads() / use parsed block.input |
| Compacting by reflex on a long conversation | Higher cost, worse recall than doing nothing | Name the constraint first; under caching, appending usually wins (see Context Engineering) |
| clear_tool_uses without clear_at_least | A full cache re-write to reclaim a few hundred tokens | Set clear_at_least so each cache break is worth taking |
Resources & Verification
This skill ships a staleness verifier and two copy-and-adapt starter assets. The model table and pricing above are the facts most likely to drift — run the verifier when you suspect they're stale.
scripts/check-model-table.py — guards the Current Models table (this file)
and the per-model prompt-cache minimum table
(references/caching-and-cost.md) against drift.
Two modes per the resource protocol §7:
# Structural (default, no network): every row well-formed, ids carry no date
# suffix, prices numeric, the two files agree on the model lineup. It also
# guards the cache-economics constants that are stated in more than one file
# (0.1x read, 1.25x/2x writes, 4 breakpoints, 20-block lookback, the
# context-management beta id), asserts each doctrine reference carries a
# "verified <ISO date>" stamp, and checks SKILL.md <-> references/ citation
# integrity in both directions. It then scans every file in the skill for model
# ids: an id in neither the table nor the Legacy list is flagged "unknown", and
# a LEGACY id sitting where a reader would copy it (model=..., "model": ...,
# --model ...) is flagged "retired" - append a `legacy-ok` comment to that line
# for a deliberate migration example. Exit 4 on any contradiction.
python skills/claude-api-ops/scripts/check-model-table.py --offline
python skills/claude-api-ops/scripts/check-model-table.py --offline --json | python -m json.tool
# Live (advisory, needs ANTHROPIC_API_KEY): curls the Models API and compares
# its id set against the documented ids. Exit 10 if a documented id is gone or a
# newer alias id is missing from the table; exit 7 (not a failure) if the key is
# unset or the API is unreachable. Live mode checks model-ID coverage ONLY — the
# API returns no pricing, so pricing/context drift stays an --offline + docs concern.
ANTHROPIC_API_KEY=sk-... python skills/claude-api-ops/scripts/check-model-table.py --live
scripts/context-budget.py — append-vs-compact calculator. Models both paths
in dollars over the turns you actually have left, checks the context ceiling
first, and exits 10 when cost favours compaction, 0 when appending wins:
# Short session — appending is cheaper (exit 0)
python skills/claude-api-ops/scripts/context-budget.py \
--history-tokens 25000 --turns-remaining 5 --base-rate 0.30
# Deep session — cost favours compaction (exit 10)
python skills/claude-api-ops/scripts/context-budget.py \
--history-tokens 120000 --turns-remaining 40 --base-rate 2.00 --json
Two results worth knowing before you trust it: the break-even turn count is scale-invariant (history size and price cancel out — it tracks the summary ratio, not how big or costly the conversation is), and it prices cost only. Recall loss is not in the model, so "compact" means cheaper, not better.
assets/cached-agent-loop.py — the cache-aware sibling of the minimal loop
below, and the executable form of the Context Engineering section: breakpoint at
the end of the static prefix, a rolling breakpoint on the newest turn, an
intermediate anchor every ~15 blocks so long tool-heavy turns don't jump the
20-block lookback, tool output capped at the boundary, and a per-turn
cache_read_input_tokens check that warns when the prefix silently changed.
Copy it when the agent is long-running; copy agentic-loop.py when it isn't.
The footgun it encodes: cache_control is a key on a content block, so it can
only be set on a dict. Appending response.content verbatim (SDK block
objects) or using the "content": "a string" shorthand leaves nowhere to put a
marker — every breakpoint aimed at those turns is discarded with no error and no
warning. Normalise content to dict blocks before placing breakpoints.
assets/recall-probe.py — the "measure it on your workload" harness:
plants a fact, buries it under N turns, probes for it, and reports recall, cost
per turn and TTFT for append vs compact. Makes real API calls, so start
small (--turns 6 --trials 1). Replace the synthetic filler turns with traffic
from your own logs — that is the point of running it.
assets/agentic-loop.py — a minimal, runnable tool-use loop (define a tool,
call messages.create, loop while stop_reason == "tool_use", append
tool_result, re-request until end_turn). Copy it as the starting point when
building a manual agent loop; the >>> ADAPT marks show what to change.
assets/output-schema.json — a known-good structured-outputs request body in
the canonical output_config.format shape (with additionalProperties: false
and a required array). Copy and reshape schema.properties when adding JSON
outputs; see references/structured-outputs.md
for the rules. (Supported on every current model — Fable 5, Opus 5, Sonnet 5,
Haiku 4.5 — and the legacy 4.5–4.8 line.)
Reference Files
| File | Covers | |---|---| | references/messages-api.md | Params, response shape, streaming events, stop reasons, error handling, retries, rate limits | | references/tool-use.md | Tool definitions, tool_choice, parallel tools, agentic loop, tool results, server tools, tool runners | | references/caching-and-cost.md | Prompt caching mechanics, Batches API, token counting, model tiering economics | | references/structured-outputs.md | output_config.format, schema rules/limits, strict tools, parse() helpers, thinking interplay | | references/agent-sdk.md | Python + TS Agent SDK, ClaudeAgentOptions, hooks, MCP, sessions, SDK vs raw API | | references/context-engineering.md | Context budget, the three tiers, progressive disclosure, cache-aware ordering, tool-result bloat, sub-agents as isolation, instrumentation | | references/compaction.md | When compaction is justified, break-even arithmetic, context_management edits, memory tool, how to compact well |
Live Documentation
When cached facts may be stale, WebFetch (append .md for clean markdown):
- Models/pricing:
https://platform.claude.com/docs/en/about-claude/models/overview.md - Messages API:
https://platform.claude.com/docs/en/api/messages - Tool use:
https://platform.claude.com/docs/en/agents-and-tools/tool-use/overview.md - Prompt caching:
https://platform.claude.com/docs/en/build-with-claude/prompt-caching.md - Structured outputs:
https://platform.claude.com/docs/en/build-with-claude/structured-outputs.md - Batches:
https://platform.claude.com/docs/en/build-with-claude/batch-processing.md - Agent SDK:
https://code.claude.com/docs/en/agent-sdk/overview - Context editing:
https://platform.claude.com/docs/en/build-with-claude/context-editing - Context engineering:
https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents