Thinking is adaptive by default
On the current models you do not pass a thinking budget; the model decides how much to reason and the effort setting (low to max) is the dial. Opus 5.5 defaults to medium effort, so set it explicitly if you want more.
API quickstart · checked October 2026
Enough to make a correct first request and avoid the six mistakes that old tutorials teach. Everything here is checked against the official documentation on the date above; when they disagree, the documentation wins.
| Model ID | Context | Max output | Input / output per 1M tokens | Cache read | Use it for |
|---|---|---|---|---|---|
claude-fable-5-1Claude Fable 5.1 | 1M | 128K | $10 / $50 | $0.25 | Hardest reasoning, long agent runs; thinking always on |
claude-opus-5-5Claude Opus 5.5 | 1M | 128K | $4 / $20 | $0.2 | Default for new work |
claude-sonnet-5-5Claude Sonnet 5.5 | 1M | 128K | $2 / $10 | $0.2 | Everyday calls, speed |
claude-haiku-4-5Claude Haiku 4.5 | 200K | 64K | $1 / $5 | $0.1 | High volume, simple jobs |
Batch requests cost half. Older models (Opus 5, 4.8, 4.7, 4.6; Sonnet 5, 4.6) are still served and cost the same or more; there is no reason to start new work on them.
from anthropic import Anthropic
client = Anthropic() # reads ANTHROPIC_API_KEY from the environment
message = client.messages.create(
model="claude-opus-5-5",
max_tokens=16000,
output_config={"effort": "medium"},
messages=[{"role": "user", "content": "Summarize the attached contract in five bullet points."}],
)
if message.stop_reason == "refusal":
print("declined:", message.stop_details)
else:
print(message.content[0].text)
Install with pip install anthropic. The client reads the key from the environment; never paste it into code. max_tokens is a hard ceiling, so do not set it low for anything that may run long; for very long replies use the streaming helper.
curl https://api.anthropic.com/v1/messages \
-H "x-api-key: $ANTHROPIC_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "content-type: application/json" \
-d '{"model": "claude-sonnet-5-5", "max_tokens": 1024,
"messages": [{"role": "user", "content": "Hello, Claude"}]}'
On the current models you do not pass a thinking budget; the model decides how much to reason and the effort setting (low to max) is the dial. Opus 5.5 defaults to medium effort, so set it explicitly if you want more.
Pre-filling the assistant turn returns an error on the 4.6-and-later models. To get JSON, use the structured output format option or a strict tool definition.
Replies can run to 128K tokens; the SDKs require streaming for large max_tokens so requests do not time out. Use the stream helper and read the final message.
A safety classifier can decline a request with HTTP 200 and stop_reason "refusal". Check it before reading content, and consider the server-side fallback option for production traffic.
Prompt caching bills repeated prefixes at a fraction of the input price. Keep the system prompt and tool list stable and put the volatile part last.
The Batch API runs non-urgent requests asynchronously at half price. Nightly classification and bulk extraction belong there.
