Skip to content

API quickstart · checked October 2026

The Claude API in one page

Enough to make a correct first request and avoid the six mistakes that old tutorials teach. Everything here is checked against the official documentation on the date above; when they disagree, the documentation wins.

Current models

Model IDContextMax outputInput / output per 1M tokensCache readUse it for
claude-fable-5-1Claude Fable 5.11M128K$10 / $50$0.25Hardest reasoning, long agent runs; thinking always on
claude-opus-5-5Claude Opus 5.51M128K$4 / $20$0.2Default for new work
claude-sonnet-5-5Claude Sonnet 5.51M128K$2 / $10$0.2Everyday calls, speed
claude-haiku-4-5Claude Haiku 4.5200K64K$1 / $5$0.1High volume, simple jobs

Batch requests cost half. Older models (Opus 5, 4.8, 4.7, 4.6; Sonnet 5, 4.6) are still served and cost the same or more; there is no reason to start new work on them.

Python

from anthropic import Anthropic

client = Anthropic()  # reads ANTHROPIC_API_KEY from the environment

message = client.messages.create(
    model="claude-opus-5-5",
    max_tokens=16000,
    output_config={"effort": "medium"},
    messages=[{"role": "user", "content": "Summarize the attached contract in five bullet points."}],
)
if message.stop_reason == "refusal":
    print("declined:", message.stop_details)
else:
    print(message.content[0].text)

Install with pip install anthropic. The client reads the key from the environment; never paste it into code. max_tokens is a hard ceiling, so do not set it low for anything that may run long; for very long replies use the streaming helper.

curl

curl https://api.anthropic.com/v1/messages \
  -H "x-api-key: $ANTHROPIC_API_KEY" \
  -H "anthropic-version: 2023-06-01" \
  -H "content-type: application/json" \
  -d '{"model": "claude-sonnet-5-5", "max_tokens": 1024,
       "messages": [{"role": "user", "content": "Hello, Claude"}]}'

Six things that changed

Thinking is adaptive by default

On the current models you do not pass a thinking budget; the model decides how much to reason and the effort setting (low to max) is the dial. Opus 5.5 defaults to medium effort, so set it explicitly if you want more.

No prefill, use structured outputs

Pre-filling the assistant turn returns an error on the 4.6-and-later models. To get JSON, use the structured output format option or a strict tool definition.

Stream anything long

Replies can run to 128K tokens; the SDKs require streaming for large max_tokens so requests do not time out. Use the stream helper and read the final message.

Check stop_reason

A safety classifier can decline a request with HTTP 200 and stop_reason "refusal". Check it before reading content, and consider the server-side fallback option for production traffic.

Cache the stable prefix

Prompt caching bills repeated prefixes at a fraction of the input price. Keep the system prompt and tool list stable and put the volatile part last.

Batch what can wait

The Batch API runs non-urgent requests asynchronously at half price. Nightly classification and bulk extraction belong there.

Source code on a screen, close up
The error that teaches you most: a 400 on a parameter that worked last year.