Providers¶
Purpose¶
Connect an agent to a real generation API through a stdlib-only adapter, with typed, retry-classified errors.
When to use¶
Whenever a run should reach a live model. Anything implementing
model.Model works; the shipped adapters cover OpenAI-compatible APIs
and Anthropic.
How it works¶
Adapters translate the provider-neutral model.Request to their wire
format and normalize responses back. They never retry on their own:
failures are classified through model.RetryableError (408, 429, 5xx,
transport faults — never context cancellation) and the runner's retry
policy decides. All shipped adapters implement model.StreamingModel
— over SSE everywhere except Bedrock, whose ConverseStream responses
are AWS binary event-stream frames decoded with the standard library.
Every adapter also translates media
parts on user messages — images, documents, audio, video — to its
native multimodal form, and rejects unsupported combinations before
any request; the per-provider matrix lives in
Multimodal input.
Every adapter accepts optional sampling and length controls on its
Config — Temperature, TopP, and MaxTokens — validated against the
provider's documented ranges at construction. A nil sampling field
leaves the provider default, and an unset MaxTokens omits the bound,
so unset controls never appear on the wire. Temperature 0 is a meaningful
value, not a default: set it with providers.Ptr(0).
Every response also carries the provider's terminal cause, normalized
onto model.FinishReason — identical on streamed and plain runs:
| Adapter | Wire field | Stop | Length | Tool call | Content filter |
|---|---|---|---|---|---|
| OpenAI / Azure | finish_reason |
stop |
length |
tool_calls, function_call |
content_filter |
| Anthropic | stop_reason |
end_turn, stop_sequence |
max_tokens |
tool_use |
refusal |
| Bedrock | stopReason |
end_turn, stop_sequence |
max_tokens |
tool_use |
guardrail_intervened, content_filtered |
| Gemini | finishReason |
STOP |
MAX_TOKENS |
— | SAFETY, RECITATION, BLOCKLIST, PROHIBITED_CONTENT, SPII |
Prompt caching¶
Every provider can serve a repeated prompt prefix from cache instead of
re-processing it; the adapters expose the control where the provider
offers one, and every adapter reports the outcome on model.Usage's
cache token fields, so a run can compute what caching saved.
- Anthropic —
anthropic.Config.CacheControlrequests the provider's automatic caching: a top-levelcache_controlfield, and the provider applies the breakpoint to the last cacheable block of each request, moving it forward as the conversation grows. A zero TTL selects the five-minute default;anthropic.CacheOneHourasks for the one-hour entry at a premium. Anthropic-compatible gateways reached throughBaseURLmay not support the parameter — the provider's own API does. - Bedrock —
bedrock.Config.CacheControlplaces an explicitcachePointcheckpoint at each request's conversation frontier, so the prefix it covers — tool declarations, system guidance, and history — is cached and re-read on later requests. TTL semantics are the same (bedrock.CacheOneHourwhere the model supports it), and Anthropic models on Bedrock also cache implicitly without any opt-in. - OpenAI and Azure — caching is automatic for eligible prefixes
(about a thousand tokens and up); there is nothing to send, and the
cached share arrives on
Usage.CacheReadTokens. - Gemini — caching is implicit for current models; nothing to
send, and
Usage.CacheReadTokensreports hits. The explicit context-caching API (a separate cached-content resource with its own lifecycle) stays out of scope.
Usage carries detail where providers break it out, normalized onto
three model.Usage fields beside the token totals:
| Adapter | CacheReadTokens |
CacheWriteTokens |
ReasoningTokens |
|---|---|---|---|
| OpenAI / Azure | prompt_tokens_details.cached_tokens — inside input |
— (implicit caching) | completion_tokens_details.reasoning_tokens — inside output |
| Anthropic | cache_read_input_tokens — beside input |
cache_creation_input_tokens — beside input |
— |
| Bedrock | cacheReadInputTokens — beside input |
cacheWriteInputTokens — beside input |
— |
| Gemini | cachedContentTokenCount — inside input |
— (implicit caching) | thoughtsTokenCount — inside output |
"Inside" and "beside" matter for pricing: a subset of the input total discounts part of it, an exclusive figure replaces it. Detail fields are zero when a provider (or a proxied endpoint) does not report them, and streamed runs carry the same figures as plain ones — every adapter reports streamed usage through its terminal chunk or event.
Unlisted values — Anthropic's pause_turn, Gemini's
MALFORMED_FUNCTION_CALL, anything a provider adds later — map to
FinishOther, and a missing field maps to the empty FinishReason;
unrecognized causes are evidence, never failures. Result.FinishReason
holds the final turn's cause, and RunError.Partial.FinishReason the
last completed turn's when a run failed: a length cause on a decode
failure means the provider truncated the output, not that the model
wrote bad JSON.
Pricing¶
Each generation adapter ships a Price struct implementing
model.Price: per-million-token USD rates you supply, combined by the
rule that adapter's own usage semantics require. Wire one with
golem.WithPrice and runs report Result.Cost (and enforce
UsageLimit.Cost); see Cost for the full contract. Golem
ships no price table — the rates are your snapshot of the provider's
price list.
| Price | Cached input | Cache writes | Rates |
|---|---|---|---|
openai.Price |
inside the input total — Input − CacheRead bills at the input rate |
— (implicit caching) | input, output, cache read |
azure.Price |
inside the input total | — (implicit caching) | input, output, cache read |
anthropic.Price |
beside the input total — billed at its own rate | beside input, premium rate | input, output, cache read, cache write |
bedrock.Price |
beside the input total | beside input | input, output, cache read, cache write |
gemini.Price |
inside the input total | — (implicit caching) | input, output, cache read |
With no cache-read rate set, the inside-input prices bill cached input at the input rate rather than free; the beside-input prices bill the input total in full by construction. Reasoning tokens are part of billed output on every adapter and take no separate rate.
Embeddings¶
Generation adapters are not the only port: the same providers serve
text embeddings through the embedding.Embedder port, covered end to
end in Embeddings. Query and document calls share one
Result shape — vectors in input order plus model.Usage input-token
evidence — and the split maps to the provider's task type where one
exists:
| Embedder | Wire | Query / documents | Dimensions |
Usage |
|---|---|---|---|---|
openai.Embedder |
POST {BaseURL}/embeddings |
one shape, split ignored | dimensions (text-embedding-3+) |
usage.prompt_tokens |
azure.Embedder |
POST {Endpoint}/openai/deployments/{deployment}/embeddings |
one shape, split ignored | dimensions (text-embedding-3+) |
usage.prompt_tokens |
gemini.Embedder |
:embedContent / :batchEmbedContents |
RETRIEVAL_QUERY / RETRIEVAL_DOCUMENT |
outputDimensionality (newer models) |
not reported today — stays zero |
Anthropic has no embeddings API, so there is no Anthropic embedder. Bedrock's embedding models (Titan, Cohere on Bedrock) are likewise not covered yet; say the word if your index runs there.
Token counting¶
The same asymmetry runs the other way for counting: where the provider
offers a counting endpoint, an adapter implements the tokens.Counter
port — covered end to end in Token counting — and
prices exactly what a generation request with the same builders would
send:
| Counter | Wire | Counts tools | Notes |
|---|---|---|---|
anthropic.Counter |
POST {BaseURL}/v1/messages/count_tokens |
yes | same request schema as inference |
gemini.Counter |
POST {BaseURL}/v1beta/models/{model}:countTokens |
yes | |
bedrock.Counter |
POST {BaseURL}/model/{model}/count-tokens |
no — lower bound | free; base foundation-model IDs only |
OpenAI, Azure OpenAI, and OpenAI-compatible endpoints have no counting
API, so there is no counter for them; implement tokens.Counter with
your own tokenizer there.
Example¶
examples/minimal— OpenAI-compatible.examples/anthropic— Anthropic Messages API.examples/gemini— Google Gemini GenerateContent API.examples/azure— Azure OpenAI deployments.examples/bedrock— AWS Bedrock Converse with SigV4.examples/local-models— Ollama or LM Studio through the OpenAI adapter.
OPENAI_API_KEY=sk-... go run ./examples/minimal
ANTHROPIC_API_KEY=sk-... go run ./examples/anthropic
GEMINI_API_KEY=... go run ./examples/gemini
AZURE_OPENAI_API_KEY=... AZURE_OPENAI_ENDPOINT=https://my-resource.openai.azure.com \
AZURE_OPENAI_DEPLOYMENT=gpt-4o AZURE_OPENAI_API_VERSION=2024-10-21 \
go run ./examples/azure
AWS_ACCESS_KEY_ID=... AWS_SECRET_ACCESS_KEY=... AWS_REGION=us-east-1 \
go run ./examples/bedrock
GOLEM_LOCAL_BASE_URL=http://localhost:11434/v1 go run ./examples/local-models
openaiClient, _ := openai.New(openai.Config{APIKey: key, Model: "gpt-4o-mini"})
anthropicClient, _ := anthropic.New(anthropic.Config{
APIKey: key, Model: "claude-sonnet-4-5", MaxTokens: 1024,
})
geminiClient, _ := gemini.New(gemini.Config{APIKey: key, Model: "gemini-2.5-flash"})
azureClient, _ := azure.New(azure.Config{
APIKey: key, Endpoint: "https://my-resource.openai.azure.com",
Deployment: "gpt-4o", APIVersion: "2024-10-21",
})
bedrockClient, _ := bedrock.New(bedrock.Config{
Credentials: bedrock.Credentials{AccessKeyID: id, SecretAccessKey: secret},
Region: "us-east-1", Model: "us.anthropic.claude-sonnet-4-5-20250929-v1:0",
})
// Optional sampling controls: nil leaves the provider default, and
// providers.Ptr sets the value, including 0.
focusedClient, _ := openai.New(openai.Config{
APIKey: key, Model: "gpt-4o-mini",
Temperature: providers.Ptr(0.2), TopP: providers.Ptr(0.9), MaxTokens: 512,
})
API surface¶
openai.New(openai.Config{APIKey, BaseURL, Model, Temperature, TopP, MaxTokens, HTTPClient})anthropic.New(anthropic.Config{APIKey, BaseURL, Model, MaxTokens, Temperature, TopP, Effort, CacheControl, HTTPClient})anthropic.CacheControl{TTL}— automatic prompt caching;anthropic.CacheOneHourgemini.New(gemini.Config{APIKey, BaseURL, Model, Temperature, TopP, MaxTokens, HTTPClient})azure.New(azure.Config{APIKey, Endpoint, Deployment, APIVersion, Temperature, TopP, MaxTokens, HTTPClient})bedrock.New(bedrock.Config{Credentials, Region, Model, MaxTokens, Temperature, TopP, BaseURL, HTTPClient})providers.Ptr(v)builds the optional*float64sampling fields.model.FinishReason—FinishStop | FinishLength | FinishToolCall | FinishContentFilter | FinishOther; onmodel.Response,golem.Result, andgolem.PartialResult(decision in ADR 0020).openai.Embedder/azure.Embedder/gemini.Embedder— text embeddings beside generation; see the Embeddings API surface (decision in ADR 0021).anthropic.Counter/gemini.Counter/bedrock.Counter— input-token counting beside generation; see the Token counting API surface (decision in ADR 0023).- Errors per adapter:
APIError|TransportError|DecodeError
Gotchas¶
- Many providers speak the OpenAI chat-completions wire format; the
OpenAI adapter serves them all through
BaseURL:
| Provider | BaseURL |
|---|---|
| Groq | https://api.groq.com/openai/v1 |
| OpenRouter | https://openrouter.ai/api/v1 |
| DeepSeek | https://api.deepseek.com/v1 |
| Mistral | https://api.mistral.ai/v1 |
| xAI (Grok) | https://api.x.ai/v1 |
| Perplexity | https://api.perplexity.ai |
| Cerebras | https://api.cerebras.ai/v1 |
| Fireworks | https://api.fireworks.ai/inference/v1 |
| Together | https://api.together.xyz/v1 |
| Cohere (compatibility) | https://api.cohere.com/compatibility/v1 |
| Ollama (local) | http://localhost:11434/v1 |
| LM Studio (local) | http://localhost:1234/v1 |
| vLLM (local) | http://localhost:8000/v1 |
Compatibility is the providers' own promise — verify structured output
and streaming support against their docs; some gate json_schema
responses or stream_options behind specific model versions.
- Local runtimes work with zero code differences — point BaseURL at
the server and pick a model it has loaded. Ollama serves after
ollama pull <model>; LM Studio serves with lms server start
(or the Developer tab). Local servers ignore the Authorization
header; when pointing BaseURL at a custom endpoint, APIKey is
optional and golem omits the Authorization header. With a capable model loaded, the full core contract
holds over both: streaming (Ollama supports the include_usage
stream option golem always sends), tool calling (each call may arrive
complete in one chunk instead of argument fragments — the adapter
accepts both shapes), structured output via response_format
json_schema, image parts (Ollama), and ReasoningEffort (Ollama).
Tool support and quality depend on the loaded model, not on golem;
check the runtime's own capability matrix for the model you pull.
Reasoning text is best-effort on these surfaces: golem captures the
reasoning_content field some endpoints return, while Ollama exposes
its trace under a different name, so runs work but Thinking blocks
may stay empty. WithToolChoice is honored only by servers that
support tool_choice (Ollama's published docs list it; older
releases ignore it). Ollama has no API for context length — raise it
in the model's Modelfile, not per request.
- Anthropic requires a positive max_tokens; zero selects
anthropic.DefaultMaxTokens (1024).
- anthropic merges consecutive user-side messages (tool results plus a
following prompt) into one turn — the Messages API expects alternating
roles.
- Gemini function calls carry no provider ID: the adapter generates
stable ones and correlates function responses by tool name.
- Azure OpenAI shares the OpenAI wire format but addresses models by
deployment URL with an explicit api-version; there is no default
version, because versions gate feature support — structured output
needs one that supports response_format.
- Gemini structured output maps to generationConfig JSON responses
(responseMimeType + responseSchema); Gemini accepts a JSON-Schema
subset — no additionalProperties — and rejects JSON response mode
combined with function calling, so a request carrying both an output
schema and tool declarations fails at request encoding with a
DecodeError pointing at tool-mode output (golem.WithOutputTool).
- The Gemini SSE stream has no terminal sentinel beyond a chunk's
finishReason: a stream that ends at EOF without one is a decode
failure, not a silent partial response — mirroring the
OpenAI-compatible ([DONE]) and Anthropic (message_stop) sentinels.
- Structured output maps to OpenAI response_format, Anthropic
output_config, and Gemini generationConfig; OpenAI and Anthropic
require strict-conformant schemas (see
Structured output).
- Bedrock credentials are wired in explicitly — the adapter never reads
the AWS environment or credential chain; requests are SigV4-signed
with the standard library. Streaming decodes ConverseStream's AWS
binary event-stream framing with the standard library (no AWS SDK):
text and tool-use fragments stream as deltas, mid-stream exception
frames classify like their HTTP equivalents, and a stream that ends
without a messageStop event is a decode failure, not a silent
partial response. OutputSchema maps to the Converse outputConfig
json_schema format with the schema passed as a string.
- Sampling ranges differ per provider and are validated at
construction: temperature is [0, 2] on OpenAI, Azure, and Gemini but
[0, 1] on Anthropic and Bedrock; top P is [0, 1] everywhere. Anthropic
requires max_tokens on every request, so its zero selects
anthropic.DefaultMaxTokens; on the other adapters zero omits the
bound.
- Deciding conventions live in docs/adr/0003-provider-adapter-conventions.md.