Skip to content

Providers

Purpose

Connect an agent to a real generation API through a stdlib-only adapter, with typed, retry-classified errors.

When to use

Whenever a run should reach a live model. Anything implementing model.Model works; the shipped adapters cover OpenAI-compatible APIs and Anthropic.

How it works

Adapters translate the provider-neutral model.Request to their wire format and normalize responses back. They never retry on their own: failures are classified through model.RetryableError (408, 429, 5xx, transport faults — never context cancellation) and the runner's retry policy decides. All shipped adapters implement model.StreamingModel — over SSE everywhere except Bedrock, whose ConverseStream responses are AWS binary event-stream frames decoded with the standard library. Every adapter also translates media parts on user messages — images, documents, audio, video — to its native multimodal form, and rejects unsupported combinations before any request; the per-provider matrix lives in Multimodal input.

Every adapter accepts optional sampling and length controls on its Config — Temperature, TopP, and MaxTokens — validated against the provider's documented ranges at construction. A nil sampling field leaves the provider default, and an unset MaxTokens omits the bound, so unset controls never appear on the wire. Temperature 0 is a meaningful value, not a default: set it with providers.Ptr(0).

Every response also carries the provider's terminal cause, normalized onto model.FinishReason — identical on streamed and plain runs:

Adapter Wire field Stop Length Tool call Content filter
OpenAI / Azure finish_reason stop length tool_calls, function_call content_filter
Anthropic stop_reason end_turn, stop_sequence max_tokens tool_use refusal
Bedrock stopReason end_turn, stop_sequence max_tokens tool_use guardrail_intervened, content_filtered
Gemini finishReason STOP MAX_TOKENS — SAFETY, RECITATION, BLOCKLIST, PROHIBITED_CONTENT, SPII

Prompt caching

Every provider can serve a repeated prompt prefix from cache instead of re-processing it; the adapters expose the control where the provider offers one, and every adapter reports the outcome on model.Usage's cache token fields, so a run can compute what caching saved.

  • Anthropic — anthropic.Config.CacheControl requests the provider's automatic caching: a top-level cache_control field, and the provider applies the breakpoint to the last cacheable block of each request, moving it forward as the conversation grows. A zero TTL selects the five-minute default; anthropic.CacheOneHour asks for the one-hour entry at a premium. Anthropic-compatible gateways reached through BaseURL may not support the parameter — the provider's own API does.
  • Bedrock — bedrock.Config.CacheControl places an explicit cachePoint checkpoint at each request's conversation frontier, so the prefix it covers — tool declarations, system guidance, and history — is cached and re-read on later requests. TTL semantics are the same (bedrock.CacheOneHour where the model supports it), and Anthropic models on Bedrock also cache implicitly without any opt-in.
  • OpenAI and Azure — caching is automatic for eligible prefixes (about a thousand tokens and up); there is nothing to send, and the cached share arrives on Usage.CacheReadTokens.
  • Gemini — caching is implicit for current models; nothing to send, and Usage.CacheReadTokens reports hits. The explicit context-caching API (a separate cached-content resource with its own lifecycle) stays out of scope.

Usage carries detail where providers break it out, normalized onto three model.Usage fields beside the token totals:

Adapter CacheReadTokens CacheWriteTokens ReasoningTokens
OpenAI / Azure prompt_tokens_details.cached_tokens — inside input — (implicit caching) completion_tokens_details.reasoning_tokens — inside output
Anthropic cache_read_input_tokens — beside input cache_creation_input_tokens — beside input —
Bedrock cacheReadInputTokens — beside input cacheWriteInputTokens — beside input —
Gemini cachedContentTokenCount — inside input — (implicit caching) thoughtsTokenCount — inside output

"Inside" and "beside" matter for pricing: a subset of the input total discounts part of it, an exclusive figure replaces it. Detail fields are zero when a provider (or a proxied endpoint) does not report them, and streamed runs carry the same figures as plain ones — every adapter reports streamed usage through its terminal chunk or event.

Unlisted values — Anthropic's pause_turn, Gemini's MALFORMED_FUNCTION_CALL, anything a provider adds later — map to FinishOther, and a missing field maps to the empty FinishReason; unrecognized causes are evidence, never failures. Result.FinishReason holds the final turn's cause, and RunError.Partial.FinishReason the last completed turn's when a run failed: a length cause on a decode failure means the provider truncated the output, not that the model wrote bad JSON.

Pricing

Each generation adapter ships a Price struct implementing model.Price: per-million-token USD rates you supply, combined by the rule that adapter's own usage semantics require. Wire one with golem.WithPrice and runs report Result.Cost (and enforce UsageLimit.Cost); see Cost for the full contract. Golem ships no price table — the rates are your snapshot of the provider's price list.

Price Cached input Cache writes Rates
openai.Price inside the input total — Input − CacheRead bills at the input rate — (implicit caching) input, output, cache read
azure.Price inside the input total — (implicit caching) input, output, cache read
anthropic.Price beside the input total — billed at its own rate beside input, premium rate input, output, cache read, cache write
bedrock.Price beside the input total beside input input, output, cache read, cache write
gemini.Price inside the input total — (implicit caching) input, output, cache read

With no cache-read rate set, the inside-input prices bill cached input at the input rate rather than free; the beside-input prices bill the input total in full by construction. Reasoning tokens are part of billed output on every adapter and take no separate rate.

Embeddings

Generation adapters are not the only port: the same providers serve text embeddings through the embedding.Embedder port, covered end to end in Embeddings. Query and document calls share one Result shape — vectors in input order plus model.Usage input-token evidence — and the split maps to the provider's task type where one exists:

Embedder Wire Query / documents Dimensions Usage
openai.Embedder POST {BaseURL}/embeddings one shape, split ignored dimensions (text-embedding-3+) usage.prompt_tokens
azure.Embedder POST {Endpoint}/openai/deployments/{deployment}/embeddings one shape, split ignored dimensions (text-embedding-3+) usage.prompt_tokens
gemini.Embedder :embedContent / :batchEmbedContents RETRIEVAL_QUERY / RETRIEVAL_DOCUMENT outputDimensionality (newer models) not reported today — stays zero

Anthropic has no embeddings API, so there is no Anthropic embedder. Bedrock's embedding models (Titan, Cohere on Bedrock) are likewise not covered yet; say the word if your index runs there.

Token counting

The same asymmetry runs the other way for counting: where the provider offers a counting endpoint, an adapter implements the tokens.Counter port — covered end to end in Token counting — and prices exactly what a generation request with the same builders would send:

Counter Wire Counts tools Notes
anthropic.Counter POST {BaseURL}/v1/messages/count_tokens yes same request schema as inference
gemini.Counter POST {BaseURL}/v1beta/models/{model}:countTokens yes
bedrock.Counter POST {BaseURL}/model/{model}/count-tokens no — lower bound free; base foundation-model IDs only

OpenAI, Azure OpenAI, and OpenAI-compatible endpoints have no counting API, so there is no counter for them; implement tokens.Counter with your own tokenizer there.

Example

  • examples/minimal — OpenAI-compatible.
  • examples/anthropic — Anthropic Messages API.
  • examples/gemini — Google Gemini GenerateContent API.
  • examples/azure — Azure OpenAI deployments.
  • examples/bedrock — AWS Bedrock Converse with SigV4.
  • examples/local-models — Ollama or LM Studio through the OpenAI adapter.
OPENAI_API_KEY=sk-... go run ./examples/minimal
ANTHROPIC_API_KEY=sk-... go run ./examples/anthropic
GEMINI_API_KEY=... go run ./examples/gemini
AZURE_OPENAI_API_KEY=... AZURE_OPENAI_ENDPOINT=https://my-resource.openai.azure.com \
    AZURE_OPENAI_DEPLOYMENT=gpt-4o AZURE_OPENAI_API_VERSION=2024-10-21 \
    go run ./examples/azure
AWS_ACCESS_KEY_ID=... AWS_SECRET_ACCESS_KEY=... AWS_REGION=us-east-1 \
    go run ./examples/bedrock
GOLEM_LOCAL_BASE_URL=http://localhost:11434/v1 go run ./examples/local-models
openaiClient, _ := openai.New(openai.Config{APIKey: key, Model: "gpt-4o-mini"})
anthropicClient, _ := anthropic.New(anthropic.Config{
    APIKey: key, Model: "claude-sonnet-4-5", MaxTokens: 1024,
})
geminiClient, _ := gemini.New(gemini.Config{APIKey: key, Model: "gemini-2.5-flash"})
azureClient, _ := azure.New(azure.Config{
    APIKey: key, Endpoint: "https://my-resource.openai.azure.com",
    Deployment: "gpt-4o", APIVersion: "2024-10-21",
})
bedrockClient, _ := bedrock.New(bedrock.Config{
    Credentials: bedrock.Credentials{AccessKeyID: id, SecretAccessKey: secret},
    Region: "us-east-1", Model: "us.anthropic.claude-sonnet-4-5-20250929-v1:0",
})

// Optional sampling controls: nil leaves the provider default, and
// providers.Ptr sets the value, including 0.
focusedClient, _ := openai.New(openai.Config{
    APIKey: key, Model: "gpt-4o-mini",
    Temperature: providers.Ptr(0.2), TopP: providers.Ptr(0.9), MaxTokens: 512,
})

API surface

  • openai.New(openai.Config{APIKey, BaseURL, Model, Temperature, TopP, MaxTokens, HTTPClient})
  • anthropic.New(anthropic.Config{APIKey, BaseURL, Model, MaxTokens, Temperature, TopP, Effort, CacheControl, HTTPClient})
  • anthropic.CacheControl{TTL} — automatic prompt caching; anthropic.CacheOneHour
  • gemini.New(gemini.Config{APIKey, BaseURL, Model, Temperature, TopP, MaxTokens, HTTPClient})
  • azure.New(azure.Config{APIKey, Endpoint, Deployment, APIVersion, Temperature, TopP, MaxTokens, HTTPClient})
  • bedrock.New(bedrock.Config{Credentials, Region, Model, MaxTokens, Temperature, TopP, BaseURL, HTTPClient})
  • providers.Ptr(v) builds the optional *float64 sampling fields.
  • model.FinishReason — FinishStop | FinishLength | FinishToolCall | FinishContentFilter | FinishOther; on model.Response, golem.Result, and golem.PartialResult (decision in ADR 0020).
  • openai.Embedder / azure.Embedder / gemini.Embedder — text embeddings beside generation; see the Embeddings API surface (decision in ADR 0021).
  • anthropic.Counter / gemini.Counter / bedrock.Counter — input-token counting beside generation; see the Token counting API surface (decision in ADR 0023).
  • Errors per adapter: APIError|TransportError|DecodeError

Gotchas

  • Many providers speak the OpenAI chat-completions wire format; the OpenAI adapter serves them all through BaseURL:
Provider BaseURL
Groq https://api.groq.com/openai/v1
OpenRouter https://openrouter.ai/api/v1
DeepSeek https://api.deepseek.com/v1
Mistral https://api.mistral.ai/v1
xAI (Grok) https://api.x.ai/v1
Perplexity https://api.perplexity.ai
Cerebras https://api.cerebras.ai/v1
Fireworks https://api.fireworks.ai/inference/v1
Together https://api.together.xyz/v1
Cohere (compatibility) https://api.cohere.com/compatibility/v1
Ollama (local) http://localhost:11434/v1
LM Studio (local) http://localhost:1234/v1
vLLM (local) http://localhost:8000/v1

Compatibility is the providers' own promise — verify structured output and streaming support against their docs; some gate json_schema responses or stream_options behind specific model versions. - Local runtimes work with zero code differences — point BaseURL at the server and pick a model it has loaded. Ollama serves after ollama pull <model>; LM Studio serves with lms server start (or the Developer tab). Local servers ignore the Authorization header; when pointing BaseURL at a custom endpoint, APIKey is optional and golem omits the Authorization header. With a capable model loaded, the full core contract holds over both: streaming (Ollama supports the include_usage stream option golem always sends), tool calling (each call may arrive complete in one chunk instead of argument fragments — the adapter accepts both shapes), structured output via response_format json_schema, image parts (Ollama), and ReasoningEffort (Ollama). Tool support and quality depend on the loaded model, not on golem; check the runtime's own capability matrix for the model you pull. Reasoning text is best-effort on these surfaces: golem captures the reasoning_content field some endpoints return, while Ollama exposes its trace under a different name, so runs work but Thinking blocks may stay empty. WithToolChoice is honored only by servers that support tool_choice (Ollama's published docs list it; older releases ignore it). Ollama has no API for context length — raise it in the model's Modelfile, not per request. - Anthropic requires a positive max_tokens; zero selects anthropic.DefaultMaxTokens (1024). - anthropic merges consecutive user-side messages (tool results plus a following prompt) into one turn — the Messages API expects alternating roles. - Gemini function calls carry no provider ID: the adapter generates stable ones and correlates function responses by tool name. - Azure OpenAI shares the OpenAI wire format but addresses models by deployment URL with an explicit api-version; there is no default version, because versions gate feature support — structured output needs one that supports response_format. - Gemini structured output maps to generationConfig JSON responses (responseMimeType + responseSchema); Gemini accepts a JSON-Schema subset — no additionalProperties — and rejects JSON response mode combined with function calling, so a request carrying both an output schema and tool declarations fails at request encoding with a DecodeError pointing at tool-mode output (golem.WithOutputTool). - The Gemini SSE stream has no terminal sentinel beyond a chunk's finishReason: a stream that ends at EOF without one is a decode failure, not a silent partial response — mirroring the OpenAI-compatible ([DONE]) and Anthropic (message_stop) sentinels. - Structured output maps to OpenAI response_format, Anthropic output_config, and Gemini generationConfig; OpenAI and Anthropic require strict-conformant schemas (see Structured output). - Bedrock credentials are wired in explicitly — the adapter never reads the AWS environment or credential chain; requests are SigV4-signed with the standard library. Streaming decodes ConverseStream's AWS binary event-stream framing with the standard library (no AWS SDK): text and tool-use fragments stream as deltas, mid-stream exception frames classify like their HTTP equivalents, and a stream that ends without a messageStop event is a decode failure, not a silent partial response. OutputSchema maps to the Converse outputConfig json_schema format with the schema passed as a string. - Sampling ranges differ per provider and are validated at construction: temperature is [0, 2] on OpenAI, Azure, and Gemini but [0, 1] on Anthropic and Bedrock; top P is [0, 1] everywhere. Anthropic requires max_tokens on every request, so its zero selects anthropic.DefaultMaxTokens; on the other adapters zero omits the bound. - Deciding conventions live in docs/adr/0003-provider-adapter-conventions.md.