All notable changes to this project will be documented in this file.
Source-breaking for downstream code that constructs these records directly or pattern-matches without ; _. Code that goes through the documented constructors (Core_tool.create, Prompt_builder.resolve_messages, Cache_control.ephemeral / ephemeral_1h) is unaffected.
Ai_provider.Prompt.System gained provider_options : Provider_options.t alongside content. Callers building System { content } literally must add provider_options = Ai_provider.Provider_options.empty.Ai_provider.Tool.t gained provider_options : Provider_options.t. The in-tree OpenAI / OpenRouter providers were updated; external providers constructing this record need the same.Ai_provider.Stream_part.Finish gained provider_metadata. The standalone Provider_metadata constructor is removed β its data now rides on Finish.Ai_core.Core_tool.t gained provider_options. The create / create_with_approval / create_client_tool helpers take an optional ?provider_options defaulting to empty, so call sites that use them are unaffected.Ai_provider_anthropic.Cache_control.t gained ttl : ttl option. Construct via Cache_control.ephemeral (5m, default) or Cache_control.ephemeral_1h.Ai_provider_anthropic.Cache_control.breakpoint_to_json / breakpoint_of_json from the public mli β they silently dropped the ttl field and had no in-tree callers.ai_provider_anthropic)cache_control plumbing landed in 0.3 was only usable through Ai_provider.Language_model directly. Ai_core.Generate_text.generate_text, Ai_core.Stream_text.stream_text, and Ai_core.Server_handler.handle_chat now accept ?system_provider_options for the prepended system prompt, and Core_tool.t gains a provider_options field that flows through to the provider Tool record. The runnable examples/prompt_caching demo exercises the new path.Stream_part.Finish carries provider_metadata matching upstream LanguageModelV4StreamPart. The standalone Provider_metadata chunk is removed; Anthropic cache token metrics now ride on the terminal Finish chunk on cached requests, and Stream_text_result.provider_metadata / Generate_text_result.step.provider_metadata expose them to callers. Cache fields are read from message_start.message.usage (where Anthropic actually emits them) with a fallback to message_delta.usage.@ai-sdk/anthropic. The previous "joined string when no cache_control, array when cache_control is set" behavior changed the model's input based on a feature flag.cache_creation fields, so new provider-side usage metadata does not break response parsing.ai_provider_openrouter)Cache_control / Cache_control_options modules for explicit per-block cache breakpoints, including Anthropic-compatible ttl values (5m / 1h) and fallback support for prompts that already use the Anthropic cache-control key.cache_control placement for system, user, assistant, and tool messages, while preserving the no-cache shape for non-system messages.usage.prompt_tokens_details is mapped into provider metadata as cache_read_tokens and cache_write_tokens. Added examples/openrouter_prompt_caching to demonstrate both top-level automatic caching and explicit breakpoint mode.ai_provider_openrouter)finish_reason = "error" now map to Finish_reason.Error, and error-shaped completion bodies (top-level or per-choice error objects) raise a Provider_error instead of being reduced to empty output. The error body is a human-readable message (upstream provider name, the most specific upstream message from error.metadata.raw, and the error_type suffix) rather than a raw JSON dump. Retryability now keys off the real HTTP status on the transport error path (an inner provider error.code no longer overrides a 5xx gateway status), and 200-embedded / streaming errors derive their status from error.code, tolerating string and float encodings.ai_provider_anthropic)Object_json mode now uses Anthropic's native output_config.format = { type: "json_schema", schema } field on capable models (Haiku 4.5, Sonnet 4.5/4.6, Opus 4.5/4.6/4.7), matching upstream @ai-sdk/anthropic. Schema enforcement is handled by the provider, not by appending instructions to the system prompt.Custom model ids, the provider synthesises a tool named json carrying the schema as input_schema and forces tool_choice = { type: "tool", name: "json" }. The caller's system prompt is left untouched.Object_json None (no schema) now receive an Unsupported_feature warning because Anthropic cannot enforce JSON without a schema.Claude_opus_4_7. The supports_structured_output capability flag is now accurate per model (previously defaulted to true for all known models).ai_core)Output.parse_output β when a step has no assistant text, falls back to decoding the json tool call's args. Enables end-to-end structured output on the Anthropic fallback path and on any future provider that adopts the same convention.Stream_text β Tool_call_delta events for the json tool drive the partial-output parser, so streaming callers see incremental JSON on the fallback path with the same UX as the native path.ai_provider)Ai_provider.Http_timeouts module and Ai_provider.Http_client wrapper. Defaults: 600s for response headers (request_timeout) and 300s for silence between streaming chunks (stream_idle_timeout). Override per-provider via Config.create ?timeouts. Conservative values chosen to catch stuck connections and bugs, not bound legitimate workloads β a 20-minute streaming response completes fine as long as chunks keep flowing.Provider_error.Timeout kind with phase (Request_headers | Stream_idle), elapsed_s, and limit_s. is_retryable is derived: Stream_idle is retryable (connection is dead); Request_headers is not (server may already be processing the request).Sse.parse_events no longer hangs consumers on upstream errors. Previously, an exception from the upstream line stream left the output stream pending forever. It now closes cleanly (via push None) and re-raises to Lwt.async_exception_hook so the underlying bug stays visible.Mode.fallback_json_tool_name β exported constant ("json") naming the synthetic tool used by the structured-output tool-use fallback convention. Shared between ai_core and providers so the convention has a single source of truth.ai_provider_openai, ai_provider_anthropic, ai_provider_openrouter)Config.t gains a timeouts : Http_timeouts.t field. All HTTP traffic now routes through Http_client, removing three copies of the unguarded body_to_line_stream helper.structured_output β live-API smoke test exercising both the native and tool-fallback paths with ppx_deriving_jsonschema for schema derivation and melange-json-native's of_json deriver for typed response decoding.ai_core)Smooth_stream β stream transformer that buffers Text_delta and Reasoning_delta chunks and re-emits them in controlled pieces with configurable inter-chunk delays. Five chunking modes: Word (default), Line, Regex (custom Re2 pattern), Segmenter (Unicode UAX#29 word boundaries via uuseg, recommended for CJK), and Custom (user function). Matches the upstream AI SDK's smoothStream transform.?transform parameter on stream_text and server_handler.handle_chat β generic stream transformer (Text_stream_part.t Lwt_stream.t -> Text_stream_part.t Lwt_stream.t) applied between the raw event stream and consumer-facing streams. Both full_stream and text_stream reflect the transformed output.Retry module with jitter, configurable initial delay and backoff factor, and parameter validation. ?max_retries threaded through generate_text, stream_text, and server_handler.handle_chat. Retries only on errors marked retryable.Telemetry module with OpenTelemetry-compatible span instrumentation via the trace library (ocaml-trace). Configurable Telemetry.t settings control enable/disable, input/output recording privacy, function ID, custom metadata, and lifecycle integration callbacks (on_start, on_step_finish, on_tool_call_start, on_tool_call_finish, on_finish). Span hierarchy matches upstream AI SDK: ai.generateText / ai.streamText root spans, *.doGenerate / *.doStream step spans, and ai.toolCall tool execution spans. ?telemetry parameter threaded through generate_text, stream_text, and server_handler.handle_chat.ai_provider)is_retryable field on Provider_error.t β defaults from HTTP status code (429, 5xx are retryable). Anthropic and OpenAI providers set it explicitly based on error classification.smooth_streaming β demonstrates all five chunking modestelemetry_logging β demonstrates integration callbacks for lifecycle loggingre2 (>= 0.16) and uuseg (>= 17.0) to ai_coretrace (>= 0.12) to ai_coreInitial release of the OCaml AI SDK β a type-safe, provider-agnostic AI model abstraction inspired by the Vercel AI SDK, targeting AI SDK v6 wire compatibility.
ai_provider)Provider_options for compile-time type-safe provider-specific settingsPrompt types (System = string only, User = text + files, etc.)Language_model.S module type with first-class module wrapperTool, Tool_choice, Mode, Content foundation typesFinish_reason, Usage, Warning, Provider_error typesProvider.S and Middleware.S module type signaturesCall_options, Generate_result, Stream_part, Stream_result typesai_provider_anthropic)Thinking support with budget_tokens smart constructor (>= 1024)Cache_control for prompt cachingAnthropic_options via the extensible GADT systemmax_tokensai_provider_openai)ai_core)generate_text β synchronous text generation with multi-step tool loopstream_text β streaming text generation with multi-step tool loop, returns synchronously with streams filled by background Lwt taskOutput.text, Output.object_, Output.enum, Output.array, Output.choice with JSON Schema validationdata: {json}\n\n encoding with x-vercel-ai-ui-message-stream: v1 header, all v6 chunk typesUi_message_stream_writer β composable stream builder with write (synchronous) and merge (non-blocking via Lwt.async), lifecycle management, ref-counted in-flight merge tracking, on_finish callbackneeds_approval predicate on Core_tool.t, step loop partitioning, Tool_approval_request chunk type, stateless re-submission with approved_tool_call_idsStop_condition β step loop termination predicates matching upstream stopWhen: step_count_is, has_tool_call, is_met (OR semantics with short-circuit); wired through generate_text, stream_text, and server_handler; max_steps remains as independent hard safety capai-sdk-react)useChat and useCompletion hook bindings for @ai-sdk/reactdata_ui_partclassify function for part type dispatchone_shot, streaming, tool_use, thinking, generate, stream_chat, agent_loop β standalone CLI exampleschat_server β cohttp chat server with React frontend, tool approval, structured outputcustom_stream β custom data streaming with Melange frontendai-e2e β end-to-end Melange app with 11 demos (basic chat, reasoning, tool use, tool approval, client tools, file attachments, structured output, completion, web search, retry/regenerate)generate_opam_files for automated opam file generationmlx-pp / ocamlformat-mlx)