kit-chat-completions

    Posoco chat-completions kit — shared OpenAI chat-completions dialect parsing (SSE streaming) with a uniform real-world tolerance policy

    posoco
    sse
    streaming
    chat-completions
    model-port
    Download zip
    Author
    Version
    0.1.0
    License
    Apache-2.0
    Last updated
    9 hours ago
    Downloads
    5

    Dependencies

    EncodedChatContent

    pub(all) enum EncodedChatContent {
    TextContent(String)
    PartsContent(Array[Json])
    }

    Result of encoding one message's content blocks.

    cached_input_tokens_from_usage

    fn cached_input_tokens_from_usage(usage : Map[String, Json]) -> Int?

    classify_chat_completions_http_error

    fn classify_chat_completions_http_error(status : Int, body_text : String, default_tz_offset_minutes? : Int, retry_after_header? : String?, now_ms~ : Int64) ->
    ModelError
    ?

    Classify an HTTP status + response body pair. Returns Some(ModelError::RateLimited(info)) for 429 verdicts and, like Kimi's coding plan, 403 verdicts whose body phrasing reads as a quota wall; None otherwise — the caller keeps its existing Transport formatting for every other non-2xx status. reset_at_ms inside the info is None when the provider stated no schedulable reset time (transient concurrency limits); such verdicts are not schedulable and the caller should back off or fail.

    Extraction order (first hit wins):
    1. resets_in_seconds — relative seconds from now (OpenAI/Codex shape)
    2. resets_at / reset_at — epoch seconds (or milliseconds)
    3. retry_after / retry-after — seconds, or a timestamp string
    4. the Retry-After HTTP header, when the caller can read one (the documented wait signal for Kimi/Moonshot verdicts, whose message text carries no timestamp)
    5. the provider message text (z.ai "{next_flush_time}", Codex English dates, relative "in N hours")
    6. the named quota-window length ("5-hour usage limit") — an upper bound on the true reset time

    Lookup covers both the nested error object and the top level. When the body is not JSON at all, only paths 4-6 apply against the raw text.

    derive_prompt_cache_key

    fn derive_prompt_cache_key(session_id : String) -> String

    Stable provider-side prompt_cache_key for a posoco session id.

    Normal session ids keep the historical "posoco-" + session_id shape so existing provider affinity is not churned unnecessarily. Oversized ids are replaced with a versioned SHA-256-derived key instead of being truncated: posoco-v1-<48 lowercase hex chars>. The hashed form is 58 chars, stays below the 64-char provider limit, and two ids that share a long prefix do not collapse to the same cache key.

    encode_chat_message_content

    fn encode_chat_message_content(blocks : Array[
    Content
    ], supports_images? : Bool) -> EncodedChatContent

    Encode content blocks for a chat-completions message content field.

    • No image blocks → TextContent (concatenated text).
    • Image blocks with supports_images=true → PartsContent.
    • Image blocks with supports_images=false → TextContent with each image downgraded to an explicit placeholder (never a silent drop).

    fold_tool_attachments

    fn fold_tool_attachments(messages : Array[
    Message
    ], supports_images : Bool) -> Array[
    Message
    ]

    Fold tool-result attachments into the chat-completions wire shape.

    The tool role carries only a text content string on this protocol, so attachment blocks cannot ride the tool message itself:
    • supports_images=true: a contiguous run of tool messages keeps its text, and the run's attachments collapse into ONE synthetic user message emitted right after the run. The fold is a pure function of the message array, so provider prefix caches replay identically.
    • supports_images=false: every attachment degrades to the explicit [Attached …, omitted] placeholder appended to its tool text — never a silent drop.

    image_placeholder

    fn image_placeholder(media_type : String, data : String) -> String

    Placeholder for an image the endpoint cannot accept. The size is the decoded-byte estimate (base64 length × 3/4) so the model can gauge how much was attached without seeing it.

    is_prompt_cache_key_rejection

    fn is_prompt_cache_key_rejection(body_text : String) -> Bool

    Whether a provider error body explicitly rejects the prompt_cache_key parameter: true only when the text contains BOTH "unsupported parameter" AND "prompt_cache_key", case-insensitively. Empty bodies, bodies naming only one phrase, and unrelated 400s never match.

    is_transient_model_error

    fn is_transient_model_error(error :
    ModelError
    ) -> Bool

    Transient-failure classification shared by every chat-completions provider: the HTTP/SSE message shapes this kit's http_utils and SSE processors produce. The router-level retry policy (posoco-ext-llm) keys off this predicate, so a provider inherits retrying by emitting these shapes — no per-provider retry wrapper.

    Conservative by design: only enumerated shapes retry.
    • ResponseParse "SSE truncated" — marker-less stream EOF.
    • Transport "stage=read_stream" — mid-stream connection reset.
    • Transport "status=408" / "status=5xx" — gateway timeouts and 5xx.

    Never retryable: RateLimited (429 belongs to the rate-limit resume domain), 401/403 (auth renewal domain), other 4xx, RequestBuild, and unmarked Transport messages.

    outcome_wire_text

    fn outcome_wire_text(outcome :
    ToolOutcome
    ) -> String

    Plain-text projection of a tool outcome (attachments contribute nothing here; they travel the message-level strategy below).

    parse_reset_timestamp

    fn parse_reset_timestamp(text : String, default_tz_offset_minutes? : Int, now_ms~ : Int64) -> Int64?

    Parse a provider-stated reset time out of free text. Supported shapes:
    • 2026-08-21T15:04:16Z, 2026-08-21 15:04:16, 2026-08-21 15:04:16(UTC+8)
    • Aug 13, 2026 at 15:52, Aug 13, 2026 at 3:45pm
    • relative: in 8 hours 30 minutes, in 13 minutes, in 45 seconds

    default_tz_offset_minutes~ applies only when the text carries no explicit timezone (z.ai flush times are UTC+8; a generic gateway's are usually UTC). Returns epoch milliseconds.

    process_chat_completions_sse

    fn process_chat_completions_sse(provider~ : String, usage_fields_required? : Bool, data : String, acc :
    StreamAccumulator
    , on_chunk : (
    StreamChunk
    ) -> Unit) -> Bool raise
    ModelError

    Parse one SSE data line into StreamChunk(s) and feed to accumulator. Returns true when the stream is finished ([DONE] or finish_reason present). provider~ labels every raised message (e.g. "openai-compatible", "DeepSeek", "kimi"). usage_fields_required~ selects the streamed-usage policy: DeepSeek/kimi require every token field (loud on missing), the generic compatible adapter maps missing fields to None.

    uncached_input_tokens_from_usage

    fn uncached_input_tokens_from_usage(usage : Map[String, Json]) -> Int?

    Extract the cached input-token count from one chat-completions usage object. NO openai-compatible endpoint is obliged to report caching at all — OpenAI itself only populates the field for prompts at or above the ~1024-token caching threshold — so absence is a normal outcome, mapped to None, never a parse failure. The spellings, each tied to a documented source:
    • OpenAI (and vLLM following its usage schema): prompt_tokens_details.cached_tokens (developers.openai.com/api/docs/guides/prompt-caching)
    • DeepSeek: prompt_cache_hit_tokens (hit + miss = prompt_tokens) (api-docs.deepseek.com/guides/kv_cache)
    • Moonshot: usage-top-level cached_tokens — the spelling this repo's kimi extension has always parsed (platform.kimi.ai context caching)
    • Anthropic-origin field surfaced by OpenAI-compatible gateways (Portkey, LiteLLM, TrueFoundry, Zenlayer, LangWatch): cache_read_input_tokens An endpoint reporting some yet-unknown spelling also reads None — the consumer degrades to "no cache facts", never a fabricated rate.