ext-llm

    Posoco LLM Extension - host-injected model slot router

    posoco
    llm
    router
    model-port
    Download zip
    Author
    Version
    0.2.0
    License
    Apache-2.0
    Last updated
    9 hours ago
    Downloads
    3

    #posoco-ext-llm

    posoco-ext-llm is a provider-agnostic model router. It does not construct HTTP adapters and it does not depend on DeepSeek, Kimi, or OpenAI. Provider extensions construct their own ProviderModelCatalog values; the host only assembles those catalogs.

    let provider = make_provider_port()
    let catalog = provider.model_catalog()
    let router = @llm.RouterModelPort::from_catalogs(catalogs=[catalog])

    RouterModelPort delegates ModelPort calls to the active slot. Provider catalogs advertise effort choices and provide a rebuild hook, so /model {slot, effort} validates and applies effort without the host knowing provider request fields.

    When a composed DecisionPort is available and the active slot advertises multiple reasoning efforts, the router may classify the current user request and choose a temporary per-call effort. This semantic judgement does not choose a provider or model and does not mutate the router's active selection. An explicit user-selected effort always wins; missing/failed/invalid decision results fall back to the existing active slot unchanged.

    The router also declares these commands:

    • /model [slot] lists the slot catalog or switches the active slot. The no-argument catalog enriches each slot-contributing provider's entries with its devkit-registered pricing phase (pricing: {tier, multiplier, window} facts stated by the provider extension; omitted when none is registered — src/command_model_pricing.mbt:13-41); the quick-pick, slot-switch, and effort payloads carry no pricing.
    • /model {slot, effort} additionally selects a provider-advertised effort.
    • /model quick-pick returns the slot catalog enriched with each slot-contributing provider's quota readings (host-injected sources first, then the devkit registry; sources are read concurrently, each under a 2500 ms cap; a failed or slow provider is omitted, never estimated — src/command_model_quick.mbt:23-73), and refreshes the active provider's status bar when its pull succeeds (src/command_model.mbt:349-359). The keyword is reserved and wins over a real slot of the same name; an effort argument is rejected (src/command_model.mbt:342-346).
    • /model <provider> pick auto-picks a cost-efficient model+effort from the provider's devkit-registered metrics source: points whose model names one of the provider's slots and whose effort that slot advertises, guarded by a minimum benchmark sample (total >= 30) and required to state every scored dimension (token total, duration, positive pass count; token efficiency is amortized tokens per solved problem — tokens * total / passed — because failed attempts burn tokens too). Points dominated on all four dimensions (IQ, token efficiency, cost, duration) drop out before ranking, then the survivors rank by iq - 1.0 * (tokens_per_pass / 1M) - 2.0 * cost_usd - 0.5 *(minutes / 10) — the priority gradient 分数 > tokens > 金钱 > 时间, calibrated so the expensive-efficient pick beats a cheap-but-verbose one (107.81 IQ @ $2.26 / 1.53M tokens-per-pass / 9 min wins over 102.23 IQ @ $0.54 / 26.34M / 37 min). The winner switches exactly like a manual /model {slot, effort}; ranks 2-3 ride along in the structured payload as {slot_id, model, effort, iq,cost_usd} candidates for host-side notices. A failed read, an unknown provider, or an empty eligible pool fails the command — never a fallback pick (src/command_model_pick.mbt, intercept at src/command_model.mbt:367-390).
    • /login [provider] lists provider authentication capabilities or runs the selected injected provider login flow. Use /login {provider: "kimi", method: "oauth"} or /login {provider: "deepseek", method: "api_key"}.
    • /status reports the active provider's facts as multi-line feedback (plus a structured JSON twin): provider/model ids, selected effort, context window (provider-stated when the slot carries one), reasoning efforts, and the 5h/weekly quota windows when a provider extension has published them on the shared bus (Codex's quota publisher does). It also pulls the active provider's registered quota source live at invocation through the devkit quota registry (DeepSeek's balance, Z.ai Coding Plan's windows); bus windows win over registry duplicates, and a provider Err reading is skipped, never estimated. With no bus and no pull readings it states quota: no readings yet instead of estimating. Quota windows render codex CLI-style as 5h: 66% left: registry readings convert the provider-declared used percent with exact arithmetic (100 − used, not an estimate), and bus-published window values — already published in N%left shape by their provider — are passed through unchanged.

    #Status facts

    With an EventBus injected (bus? = None by default — without one every publication below is a silent no-op), the router publishes three status-bar segments through the devkit status protocol for a status-bar bridge to render:

    • model (priority 20) — "<model-id>:<effort>": the active slot's model id plus the selected effort, falling back to the slot's declared default effort and then "default".
    • ctx (priority 30) — ctx: <occ>/<window> · <pct>: the core-projected context state of the active session (src/observer.mbt:56-145). Occupancy is the last measured reading plus the core-reported estimate, marked ~ whenever an estimated component is included; an absent or untrusted reading renders ? — never zero. The window renders ? when the slot carries none, and the percentage appears only when occupancy and a positive window are both known. State is per session: TurnStarted re-publishes the scoped session's own state and never resets occupancy, and compact/operation-lifecycle events never touch the segment (core projects the state around them; post-compact it is explicitly awaiting measurement; src/router_wbtest.mbt:411-663).
    • cache (priority 40) — "<n>%": the last committed round's cached-input percentage (cached × 100 / input). When the round reports no usable cache counters (or zero input) the segment is unregistered so a stale percentage never lingers.

    Publication triggers: ContextStateUpdated (store the session's state and publish ctx), TurnStarted (re-publish the scoped session's ctx — no reset), and ModelResponseReceived (republish cache and model). A successful /model or /effort switch re-publishes model immediately — slot switches go through switch_slot, which publishes too, so the bar reflects the choice before the next turn.

    Whenever the active slot's provider can change — on_compose, a /model switch, /effort, switch_slot, or the replace_provider_slots fallback — the router also publishes a provider event on the bus: source "posoco_ext_llm", topic "provider", payload {"provider_id": "<active provider id>"}. Provider-scoped status publishers (Codex's quota segments, for example) subscribe to it to show or hide their own segments when the active provider changes; republication is idempotent. Without a bus the publication is a silent no-op.

    The router also publishes registry quota segments (priority 45) when the active provider has a source in the devkit quota registry — DeepSeek's balance, and the Z.ai Coding Plan / Kimi 5h/weekly windows: balance readings render a balance segment with "<value> <currency>", window readings render the window name with "<n>% left" (≤20 warning, ≤0 error, matching codex). Refresh points: after each chat response (the pull never delays the first token), on /status (the command's live pull doubles as the bar refresh), and immediately on a /model//effort switch; switch_slot only clears the previous provider's segments and resets the pull cache. Successful pulls throttle re-reads for 30 minutes and failed pulls back off for 5 minutes, so a bad endpoint never slows chat. Codex deliberately has no registry source — its windows are push-shaped and stay owned by CodexQuotaStatus, while the router clears its own quota segments for it (distinct source; provider scoping keeps the two publishers mutually exclusive).

    Authentication is explicit and provider-neutral. OAuth uses a CredentialStore, an AuthInteraction, an OAuthProvider on the slot, and, when credentials change adapter configuration, a rebuild_on_credential factory. API-key login uses an ApiKeyStore, an AuthPromptInteraction, an ApiKeyFactory, and rebuild_on_api_key. The router persists credentials and replaces every matching slot that declares the corresponding factory. Missing dependencies fail as CommandError::ExecutionFailed; no silent fallback is used. API-key secrets are never included in diagnostic text.

    Provider adapters and catalogs live in separate extensions, one per provider; each catalog advertises only the capabilities that provider supports (for example chat/streaming/reasoning only for a provider without FIM support).

    Providers with authenticated model discovery may additionally implement the optional async RefreshableProviderFactory seam. Hosts invoke it explicitly (for example at startup or immediately after login); normal catalog composition remains a pure snapshot build and /model never performs an implicit network refresh. A refresh failure is typed and observable rather than hidden behind a stale/static fallback.

    ApiKeyFactory

    pub(open) trait ApiKeyFactory {
    fn provider_id(Self) -> String
    fn credential_id(Self) -> String
    async fn login(Self, interaction : &
    AuthPromptInteraction
    ) ->
    ApiKeyCredential
    raise
    OAuthError

    }

    Provider-owned API-key login contribution. The provider prompts and validates its secret; the router only persists it and rebuilds slots.

    OAuthFactory

    pub(open) trait OAuthFactory {
    fn provider_id(Self) -> String
    fn credential_id(Self) -> String
    async fn login(Self, interaction : &
    AuthInteraction
    ) ->
    Credential
    raise
    OAuthError

    }

    Optional OAuth contribution paired with a ProviderFactory. Keeping login as a separate trait means non-OAuth providers do not implement fake login paths or know about host interaction details.

    ProviderFactory

    pub(open) trait ProviderFactory {
    fn provider_id(Self) -> String
    fn credential_id(Self) -> String
    fn build(Self, source : ProviderConfigSource) -> ProviderBuildResult raise
    CompositionError

    }

    Public extension seam for provider-owned configuration and catalog construction. Implementations never expose endpoint or protocol details to the host/router.

    RefreshableProviderFactory

    pub(open) trait RefreshableProviderFactory {
    fn provider_id(Self) -> String
    async fn refresh(Self, source : ProviderConfigSource) -> ProviderBuildResult raise
    CompositionError

    }

    Optional provider-owned model discovery seam.

    A provider implements this trait only when its authenticated endpoint can return a live model catalog. Hosts must call refresh explicitly (for example during first startup or immediately after a successful login); a normal ProviderFactory::build remains a pure snapshot reconstruction and never performs an implicit network request. Transport and schema failures are raised as CompositionError by the provider implementation, so a host cannot accidentally fall back to a stale or static catalog.

    AuthMethod

    pub(all) enum AuthMethod {
    ApiKey
    OAuth
    } derive(Eq,
    Debug
    )

    Provider-neutral authentication method advertised by a model catalog.

    AuthMethod::equal

    fn AuthMethod::equal(AuthMethod, AuthMethod) -> Bool

    AuthMethod::not_equal

    fn AuthMethod::not_equal(x : AuthMethod, y : AuthMethod) -> Bool

    AuthMethod::to_id

    fn AuthMethod::to_id(self : AuthMethod) -> String

    ModelSlot

    pub(all) struct ModelSlot {
    id : String
    label : String
    provider_id : String
    model_id : String
    port : &
    ModelPort

    oauth : &
    OAuthProvider
    ?
    api_key : &ApiKeyFactory?
    thinking_efforts : Array[String]
    context_window : Int?
    capabilities : Array[String]
    default_effort : String?
    display_group : String?
    rebuild_on_credential : (
    Credential
    ) -> ModelSlot?
    rebuild_on_api_key : (
    ApiKeyCredential
    ) -> ModelSlot?
    rebuild_on_effort : (String) -> ModelSlot?
    }

    A selectable model slot.

    The router owns selection state only. Hosts and provider extensions build the ModelPort and inject it here, so the router does not know how a port is configured, which protocol it speaks, or how credentials are obtained.

    ModelSlot::ModelSlot

    fn ModelSlot::ModelSlot(id~ : String, label~ : String, provider_id~ : String, model_id~ : String, port~ : &
    ModelPort
    , oauth? : &
    OAuthProvider
    ?, api_key? : &ApiKeyFactory?, thinking_efforts? : Array[String], context_window? : Int?, capabilities? : Array[String], default_effort? : String?, display_group? : String?, rebuild_on_credential? : (
    Credential
    ) -> ModelSlot?, rebuild_on_api_key? : (
    ApiKeyCredential
    ) -> ModelSlot?, rebuild_on_effort? : (String) -> ModelSlot?) -> ModelSlot

    Construct a selectable model slot.

    ProviderBuildResult

    pub(all) enum ProviderBuildResult {
    Ready(ProviderModelCatalog)
    Unconfigured
    }

    Result of asking a provider extension to resolve a catalog. Missing credentials/settings are a normal unconfigured state; malformed configured values remain composition failures and must not be silently skipped.

    ProviderConfigSource

    pub(all) struct ProviderConfigSource {
    settings : Json?
    provider_credential :
    ProviderCredential
    ?
    credential :
    Credential
    ?
    api_key_credential :
    ApiKeyCredential
    ?
    cached_models : Json?
    credential_store : &
    ProviderCredentialStore
    ?
    }

    ProviderConfigSource::ProviderConfigSource

    ProviderConfigSource::effective_credential

    Resolve the canonical tagged credential for a provider. The legacy fields remain source-compatible for existing hosts, but a host that exposes both records is an observable composition failure: silently choosing OAuth would make an explicit API-key login disappear after recomposition.

    ProviderConfigSource::has_settings

    fn ProviderConfigSource::has_settings(self : ProviderConfigSource) -> Bool

    ProviderConfigSource::setting

    Return one opaque setting without imposing provider-specific names on the host. A missing/non-object source simply has no setting; factories decide which keys are required and raise a typed composition error for malformed values.

    ProviderModelCatalog

    pub(all) struct ProviderModelCatalog {
    provider_id : String
    slots : Array[ModelSlot]
    capabilities : Array[String]
    }

    Provider-owned model catalog.

    A provider extension builds this value after resolving its own endpoint, protocol, credentials, model ids, and capability rules. The router only validates composition and aggregates the resulting slots; it never parses provider configuration.

    ProviderModelCatalog::ProviderModelCatalog

    fn ProviderModelCatalog::ProviderModelCatalog(provider_id~ : String, slots~ : Array[ModelSlot], capabilities? : Array[String]) -> ProviderModelCatalog raise
    CompositionError

    Construct and validate a provider catalog atomically.

    ProviderModelCatalog::auth_methods

    Authentication methods advertised by the provider's slots. This is derived from extension-owned factories so a host receives a generic capability list without inspecting provider configuration.

    ProviderModelCatalog::capabilities

    fn ProviderModelCatalog::capabilities(self : ProviderModelCatalog) -> Array[String]

    Provider capabilities such as chat, streaming, or provider-native operations. Cetas treats these as opaque capability identifiers.

    ProviderModelCatalog::provider_id

    fn ProviderModelCatalog::provider_id(self : ProviderModelCatalog) -> String

    Provider id for diagnostics and UI labels.

    ProviderModelCatalog::slots

    Provider-owned model slots in declaration order.

    RouterModelPort

    pub(all) struct RouterModelPort {
    slots : Map[String, ModelSlot]
    slot_order : Array[String]
    active_id : String
    active_effort : String?
    bus :
    EventBus
    ?
    context_states : Map[String,
    ContextState
    ]
    provider_credential_store : &
    ProviderCredentialStore
    ?
    cred_store : &
    CredentialStore
    ?
    auth_interaction : &
    AuthInteraction
    ?
    api_key_store : &
    ApiKeyStore
    ?
    auth_prompt : &
    AuthPromptInteraction
    ?
    credentials : Map[String,
    Credential
    ]
    api_key_credentials : Map[String,
    ApiKeyCredential
    ]
    extra_quota_sources : Map[String, &
    QuotaSource
    ]
    // private fields
    }

    RouterModelPort::RouterModelPort

    Construct a router from a non-empty, uniquely keyed slot list.

    An empty default id selects the first slot in insertion order. A non-empty default that is absent from slots raises, unless allow_missing_default is set: hosts then keep the persisted selection as a pending active id — requests to it fail visibly until the slot appears (login/refresh) instead of silently moving to another model. Composition failures are raised before the router is returned, so callers never see a partially populated routing table.

    RouterModelPort::active_effort

    fn RouterModelPort::active_effort(self : RouterModelPort) -> String?

    RouterModelPort::active_slot_id

    fn RouterModelPort::active_slot_id(self : RouterModelPort) -> String

    RouterModelPort::current_model_id

    fn RouterModelPort::current_model_id(self : RouterModelPort) -> String

    RouterModelPort::current_provider_id

    fn RouterModelPort::current_provider_id(self : RouterModelPort) -> String

    RouterModelPort::extension_id

    fn RouterModelPort::extension_id(_self : RouterModelPort) -> String

    RouterModelPort::from_catalogs

    Construct a router from provider-owned catalogs. Each provider extension builds its own slots and capability metadata; this function only flattens catalogs and lets new perform the final collision check.

    RouterModelPort::list_auth_capabilities

    fn RouterModelPort::list_auth_capabilities(self : RouterModelPort) -> Json

    List each provider once with its provider-owned auth methods.

    RouterModelPort::list_oauth_provider_ids

    fn RouterModelPort::list_oauth_provider_ids(self : RouterModelPort) -> Array[String]

    Backward-compatible OAuth-only provider id list.

    RouterModelPort::list_slot_ids

    fn RouterModelPort::list_slot_ids(self : RouterModelPort) -> Array[String]

    Return all slot ids in composition order.

    RouterModelPort::list_slots

    fn RouterModelPort::list_slots(self : RouterModelPort) -> Array[ModelSlot]

    Return slots in the same order in which they were composed.

    RouterModelPort::on_bus_event

    fn RouterModelPort::on_bus_event(self : RouterModelPort, event :
    BusEvent
    ) -> Unit

    RouterModelPort::on_event

    RouterModelPort::on_event_at

    RouterModelPort::on_shutdown

    async fn RouterModelPort::on_shutdown(self : RouterModelPort) -> Unit

    RouterModelPort::on_start

    async fn RouterModelPort::on_start(_self : RouterModelPort) -> Unit

    RouterModelPort::provider_config

    RouterModelPort::replace_provider_slots

    fn RouterModelPort::replace_provider_slots(self : RouterModelPort, provider_id : String, slots : Array[ModelSlot]) -> Unit raise
    CompositionError

    Replace every slot of one provider in place, preserving composition order for all other providers. Hosts use this after a live model-discovery refresh: the refreshed catalog's slots replace the provider's previous slots without recomposing the Agent.

    Validation mirrors new (non-empty, unique ids, provider id match) and fails before any mutation, so a rejected replacement leaves the routing table untouched. The active selection survives when its slot id still exists afterwards (the slot object is swapped, selection and effort stay); when a previously present active slot disappears, selection falls back to the provider's first new slot and the effort selection is cleared. A pending default is untouched: if the replacement provides its slot the selection heals in place, otherwise it stays pending.

    Every success path ends with a model-status republication, so the status bar reflects the swapped catalog; the provider event rides only the fallback path, where the active slot context changed.

    RouterModelPort::select_effort

    fn RouterModelPort::select_effort(self : RouterModelPort, slot_id : String, effort : String) -> Unit raise
    CommandError

    Select a provider-advertised reasoning effort on one slot through its rebuild hook. Public so hosts can restore a persisted effort at composition and apply ACP/config-driven effort switches; the /effort command remains the interactive entry point.

    RouterModelPort::switch_slot

    fn RouterModelPort::switch_slot(self : RouterModelPort, slot_id : String) -> Unit raise
    CompositionError

    Select a slot without rebuilding its ModelPort.

    TransientRetryModelPort

    type TransientRetryModelPort

    TransientRetryModelPort::provider_config

    TransientRetryPolicy

    pub(all) struct TransientRetryPolicy {
    max_retries : Int
    initial_delay_ms : Int
    } derive(
    Debug
    )

    Router-level transient-failure retry policy. Applied by RouterModelPort::from_catalogs to every ingested slot port so all providers inherit the same exception-retry semantics (sudden 5xx, mid-stream resets, truncated SSE) without per-provider wrappers. Rate-limit (429) verdicts never retry here — they belong to the rate-limit resume domain.

    TransientRetryPolicy::default

    Mirrors the OpenCode session retry: 3 retries, 2s initial doubling delay.

    transient_retry_port