03: Chunk Normalization
Converts a provider-native RawStreamChunk into the uniform ProcessedChunk shape the orchestrator routes.
Contract: BaseStreamAdapter.processChunk — src/types/stream/interfaces.ts:247
Canonical implementation: GoogleStreamAdapter.processChunk — src/providers/google/googleStreamAdapter.ts:649-760
Mission
Section titled “Mission”Each RawStreamChunk yielded by stage 02 is immediately passed to processChunk() inside the
stage 04 orchestrator loop. This method is the boundary at which all provider-specific chunk
formats collapse into the single ProcessedChunk union that the orchestrator understands:
interface ProcessedChunk { type: "text" | "function_call" | "error" | "done"; content?: string; // type="text" functionCall?: FunctionCall; // type="function_call" error?: ProviderError; // type="error" thoughts?: ThoughtLogEntry[]; // any type — thought log capture metadata?: Record<string, unknown>; // terminal metadata (finish reason, thought signature)}The method also handles two additional responsibilities:
-
Error normalisation — raw SDK errors (HTTP status codes, provider-specific error objects) are converted to the shared
ProviderErrorshape viahandleProviderError(). This includes classifying the error type (api_error,rate_limit,content_blocked,timeout,provider_overloaded,model_error) and settingretryableso the stage 04 orchestrator and the upstream key-rotation logic inrunGenerationTurncan make retry decisions without inspecting provider-specific error objects.model_erroris reserved for model-selection or model-availability failures such as “unsupported model” or “model not found”; the stream UI surfaces the provider’s raw supported-model details instead of hiding them behind a generic 400. -
Thought log extraction — for providers that emit reasoning fields, thought summaries, or thought signatures (for example Google/Gemini
part.thought,thoughtSummary, andthoughtSignaturefields), these are extracted intoThoughtLogEntry[]on the returned chunk so the orchestrator can accumulate them intostate.thoughtSummarySegments/state.thoughtRawSegmentsindependently of visible text. For OpenRouter, the upstream serving backend (the chunk-levelproviderfield, e.g.minimax-cn) is also carried onProcessedChunk.servingProvider, recorded intostate.servingProvider(first non-empty wins), and rendered in the thought-log footer asProvider: openrouter via <backend>— so a backend that bleeds reasoning into content can be identified and pinned/avoided. -
Leaked reasoning-tag guardrails — the clean path is the provider (or OpenRouter) returning reasoning in a dedicated field. When a backend instead leaks reasoning into
delta.content, OpenAI-compatible adapters (openrouter,openaiCompatible) run two worst-case guards over the text delta:ThinkBlockContentStripper(reroutes<think>…</think>blocks — including stray closers split across chunks — intodelta.reasoning) andReasoningContentSpillGuard(catches a tagless reasoning tail glued to the first visible delta — e.g.must do.Hello!). The spill guard only fires on the first visible content after reasoning, when that content starts lowercase and a sentence boundary is glued (no following whitespace). It strips when EITHER the text after the boundary looks like an answer start (uppercase / caseless letter, emoji, or quote/bracket — catcheswait.Actually) OR the fragment before the boundary is a multi-word clause (catches a casual lowercase reply glued onto a reasoning tail, e.g.g it out.hey master 👋, where capitalization is blind because the real reply is also lowercase). A spaced boundary is treated as the model’s own prose and emitted untouched; the missing space is the fingerprint of a backend concatenating its reasoning buffer onto content with no separator (sowait. Actuallyis two real sentences, butwait.Actuallyis a seam). A single-word fragment glued to a lowercase continuation is left alone. Dotted code identifiers are protected by an inline-code guard: a boundary preceded by an odd number of backticks (inside a`obj.Method`span, or a still-open one mid-stream) is left intact, as is code at the very start of content (a leading backtick is not lowercase). Bare, non-backticked identifiers in prose are accepted collateral of the aggressiveness. The shape of a think tag is defined once insrc/providers/utils/reasoningTags.tsand is namespace-aware (<think>,<mm:think>,<ns:think>); the stripper, the Discord-layerbufferManager, and the finalcleanLLMOutputsweep all consume it. Adding a new vendor namespace is a one-line change there. Note: API stop strings (stopStrings.ts) are matched literally by the provider, so namespaced close tags must be added per model rule explicitly rather than via the shared pattern. -
Custom verbatim tool-call fallback — when
server_capabilities_configs.verbatim_tool_calling_enabledis true, the active Custom text model has tools, and the request includes OpenAI-compatible tool schemas,CustomStreamAdapterrunsVerbatimToolCallParserover visibledelta.contentafter existing Custom/Gemma cleanup. It scans the stream for an anchor<knownToolName>(— only names from the exposed tool set trigger — then accumulates from that name until the parentheses balance (quote-aware, so a)inside a JSON string does not close early) and parses thename(...)body. The call may be bare or wrapped in an inline code span / fenced block, and prose before it is allowed (chat models narrate before they act): leading narration is emitted as normal text and the call is recovered after it, e.g.Fine. `generate_image({"prompt":"a cat","mode":"txt2img"})`or the same call with no backticks. The parse step is the false-positive guard — a tool name merely mentioned in prose (generate_image (it makes art)) fails JSON/arity validation and is released as text. Successful parses are emitted astype: "function_call"before the literal text reaches Discord; rejected or incomplete text is released normally. Because the stream adapter dropsvisibleTextwhenever afunctionCallis present, the parser emits any preceding prose first (call unresolved) and resolves the call on a later chunk or at stream flush.
chunk: RawStreamChunk — the provider-native envelope yielded by stage 02.
Output
Section titled “Output”ProcessedChunk — one of four variants:
type |
Carries | Orchestrator action |
|---|---|---|
"text" |
content: string |
Route to stage 05 buffer flusher |
"function_call" |
functionCall: FunctionCall |
Flush pending buffer → return { status: "function_call" } |
"error" |
error: ProviderError |
Flush pending buffer → show error embed → return { status: "error" } |
"done" |
metadata?: { finishReason } |
Record terminal metadata; continue loop (generator will exhaust next) |
Any variant may additionally carry thoughts?: ThoughtLogEntry[] when the provider emits
reasoning content in-band.
Side effects
Section titled “Side effects”- None.
processChunkis a pure transformation — it does not mutateStreamState, call Discord APIs, or trigger any timer. All side effects are owned by the orchestrator and downstream stages.
Invariants
Section titled “Invariants”After this stage:
- The returned chunk’s
typeis one of exactly"text","function_call","error","done". - If
type === "error",chunk.erroris a fully-formedProviderErrorwithtype,message,retryable, andcodeset.originalErrorpreserves the raw SDK error for logging. - If
type === "function_call",chunk.functionCallis a provider-agnosticFunctionCall{ name, args, thoughtSignature? }. - Content-blocked responses (e.g., Gemini
promptFeedback.blockReason, safetyfinishReason) are normalised totype: "error"witherror.type === "content_blocked"andretryable: false.
Message-based classifiers
Section titled “Message-based classifiers”HTTP status codes are too coarse to pick the right advice, so utils/provider/providerErrorClassification.ts
re-classifies a normalized ProviderError by matching its message text (and originalError payload).
errorUi.ts consults these before falling back to ProviderError.type:
| Classifier | Recognizes | Why it is separate |
|---|---|---|
isProviderModelError |
unsupported / unknown / deprecated model IDs | Steer to model selection, not to the key |
isAccountBalanceExhaustedError |
zero spendable balance (DeepSeek 402 Insufficient Balance, Z.ai no resource package) |
No request of any size can succeed, so trimming tips are useless |
isCreditAffordabilityError |
affordability ceiling (OpenRouter 402 can only afford N) |
A smaller max_tokens still fits the remaining credit |
isContextLengthError |
hard context-window overflow | Trimming history genuinely resolves it |
The two credit classifiers must stay mutually exclusive. They map to opposite advice: an exhausted
balance needs a top-up, an affordability ceiling needs fewer output tokens. Ordering in
resolveProviderErrorPresentation() runs exhausted-balance first because it is the stricter
condition. Both are checked before the generic api_error default, whose verify_api_key tip
misdiagnoses a billing failure as a bad credential.
A provider that reports billing denial under a non-402 status (Z.ai uses 429) is re-coded in its error formatter so it does not present as a rate limit.
Credential-scoped recovery tips
Section titled “Credential-scoped recovery tips”Which classifier fires decides what to advise; StreamContext.textCredentialSource decides
which commands the advice names. It is the resolved credential policy of the turn
("server" | "personal"), threaded from ChatTurn through StreamingContext and
buildStreamContext(). Scope is never inferred from the provider or model name, because both
scopes can run the same provider with the same model.
resolveProviderErrorPresentation() resolves every command-bearing tip through a scoped()
helper that appends _personal to the locale key when the failing request used a personal key,
so genai.tips.model_fallback becomes genai.tips.model_fallback_personal. Two rules fall out
of that and are easy to regress:
genai.tips.api_key_rotationis suppressed entirely on personal failures. Rotation pools are a server-scoped, manager-only feature, so the tip would point a normal member at a command they cannot run and that could not repair their own key anyway.modelFallbackTipgates on!isPersonal. A personal text override reads the user’s fallback chain, sotomoriState.fallback_chainsays nothing about whether that request had a fallback available.
Personal text failures additionally get genai.tips.disable_personal_text_override, which offers
the fastest recovery (hand the turn back to the server default) and carries the User BYOK caveat,
since a server in that mode rejects server fallback and will still require a personal provider.
Non-chat provider calls may leave textCredentialSource unset; the tips then resolve to their
server-scoped form, which is the correct default for configuration a member does not own.
Extension points
Section titled “Extension points”| Surface | Plugin-relevance |
|---|---|
BaseStreamAdapter.processChunk() abstract method |
A new provider adapter implements this to map its SDK chunk shapes to ProcessedChunk. The contract is at src/types/stream/interfaces.ts:184. The implementation must be synchronous. |
BaseStreamAdapter.handleProviderError() abstract method |
A new provider adapter implements this to classify its SDK errors. The ProviderError.retryable flag is consumed by the key-rotation loop in runGenerationTurn; the type field drives user-facing error embed formatting. Model-name and model-availability failures should become model_error or carry a message that the shared model-error classifier can recognize. Contract at src/types/stream/interfaces.ts:201. |
BaseStreamAdapter.createErrorDescription() abstract method |
A new provider adapter implements this to produce localized, provider-specific error text for the error embed shown in Discord when retryable: false and user errors are not suppressed. For model_error, preserve the provider’s actionable details, such as supported model IDs. Contract at src/types/stream/interfaces.ts:207. |
FunctionCall shape (name, args, thoughtSignature) |
The provider-agnostic function call format — src/types/provider/interfaces.ts:145. Fields like thoughtSignature, reasoning_details, and deepseekReasoningContent are provider-specific optional fields that must be preserved when passing tool results back to the provider in stage 01. |
Related docs
Section titled “Related docs”- Stage 02 (produces
RawStreamChunkconsumed here): →02-raw-chunk-generation.md - Stage 04 (routes the
ProcessedChunkproduced here): →04-orchestrator-state-machine.md ProcessedChunktype:src/types/stream/interfaces.ts:36ProviderErrortype:src/types/stream/interfaces.ts:47FunctionCalltype:src/types/provider/interfaces.ts:145- Error embed formatting:
src/utils/discord/stream/errorUi.ts