* fix(elevenlabs): fall back to default base URL when config value is malformed
* fix(elevenlabs): reject malformed and non-http(s) base URL overrides
* fix(elevenlabs): redact configured URL from base URL validation errors
* fix(elevenlabs): preserve realtime WebSocket overrides
* test(elevenlabs): avoid stale realtime test overlap
* style(elevenlabs): format realtime URL assertion
---------
Co-authored-by: Peter Steinberger <steipete@gmail.com>
* fix(mattermost): adopt durable ingress drain at the websocket chokepoint
Posted events processed detached from the websocket receive with only a
5-minute in-memory guard; a crash lost the post and reconnect never replays
it. Raw posted envelopes now journal durably (event_id = post.id per the
upstream Post model, lane per channel_id, one row per post) before handler
scheduling; dispatch runs through the core drain with deferred claims through
debounce, merged-flush fan-out adoption, gated-turn settlement, and 30d/20k
tombstones covering the old 5min/2k guard, which is deleted after its parity
test. post_edited stays excluded and cannot be swallowed by posted
tombstones. Cold-gap limitation stated: Mattermost cannot replay posts missed
while disconnected.
Autoreview blocked by codex sandbox network in the build stage; full manual
review performed (updated one websocket test asserting the pre-adoption
parsed-post contract to the raw-envelope contract). Part of #109657 wave 2.
* style(mattermost): keep ingress monitor type internal
* fix(mattermost): retry then loudly escalate a failed durable append
Landing autoreview caught a real loss path: a durable enqueue failure at the
websocket chokepoint was logged and swallowed — the raw envelope discarded,
the connection kept running against a broken store, and reconnect never
replays, so a transient SQLite failure silently lost the post. The append now
retries with short backoff for transient blips; a persistent failure
propagates and the websocket terminates loudly so the outage is
operator-visible instead of silently dropping every subsequent post.
Regression test covers both the absorbed-transient and escalation paths.
* style(mattermost): format rebased ingress handler
* docs(mattermost): document bounded auth-failure retries under deferred claims
* fix(mattermost): guard the drain pump against stop racing the async prune
stop() disposing before the startup pump finished pruning let the pump
lazily create a fresh undisposed drain and dispatch after shutdown. The pump
now re-checks running after the prune, and stop() disposes again after
awaiting the pump so a drain created mid-race is torn down. Regression test
blocks the prune across stop and asserts no dispatch.
* fix(mattermost): serialize durable admissions to preserve lane order
Concurrent websocket callbacks let a post in append-retry backoff be
overtaken by its successor, inverting same-channel arrival order in the
queue. Admissions now chain (order over latency, mirroring the iMessage
admission tail); regression proves a retried post still lands ahead of a
concurrently received one.
* test(mattermost): assert lane order via dispatch sequence
* fix(mattermost): stop() awaits in-flight admissions before disposal
* fix(mattermost): satisfy ingress lint checks
* fix(mattermost): honor envelope-level channel ids in the durable inspector
Posts can carry their channel id on the post, the event data, or the
broadcast envelope — the monitor dispatch honors all three, but the ingress
inspector and claim-side validator required the nested field, rejecting valid
posts as permanent and (via the storage-failure escalation) tearing down the
socket for a failure that never happened, losing posts reconnect cannot
replay. Both sites accept the three shapes; regression proves an
envelope-level post dispatches.
* fix(anthropic): add claude-fable-5 to CLI allowlist, aliases, labels, and context window
CLI path was missing claude-fable-5 across all four metadata layers while
the direct Anthropic provider path in register.runtime.ts already had full
support via isAnthropicFable5Model() / resolveAnthropicFixedContextWindow().
- CLAUDE_CLI_DEFAULT_ALLOWLIST_REFS: add claude-cli/claude-fable-5
- CLAUDE_CLI_MODEL_ALIASES: add fable/fable-5/claude-fable-5 mappings
- CLAUDE_CLI_CONTEXT_WINDOWS: add claude-fable-5 -> 1_000_000 (was falling back to 200K)
- CLAUDE_CLI_MODEL_LABELS: add Claude Fable 5 (Claude CLI)
- resolveClaudeCliImageMediaInput: add fable-5 to 2576 max-side tier
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(anthropic): route fable alias through context/model-ref canonicalization paths
* fix(agent): include claude-cli in fable-5/mythos-5 fixed context window resolution
* fix(anthropic): add claude-fable-5 to static CLI manifest catalog
* test(anthropic): cover fable-5 alias canonicalization in bare and provider-qualified forms
* style(anthropic): fix indentation in plugin manifest
* fix(anthropic): publish Fable CLI output limit
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Peter Steinberger <steipete@gmail.com>
* refactor(whatsapp): replay durable inbound through the shared drain
Accept-before-processing stays; the private startup replay loop and inline
complete/release bookkeeping move to the core drain. Drain-owned entries keep
their durable id so downstream failures release for replay, dedupe-claim
contention stays retryable instead of tombstoning, and undeliverable claims
still send configured read receipts. Rows tombstone at dispatch return
(pre-migration timing), not at adoption (#108656).
* fix(whatsapp): finish durable ingress drain adoption with serialized conversation lanes
Core durable ingress intentionally serializes claims by remote JID. Same-conversation bursts now merge as queued followups in the core reply lane after adoption; channel debouncing remains for control and retry paths, with merged fan-out and gated settlement.
Built on the durable drain branch base by @obviyus.
* fix(whatsapp): fail closed on persistence failure and guard prepared-inbound races
Landing autoreview caught two holes: the persistence-failure fallback
dispatched live, bypassing drain dedupe and lane serialization now that the
replay guard is deleted (duplicate replies / session races); and duplicate
pending rows could orphan or clobber preparedInboundByDurableId entries. The
append now retries transient failures with short backoff then drops loudly
(fail closed — the fallback traded a correctness race for availability
against an already-broken store); redeliveries neither clobber the first
delivery's in-flight preparation nor keep orphan entries. The old
fallback-contract test is repurposed to prove retry-then-durable-delivery.
* fix(whatsapp): dispose the durable drain when the close drain times out
The close timeout abandoned the graceful drain path before its finally could
dispose, so a successor monitor could pump the same account queue against a
still-live drain. The timeout wrapper now owns disposal; dispose is
idempotent so the graceful path's cleanup is unaffected.
* style(whatsapp): keep ingress internals unexported and drop a dead harness helper
* refactor(whatsapp): move durable payload serialization to its own module
The dead-export gate refused serializer/error exports whose only external
consumers were tests. The payload contract (Long-timestamp-preserving
serialize/deserialize + the permanent-error class) now lives in
durable-payload.ts with durable-receive as its production consumer; the dead
WhatsAppRetryableInboundError class is deleted (any non-permanent error is
retryable to the drain classifier, so test-support throws a plain Error).
* fix(whatsapp): guard prepared-map ownership and bound its growth
A duplicate pending delivery's finishPreparation deleted the first
delivery's kept entry (ownership guard now requires the deleter to be the
installer), and queue pruning could evict pending rows whose prepared-map
entries then lingered forever on a blocked lane (map now evicts oldest-first
at a bound above the queue's pending cap; dispatch already re-normalizes the
journaled payload when its entry is gone). Regression covers the duplicate
case.
* style(whatsapp): prune split leftovers and satisfy consistent-return lint
* style(whatsapp): satisfy type-aware lint across the drain monitor
Closed verdict shape for the retry result (queue enqueue's metadata generic
leaks any into the union), Promise.resolve-wrapped abandonment aggregators,
braced timeout executor, and the drain merged to a single const (lazy
closures above only run post-start, so const-after-use is TDZ-safe).
---------
Co-authored-by: Ayaan Zaidi <hi@obviy.us>
* fix(imessage): adopt durable ingress drain with cursor advanced after append
The chat.db watcher advanced its ROWID cursor after yielding rows to an
in-memory pipeline; a crash between read and dispatch skipped those messages
forever. Raw rows now journal durably (event_id = message GUID, lane per
chat) at the monitor chokepoint, and the cursor advances only after the
durable append — mirroring Telegram's offset-after-spool contract. Dispatch
runs through the core drain with deferred claims, merged fan-out adoption,
and gated settlement; legacy catchup enters the same raw path. The
transitional GUID replay guard is deleted per the layering contract with
4h/10k tombstone parity; the permanent age fence and recovery caps are
untouched. Twin verdict: tapbacks are distinct rows with their own
ROWID/GUID; fresh-GUID stale backlog remains age-fence territory (upstream
imsg source cited in PR).
Autoreview blocked by codex sandbox network in the build stage; full manual
review performed. Part of the #109657 fleet adoption program (wave 2).
* fix(imessage): exempt operator-requested catchup rows from the live age fence
Landing review caught real message loss: legacy catchup accepts rows up to
its configured maxAge (120min default), but the unified durable path ran them
through the 15-minute live Push-flush fence — suppressing AND tombstoning
any catchup row older than 15 minutes, permanently losing history the
operator explicitly asked to replay. Catchup rows now carry provenance
through the journal payload and skip the live fence; the catchup query's own
maxAge window is their age gate. Regression test proves provenance survives
the durable round trip.
* test(imessage): resolve ingress rebase mocks
* docs(imessage): document the fence-suppression vs catchup overlap tradeoff
* fix(imessage): satisfy durable ingress lint gates
* fix(msteams): adopt durable ingress drain with ack gated on activity append
Bot Framework webhooks acked before detached processing with no platform
retry, silently losing inbound activities on a crash. Dispatchable activities
(message + adaptiveCard/action invokes) now persist their raw JSON durably
(event_id = activity.id per the @microsoft/teams.api uniqueness contract,
lane = conversation.id) before the webhook 200; dispatch runs through the
core drain with deferred claims through debounce, merged-flush fan-out
adoption, gated-turn settlement, retry/dead-letter classification, and
30d/20k tombstones. Restart replay reconstructs proactive routing context.
No inbound replay guard existed; the outbound echo cache is untouched. Twin
check: messageUpdate/votes/consent/reactions run dedicated non-agent
workflows and are excluded from the journal.
Autoreview blocked by codex sandbox network in the build stage; full manual
review performed. Part of the #109657 fleet adoption program (wave 2).
* fix(msteams): close live-context races and double next() in drain dispatch
Three defects found in landing review: the live turn context could swap a
duplicate delivery's mutated activity in place of the journaled payload
(dispatch now always uses the journaled activity, live context is transport
surface only); the context registry installed entries after the durable
append, leaking them when the drain consumed the claim first (install now
precedes enqueue, tombstoned duplicates clean up, first delivery's context
wins); and a failed next() in the message handler was invoked a second time
from the catch fall-through (ran-flag guards it).
* fix(msteams): uninstall a failed append's live context before retry
Follow-up to the install-before-enqueue race fix: an enqueue rejection
bypassed cleanup, so a retry with the same activity id dispatched the failed
request's stale context. Uninstall is identity-guarded (only our context is
removed, never a concurrent redelivery's fresh install) and covers both the
throw and tombstoned-duplicate paths. Regression proves the retry's own
context dispatches.
* test(msteams): track the buffered dispatcher seam after main retired the settled variant
* test(msteams): resolve dispatch union before catch for promise lint
* test(msteams): explicit promise-returning mock for the promise-misuse lint
* test(msteams): type the ingress accept mock promise-returning
ReturnType<typeof vi.fn> types the implementation callback void-returning, so
every promise-returning accept mock tripped typescript(no-misused-promises)
in CI regardless of call-site shape.
* test(msteams): restore compact async mock now that its type permits promises
* test(msteams): extract gated-accept helper to stay under max-lines
* test(imessage): mock core turn dispatch for media policy coverage
PR #110095 missed this iMessage media-policy test file.
* test(imessage): mock core turn dispatch for last-route coverage
* test(whatsapp): mock core turn dispatch for broadcast coverage
* test(imessage): drop type-dead terminal-admission branch in media harness
* fix(qa-lab): close transport before gateway teardown
* test(qa-lab): bind teardown order to production plan
---------
Co-authored-by: Peter Steinberger <steipete@gmail.com>
* fix(whatsapp): resolve abortPromise when stop signal already fired
waitForClose() races on abortPromise, but a pre-aborted gateway stop
signal never settled that promise because the abort listener was skipped.
Resolve abortPromise immediately when abortSignal.aborted is already true.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(whatsapp): resolve tsgo TS18048 and test mock in pre-aborted abortPromise fix
- Extract params.abortSignal to local const so tsgo can narrow the type
through else-if control flow, fixing TS18048: possibly undefined
- Remove setupAbortController/ownerAcquireAbortController.abort() from
the constructor's pre-aborted branch so openConnection() still works
when the stop signal was already fired at construction time
- Use createSocketWithTransportEmitter() mock in the new test so
shutdown() cleanup has sock.end() and ws.removeListener() available
* fix(whatsapp): preserve pre-aborted setup cancellation
* fix(whatsapp): unify connection abort handling
---------
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Peter Steinberger <steipete@gmail.com>
* fix(discord): adopt durable ingress drain for gateway messages
Gateway MESSAGE_CREATE events dispatched with only an in-memory/persistent
5-minute dedupe; a crash between receive and dispatch lost the message (live
sessions replay via RESUME, but session state is memory-only so a cold
restart re-IDENTIFYs — stated platform limitation: the drain closes the
receive->dispatch window, not cold downtime). Raw API messages now journal
durably (event_id = message snowflake, lane per channel/thread) before
normalization; dispatch runs through the core drain with deferred claims,
fan-out adoption, gated settlement, and 30d/5k tombstones. The persistent
guard is deleted after its RESUME duplicate-delivery parity test. Edits stay
outside the queue: only MessageCreateListener is registered, MESSAGE_UPDATE
never dispatches, so snowflake tombstones cannot swallow edit turns; a future
edit dispatcher must namespace its event ids.
Autoreview blocked by codex sandbox network in the build stage; full manual
review performed. Part of the #109657 fleet adoption program (wave 2).
* style(discord): keep drain-internal types and dispatcher unexported
* refactor(discord): move message dispatcher to its own module for the dead-export gate
* fix(discord): shutdown dispatch releases claims instead of tombstoning
Landing autoreview reproduced a real loss race: a dispatch entering after
dispatcher shutdown (or with an aborted signal) returned completed, so the
drain tombstoned a message that never ran and a restarted gateway skipped it
forever. Aborted-before-dispatch now returns failed-retryable so the claim
releases for replay. Test type imports of drain-internal types converted to
factory-derived types so the dead-export gate and typecheck agree.
* fix(msteams): bound probe token acquisition to request deadline
probeMSTeams() at extensions/msteams/src/probe.ts:75 and :89 awaited
tokenProvider.getAccessToken(...) for the Bot Framework and Microsoft
Graph token endpoints with no surrounding deadline. The Microsoft
Teams SDK does not carry an inherent timeout on these calls, so a
stalled Azure AD token endpoint pinned the probe indefinitely.
Wrap both awaits with withMSTeamsRequestDeadline (default
MSTEAMS_REQUEST_TIMEOUT_MS = 30_000), matching the pattern already
used by six other MS Teams call sites: attachments/bot-framework.ts:252,
attachments/graph.ts:258, monitor-handler/message-handler.ts:594/654/685/692,
attachments/download.ts:167, team-identity.ts:37.
The probe was the one missing site. No new helper, no SDK change.
The existing outer catch at probe.ts:138 and inner catch at probe.ts:110
convert the timeout into a ProbeMSTeamsResult with ok: false and a
structured error field.
Added probe.timeout.test.ts: real probeMSTeams() with vi.mock
injected never-resolving getBotToken/getGraphToken; asserts the call
returns within the 30s bound instead of hanging to the proof budget.
* test(msteams): drive probe timeout test with vi.useFakeTimers
The original probe.timeout.test.ts waited 90 seconds of wall-clock per
focused run (3 stalled cases racing against a real setTimeout budget).
Per ClawSweeper P2 (automation), this material deterministic CI cost can
slow or time out test shards.
Drive the withTimeout race (from @openclaw/fs-safe/dist/timing.js, uses
setTimeout + clearTimeout) via vi.useFakeTimers() so each stalled case
resolves in milliseconds. Add one new case that spies on withTimeout's
timeoutMs argument to assert the production default deadline is exactly
MSTEAMS_REQUEST_TIMEOUT_MS = 30_000, so the production contract is not
silently weakened by the fake-timer change.
Per-case wall-clock: 25ms / 3ms / 2ms / 1ms / 2ms (was: 30s / 30s / 30s /
2ms / n/a).
Co-Authored-By: Claude <noreply@anthropic.com>
* fix(msteams): bound remaining token acquisition
* test(msteams): keep credential fixture unchanged
---------
Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: Peter Steinberger <steipete@gmail.com>
* fix(test): converge Slack harness state across module reloads
Five reaction tests in monitor.tool-result.test.ts failed whenever a sibling
file's vi.resetModules() ran earlier in the same non-isolated worker: the
cached globalThis __slackClient kept routing reactions to the OLD test-helpers
module's mocks while tests asserted on the recreated module's fresh mocks —
reactions 'never fired' (replies survived via freshly injected runtime).
slackTestState is now a globalThis-backed singleton so every module
incarnation and the cached client share one state object. Also: reaction
assertions wait on the mock (reactions apply via detached debounce/queue
work), and stale Bolt handler registrations clear per test/stop so
waitForSlackEvent cannot match a previous provider's handler. Proven with six
consecutive full extensions/slack runs (was ~50% failure).
* style(slack): bracket global test-state access for underscore lint
* fix(perplexity): send Search API date filters with official field names
* test(perplexity): cover date filter request body
* test(perplexity): simplify request capture
* style(perplexity): format request test
---------
Co-authored-by: Peter Steinberger <steipete@gmail.com>
* fix(discord): avoid Activity OAuth stalls after API rejection
* test(discord): preserve status on cancel failure
---------
Co-authored-by: Peter Steinberger <steipete@gmail.com>
* fix(discord): honor caller abortSignal during 429 retry backoff
* test(discord): prove 429 backoff abort through a real loopback server
* test(discord): make retry abort proof deterministic
---------
Co-authored-by: Peter Steinberger <steipete@gmail.com>