* Strip liveness gate from AI Vault session delete
Delete now requires only path validation + user confirmation — no process
roster, no liveness check, no quiescence, no ownership ledger.
Co-authored-by: Orca <help@stably.ai>
* Remove obsolete AI Vault liveness delete reliability gate
Session delete no longer checks process liveness, so drop the
manifest entry that still referenced the deleted test files.
* minor fix
---------
Co-authored-by: Orca <help@stably.ai>
Prevent older paired hosts from discarding the entire taskResumeState
when they encounter the strict linearIssueView field, which was causing
silent loss of github and jira queries during remote pairing. Layout,
grouping, ordering, and per-workspace filters are now device-local.
* fix(terminal): show a preedit the IME resumes without a compositionstart
Typing 2-Set Korean shows committed syllables but not the in-progress jamo, so
the user composes each syllable blind. Long-standing hole in the vendored
terminal library, not a regression: the same test fails identically against the
bundle this branch starts from.
The `.active` class that CSS keys `display: block` off is added only in
`compositionstart` and dropped in `_finalizeComposition`. Some IMEs (observed on
Windows/WSL Korean) resume a composition with a bare `compositionupdate` and no
second `compositionstart`, by which point `compositionend` has already hidden the
overlay, so the resumed preedit is written into a hidden element and never
positioned. `updateCompositionElements` also early-returned on `!_isComposing`,
so it would not lay the overlay out either.
Re-show the overlay on an update that carries data, and key the layout guard on
the shown overlay instead. `_isComposing` is deliberately left alone, so no
commit bookkeeping changes and `onData` stays byte-identical. The two guards are
equivalent on every pre-existing path: `compositionstart` sets both,
`_finalizeComposition` clears both.
The bundle hunks are the same two edits applied to the shipped minified output;
the sourcemaps are carried through unchanged.
* test(terminal): prove the resumed-preedit fix against a recorded Windows capture
The synthetic test pins the shape; this replays events a real Microsoft Korean
IME emitted on Windows/WSL. The capture holds three compositionupdates that
resume a composition with no second compositionstart — the exact ordering that
wrote the preedit into a hidden overlay.
Without the fix all three report shown:false; with it all three are visible.
Fixture derived from the sealed 11919-windows-wsl-current capture, which is
read-only and unmodified.
Co-authored-by: Orca <help@stably.ai>
* test(terminal): stop the recorded Hangul fixture pinning a derivation artifact
The capture logs each event twice — a dispatch record and a batched next-frame
re-log. Deriving from both replayed every event twice, which made three
compositionupdates appear to land after a session had ended. Filtered to
dispatch records the capture holds zero resumes and 11 balanced sessions, so
the previous toHaveLength(3) was pinning an artifact of the derivation.
Re-scoped to what the capture does prove: the preedit stays visible across all
37 real updates. Verified by reverting the patch that this passes either way,
so it is coverage and the synthetic test remains the discriminator. Both facts
are now stated in the file.
Co-authored-by: Orca <help@stably.ai>
* fix(terminal): restore the preedit visibility patch onto its own branch
The previous commit accidentally reverted it: checking main's patch and lockfile
into the worktree to test whether a test discriminates also stages them, so the
commit that followed swept them up.
Co-authored-by: Orca <help@stably.ai>
* fix(terminal): claim printable keydowns structurally so committed text survives
Co-authored-by: Orca <help@stably.ai>
* chore(reliability-gates): retarget the IME forwarding gate after the allowlist removal
The gate listed terminal-ime-input-source.test.ts, which went with the
input-source allowlist. Points at the substituted-text commit test instead,
which covers what the gate is actually protecting: text committed outside a
composition session reaching the pty exactly once.
Co-authored-by: Orca <help@stably.ai>
* docs(terminal): record why withholding a claimed keydown needs no timer
The predicate withholds a keydown's byte until the commit arrives, so a key the
IME eats without committing would be dropped. Measured across the recorded
corpus that case does not occur, and the browser marks IME-owned presses on the
keydown itself. Both facts belong next to the predicate rather than only in a
handoff note, since the obvious fix for the imagined gap is a timer, and a timer
here once wrote a newline the user never typed.
Co-authored-by: Orca <help@stably.ai>
* test(terminal): pin the kitty all-keys-as-escape-codes hole explicitly
Flag 8 asks for every printable key as an escape code; this path sends the
committed text raw instead. That is a deliberate trade, not an oversight, but it
was untested — the suite only covered the disambiguate flag. Pinning it makes
the choice visible and records the gate to use if it ever needs closing.
Co-authored-by: Orca <help@stably.ai>
* fix(terminal): keep the kitty key-release report for presses that reached the pty
Claiming the keyup unconditionally suppressed xterm's release report. That was
sized for the old design, which claimed only a short punctuation list; the
structural claim takes every printable keydown, so on macOS an app that
negotiated kitty report_event_types stopped seeing releases for ordinary typing
and would treat every printable key as held down.
Suppress the release only when the press put nothing on the wire — swallowed by
the input source, or owned by a composition transaction. xterm emits nothing
from keyup unless kitty report_event_types (or win32 input mode) is on, so
letting it through is inert everywhere else.
Co-authored-by: Orca <help@stably.ai>
* test(e2e): assert IME preedit geometry headlessly for Korean and CJK input
Both IME defects that shipped and were reverted walked through a suite of ~3000
passing assertions, because every one of them was about bytes reaching the PTY.
A preedit rendered into a hidden overlay satisfies all of them while the user
composes blind. The real-geometry coverage that would have caught it existed but
was headful, env-gated and macOS-only, so it never ran in CI.
Drives composition through CDP Input.imeSetComposition instead of a native input
source, which removes the accessibility grant and the system input source that
forced the headful gate. The suite runs in the normal headless project in about
55s serially, and asserts the composition overlay's real bounding rect — the one
property an overlay clipped to max-width:0 cannot fake and a DOM emulator cannot
produce.
Three tests are red on main and marked test.fail() so they stay visible in CI and
flip loud when their fix lands: the preedit resumed by a bare compositionupdate,
and full-width punctuation and digits committed from a keydown that still carries
the ASCII layout key.
Co-authored-by: Orca <help@stably.ai>
* test(e2e): drop the known-broken markers now the stack closes all three
Validated on real macOS hardware: with the two fixes below this layer, all
three report "Expected to fail, but passed". Korean preedit renders at
non-zero geometry through every jamo, and an Apple pinyin source sends
ef bc 8c e3 80 82 to the pty where main sends ASCII.
Worth recording why the punctuation case looked green on main once: an input
source whose id happens to contain an allowlist term, as Sogou's does, satisfies
the old gate. Correctness there depended on which IME the user had selected.
Co-authored-by: Orca <help@stably.ai>
* test(e2e): cover the Linux and Windows IME ownership shapes headlessly
The ten headless IME specs on this branch all decided ownership through a macOS
user-agent override, so the two platforms whose failure mode is a *dropped*
character rather than a downgraded one had no coverage at all, and the one
Windows-recorded trace already in the suite was replayed under whichever policy
the runner happened to report — macOS locally, Linux on the CI shards.
Adds four Linux specs and two Windows specs, every one of them replaying a
native capture rather than a hand-authored ordering:
- IBus/X11 Hangul mixed with literal ASCII. Its `compositionend` is EMPTY and
the syllable arrives afterwards as a bare `insertText`, so reading the commit
off `compositionend.data` — which the Windows capture rewards — drops every
syllable on this framework.
- fcitx5/Wayland Hangul. No keydown at all for a composing key, not even 229,
and physically wrong `code` values on the literal keys. Any ownership rule
reading 229 or `code` fails here.
- Numeric pinyin candidate selection under both frameworks, with the ordinary
digit kept as the negative control, so the two directions are pinned against
each other rather than separately.
- Windows Microsoft Korean captured with real scan codes, including the two
lines committed with Shift held.
Each asserts both sides of the boundary: the preedit's real geometry at every
frame the user would see, and the exact byte stream the native run put on the
PTY. The recorded `onData` the Windows/WSL fixture already carried is now
asserted instead of sitting unused.
Chinese moves up to first-class alongside Korean: pinyin preedit width is now
pinned the way the Japanese phrase already was, and full-width punctuation is
covered in the composition-session shape the Windows and Linux frameworks use,
not only the macOS insertText shape.
Two harness fixes fell out of the recorded traces and are why the IBus one
passes. The replay applied each event's recorded textarea state *after*
dispatch, one event too late for the handlers that read `textarea.value`; and
it left a task boundary between `compositionend` and the `input` carrying the
commit, which Chromium never inserts, letting xterm's deferred finalizer settle
against a textarea the committed text had not reached yet.
Co-authored-by: Orca <help@stably.ai>
* test(e2e): replay the recorded macOS IME shapes instead of only synthesising them
The macOS coverage on this branch drives Chromium composition through CDP, which
is genuine but hand-ordered, so it could not assert the one property that
decides the macOS rule: a composing keydown arrives with keyCode 229 while `key`
is still the single translated character the input source produced — `ㅎ`, not
`Process` — which is indistinguishable by length from an ordinary printable key.
Three native captures were sitting unused in the evidence set.
Adds a recorded 2-Set Korean session, including the syllable boundary where one
composition closes and the next opens with no keydown between them, asserted
against its own recorded byte stream.
Adds the third failure mode, which had no coverage in any shape: an abandoned
preedit leaking to the shell. Pinyin and Cangjie both backspace a composition
away to nothing, and the assertion is not "the right bytes" but "no bytes".
Both cancellation captures continue with a literal `ordinary` typed as bare
keydowns, the recorder's own negative control. That tail carries no `input`
events because the build it was captured on produced the byte from the keydown
itself, so replaying it would measure the recorder rather than the product; the
specs cut at the `compositionend` and say so.
Co-authored-by: Orca <help@stably.ai>
* test(e2e): promote the non-allowlist input-source punctuation spec to the suite
Qingg matches none of the terms the pre-structural build enumerated, so on that
build its bypass never installs. It is the one arm no headless spec can express,
and the only test here that the pre-structural build cannot pass.
Promoted from scratch with four changes, each forced by a measurement rather than
by taste:
- A non-attached input method is a refusal, not a negative result. macOS attaches
per app instance and the attach can simply fail — 3 of 6 instances under
exclusive host access, and re-selecting the source did not recover one of them
across 9 keystrokes. That is now `test.skip()` with a reason naming the rerun,
not a thrown error, and the suite does not gate on a fully green session.
- Attachment is probed with a LETTER. Punctuation substitution emits no
compositionstart and no keyCode 229 even under a fully attached source, so at
the keydown it is indistinguishable from having no source at all. The
punctuation arms are judged on PTY bytes alone.
- The ASCII-layout control is now part of the spec rather than a side experiment.
Without it a build that rewrote every `.` into `。` unconditionally would pass
the Qingg arm and be badly wrong.
- The assertion runs by default instead of behind a strict-mode flag, and gained
a non-vacuity check: the input source must have committed something. That is
the sharp end of the mechanism — on the old build the DOM carries only keydown
and keyup, so nothing is committed at all and the character is destroyed before
the source is asked.
The verdict stays an equality between two measurements, never a comparison
against a hardcoded glyph, so it holds whatever punctuation mode the operator's
input source happens to be in. It reads `beforeinput`, not `input`: the forwarder
consumes `input` in the capture phase on the pane element, so a probe on the
helper textarea never sees it and a strict run fails with correct bytes
underneath.
Co-authored-by: Orca <help@stably.ai>
* feat(terminal): encode IME commits as CSI-u under the all-keys kitty flag
A pane that negotiates `report_all_keys_as_escape_codes` (bit 3) asked for every
printable key as a CSI-u report. The commit path wrote IME-committed text raw,
so such a pane got a legacy byte stream it had declined. That predicate has no
IME-specific condition, so it affected every macOS user in such a pane, not just
CJK users.
Encode the press that produced the commit instead, reusing xterm's own kitty
encoder rather than hand-rolling CSI-u.
`claimKeyEvent` is untouched: still unconditional, still structural, still no
kitty read on the keydown. The flag read happens once per commit.
The gate is bit 3 alone. Flags 1/2/4/16 leave printable keys as text, so panes
negotiating only those keep receiving substituted characters; gating on "kitty
active" would strip the substitution from every pane that negotiates anything.
Known limit, pinned by test: the report carries the physical key's codepoint,
not the committed glyph. Bit 3 is the app declaring it does not want text, and
bit 4 is how it asks for text back — but xterm's encoder derives that text field
from the same `key` it derives the keycode from, so carrying the committed glyph
needs an encoder change, not a wider gate.
* fix(terminal): report a held key's repeats as REPEAT under the kitty flags
The commit encoder never passed an event type, so xterm's encoder applied its
PRESS default to every auto-repeat keydown. A pane negotiating report_event_types
alongside bit 3 saw one held key as N separate strikes.
Carry the keydown's `repeat` on the claimed press and map it to the protocol's
REPEAT. The event type only reaches the wire when report_event_types is
negotiated, so this is inert for panes that asked only for bit 3.
Co-authored-by: Orca <help@stably.ai>
* test(e2e): cover a macOS system key remap reaching the terminal (#11170)
Co-authored-by: Orca <help@stably.ai>
* chore: drop non-mergeable IME e2e scratch files
---------
Co-authored-by: Orca <help@stably.ai>
* fix(terminal): show a preedit the IME resumes without a compositionstart
Typing 2-Set Korean shows committed syllables but not the in-progress jamo, so
the user composes each syllable blind. Long-standing hole in the vendored
terminal library, not a regression: the same test fails identically against the
bundle this branch starts from.
The `.active` class that CSS keys `display: block` off is added only in
`compositionstart` and dropped in `_finalizeComposition`. Some IMEs (observed on
Windows/WSL Korean) resume a composition with a bare `compositionupdate` and no
second `compositionstart`, by which point `compositionend` has already hidden the
overlay, so the resumed preedit is written into a hidden element and never
positioned. `updateCompositionElements` also early-returned on `!_isComposing`,
so it would not lay the overlay out either.
Re-show the overlay on an update that carries data, and key the layout guard on
the shown overlay instead. `_isComposing` is deliberately left alone, so no
commit bookkeeping changes and `onData` stays byte-identical. The two guards are
equivalent on every pre-existing path: `compositionstart` sets both,
`_finalizeComposition` clears both.
The bundle hunks are the same two edits applied to the shipped minified output;
the sourcemaps are carried through unchanged.
* test(terminal): prove the resumed-preedit fix against a recorded Windows capture
The synthetic test pins the shape; this replays events a real Microsoft Korean
IME emitted on Windows/WSL. The capture holds three compositionupdates that
resume a composition with no second compositionstart — the exact ordering that
wrote the preedit into a hidden overlay.
Without the fix all three report shown:false; with it all three are visible.
Fixture derived from the sealed 11919-windows-wsl-current capture, which is
read-only and unmodified.
Co-authored-by: Orca <help@stably.ai>
* test(terminal): stop the recorded Hangul fixture pinning a derivation artifact
The capture logs each event twice — a dispatch record and a batched next-frame
re-log. Deriving from both replayed every event twice, which made three
compositionupdates appear to land after a session had ended. Filtered to
dispatch records the capture holds zero resumes and 11 balanced sessions, so
the previous toHaveLength(3) was pinning an artifact of the derivation.
Re-scoped to what the capture does prove: the preedit stays visible across all
37 real updates. Verified by reverting the patch that this passes either way,
so it is coverage and the synthetic test remains the discriminator. Both facts
are now stated in the file.
Co-authored-by: Orca <help@stably.ai>
* fix(terminal): restore the preedit visibility patch onto its own branch
The previous commit accidentally reverted it: checking main's patch and lockfile
into the worktree to test whether a test discriminates also stages them, so the
commit that followed swept them up.
Co-authored-by: Orca <help@stably.ai>
* fix(terminal): claim printable keydowns structurally so committed text survives
Co-authored-by: Orca <help@stably.ai>
* chore(reliability-gates): retarget the IME forwarding gate after the allowlist removal
The gate listed terminal-ime-input-source.test.ts, which went with the
input-source allowlist. Points at the substituted-text commit test instead,
which covers what the gate is actually protecting: text committed outside a
composition session reaching the pty exactly once.
Co-authored-by: Orca <help@stably.ai>
* docs(terminal): record why withholding a claimed keydown needs no timer
The predicate withholds a keydown's byte until the commit arrives, so a key the
IME eats without committing would be dropped. Measured across the recorded
corpus that case does not occur, and the browser marks IME-owned presses on the
keydown itself. Both facts belong next to the predicate rather than only in a
handoff note, since the obvious fix for the imagined gap is a timer, and a timer
here once wrote a newline the user never typed.
Co-authored-by: Orca <help@stably.ai>
* test(terminal): pin the kitty all-keys-as-escape-codes hole explicitly
Flag 8 asks for every printable key as an escape code; this path sends the
committed text raw instead. That is a deliberate trade, not an oversight, but it
was untested — the suite only covered the disambiguate flag. Pinning it makes
the choice visible and records the gate to use if it ever needs closing.
Co-authored-by: Orca <help@stably.ai>
* fix(terminal): keep the kitty key-release report for presses that reached the pty
Claiming the keyup unconditionally suppressed xterm's release report. That was
sized for the old design, which claimed only a short punctuation list; the
structural claim takes every printable keydown, so on macOS an app that
negotiated kitty report_event_types stopped seeing releases for ordinary typing
and would treat every printable key as held down.
Suppress the release only when the press put nothing on the wire — swallowed by
the input source, or owned by a composition transaction. xterm emits nothing
from keyup unless kitty report_event_types (or win32 input mode) is on, so
letting it through is inert everywhere else.
Co-authored-by: Orca <help@stably.ai>
* test(e2e): assert IME preedit geometry headlessly for Korean and CJK input
Both IME defects that shipped and were reverted walked through a suite of ~3000
passing assertions, because every one of them was about bytes reaching the PTY.
A preedit rendered into a hidden overlay satisfies all of them while the user
composes blind. The real-geometry coverage that would have caught it existed but
was headful, env-gated and macOS-only, so it never ran in CI.
Drives composition through CDP Input.imeSetComposition instead of a native input
source, which removes the accessibility grant and the system input source that
forced the headful gate. The suite runs in the normal headless project in about
55s serially, and asserts the composition overlay's real bounding rect — the one
property an overlay clipped to max-width:0 cannot fake and a DOM emulator cannot
produce.
Three tests are red on main and marked test.fail() so they stay visible in CI and
flip loud when their fix lands: the preedit resumed by a bare compositionupdate,
and full-width punctuation and digits committed from a keydown that still carries
the ASCII layout key.
Co-authored-by: Orca <help@stably.ai>
* test(e2e): drop the known-broken markers now the stack closes all three
Validated on real macOS hardware: with the two fixes below this layer, all
three report "Expected to fail, but passed". Korean preedit renders at
non-zero geometry through every jamo, and an Apple pinyin source sends
ef bc 8c e3 80 82 to the pty where main sends ASCII.
Worth recording why the punctuation case looked green on main once: an input
source whose id happens to contain an allowlist term, as Sogou's does, satisfies
the old gate. Correctness there depended on which IME the user had selected.
Co-authored-by: Orca <help@stably.ai>
* test(e2e): cover the Linux and Windows IME ownership shapes headlessly
The ten headless IME specs on this branch all decided ownership through a macOS
user-agent override, so the two platforms whose failure mode is a *dropped*
character rather than a downgraded one had no coverage at all, and the one
Windows-recorded trace already in the suite was replayed under whichever policy
the runner happened to report — macOS locally, Linux on the CI shards.
Adds four Linux specs and two Windows specs, every one of them replaying a
native capture rather than a hand-authored ordering:
- IBus/X11 Hangul mixed with literal ASCII. Its `compositionend` is EMPTY and
the syllable arrives afterwards as a bare `insertText`, so reading the commit
off `compositionend.data` — which the Windows capture rewards — drops every
syllable on this framework.
- fcitx5/Wayland Hangul. No keydown at all for a composing key, not even 229,
and physically wrong `code` values on the literal keys. Any ownership rule
reading 229 or `code` fails here.
- Numeric pinyin candidate selection under both frameworks, with the ordinary
digit kept as the negative control, so the two directions are pinned against
each other rather than separately.
- Windows Microsoft Korean captured with real scan codes, including the two
lines committed with Shift held.
Each asserts both sides of the boundary: the preedit's real geometry at every
frame the user would see, and the exact byte stream the native run put on the
PTY. The recorded `onData` the Windows/WSL fixture already carried is now
asserted instead of sitting unused.
Chinese moves up to first-class alongside Korean: pinyin preedit width is now
pinned the way the Japanese phrase already was, and full-width punctuation is
covered in the composition-session shape the Windows and Linux frameworks use,
not only the macOS insertText shape.
Two harness fixes fell out of the recorded traces and are why the IBus one
passes. The replay applied each event's recorded textarea state *after*
dispatch, one event too late for the handlers that read `textarea.value`; and
it left a task boundary between `compositionend` and the `input` carrying the
commit, which Chromium never inserts, letting xterm's deferred finalizer settle
against a textarea the committed text had not reached yet.
Co-authored-by: Orca <help@stably.ai>
* test(e2e): replay the recorded macOS IME shapes instead of only synthesising them
The macOS coverage on this branch drives Chromium composition through CDP, which
is genuine but hand-ordered, so it could not assert the one property that
decides the macOS rule: a composing keydown arrives with keyCode 229 while `key`
is still the single translated character the input source produced — `ㅎ`, not
`Process` — which is indistinguishable by length from an ordinary printable key.
Three native captures were sitting unused in the evidence set.
Adds a recorded 2-Set Korean session, including the syllable boundary where one
composition closes and the next opens with no keydown between them, asserted
against its own recorded byte stream.
Adds the third failure mode, which had no coverage in any shape: an abandoned
preedit leaking to the shell. Pinyin and Cangjie both backspace a composition
away to nothing, and the assertion is not "the right bytes" but "no bytes".
Both cancellation captures continue with a literal `ordinary` typed as bare
keydowns, the recorder's own negative control. That tail carries no `input`
events because the build it was captured on produced the byte from the keydown
itself, so replaying it would measure the recorder rather than the product; the
specs cut at the `compositionend` and say so.
Co-authored-by: Orca <help@stably.ai>
* test(e2e): promote the non-allowlist input-source punctuation spec to the suite
Qingg matches none of the terms the pre-structural build enumerated, so on that
build its bypass never installs. It is the one arm no headless spec can express,
and the only test here that the pre-structural build cannot pass.
Promoted from scratch with four changes, each forced by a measurement rather than
by taste:
- A non-attached input method is a refusal, not a negative result. macOS attaches
per app instance and the attach can simply fail — 3 of 6 instances under
exclusive host access, and re-selecting the source did not recover one of them
across 9 keystrokes. That is now `test.skip()` with a reason naming the rerun,
not a thrown error, and the suite does not gate on a fully green session.
- Attachment is probed with a LETTER. Punctuation substitution emits no
compositionstart and no keyCode 229 even under a fully attached source, so at
the keydown it is indistinguishable from having no source at all. The
punctuation arms are judged on PTY bytes alone.
- The ASCII-layout control is now part of the spec rather than a side experiment.
Without it a build that rewrote every `.` into `。` unconditionally would pass
the Qingg arm and be badly wrong.
- The assertion runs by default instead of behind a strict-mode flag, and gained
a non-vacuity check: the input source must have committed something. That is
the sharp end of the mechanism — on the old build the DOM carries only keydown
and keyup, so nothing is committed at all and the character is destroyed before
the source is asked.
The verdict stays an equality between two measurements, never a comparison
against a hardcoded glyph, so it holds whatever punctuation mode the operator's
input source happens to be in. It reads `beforeinput`, not `input`: the forwarder
consumes `input` in the capture phase on the pane element, so a probe on the
helper textarea never sees it and a strict run fails with correct bytes
underneath.
Co-authored-by: Orca <help@stably.ai>
* test(e2e): normalise the pty line terminator so IME specs can run on Windows
A Unix pty's line discipline turns the terminal's CR into a bare LF, but
Windows ConPTY hands the reading process CRLF. Every spec compares against a
recorded line ending in LF, so on Windows 16 of 21 failed on the terminator
alone while the IME payload bytes ahead of it were byte-for-byte correct.
Verified on real Windows: the renderer-to-pty boundary assertion passed there,
so the app writes a bare CR and the LF is added by the console downstream of
anything we control. Normalising in the reader keeps the specs asserting the
IME bytes unchanged rather than loosening them.
Co-authored-by: Orca <help@stably.ai>
* chore: drop non-mergeable IME e2e scratch files
---------
Co-authored-by: Orca <help@stably.ai>
* fix(sidebar): show Cursor rows and stop a stray "claude" title hijacking OpenCode
Two defects in the same title-resolution path.
**#10258** — Cursor's only native OSC title is the literal `cursor agent`, which both title trackers dropped unconditionally. A hookless Cursor pane therefore had neither a status entry nor any title carrying Cursor identity, so the worktree card showed nothing at all.
**#8940** — two owner-blind paths let an incidental `claude` token anywhere in an OpenCode session or task title outrank the pane's known owner, so the tab icon and sidebar row flipped to Claude Code.
#10258: let the literal through exactly once as identity, so a restored or mobile tab keeps its Cursor row instead of vanishing. #8940: require an *identity frame* — after stripping status decoration the title must PRESENT Claude, not merely mention it — before a Claude title may reclaim a pane from its prior identity, and make the sidebar row builder owner-aware.
> These two are in one PR because they share the `ownerAgentType` plumbing through `buildTitleDerivedAgentRow` — split apart, neither half compiles on its own.
Fixes#10258Fixes#8940
Co-authored-by: Orca <help@stably.ai>
* test(e2e): add recordable proof for sidebar-agent-row-identity
Fails on origin/main, passes on this branch.
Test: sidebar keeps a Cursor pane visible and an OpenCode pane out of Claude Code hands
Co-authored-by: Orca <help@stably.ai>
* fix(terminal): preserve restored Cursor identity
* test(terminal): cover restored Cursor redraw suppression
* refactor(terminal): tighten Cursor identity handling and Claude frame matching
Review follow-ups on the title-resolution path:
- pty-transport dropped a native Cursor literal that main emits whenever a
non-Cursor title preceded it, re-introducing the #10258 blank row in the
renderer path. The pre-filter now projects the predecessor the drain will
actually see, and defers to the drain gate while facts are still queued.
- applyTrackedPtyTitle threaded the cursor flag through 12 sites, including
ptyRecordChanged bookkeeping the sole caller ignores. Force the status null
once, and the activity-gated effects fall out unchanged.
- isClaudeIdentityFrameTitle missed a multiplexer-wrapped Claude title
("zsh | Claude Code"), costing a genuine Claude pane its identity. Reuse
the ' | ' segment split that agent-title-owner already had inline.
- Keep title normalization on launchAgent: it only rewrites within an
identity group (OMP wraps Pi), so a split does not make it wrong, and
the hook-row path normalizes the same way.
- Drop the tab.ptyId tracker fallback, which read a pty that the pane
identity check had just rejected.
Co-authored-by: Orca <help@stably.ai>
---------
Co-authored-by: Orca <help@stably.ai>
* Revert "test(ime): restore coverage the composition-ownership change removed (#13168)"
This reverts commit 25a8c517e1.
* Revert "refactor(terminal): return IME composition ownership to xterm (#13128)"
This reverts commit 17b3dff3c4.
* test(ime): keep the architecture-neutral Korean trace coverage
The recorded IBus/fcitx5 and Windows MS-Korean traces from #13168 assert PTY
byte order, not composition ownership, so they still hold once the terminal
composition layer is restored. The mobile accessory-order test pinned the new
handleLiveInputChange signature and does not.
Co-authored-by: Orca <help@stably.ai>
* fix(terminal): keep the macOS Backslash bypass through the revert
The restored native-text forwarder only claims keys for input sources in its
hardcoded CJK allowlist, so third-party IMEs off that list (Qingg, #10896) still
get a raw backslash. #13128 added this bypass as a partial replacement; keep it
rather than trade the open issue back.
Scoped to the bare backslash key. The rest of shouldBypassXtermForMacNativeText
bypassed all unmodified non-ASCII text, which would race the restored forwarder.
Co-authored-by: Orca <help@stably.ai>
* fix(mobile): move the mirror-step ref write out of render
The restored hook assigned runMirrorStepRef during render, which is not
replay-safe — React can discard render work, so the mutation can leak from UI
that never commits. Its only read is inside the held-commit timer, which fires
long after commit, and the ref has a safe default, so an effect is soon enough.
Surfaced by the changed-lines React Doctor gate: the rule postdates this code,
so restoring the file re-introduced it as a new violation.
Co-authored-by: Orca <help@stably.ai>
---------
Co-authored-by: Orca <help@stably.ai>
* Focus search inputs for immediate typing
- Autofocus inputs in AutomationListSearchField, SettingsSidebar, and WorktreeParentPickerPopover
- Only autofocus Settings search when opening directly, not via deep-link
- Use modal mode and explicit focus management in popover for proper restoration
- Forward CommandInput ref and add autofocus test coverage
* Restore focus when closing worktree parent picker popover
- Find the nearest focusable ancestor of the anchor row to restore focus
to instead of letting it drop on the detached input element
- Simplify focus assertion in AutomationListSearchField test to verify
actual focus behavior rather than autofocus attribute presence
* fix(terminal): return IME composition ownership to xterm
* fix(mobile): derive terminal input from native replacement ranges
* test(mobile): record iOS Japanese IME traces
* fix(mobile): preserve native IME replacement ranges
* fix(xterm): flush queued application input after IME commit
* test(terminal): pin Korean intermediate commit
* test: pin Windows IME shortcut ownership
* test: replay IBus number candidate commit
* fix: preserve native macOS input-method punctuation
* refactor(terminal): remove stale mac focus override
* fix(mobile): preserve soft keyboard deletion ranges
* fix: keep IME-owned palette chords in renderer
* fix: stop carried IME shortcuts at renderer owner
* fix: preserve carried IME shortcut dispatch
* fix: narrow main-owned shortcut actions
* test(mobile): pin Japanese IME replacement traces
* test(terminal): retain paired native IME trace
* fix(chat): preserve browser IME composition ownership
* fix(chat): retain macOS IME confirm gesture
* fix(chat): expire unmatched IME confirm carry
* fix(chat): isolate IME confirmation expiry
* fix(chat): retain active IME confirmation
* refactor(terminal): remove dead composition handler
* feat(ime): add shared Enter-ownership seams for CJK composition
The confirming Enter of a CJK composition arrives as two keydowns and the
orderings differ by platform: Windows/Linux redispatch the unmarked Enter/13
before keyup, macOS delivers keyup first. A guard reading only isComposing or
keyCode 229 misses the redispatch, so surfaces submitted on a confirm.
Adds useImeEnterGestureOwnership (carry token, next-frame expiry), a shared
ImeEnterGuardedForm for native implicit submission, and the cmdk seam covering
18 CommandInput surfaces at one site.
A chorded Enter arms the carry but is never swallowed — the reverse would eat a
user's deliberate Cmd/Ctrl+Enter. Both failure modes are pinned by
ime-enter-gesture-ownership-contract.test.ts.
Co-authored-by: Orca <help@stably.ai>
* refactor(terminal): consolidate native input listeners and parked-screen owner
Extracts the shared native-input listener installer and renames the parked-screen
detector for what it actually does, replacing per-call-site duplication. The
listener installer keeps a forgetOptionKeyLocationOnBlur flag so per-window
semantics are preserved rather than flattened.
Net deletion; no behaviour change intended.
Co-authored-by: Orca <help@stably.ai>
* test(terminal): pin recorded IME shapes as regression tests
Nine regression tests built from hashed affected-platform captures, each with a
paired ordinary negative and a discriminating mutation verified to take the file
from all-passing to exactly one failure.
Covers the Windows MS-Korean Shift family (#12179, #11878, #12151, #11946,
#12152) and the Korean TUI line-break rows (STA-3237, STA-3222, STA-3129).
STA-3237 pins the empirical 3-Shift / 2-active-composition / 2-newline ratio the
device run established — the third Shift produces nothing because Space has
already committed. That ratio is not derivable from a static capture.
Co-authored-by: Orca <help@stably.ai>
* fix(ime): guard Enter-commit surfaces against CJK confirm
Applies the Enter-ownership guards across the surfaces whose Enter commits
something: publishes, clones, pairs, installs, posts, or persists.
Tiered deliberately rather than uniformly. Irreversible and remote-effect sites
take the carry token, which also blocks the unmarked redispatch. Locally
reversible sites take the oracle check with a one-line comment naming the
residual, because a spurious commit there costs one undo.
Three numeric fields are left unguarded with the reason in-code: Chromium blanks
number inputs at compositionstart, so a confirm-Enter only ever reaches an
empty-draft reset. Measured with a CDP probe rather than assumed — a guard that
cannot fire is noise.
Co-authored-by: Orca <help@stably.ai>
* test(ime): teeth-check the Enter guards on every guarded surface
One suite per guarded surface, each verified by deleting the guard and
confirming the test fails. A green guard test without that check is unverified,
not verified.
Two shapes pass vacuously in happy-dom and are avoided here: native implicit
form submission never fires, and blur() is inert on an unfocused element. Both
made "the commit did not happen" assertions pass with the guard removed, so the
suites assert the guard's contract directly instead.
Co-authored-by: Orca <help@stably.ai>
* fix(mobile): keep iOS Korean commits whole through the live-input path
iOS Korean reports isComposing: false on every event, so it bypasses the
composition guard entirely. The strict owner rejected UIKit's transformed
post-change field and sent only the leading jamo — the reported symptom.
Prefers the authoritative same-event field text over the predicted text when the
supplied operation cannot produce it. Generic: no Korean special-case, no locale
classifier, no normalization. Adds the RN-target-keyed submit carry alongside it.
Co-authored-by: Orca <help@stably.ai>
* test(e2e): make IME capture harnesses fail loudly instead of silently
Four instruments recorded silence as success, so a void run scored as a clean
one:
- readTerminalImeBoundaryTrace returned an empty trace when the probe never
installed, making every "nothing leaked" negative pass vacuously
- summarizeLatencies([]) returned a perfect zero distribution that passed all
three latency thresholds
- the macOS Vietnamese spec pinned an input-source ID that does not exist, and
failed as though the operator had chosen the wrong source
- the expectedLineCount=1 prefix property was undocumented and one edit from
silently downgrading a PTY assertion
Input sources now resolve by enumeration and name the near-matches on failure.
Co-authored-by: Orca <help@stably.ai>
* test(terminal): cover Cangjie cancellation and fix a cross-namespace assertion
Adds #11951's recorded Cangjie cancel shape to the existing cancellation suite,
which covered Pinyin and Sogou but not Cangjie. One keystroke then Backspace
arriving as deleteContentBackward with data: null, so the stale preedit is the
only thing a fallback could replay.
Verified against the historical pre-6cd944c62b3 bundle: the positive fails with
['尸'] where [] is expected, while the ordinary negative stays green.
Also fixes the Vietnamese spec, which asserted a TIS-space input-source ID
against getKeyboardInputSourceId(). Those two Orca APIs report the same source
in different namespaces — TIS nests it under VietnameseIM, the app API does not.
The resolver stays as an installation precondition; the assertion matches the
leaf.
Co-authored-by: Orca <help@stably.ai>
* test(e2e): add a real-IME macOS arm for the Korean chord commit
The existing korean-ime-terminal-shift-enter-commit spec synthesizes composition
over CDP: Input.imeSetComposition sets the preedit directly and Input.insertText
performs the commit. Asserting the IME produced events you injected yourself is
circular, so that spec cannot certify real-IME behaviour.
This arm selects 2-Set Korean via TIS, reads it back live, and injects through
System Events key codes, so the OS owns the preedit, the commit instant, and
isComposing. PTY byte expectations are preserved verbatim.
Covers 2 of the original 4 cases by design. The other two are the Windows/Linux
redispatch-before-keyup ordering, which macOS cannot produce and which cannot be
selected -- the OS decides it. Reintroducing synthesis to "restore coverage"
would reintroduce the circularity.
Co-authored-by: Orca <help@stably.ai>
* test(e2e): assert the macOS chord arm at the PTY boundary, not the renderer
The byte expectations were transcribed from korean-ime-terminal-shift-enter-commit
:364/:383, which assert against onData -- a renderer boundary where the terminator
is CR. This spec reads the PTY child, where the tty has already converted CR to LF.
Names both forms per row rather than swapping the constant, so the conversion reads
as evidence that the capture reached past the renderer, as #11936 and #11951 record.
Ctrl+Enter's CSI-u sequence is unaffected and is identical at both boundaries.
Co-authored-by: Orca <help@stably.ai>
* test(e2e): measure composer-to-onData latency and stop dropping IME keystrokes
Two defects in the echo latency probe.
It hooked onWriteParsed and onRender but never onData, so it measured
key->parse->render echo rather than the composer-vs-onData delta the latency rows
need. Adds a third hook feeding its own sample set.
And `event.key.length !== 1` silently dropped IME keystrokes: Pinyin and Cangjie
keydowns arrive as key:'Process' (length 7). Replayed over the captured corpus,
the old filter accepted 580 of 4137 Chinese IME keydowns -- it was discarding 80%
of them. The new filter matches the shape the owner itself branches on.
Attribution charges each onData to the latest keydown rather than a FIFO head,
because composing jamo emit no onData at all and a queue would credit a whole
composition to its first keystroke. The consumer now asserts sample count before
any percentile, so a zero-sample run cannot render as a flawless distribution.
Co-authored-by: Orca <help@stably.ai>
* test(terminal): pin the WSL shifted-jamo newline shape for #11919
In Korean 2-set, Shift types ordinary letters -- the double consonants and the
compound vowels. Each such keystroke reaches Chromium as key='Process',
keyCode=229, shiftKey=true.
The v1.4.163 classifier matched exactly that pattern with no code guard, so it
called those keystrokes Enter, rewrote them to a synthetic Shift+Enter, and
injected a newline into the middle of the word -- with no Enter key pressed.
That is why the reporters said "no modifier key pressed": they had not chorded
Shift+Enter, but they had pressed Shift, to type the double consonant.
Asserts the row's own recorded capture: 40 immediate keydowns, exactly 3 of them
Shift-carrying inside a single syllable, and an onData stream with one newline
per Enter press and none mid-word. Two ordinary negatives keep it from being a
blanket mute -- the same session's non-IME keydowns still reach shortcut policy,
and an ordinary Shift+Enter still resolves through the real policy.
Co-authored-by: Orca <help@stably.ai>
* test(terminal): pin the composition commit lag that made Korean type one behind
macOS Korean 2-Set commits syllable N only when the first jamo of N+1 arrives, so
compositionend and compositionstart land in the same task. A composition-start
handler cancelled the pending finalizer that was the only path to triggerDataEvent
and ended the session without emitting bytes, so every committed syllable reached
onData exactly one syllable late and the backlog cleared only at a Space or Enter.
Types continuously with no Enter and no Space -- either would flush the backlog and
hide it -- and samples onData at every syllable boundary. Paired with a
length-matched ASCII arm that stays green throughout, so the positive is a fact
about composition rather than about timing in general.
Bisected to a single call site across five builds: pristine, 1.4.155 and 1.4.162
pass, 1.4.163 fails, removing the one call repairs it, restoring it fails
identically. That window is exactly the reporter's "started immediately after
updating".
Co-authored-by: Orca <help@stably.ai>
* test(mobile): cover the send-queue abort that silently drops queued keystrokes
One failed send in use-terminal-live-input-commit aborts every keystroke
queued behind it, with the error swallowed by .catch(() => false). The
existing test resolves(true) on every send, so the failure branch was
uncovered.
Four arms: the abort itself, an ordinary negative on the healthy path, a
throwing sender, and a liveness control proving the queue recovers once
the chain settles. Deleting the abort takes 4 passed to 3 failed, with the
ordinary negative correctly surviving.
Scope is stated in the docblock: this is a transport send-queue abort,
reachable only via a real disconnect or RPC error. REQUEST_TIMEOUT_MS is
30s, so latency alone cannot reach the branch — consistent with #7094's
symptom class, not proven to be its cause.
* test(terminal): pin that daemon snapshot/restore cannot disturb a composition
Two independent reporters attributed broken Korean composition to the
always-on PTY daemon repainting terminal state over the preedit. The
attribution is wrong on ancestry — the daemon shipped three months before
the version both call good — but the boundary was never actually tested.
Runs the real applyMainBufferSnapshot choreography against a live
composition, including the full 2J/3J/H wipe plus the resize and
alt-screen branches. textarea.value, selectionStart/End,
compositionView.textContent and .active all survive byte-identical, and
interleaving a restore between every jamo of 문제 still commits 문제 at
onData. Also pins that the uncommitted preedit is absent from the captured
snapshot: it lives in the textarea, never the buffer, so a restore has
nothing stale to echo back.
Injecting one textarea.value = '' into the restore fails exactly the three
restore-boundary tests.
* test(terminal): pin that Cmd tears down a composition where Ctrl and Shift do not
xterm's composition keydown exempts only keyCode 16/17/18 (Shift/Ctrl/Alt)
plus 20/229. macOS Meta — 91/93/224 — is absent, so a Cmd press mid-composition
takes _finalizeComposition(false): the overlay goes dark and never recovers,
because compositionstart is not re-fired. The user composes the rest of the
word blind. Linux and Windows users press Ctrl and are exempt.
xterm already has a Meta-aware modifier predicate in wasModifierKeyOnlyEvent,
so this is an internal inconsistency rather than a deliberate choice.
Owns no reported row and is version-neutral: 5/5 on both 1.4.162 and 1.4.163.
The branch is unexercised in all 328 recorded traces, so this is a hazard pin,
not a regression guard. Only the teardown is asserted; the likely duplicated
commit needs a compositionend the IME kept alive across the Cmd, which no
capture contains.
Deleting the exemption fails exactly the three paired negatives; adding Meta
to it fails exactly the two Cmd arms.
* test(native-chat): characterize preedit loss when a question card replaces the composer
An AskUserQuestion card fully replaces the composer by design, but the
in-flight composition goes with it: the composer unmounts before
compositionend reaches it, so the preedit is never committed to the draft.
The committed text survives only because the draft is cached and restored
via defaultValue. Node identity changes, value 'abc' is preserved, the 가
is gone.
Drives the real NativeChatView -> SessionGate -> InteractiveCard ->
questionActive swap -> Composer -> ComposerField, flipped by writing the
same store field an AskUserQuestion hook event writes. Flipping
questionActive to false fails exactly this test and nothing else across
639 native-chat tests, so the path was entirely unguarded.
CHARACTERIZATION TEST: it asserts the loss. Fixing the defect — committing
the preedit before the swap, or keeping the composer mounted — will make
this file fail. Update the expectations to the new contract rather than
working around them.
Owns no reported row. #12118/STA-3219 flicker is keyed to token counters,
which provably do not remount, and a question card arrives once per
question.
* test(terminal): pin the duplicated commit when Meta interrupts a composition
_finalizeComposition(false) sends textarea.value.substring(start, end) but
cannot clear the IME-owned textarea, so a later compositionend re-sends the
same range. Meta reaches that path because CompositionHelper exempts only
Shift/Ctrl/Alt; xterm's own wasModifierKeyOnlyEvent covers Meta four ways,
so the omission is an internal inconsistency rather than a choice.
Companion to the modifier-exemption guard, which deliberately pins only the
overlay teardown. This pins the data consequence.
HAZARD PIN: owns no reported row. The trigger is unverified on hardware —
no capture in the corpus contains a Meta-during-composition gesture, and
whether macOS keeps the composition alive across it is unmeasured. The
duplication follows from the code given that sequence; whether users reach
the sequence is the open half.
An earlier premise that Space (keyCode 32) reaches this path was refuted by
a corpus scan: 0 of 731 evidence files carry a keyCode-32 Space while
composing, against 171 at 229, and 229 returns early.
* test(terminal): characterize the syllable lost when the textarea blurs mid-composition
CoreBrowserTerminal._handleTextAreaBlur clears the helper textarea
unconditionally — "Text can safely be removed on blur" — while
CompositionHelper._finalizeComposition reads the committed text back out of
that same value from a deferred timeout. By the time it runs the value is
empty, the substring is '', and triggerDataEvent never sees the syllable.
xterm checks composition state in _syncTextArea and omits the same check
here.
Six cases. Blurring mid-composition loses the syllable in every ordering,
including compositionend-before-blur, which is Chromium's real order — so
it is not an ordering artifact. A bare textarea.blur() with no Orca code
loses it too, which places the owner upstream: Orca's unguarded release on
outside pointerdown is one trigger, not the cause. Committing 한 then
blurring mid-가 yields ['한'] where ['한','가'] is correct: one syllable
gone, surrounding text intact.
Teeth checked by inverting — adding an Orca-side composition guard flips
exactly the three cases that route through the release path and leaves the
bare-blur and no-blur cases green, which is the scope split: a fix in
regular-terminal-focus-ownership alone would not close this.
HAZARD PIN, but unlike the others this one has a real production injector —
clicking outside the terminal mid-composition. Owns no reported row. The
shape matches #9738's report; the injector does not, and a shape match with
a mismatched injector is not an owner.
* test(terminal): say which arm the STA-3237 fixture came from
The recorded keydowns are wave 4's A-shift-unmarked-only — the arm that
emits no PTY bytes. Nothing in the file said so, so two readers concluded
the row's events fail the owner's predicate and that STA-3237 and STA-3222
were different defects. They share an owner; the arm that fires is
Process/229+Shift, absent from this bubble-phase trace because the owner
claims it in the capture phase.
Also corrects "code-blind": the v1.4.163 policy emits \x1b\r only for a
shift-only key:'Enter', and a jamo keydown reaches that branch solely via
the isTerminalImeProcessEnter rewrite. The mock is deliberately wider so
the ownership guard stays under test if that rewrite moves.
Comments only — no assertion, fixture value, or mock behaviour changed.
* test(e2e): track the input-source selector the macOS specs shell out to
Five tracked macOS IME specs ran `swift .tmp/select-input-source.swift`, a
file that is gitignored and existed only on one machine. Anyone else
checking out the repo — or the same machine after .tmp is cleaned — could
not run them, and they are the capture drivers for the macOS rows that are
blocked waiting for exactly those runs.
Moves it to tests/e2e/ beside its callers. The chord spec now resolves it
from __dirname rather than reaching two levels up into .tmp.
* test(terminal): pin the CJK repaint decision against the reporter's own output
#12164 comment 1 and #5921 report agent output with double-width glyphs
rendering duplicated character-by-character while ASCII in the same line
stays clean. No IME, no composition, no keystroke — the user never types
the CJK.
Segmenting all three verbatim samples into maximal same-risk-class runs
gives 33 runs and zero violations of "this run is corrupted iff the
production detector flags it": 17 wide runs all corrupted, 16 narrow runs
all byte-identical. The paired negative is co-located in the same line
rather than in a separate run — the reporter supplied it without knowing.
Doubling is asserted as present, not uniform: 자바스크립트 and 시스템 each
leave a jamo undoubled, which is a repaint-region boundary artifact rather
than a per-character transform.
The discriminating arm is in the test rather than a source mutation:
be3f30e2f8 (#6890) elects a repaint for all 17 corrupted runs when the
agent types nothing, and reverting its disjunct elects none. Both
predicates agree once the user has recently typed, which is the pre-#6890
condition.
Samples inlined with per-sample SHA-256 because .tmp is gitignored and
cannot back a landed test.
* test(terminal): pin macOS period substitution landing after the composition
#11504's reporter published a DOM trace showing insertText ". " arriving
149ms after compositionend, when two spaces are typed with a CJK input
source and NSAutomaticPeriodSubstitutionEnabled is on. This replays that
trace against a real Terminal and asserts what reaches onData — bytes to
the PTY, not anything visual.
The owner is stock upstream CoreBrowserTerminal._inputEvent, not an Orca
module, confirmed at the resolved install and in the shipped bundle Vite
loads rather than in the TypeScript source.
Three mutations against that install, predictions written before the runs,
each failing exactly the arms predicted: dropping the composed/keyDownSeen
guard fails two, dropping Orca's intercept fails the one arm where the
payload arrives before the send window drains, and flipping || to && —
the candidate-fix shape — fails the arm that pins the defect itself.
CHARACTERIZATION: arm 1 asserts the broken behaviour and will fail the
moment #11504 is fixed. Update it to the new contract rather than working
around it.
composed is absent from every recorded bundle, so composed: true is the
spec-required value rather than a captured one; the test asserts it before
dispatching so a harness that dropped the field fails loudly.
* test(terminal): replay the recorded Windows Shift sessions through the IME guard
STA-3179 reports a Shift release sending Enter; #12171 reports delayed
Hangul plus doubled newlines. Both replay their own recorded Windows
MS-Korean keydowns through resolveTerminalKeyboardShortcutAction with the
shortcut policy mocked, so the assertions are about which events reach the
policy and what reaches terminal input.
STA-3179's held-Shift gesture yields exactly one newline, from the unmarked
Enter alone; its release arms nothing for the next composition, asserted
after a precondition check that the release really is keyups with shiftKey
already dropped; and an ordinary Shift press-and-release still routes every
keydown, which is the paired non-IME negative.
Teeth, verified by mutation: bypassing the isImeOwnedKeyboardEvent guard in
keyboard-handlers takes STA-3179 from 3 passed to 2 failed / 1 passed — the
survivor being the ordinary-session negative, which is correct, since a
non-IME session should not depend on that guard — and #12171 from 2 passed
to 2 failed. Source restored byte-identical.
Recorded shapes are inlined and the bundles cited in comments; nothing is
imported from .tmp, which is gitignored.
* test(native-chat): correct 61977d45177 — the preedit survives the question card
61977d45177 claimed a composed syllable vanishes silently when an
AskUserQuestion card replaces the composer, and characterized that loss.
The claim was false. Its premise was an artifact of the harness: the test
simulated a preedit with a silent textarea.value assignment and no input
event, which no IME does.
Real composition fires input with insertCompositionText on every keystroke
— the shape this repo already records in its own observed-event capture —
and React's change handler returns on input/change with no composition
gate, so onChange runs for each frame. The draft cache is written
synchronously inside the updater, so the preedit is already committed
before the card can arrive. Driven that way, it survives.
Renamed to match the contract that actually holds, and extended: Hangul
jamo-per-frame, Japanese kana accumulation followed by per-segment
conversion asserting the candidate the user was looking at survives, and a
pin on the mechanism itself — the draft cache holds the preedit while the
card is up.
Teeth: there is no fix to revert, so the mutation is the plausible wrong
one — gating onChange on isComposing(). That takes 4 passed to 3 failed,
with the English negative correctly surviving, since it has no composition
to gate.
Two consequences remain, recorded rather than fixed: the OS aborts the
composition when the field disappears, so a lone jamo returns as a
compatibility jamo the user cannot compose onto, and the remounted
composer is unfocused because the card owned focus.
* docs(native-chat): name the corrected commit and the degraded-jamo consequence
Records in the file itself that 61977d45177 is pushed and wrong, quoting
the two claims that are false, so a reader who finds it in git log reaches
the correction from the file that replaced it.
Also states the residual as a consequence rather than a curiosity: a lone
leading jamo returns as a standalone compatibility jamo (U+3131), which is
not a composable state — the user cannot resume the syllable, only delete
and retype. Preserved, but degraded into something unusable. That is the
note to find if a reporter ever describes exactly that.
The invariant these tests pin is not "the composer commits on unmount" but
"composition input events must reach React" — which is what a future IME
change would break, and is not visible from the swap site at all.
* test(terminal): replay the recorded macOS Telex commit boundaries
#6905 reports Vietnamese composed characters breaking in the terminal.
Replays the retained macOS built-in Simple Telex capture — recorded
selection and value set before each dispatch, since that is what the
commit range reads — and asserts what reaches onData: the first commit
alone, then through the real Enter, then the ASCII tail of the same run.
A code-point count would catch NFD normalisation.
ENGINE CAVEAT, stated first in the docblock: this is macOS built-in Simple
Telex, Telex only. The reporter's three named engines cannot run on the
platform they declared, and which macOS Vietnamese engine they used is
unconfirmed. This file certifies no engine, and does not imply VNI.
The owner is upstream's — CompositionHelper._finalizeComposition's
waitForPropagation branch — so the arms are copies under .tmp aliased by a
scratch config, with node_modules verified unchanged by shasum after every
run. Collapsing the range end onto its start fails all three; collapsing
the start to zero re-emits the first word into the second commit, which is
the reporter's "duplicated" direction. Different failure sets, so the
mutants are distinguishable rather than merely detectable, and the ASCII
assertion passes under both.
Falsifiability here is by mutation, not by a defective build: #6905 does
not reproduce at HEAD, so this has never been watched going red on a real
reproduction.
* docs(terminal): lead the #6905 test with its engine caveat
Comment-only. Moves the caveat above the source line so a reader meets what
the file does NOT establish before what it does — the capture is macOS
built-in Simple Telex, the reporter's named engines cannot run on the
platform they declared, and which engine they used is what gates this row.
Co-authored-by: Orca <help@stably.ai>
* docs(terminal): record that the swallow eats a keystroke after Japanese conversion
This pin framed the swallowed insertText around Cmd interrupting a
composition. A differential through Japanese multi-segment conversion shows
it is broader: type a segment, convert, then press `a`, and the `a` is
lost. No modifier, no exotic gesture. Korean surfaced it first only because
2-Set composes on nearly every keystroke.
Also records why it cannot simply be fixed. The suppression de-duplicates
IMEs that deliver their commit a task after compositionend, which a sibling
test pins; this swallow is that dedup's false positive, and the two events
differ only in payload, so no flag-timing change separates them. Both a
smaller redesign and a content-aware variant were built and measured — the
first duplicates on IBus, the second costs a reported row's test and is
blocked while the patch cannot be regenerated.
The Japanese arrays behind this are authored, not observed: no Japanese DOM
composition trace exists in the corpus.
* docs(terminal): a Japanese capture does exist — correcting dbeecb11bee
That commit said no Japanese DOM composition trace exists in the corpus.
False. One does, filed under the Linux bundles rather than the bundle named
for Japanese: 30 DOM events, two にほんご->日本語 conversions, with full
selection state per event. It is retained byte-identically in three further
bundles — one capture copied four times, not four observations, checked by
hash rather than by counting files.
The claim came from checking the bundle named for Japanese, finding nothing,
and generalising to the corpus without querying the rest of it.
Replaying it emits 日本語日本語 under both sequencing extremes on all four
arms, matching its own recorded onData. So "repeated conversion is
undisturbed" is now captured rather than authored. It carries no
post-compositionend insertText, so it cannot speak to the swallow: the
a-after-conversion figure stays authored and unobserved.
Also rewords the paragraph opener. It claimed to broaden a Cmd framing, but
hazard 2 was never Cmd-framed — the lines above already say Cmd does not
reach it. The real gap was that hazard 2 named no trigger at all, which
reads as exotic when it is ordinary.
* build(xterm): land the patch regeneration harness
The five dependency patches under config/patches/ shipped with no tracked
way to regenerate any of them. The xterm one is the hard case: it is derived
from an upstream build, so no fix could be made without rebuilding, and the
tooling to rebuild lived only in one machine's scratch directory. That
blocked a measured fix for a live keystroke-loss bug, and the EditContext
reduction an OSS survey identified as the only real one available.
Adds the regenerator, the upstream pin, the hand-written source patch the
bundle hunks derive from, tests, docs, and a PR job that verifies the
shipped patches still match the pinned build. The job caches the shallow
clone keyed on the manifest, so a cold run is minutes and a warm one under
one. Round-trip verified: regenerating from a clean checkout reproduces the
shipped patch byte-for-byte.
Marks the emitted patch -diff -text. pnpm hashes it byte-for-byte, so a
CRLF checkout would break install on Windows, and its minified bundle lines
make a diff nobody can read — review the source patch instead.
Also rejects unknown flags. --check was the fallback for any unrecognised
argument, so a typo, or --help, silently triggered a full upstream build
instead of what the caller asked for.
* fix(xterm): stop swallowing a keystroke typed after an IME commit
Type a Japanese segment, convert it, then press a key one macrotask later
and that key was lost. No modifier, nothing exotic — every user who keeps
typing straight after converting. Korean surfaced it first only because
2-Set composes on nearly every keystroke.
handleCompositionInput discarded the payload unconditionally in the window
after the deferred send: _isSendingComposition stays true for one macrotask
after the timer cleared _pendingCompositionStart, and the branch substituted
'' for whatever arrived. The suppression is not itself wrong — it
de-duplicates IMEs that deliver their commit an event-loop turn late, which
terminal-stock-composition.test.ts pins. It just could not tell a duplicate
from new input, because the two events are identical apart from payload.
Now it compares against _sentComposition, the text the deferred send
actually emitted, and discards only a match. A flag-timing redesign was
measured first and rejected: it fixed this and duplicated on IBus, because
no timing change can separate events that differ only in content.
Edited in config/patches/xterm-src/ and regenerated through the harness, so
the emitted patch and the lockfile hash are derived, not hand-written.
The commit-overlap pin's swallow arm now asserts the repaired contract —
the value its own comment already named as correct and as what stock
beta.287 emits. #11504's arm at :184 flips too; it never covered that
report, as its own prior note recorded, and the reporter's +149ms arm is
untouched and still asserting the defect. Provenance hashes in three test
docblocks are updated, since regenerating changes the patch hash and with it
the resolved install directory.
* docs(terminal): re-measure the #6905 mutation citations against the new bundle
Regenerating the patch moved the resolved install, so this docblock's
patch_hash, line count, two line numbers and three mutation outcomes all
described a bundle that no longer exists. The deferred branch is one the
fix writes into, so the outcomes could not be re-pointed on reasoning.
Line numbers read off both files by diffing anchors rather than derived by
arithmetic: 201 to 205, 159 to 163. Outcomes re-run through the retained
rig, which re-resolves through the module loader and re-derives each arm
from a unique minified anchor: pristine 3 passed, m1 3 failed, m2 2 failed,
m3 3 passed — identical to the old bundle. Guard controls in both
directions exit 1, so the counts are falsifiable.
Comment-only; the assertions and expectations are unchanged.
* fix(xterm): size the preedit overlay to the cells its text will occupy
updateCompositionElements computed the overlay's left edge from the grid
but never its width, so the preedit rendered at the font's natural advance
while the committed text takes two cells per wide glyph. Measured in
Chromium 150: 가나다라 drew 48.45px as a preedit and 69.20px once committed
— the same characters, same font, 30% narrower, and drifting further with
each syllable. Every macOS mono font carrying Hangul measured 0.49–0.72 of
two cells; never 1.0.
Deriving the width from wcwidth and the cell measure moves Korean, Japanese
and Chinese to 1.000 and leaves ASCII at 1.000, which it already was:
한 12.125 -> 17.297 (17.30 expected)
가나다라 48.453 -> 69.188 (69.20)
안녕하세요 60.563 -> 86.500 (86.50)
日本語 42.000 -> 51.906 (51.90)
abcdefgh 69.234 -> 69.203 (69.20, unchanged)
Edited in config/patches/xterm-src/ and regenerated through the harness, so
the emitted patch and lockfile hash are derived rather than hand-written.
The unit test asserts the arithmetic, which is what CI can run. The pixel
consequence was measured on macOS with SF Mono in an Electron harness, not
on the Windows font stack STA-3232 reports from — so this demonstrates the
mechanism and does not stand as that row's platform evidence.
* test(e2e): pin the macOS Korean preedit as visible only while composing
#11914 reports the composing text invisible until Space. Its c3 was recorded
as unobtainable, and the reason on file was wrong: the boundary IS
assertable, but not in happy-dom, which reports display:block in BOTH the
active and inactive states and zeros for every rect. A test there passes
with the defect present.
Captured on real hardware instead: hidden and 0x0 before, .active with
display:block, a 15.84x16 rect and checkVisibility() true while composing
그, hidden again after. 39 DOM events, 2 composition starts, onData
["한","그","\r"].
Two mechanism findings are carried in the setup because both are invisible
in the result and fatal if removed. The input source must be selected AFTER
the app takes focus — focusing resets it to ABC. And the IME must be warmed
until an observed keyCode 229; typed cold it emits raw QWERTY (g k s r m)
with no composition at all, which is indistinguishable from an IME that is
not installed. Two runs were voided on exactly that signature before the
warm-up was found.
The has229 and compositionStarts assertions exist to make such a run fail
loudly rather than pass as a clean negative.
Gated on darwin plus ORCA_E2E_NATIVE_MACOS_KOREAN, like its siblings. The
final spec form has not itself been executed — the machine became
unavailable — so it carries the probe's measured values as literals rather
than a run of its own.
* docs(e2e): correct 19a8d133db7 — the Korean preedit spec has been executed
That commit said the landed form had never run and carried the probe's
values as literals. It has now run on real hardware: 1 passed, 9.1s, rc=0,
with the capture and log sealed under a verified hash manifest.
The teeth check was also run rather than reasoned about, and it changes
which assertion matters. Forcing the active overlay to max-width:0 with
overflow:hidden — invisible on screen — leaves the active class, the
textContent, display:block AND checkVisibility() all passing. Only
during.rect.width fails. checkVisibility() is not sufficient against this
defect; the bounding rect is the single load-bearing assertion, which the
docblock already said and this run confirms.
An earlier teeth attempt injected the CSS mid-run and tripped the
hasActiveClass poll instead, failing at the wrong assertion. It is
inconclusive and excluded from the seal rather than counted.
* test(terminal): add #12171's ordinary-English arm from a real Windows capture
c4 was recorded as unmet and the ledger sourced its control to
evidence/windows-current/, which holds 12 captures and not one English one.
The arm here comes from windows-9803-final instead — same probe, same host
geometry, same injector, en-US with no IME, replayed keydown for keydown.
Two limits are stated in the file rather than left for a reader to find. It
is a different bundle and a different run about 3.6 hours later, so it is
not a same-run arm. And it is #9803's range-active MUTANT arm: ordinary
English stays byte-exact even with that saved-range mutation live, which is
why it reads as a negative rather than as a baseline.
Bundle cited by directory with its file SHA-256; MANIFEST.sha256 verifies
21/21, rc=0. Nothing imported from .tmp.
* docs(terminal): correct #12164's grounds — the cited comments say no such thing
The rejection of #12164 from this file's family was recorded as resting on its
comment 1 (output doubling) and comment 2 (filed against 1.4.163). Checked
against the API: the issue has exactly two comments, neither of which says
that, and the string 1.4.163 appears nowhere in the thread.
The conclusion survives on better grounds. The issue BODY's repro is "Run any
CLI agent (Codex, AGY, Claude, etc.) that outputs Korean text into the Orca
terminal" — untyped output, no keystrokes, no composition — so excluding
CompositionHelper is right, and the input-path hunt was looking in the wrong
place. The body is also LLM-authored (it still contains a literal
"## 5. GitHub Submission Draft (Ready to Post)") and its Root Cause section
blames a CJK IME preedit buffer its own repro never engages, so it should not
be read as observation.
Comment-only; suite unchanged at 5/5.
* test(native-chat): make composition frames carry isComposing, not just inputType
This suite's comment claimed "Gating onChange on `isComposing` breaks here."
It did not. composeFrame() fired `input` with `inputType` but never set
`isComposing`, so a gate on `isComposing` passed all four tests untouched —
the suite asserted a discriminator it did not exercise.
Composition frames now carry both, so neither gate is exempt. Verified by
pointing the mutant at it: with an `isComposing` gate on the composer's
onChange, this suite now fails 3 of 4 (it passed 4 of 4 before), and the
ordinary-English arm correctly survives, since a composition gate should not
touch it. Production code is unchanged and stays gate-free; the mutation was
applied, measured, and reverted.
Found while excluding NativeChatView's question-card remount as the owner of
#12118 / STA-3219: the remount is real, but the preedit survives it precisely
because this write path has no composition gate.
* fix(mobile): ship the patched xterm build, matching desktop
mobile pinned @xterm/xterm 6.1.0-beta.285 while the patch is keyed to
6.1.0-beta.287, so mobile shipped stock xterm and neither IME defect fix
reached it: the swallowed keystroke after an IME commit (9506039de72) and
the preedit sized to the font rather than the grid (e04e0c88da5).
Bumps the three xterm packages to the desktop versions and adds the patch
to mobile's own pnpm.patchedDependencies. No copy of the patch: pnpm
accepts the parent-relative path and records it in the lockfile against
hash 8d63166272e9040a…, byte-identical to what desktop resolves, so the
two stay in step by construction rather than by a drift check.
The workspace separation is untouched — root pnpm-workspace.yaml still
declares `packages: []` and mobile keeps its own lockfile, which is what
keeps the root's patches from failing as ERR_PNPM_UNUSED_PATCH.
Verified in the generated webview bundle rather than at the install:
alignPreeditToGrid 0->2, sentComposition 0->3, pendingInput 0->11, and the
stock-only _handleAnyTextareaChanges 2->0 and dataAlreadySent 4->0. pnpm
applies patches during linking before postinstall regenerates the bundle,
confirmed by a revert/reinstall/re-apply cycle in both directions.
Mobile suite 2971 passed, 3 skipped — identical before and after. Bundle
+1,514 B (+0.24%). Lockfile churn is xterm-only; --frozen-lockfile passes.
mobile/src/ime/ime-submit-carry.ts is NOT made redundant and is untouched:
it handles iOS firing onSubmitEditing on a React Native native TextInput
after unmarking a composition, which is outside the WebView entirely.
Known divergence left alone: desktop also patches @xterm/addon-webgl and
mobile now runs that version unpatched. That patch is glyph/texture-atlas
rendering with nothing IME-related, so it affects neither fix.
* test(terminal): pin the preedit overlay against already-committed cells
STA-3132 (arm A), STA-3170 and STA-3232 report a Korean preedit painted on
top of text already on screen. Builds v1.4.163-v1.4.166 cancel the pending
finalizer in compositionstart, so a committed syllable reaches onData one
syllable late and buffer.x is stale — the overlay lands on the cell the
flushed syllable is about to occupy.
Replays a recorded hardware trace rather than an authored one: the ordered
DOM event stream captured on Windows + MS Korean (wave5-r2 evidence, 64
events), echoing onData back as PTY output.
The load-bearing assertion is deliberately not the obvious one. Comparing
overlay style.left against cursorX is tautological — left is computed from
buffer.x. This counts committed syllables from the compositionend events
the IME fired, so the two sides are independently derived.
Discriminated by a historical re-add across seven real bundles, since the
owner is deletion-shaped: pristine beta287, v1.4.155 and v1.4.162 pass;
v1.4.163 fails; v1.4.163 with that single call removed passes; the byte
identical baseline restored fails again; head passes. Every failing arm
fails only this case — the ordinary negative stays green in all seven.
The negative asserts its own category rather than claiming it: zero
composition events, zero isComposing, zero keyCode 229, exactly 16 events,
paired against the Korean arm's 4 starts / 3 ends / 11 updates / 64 events.
Scope: cell indices, not pixels. happy-dom has no layout, so the recorded
8x16 cell metrics are supplied to the render service. This makes no claim
about pixels visually overlapping; that is affected-OS confirmation and
stays open. Covers the overlap arm only — STA-3132's auto-line-break arm
and STA-3232's half-line-capacity and a11y arms are untouched.
* test(e2e): matrix macOS period substitution against the OS preference
#11504 reports macOS inserting ". " after a Hangul Space commit. This
sweeps six arms across both states of NSAutomaticPeriodSubstitutionEnabled,
reading the preference live per run rather than asserting a literal.
Two results worth having on record.
The reporter's stated trigger did not reproduce. Their words are "There is
no second press at all. One space is enough", but korean-single-space emits
zero insertText with the preference on or off. So does word-space-word-space.
Their timing does reproduce, with different content. korean-double-space and
korean-longer-word-double-space emit a delayed insertText at +122.5-122.7ms
after compositionend — squarely the reported +149ms — but the payload is a
space, never ". ". Consistent with the double-space rule seeing two slots
under ABC and only one under Korean, where the IME commit consumes the first.
The substitution itself is real and preference-bound: latin-double-space
gives "ab . " with the preference on and "ab " with it off, on one build
with the preference as the sole variable, reproduced across two runs.
That also refutes a claim in PR #11506, which states the substitution "is
enforced outside the renderer and never reproduces in dev builds, so changes
here must be verified against a packaged app". It reproduced in the dev build
twice and did not reproduce on the signed packaged app. That claim should not
be used as a verification gate.
Gated @headful behind ORCA_E2E_NATIVE_MACOS_PERIOD, same shape as the Korean
preedit spec, so it does not run in ordinary CI. Evidence is onData and DOM
only — the PTY-child reader aborted and no packaged-app arm was stable.
* test(terminal): actually enforce the recorded jamo progression
The preedit assertion compared sample.overlayText against sample.overlayText
— the same expression on both sides. A lane proved it by mutation: corrupting
seven of the eight recorded preedit values left the suite fully green. So the
docblock's ㄱ→가→간→나→낟→다→달→라, which the matrix also cites as this row's
recorded shape, was cited and unenforced.
The first attempt at a fix was insufficient and is worth recording. Threading
stroke.preedit through to the expectation still passed on a corrupted fixture,
because that value both drives the rig and was the expectation — corrupting it
moved both sides together. Same tautology, one level down.
The expectation is now an independent literal. Verified by mutation rather
than by reading: corrupting two recorded values fails one arm; restoring them
passes 3/3.
overlayCell was never affected — it is compared against a count derived from
the compositionend events, not from the buffer, and remains the load-bearing
assertion for the overlap.
* fix(e2e): select the selectable input source, not the first match
TISCreateInputSourceList can return several entries for one input source
id. A third-party IME publishes a non-selectable parent alongside the
selectable mode, and taking sources.first can return the parent — after
which TISSelectInputSource fails with paramErr (-50) while the caller
reports success from the enable step.
Found with Qingg (com.aodaren.inputmethod.Qingg), which exposes exactly
that pair under one id. Its mode id equals the bundle id, so filtering by
name would not have helped; selectability is the discriminator.
Now filters on kTISPropertyInputSourceIsSelectCapable and falls back to
the old behaviour when nothing advertises it, so single-entry sources are
unaffected. Also enables every entry for the id rather than only the one
being selected: selecting a mode whose parent is still disabled fails the
same way.
Compile-checked, and selecting com.apple.keylayout.ABC still exits 0.
Unrelated to the enable path: on macOS 26.5.2 third-party IMEs are gated
behind a consent sheet in System Settings. TISEnableInputSource returns
noErr immediately regardless, and the enable only lands if that sheet is
answered while the requesting process is still alive.
* test(terminal): discriminate #12171 against the real shortcut policy
The prior candidate mutation for this row was correctly refused: its suite
mocked shortcut policy so Process/229 became actionable, while the real
resolveTerminalShortcutAction has no Process branch — so the kill measured
the mock. This does not mock it.
Replays a capture of this row's own gesture (d, l, Shift+T, e, k, Space,
Enter under MS Korean, committing 있다) taken on Orca 1.4.164, through the
real useTerminalKeyboardShortcuts hook, capturing bytes at terminal.input.
The earlier capture could not discriminate at all because it recorded no
shiftKey; this one records it on 10 of 10 keydowns with code populated.
One physical Shift+T produces two shifted Process/229 keydowns. Under the
pre-#12265 classifier each synthesizes {key:'Enter', shiftKey:true}, which
the real policy resolves to sendInput '\x1b\r' — twice, giving 1b0d1b0d,
the two escapes the known-bad ed96881b0d1b0d contains.
Mutation is the retained pre-12265-process-shift.patch applied to HEAD, not
an authored one: patch -p1 applies clean and diffs identical to the mutant
copy. Arms are copies; shared source hashes the same before and after.
The English arm stays clean under both modules, so the mutation
discriminates by language rather than by harness — and a real Shift+Enter
through the same rig yields exactly ['\x1b\r'] in every arm, so a silent
pristine result means the code is quiet rather than the harness dead.
Scope: 1b0d1b0d is measured at the renderer boundary. The capture recorded
no PTY bytes — window.api.pty is frozen on shipped builds and the onData
channel needs a build-time flag — so this shows the renderer producing the
two escapes that payload contains, not a re-observation of the payload.
* docs(native-chat): narrow this file's disclaimer to what is now true
It said "THIS OWNS NO REPORTED ROW". Half of that is stale: the remount site
is now the attributed owner of #12118 and STA-3219. On real Windows TSF the
questionActive swap aborts a live composition — the old node gets only a
blur and no compositionend, the text returns as committed, and the next jamo
yields 아ㄴ rather than 안.
The other half holds. This file pins the opposite property, that the text
survives, which is the half those reporters already agree with. Mutation
shows the gap rather than asserting it: deleting the unmount entirely leaves
three of four tests green, because every substantive assertion is
after.value === … and a composer that never unmounts keeps its value.
Also records why the abort cannot be asserted here. The DOM exposes no
observable separating committed text from a live preedit — value is the same
string either way, there is no EditContext, and the only composing-ness
state is a per-instance ref discarded with the node. A test pinning "no
compositionend fires" would be an anti-guard: red the day it is fixed.
The cadence objection is kept, since it is now the open question rather than
the reason for exclusion.
* refactor(terminal): drop the unread isComposing field from XtermBypassEvent
Added by #6396 for terminal IME candidate handling that this branch has since
removed. No production or test code reads it, and the policy is safe without it:
during composition `key` is 'Process', so the non-ASCII printable checks that
would care never match.
Co-authored-by: Orca <help@stably.ai>
* fix(native-chat): keep the composer mounted through an in-flight IME composition
A question card replaced the composer outright
(`{questionActive ? null : <NativeChatComposer/>}`). Unmounting the field
mid-composition aborts the composition in the OS: the node is detached before
`compositionend` can fire, the preedit returns as committed text, and a resumed
Hangul syllable degrades — 아 then ㄴ yields `아ㄴ`, never `안`.
Confirmed in rasterised pixels on Windows with a real MS Korean IME, at both
v1.4.171 and the reporter-era v1.4.164 (the swap block is byte-identical
across them): the preedit underline present before the swap, the composer
visibly absent during it, and the same glyph back afterwards WITHOUT the
underline — committed, not composing.
The swap is now deferred while a composition is in flight, which is what
editors that survive IME do: ProseMirror gates DOM work on `view.composing`,
CodeMirror protects the composing subtree from redraws. Hiding instead of
unmounting does not work — `display:none` and `visibility:hidden` both blur the
focused element and abort the composition the same way.
The hold releases on `compositionend`, which browsers also fire on blur, so
clicking into the card's own answer input yields the input region immediately;
with nothing composing the card still replaces the composer at once, so no
stray "Send a message" appears beside a question.
The existing characterization test flips to a regression guard: it pinned the
node being destroyed, which was the defect. Node identity is the load-bearing
assertion — value-only checks are trivially satisfied by a composer that never
unmounts and cannot tell a held composition from a destroyed one.
The typing-redirect handler moves to its own hook. That is not cosmetic: both
touched files sat at the 400-line cap, and `max-lines` suppressions are
forbidden, so the room had to come from a real extraction.
* fix(macos): opt Orca out of AppKit automatic period substitution
macOS "Add period with double-space" (`NSAutomaticPeriodSubstitutionEnabled`,
on by default) is applied by AppKit's text input system. Native terminals never
join that system; Chromium text fields do, so xterm's helper textarea inherits
it and a double space arrives as `". "` — a period nobody typed, handed straight
to the PTY (#11504).
Chromium answers AppKit for quote and dash substitution and defaults both off,
but declares no period accessor at all, so AppKit applies that one without
asking. This user default is the only lever: there is no per-field or
per-webContents opt-out to prefer over it. Writing the key into Orca's own
defaults domain overrides the global value for this app alone and leaves the
user's system-wide setting untouched. It necessarily covers every Orca text
field, not only terminals — AppKit offers no narrower scope, and that tradeoff
is deliberate rather than accidental.
Measured on the reporter's own build v1.4.161: with the preference ON, typing
a,b,space,space yields `onData ["a","b"," ",". "]`; with it OFF the same arm
yields two spaces.
Note the issue's causal model is wrong and this fix does not follow it. It
claims the substitution only fires with a CJK input source and never with ABC.
The measurement is the inverse — every Korean arm is clean and the ABC arm is
the one that fires — so the fix is not conditioned on input source.
NOT YET VERIFIED ON HARDWARE. The unit tests inject the writer, so they prove
the call is made on darwin and skipped elsewhere; they do not prove AppKit
honours an app-domain override for this key. That check is outstanding.
* fix(xterm): keep a live composition across a lone Cmd press on macOS
CompositionHelper.keydown exempted keyCodes 16/17/18 from tearing a composition
down, which covers Shift/Ctrl/Alt but not macOS Meta — 91/93 in Chromium, 224 in
Firefox. A lone Cmd press mid-composition therefore reached
_finalizeComposition(false), which dropped the preedit overlay's `active` class
and committed the live syllable early. macOS keeps the marked text alive across
that press, so no later compositionstart re-arms the overlay and the rest of the
word composes invisibly.
Measured on hardware (m4air, macOS 26.5.2, Apple M4, 2-Set Korean) with the Cmd
posted as a CGEventType.flagsChanged, which is what a physical modifier emits.
AppleScript `key code 55` posts nothing a browser can see — a bare `key code 56`
for Shift is equally silent — which is why no capture in the corpus ever reached
this branch. Three arms, same build otherwise: overlay live throughout with the
exemption, dark and prematurely committed without it, live again with it
restored. Evidence under
.tmp/ime-handoff/swarm-scratch/wave31-cmd-preedit/evidence/.
The fix cannot widen past a lone modifier: only a standalone press reports these
keyCodes, and a Cmd chord during composition is reported by Chromium as 229,
which was already exempt. Cmd+A still ends the composition, via the IME's own
compositionend. Ghostty draws the same line, returning early from flagsChanged
under hasMarkedText() for every modifier including Super.
Orca's terminal pane was never affected — shouldSuppressTerminalModifierKeyboardEvent
drops a standalone Meta keydown before xterm sees it, and deleting only 'Meta'
from that set is what flipped the hardware arm to broken. The popout preview
terminal and mobile's webview install no such guard and did reach the teardown.
terminal-ime-xterm-composition-commit-overlap.test.ts asked its fixer to update
the two Cmd arms to the values it named as correct; both now emit a single ['한'].
* test(native-chat): drop two byte-identical duplicate cases
`4632b86919d` copy-pasted two cases twice into the same describe block:
`retains carry across a same-frame non-Enter keyup before redispatch` and
`expires carry before a deliberate Enter after the next frame`. Each pair is
byte-identical — same title, same body — so the copies asserted nothing the
originals did not.
This is what has been failing `static analysis` on this branch since 2026-08-06:
`oxlint vitest(no-identical-title)` reports both under `--deny-warnings`, and
`verify` fails solely because it requires static analysis to pass. Every other
gate in `verify` was already green, including typecheck, xterm patch sync, the
full test shard set, and both package jobs.
12 cases still pass in the file.
* test(e2e): skip the WebGL arm when no WebGL renderer exists
The #12164 probe runs two arms, webgl and dom, and closes by asserting the
active renderer is the requested one. That assertion is right for the dom arm —
it is what proves the pane actually left WebGL, without which the arm is
meaningless — but headless CI has no GPU, xterm falls back to DOM silently, and
the webgl arm then fails.
The failure reads as a Korean rendering defect and is not one, so the webgl arm
now skips with the active renderer named. The dom arm keeps the assertion
unchanged.
This is the third of three checks that have been red on this branch since
2026-08-06. `static analysis` and `verify` were fixed in c51c6b5837e; the CI log
shows this job as 1 failed / 1 passed, the pass being the dom arm.
* chore(lint): drop five unused no-console disable directives
`check-changed-code-quality` reports unused eslint-disable directives as errors,
and these five sat above diagnostic `console.log` calls in IME test and spec
files where `no-console` is not enabled — so each suppressed nothing.
This is the second of the two static-analysis steps. `c51c6b5837e` fixed
"Enforce focused code-quality plugins" (duplicate test titles); this fixes
"Enforce changed-code quality". Both had been red on this branch since
2026-08-06, and I mistook the first for the whole job.
The diagnostic logs themselves are kept — they are what a failing IME arm prints
for a reader to inspect.
* test(e2e): cover #12164 under fractional device scale factor
Fractional display scaling was #12164's last unexplored branch, and the reason
is worth recording: earlier attempts were BLOCKED, correctly, because they
proposed mutating the Windows display scale on a remote physical machine with no
console recovery. `--force-device-scale-factor` reaches the same renderer state
per process, so nothing outside the Electron instance changes and there is
nothing to restore.
The hypothesis was specific: `프프로로젝젝트트` is what a half-pixel cell boundary
could produce on a 2-column glyph, and nothing else in the suite varies dpr.
Measured at 1.25 and 1.5, both under WebGL: ink extents 25/21/16 with identical
ink groups, matching the scale-1 run. No doubling.
The arm self-certifies before asserting — if the flag does not take, the test
fails rather than silently measuring at dpr 1. That matters here: the sibling
spec's WebGL arm went two days reporting a missing GPU as a Korean rendering
defect precisely because a silent fallback looked like a result.
* fix(terminal): match Mod+letter shortcuts by physical key, not IME-rewritten key
With a CJK input source active, macOS and Windows report the physical key through
`code` but rewrite `key` to the layout's character: Korean 2-Set turns Cmd+C into
`{ key: "ㅊ", code: "KeyC", metaKey: true }`. Every `key.toLowerCase() === 'c'`
match misses it, so the shortcut is not recognised and xterm encodes the chord as
PTY input instead — issue #13033 reports `ESC[12618;9u` and a terminal that jumps
to the bottom, because user input scrolls the viewport.
This is the same key-vs-code confusion that owned #12171, where a `Shift+T`
typing ㅆ was read as Enter for want of a `code` guard, so the fix is the same
shape: trust `code` when it is present, fall back to `key` and then the legacy
`keyCode` when it is not (Chromium omits `code` on synthetic and some keypress
events, and `keyCode` keeps its US value even when `key` is rewritten).
Applied to the four terminal-side sites, including the dashboard pop-out, which
#13033 called out specifically as having its own key handler:
pty-connection.ts Cmd/Ctrl+C copy guard
keyboard-handlers.ts Cmd+G search navigation
agent-interrupt-inference.ts interrupt inference
preview-terminal-key-handler.ts pop-out paste
Nine further `key.toLowerCase()` letter matches exist outside the terminal
(TaskPage, editor, GitHub composer, browser markup). They have the same defect
and are deliberately left for a separate change rather than widening this one.
An existing case, `matchSearchNavigate > returns null for wrong key`, overrode
only `key` and left `code: 'KeyG'`, so it began passing for the wrong reason. It
now overrides both — which is what "wrong key" means once matching is physical —
and a companion case pins the Korean-rewritten chord still matching.
#13033 was closed NOT_PLANNED; the reporter's event shapes drive the new test.
* fix(renderer): match every Mod+letter shortcut by physical key, not IME-rewritten key
Completes the previous commit. A CJK input source rewrites `event.key` while
`event.code` keeps the physical key, so `key.toLowerCase() === 'z'` and friends
silently stop matching — the shortcut is not recognised and the keystroke falls
through to whatever handles unclaimed input.
The helper moves to `@/lib/ime-latin-shortcut-key` first: it now serves the
editor, GitHub composer and browser markup, and importing terminal-pane
internals into those would be the wrong direction. `lib/` already hosts
`ime-composition-keyboard-event` for the same reason.
Nine remaining sites, all previously unreachable under Korean/Japanese/Chinese/
Vietnamese input:
TaskPage, ActivityPrototypePage, ProjectViewWrapper Cmd+F search
useMarkupKeyboardShortcuts Cmd+Z undo
GitHubMarkdownComposer, RichMarkdownLinkBubble,
rich-markdown-link-shortcut Cmd+K link
native-chat-shortcut Cmd+J
rich-markdown-key-handler Cmd+Shift+X
Six of the nine test `!== 'letter'` as early-return guards and three test
`=== 'letter'`; the negation is applied per site, since a blind substitution
would have inverted six of them.
Full suite: 4249 files pass. Three files fail locally and none is caused by this
change — the branch touches no file under `src/main/` or `src/relay/`, all four
failures reproduce on an unmodified tree or pass in isolation (the worktree
poller passes 21/21 alone, so it is full-suite parallelism), and all 16 CI test
shards are green.
* docs(ime): scope the IME composition rules to the terminal-pane directory
#11893 proposed adding these to the root `AGENTS.md`, which every agent loads on
every task regardless of what it is doing. They only bind keyboard handling, the
composer and the terminal input path, so they belong next to that code —
`tests/e2e/AGENTS.md` already establishes the nested convention here.
Kept from #11893: range-derived commits, guarding above the key dispatch,
the `attachCustomKeyEventHandler` / `CompositionHelper` interaction, no
normalization at commit, and the recorded-trace evidence bar.
Added from defects found since it was written:
- match shortcuts on `event.code`, not `event.key` (#12171, #13033)
- `keyCode === 229` means an IME owns the press
- do not unmount a field mid-composition, and hiding is not a fix because
`display:none` blurs and aborts it too (#12118, STA-3219, #11332)
The evidence bar now also names the mutation check, since a test that survives
deleting the code it guards is guarding nothing — a failure this effort hit more
than once.
* fix(terminal): gate Ctrl+Enter CSI-u on a negotiated pane, porting #12462
Found while scoping the rebase onto `main`: #12462 landed on 2026-08-06 and
fixes a real defect this branch does not carry. Ctrl+Enter emitted
`\x1b[13;5u` unconditionally, so a pane that never negotiated the kitty
keyboard protocol — local Windows ConPTY, plain shell — printed the escape
verbatim into the prompt.
This branch deletes `terminal-ime-deferred-newline.ts`, which is one of the
files #12462 touched, so a rebase resolving those conflicts by taking our side
wholesale would silently reintroduce the defect. Porting it forward now means
the fix survives the rebase however the conflicts are resolved.
Mirrors the Shift+Enter guard already here: local ConPTY falls back to the
legacy CR every emulator sends, and a negotiated pane keeps the chord, so the
fallback is scoped to panes that cannot receive CSI-u rather than to Windows.
NARROWER THAN #12462 BY ONE CONDITION, deliberately. `main` also allows CSI-u
via `hasCtrlEnterCsiUAuthority()` (trusted consumer evidence, #12329); that
helper and its plumbing do not exist on this branch. Omitting it is the
conservative direction — an authorised pane gets `\r` instead of the chord,
rather than an unnegotiated pane printing an escape — but it should be restored
when the two histories are reconciled.
Test covers both directions and is mutation-checked: forcing the gate open
fails it, so it cannot pass by construction.
* fix: reconcile two more fixtures main moved while the stack waited
Both caught by CI, not locally, and the reason the local run missed one is
worth recording:
1. `browser-toolbar-profile-dialogs.ime-enter.test.tsx` did not pass
`useNativeUserAgent` / `onUseNativeUserAgentChange`, which `main` added to
`BrowserToolbarProfileDialogsProps`.
Local `pnpm typecheck` reported 0 errors on the same commit CI failed. The
cause was a stale `config/*.tsbuildinfo` — tsc reused an incremental cache
from before the merge. Deleting it reproduced CI's error exactly. Any
"typecheck clean" during this merge should be treated as unverified unless
the cache was cleared first.
2. Localization keys for `SshDisconnectedDialog` were absent from `en.json`:
the merge took this branch's component alongside `main`'s catalog.
Regenerated with `pnpm run sync:localization-catalog` rather than hand-added.
* fix(mobile): regenerate the lockfile the merge resolved by taking one side
CI's `verify` failed with `ERR_PNPM_OUTDATED_LOCKFILE` on `mermaid (lockfile:
11.16.0, manifest: 11.16.1)`. The mismatch was in `mobile/`, not the root — the
root lockfile was consistent throughout, which is why inspecting it (and even
GitHub's merge ref) found nothing wrong.
Cause: during the merge I resolved `mobile/pnpm-lock.yaml` by taking this
branch's side wholesale rather than merging it, so it kept `mermaid 11.16.0`
while `mobile/package.json` came from `main` at `11.16.1`. Taking one side of a
lockfile is only safe when the corresponding manifest comes from the same side.
Regenerated with `pnpm install --lockfile-only`; `--frozen-lockfile` now passes
in `mobile/`. Verified the xterm patch entry survives intact — same hash
`4f1b42d268f3964d…` and the parent-relative path into `config/patches/`, which
is the desktop/mobile coupling that would silently break the mobile build.
Two earlier diagnoses of this failure were wrong and are worth recording: it was
not the root lockfile, and it was not a stale merge ref (a rerun reproduced it
exactly).
---------
Co-authored-by: Orca <help@stably.ai>
Terminals froze app-wide several times daily, needing a manual pkill. libuv
unlinks the pathname a server bound to when it closes, with no ownership check,
so a departing daemon deleted whichever socket then sat at the canonical path —
including a live replacement's. The replacement kept hosting PTYs no client
could reach.
#12709 fixed that mechanism; this replaces the shape around it. Two invariants:
only a daemon publishing itself onto the canonical endpoint may mutate that
entry, and only by replacing one it has itself just proven dead; and no actor
removes a name it did not create.
Publish binds a private name, takes the canonical one with an exclusive link,
and on EEXIST proves the incumbent dead by connecting before replacing it in a
single rename. Only 'connected' means occupied and only refused/missing prove
death — a timeout proves nothing and declines. Deletes the claim sweeper, the
reclaim tail of killStaleDaemon, and three unfenced unlinkSync(socketPath)
calls in the launcher.
Measured: rename exposed no gap across 6,525 darwin / 8,004 linux probes of a
live handover, where unlink-then-link gapped on 200 of 200.
Verified on all three platforms: full suite on macOS and Linux, and daemon
restart e2e on a real windows-2022 host. Contract in src/main/daemon/AGENTS.md.
* Persist the Linear issue list view and per-workspace filters
Layout, grouping, ordering, columns, and attribute filters survive a restart.
Facet ids are workspace-scoped, so filters are kept per Linear workspace and the
active filter is *derived* from the selected workspace rather than reset by an
effect on switch — no ordering race can apply workspace A's facets to B, and an
unresolved or cross-workspace selection reads as unfiltered without erasing
anything.
A single shared catalog backs the renderer state, `TaskResumeState`, and the
strict `ui.set` schema, so a new view option cannot leave paired web/mobile/relay
clients rejecting the whole payload. Persisted values are normalized as untrusted
input: a corrupt preference or a single bad workspace entry is dropped without
taking the rest of the resume state with it.
Deriving the filter also removed the guard that used to make three neighbouring
behaviours safe, so they are re-scoped here:
- The primary-team facet reset now fires only on an in-workspace team change.
A workspace switch also changes the primary team, and clearing there wiped the
filter that had just been restored for the workspace being switched *to*.
- The list-read force check no longer fires on the session's first read, so a
restored filter serves warm cache instead of forcing a network round trip
behind a blocking spinner on every cold start.
- The filter dropdown derives "no single workspace" from `workspaceId` alone.
With an unresolved workspace it previously rendered the statically populated
priority section, whose clicks now have nowhere to be stored.
* Harden Linear view persistence against the failures review surfaced
Five issues, each found by a reviewer and reproduced before fixing:
- The filter dropdown's prune effect only ran when the user opened the popover,
because the filter was always empty at startup. Restoration makes it run on
mount, where `availableTeams` may still be the issue-scraped fallback rather
than the real fetch. Metadata complete for a *partial* team set passes every
R12 guard, so it pruned facets belonging to teams it simply hadn't seen — and
the write persisted, deleting them permanently. Gated on `teamsSettled`.
- `canonicalize` dedupes but enforces none of the transport bounds; only the
throwing parser does. So `serialize` could emit a 101-label filter that the
strict `ui.set` schema rejects, which drops the WHOLE taskResumeState — github,
jira and linear query included — on every subsequent write, since the renderer
resends the merged object each time. Added `boundLinearIssueAttributeFilter`
and a round-trip test built from serializer output rather than a literal, which
is the only kind that can catch renderer/schema drift.
- `linearIssueView` now carries `.catch(undefined)`: value tolerance stops at the
top level, so any future instance of the above is a cosmetic reset of the view
instead of silent loss of every other resume field.
- A workspace switch forced an uncached list read in both directions. The switch
is a later observation, so the null-baseline fix didn't cover it; the cache is
already workspace-keyed, making the force pure cost.
- Recency for the 20-workspace cap came from object key order, which is wrong
twice: re-filtering an existing workspace left it at the head (first evicted,
though just used), and an array-index-like key enumerates first regardless of
insertion, so a write could evict the very entry it added. Recency is now an
explicit ordered key list.
Also adds the nested parity assertion — the top-level one compares only
TaskResumeState's own keys, so a field added to LinearIssueViewResumeState stayed
invisible to it, which is exactly what `.strict()` rejects.
The wiring test was blind: deleting the hydration guard outright left all four
assertions green. The gate is now `shouldPersistLinearIssueView`, unit-tested
directly, and the file is renamed to the repo's `*-boundary.test.ts` convention
with an assertion that fails on that mutation.
* Log discarded Linear views and fix empty-filter serialization
- Schema now logs when linearIssueView is discarded, making validation failures visible
- Fixed serialization: filters that become empty after bounding are now omitted
- Added AssertNoExtraKeys type check for bidirectional schema/type parity
- Refactored view option catalogs to use canonical constants, preventing UI/schema drift
* Remove workspace persistence limits and LRU eviction
Stop capping persisted Linear workspace filters at 20 and evicting
least-recently-used workspaces. Simplify persistence to store all
workspace filters, gate persistence only on resume state application,
and remove tests that pinned implementation details. Users can now
persist filters for all their workspaces without arbitrary limits.
* add test for linear persistence
* Improve Linear filter test clarity and fix e2e overlay dismissal for CI
- Convert parameterized filter-pruning test to sequential assertions
- Fix dismissOverlayChrome to toggle overlay triggers instead of
force-clicking inert page elements in headless CI
* Prevent TaskPage from stealing Escape from Radix menus
- Add check to detect open Radix dropdown menus and popovers; return
early from Escape handler to respect their capture-phase ownership
- Update overlay dismissal in e2e tests to use keyboard.press('Escape'),
now that TaskPage no longer interferes
* The capture-phase Escape guard in TaskPage bailed out for open dropdown menus and popovers, but an open Radix Select matches none of those selectors: the shared SelectContent wrapper (src/renderer/src/components/ui/select.tsx:60) renders data-slot="select-content" and Radix gives its content role="listbox", not role="menu". So with a select open, the window-level capture handler ran first, called preventDefault() and closeTaskPage() — closing the whole task page instead of just the select. Added [data-slot="select-content"] to the guard, as suggested. I did not add [role="listbox"]; the reviewer explicitly notes it's too broad, and the data-slot selector covers every select rendered through the shared wrapper.
---------
Co-authored-by: m4air <m4air@MacBook-Air.localdomain>
Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
Co-authored-by: Jinjing <6427696+AmethystLiang@users.noreply.github.com>
* Extract GitHub task search commit debouncing into a hook
Move debounce logic from TaskPage into useGitHubTaskSearchCommit to
prevent excessive GitHub API calls on every keystroke. Uses a 750ms
idle window before committing search values. Add tests for the new hook.
* Keep task rows visible while typing search query
Removed premature row-hiding logic from the search input handler that was
triggering before debounced queries fire. The handler now only updates the
input state; debouncing and query timing are handled by a dedicated hook.
Added e2e test verifying search idles before fetching and Enter doesn't
double-fetch.
* Test that GitHub task search commits cancel on disable and unmount
Verify the useGitHubTaskSearchCommit hook properly cleans up pending
commits when disabled or when the component unmounts. This prevents
unnecessary GitHub API calls during normal user interaction.
Also fix e2e test instrumentation to find the active repo through the
worktree relationship rather than assuming the first repo with a path.
* feat(ai-vault): validate session-delete targets for single-file providers
Add the pure judgement layer for deleting an Agent Session History entry.
`validateAiVaultSessionDeleteTarget` decides whether a session may be removed:
the agent must be one of the nine providers where a single file is the whole
session (gemini, copilot, cursor, hermes, devin, openclaw, droid, pi, omp),
the host must be local, and the renderer-supplied path must resolve inside
that agent's own session roots and match its discovery predicate.
To keep the delete roots from drifting from the scanner's own roots, the
WSL-expansion helper moves to session-scanner-root-dirs.ts and the OpenClaw
root derivation + session predicate become shared helpers that
discoverOpenClawFiles itself consumes.
The result is path-only and never touches the filesystem; a returned
`allowed: true` still requires an lstat/realpath re-check in the executor
(S-2) before removal, documented as a caller contract on the result type.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XDLggjSAjDnaWi3Y8U622i
* feat(ai-vault): move a validated session transcript to the trash
Add the filesystem executor behind session deletion. It calls the S-1 path
validator, then performs the fs-side guards that validator documented it
could not: lstat().isFile() rejects a directory or symlink, and realpath is
re-fed through the validator so a regular file reached through a symlinked
parent that escapes the agent's roots is rejected too. Only then is the file
moved to the OS trash via shell.trashItem, with ENOENT treated as success so
a delete racing an external removal stays idempotent.
WSL UNC paths (no Recycle Bin) are delegated to tryDeleteWslUncPath before the
Windows-local fs guards, mirroring fs:deletePath. Any non-ENOENT error is
returned as a failure result rather than thrown, since IPC payloads are
untyped at runtime.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XDLggjSAjDnaWi3Y8U622i
* feat(ai-vault): delete-session IPC handler, preload bridge, cache invalidation
Wire the S-2 delete executor to an IPC endpoint and expose it on the preload
bridge. The renderer calls aiVault:deleteSession with { agent, filePath,
executionHostId }; the handler fetches WSL homes, delegates to the executor
(which re-validates and trashes), and on a real delete invalidates the caches
that could otherwise keep serving the deleted session.
Cache invalidation is generation-guarded: a scan already in flight when the
delete lands carries an older generation and must not write its pre-delete
result back into the cache. Without this, an in-flight scan resolving just
after the delete would resurrect the deleted session for the 15s TTL — and
force-refreshing the panel only masks it for the desktop, not for the paired
mobile client or runtime RPC that share the same cache module. Both the shared
local-scope cache and the desktop multi-host cache carry the guard, with
regression tests for the in-flight race.
The delete result type moves to shared/ai-vault-types.ts so the renderer can
import the same contract the executor returns. To keep ai-vault.ts within the
max-lines budget after adding the delete wiring, two cohesive pieces are
extracted to their own files: the delete orchestration (ai-vault-delete.ts)
and listAiVaultSubagentSessions (ai-vault-subagent-list.ts). The latter is the
only handler with no dependency on this module's private cache state, so it is
the one piece that moves verbatim without threading state through a seam.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XDLggjSAjDnaWi3Y8U622i
* feat(ai-vault): renderer judgement for whether Delete is offered
Add the renderer counterpart to the main-side delete validator: given a
session, decide whether the row menu shows Delete enabled, or disabled with a
reason a tooltip can render. It reuses the shared deletable-agent set and
unsupported-reason map so the two sides can never disagree about which agents
are deletable, and reuses the existing local-host / synthetic-path renderer
helpers.
This is intentionally not a security boundary — it validates neither the path
root nor the file predicate. Those are the main process's untrusted-input
defense; the renderer only picks the affordance, and the main side re-checks
on delete regardless.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XDLggjSAjDnaWi3Y8U622i
* docs(ai-vault): correct deletability parity claim; test multi-reason agent
The renderer deletability check runs host -> synthetic -> agent, while the
main validator runs agent -> host -> synthetic. The two layers agree only on
deletable-or-not (renderer-false is a subset of main-false), not on the reason
code a doubly-failing session carries. Document that explicitly instead of
implying the orders match, and add the antigravity case (two reason codes) so
the agentReasonCodes array shape is actually exercised.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XDLggjSAjDnaWi3Y8U622i
* feat(ai-vault): add Delete to the session row menu with a confirmation dialog
Wire the delete affordance into AI Vault. Both the dropdown and the context
menu gain a destructive Delete item; a session that can't be completely deleted
(remote host, synthetic OpenCode-SQLite path, or a directory/registry-backed
agent) shows the item disabled with a reason surfaced both as a tooltip and as
an aria-label so keyboard and screen-reader users learn why. Confirming opens a
dialog that names the session and states it will no longer be resumable from
the provider's own CLI, then calls the delete IPC and force-refreshes the list
for immediate feedback (the main side has already invalidated its caches).
The confirmation copy says the session "will be deleted" rather than "moved to
the trash": on Windows a WSL session is deleted with rm inside the distro (no
Recycle Bin), so promising recoverability would be a lie on that platform.
Deletability is computed once per row and shared by both menus so they can
never disagree. New pure logic — the reason-to-tooltip mapping (including the
multi-reason join) and the delete action hook's deleted/rejected/failed
branches — is covered by unit tests.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XDLggjSAjDnaWi3Y8U622i
* fix(ai-vault): state that Delete is unavailable without naming the cause
The disabled Delete item explained a provider's storage layout to the user
("Claude sessions can't be deleted here: stores sessions as a folder, not a
single file"). That is Orca's problem, not the reader's — the tooltip now says
which sessions are affected and stops there. The non-local-host string stays as
it was: it states scope, not a cause, and tells the user what would work.
The reason-code plumbing existed only to compose that tooltip, so
AI_VAULT_UNSUPPORTED_DELETE_REASONS, AiVaultUnsupportedDeleteReasonCode, and the
renderer result's agentReasonCodes field go with it. Why each agent is excluded
moves into the comment above AI_VAULT_DELETABLE_AGENTS, where a reader looking
up the deletable set will find it.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGuzChimmQ1dYX2raecrH7
* feat(ai-vault): delete claude, rovo, and grok sessions by their directory
These three were excluded only because the delete unit was one file. Their
sessions are directories — claude keeps Task subagent transcripts in a sibling
`<uuid>/subagents/`, rovo and grok keep everything under `<sessionId>/` — and
nothing in them is shared with another session, so a directory-aware delete is
still a complete delete. Supported goes from 9 agents to 12; the four that
remain (antigravity, kimi, codex, opencode) are blocked by a registry or a
SQLite row, which no delete unit fixes.
Validation now returns an ordered removal plan instead of a single path. Each
removal carries the kind it must be on disk and the roots its realpath must
stay inside, so the executor's guard is the same shape for a file and for a
directory. Companions come first and the transcript last: the transcript is
what puts the row on screen, so a part-way failure leaves the row to retry
from rather than dropping it and stranding the rest on disk.
Claude's `session-env/<uuid>/` goes with the transcript — it holds that
session's generated shell exports and nothing else. Its sibling
`file-history/<uuid>/` deliberately does not: it is the rewind buffer holding
earlier versions of the user's own files, and retiring a session is no reason
to take away the only copy that can restore them.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGuzChimmQ1dYX2raecrH7
* fix(ai-vault): remove a claude session's own directory, not just its subagents
Deleting a claude session trashed `<uuid>/subagents/` and left `<uuid>/` behind
as an empty directory — one per deleted session, accumulating under every
project. The directory is named after the transcript, so it belongs to that
session as a whole; take it rather than the one subdirectory inside it. Still
derived from the scanner's own subagents path, so the two cannot drift.
Reaching the parent means a degenerate stem now matters: `..jsonl` passes the
extension check and its stem is `.`, which would resolve the session directory
to the project directory holding every session. Reject it instead.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGuzChimmQ1dYX2raecrH7
* fix(ai-vault): keep a session row collapsed when a menu action is chosen
Radix portals the row's dropdown and context menus out of its DOM, but React
still bubbles their clicks back through the component tree, so every menu
selection also hit the row's own click handler and expanded it. The trigger
button already stopped propagation, which is why opening the menu looked fine
and only choosing an item misbehaved.
It shows worst on Delete: the row expands behind the confirm dialog, so
cancelling leaves the list rearranged under a dialog the user just backed out
of. Toggle details only for clicks that land in the row's own subtree — that
covers the context menu and any future portalled surface, not just this one.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGuzChimmQ1dYX2raecrH7
* fix(ai-vault): harden the delete-confirmation flow against IPC rejection and mid-delete dismissal
Two robustness gaps flagged in review:
- handleConfirmDelete only branched on result.outcome. The main handler
resolves with a 'failed'/'rejected' outcome rather than throwing, but the
IPC invoke itself can still reject on a transport/serialization error, and
the caller fires it with `void`. That reject would surface as an unhandled
rejection with no toast. Catch it and show the same generic failure toast.
- handleDialogOpenChange cleared sessionPendingDelete on every open=false.
The Cancel button is disabled mid-delete, but Radix still fires its
Escape/outside-click/X close, which could dismiss an in-flight delete out
from under itself. Ignore close requests while deletingSession is true.
Both covered by regression tests (verified failing without the fix).
* fix(ai-vault): route WSL UNC directory removals through the WSL rm branch
Directory-shaped deletes (claude's subagents/session-env dirs, rovo/grok's
session dir) gated the WSL branch on kind === 'file', so on Windows a session
under a WSL distro home fell through to shell.trashItem — which can't trash a
WSL-volume item (no Recycle Bin) and throws, or worse is silently stranded when
the 9P filesystem's unreliable lstat false-reports ENOENT and the executor
treats that as success. Single-file deletes predate the directory kinds, so the
file-only gate was correct until directory removals were added.
tryDeleteWslUncPath already supports recursive removal; pass recursive for
directory removals so they take the same WSL rm path as files instead of
shell.trashItem. Covered by two regression tests (file: non-recursive,
directory: recursive), verified failing without the fix.
Also drops the internal ledger-ID references (D-*, S-*) from comments in these
two files; they pointed at a private design doc a reader can't see.
* docs(ai-vault): drop internal design-ledger IDs from shipped comments
Comments across the session-delete feature cited decision/slice IDs (D-1..D-7,
S-1..S-5) from a private design document. Those references are meaningless to
anyone reading the code without that doc, so remove the IDs while keeping the
reasoning each comment carried. No behavior change.
* test(ai-vault): e2e-cover the real on-disk session delete
The unit tests mock lstat/realpath/trashItem, so nothing proved the whole IPC
path actually removes files. This spec seeds sessions into the E2E harness's
isolated HOME and deletes them through window.api.aiVault.deleteSession:
- a single-file session (gemini): the transcript is gone from disk and drops
out of the list.
- a directory-shaped session (claude): the transcript, the <uuid>/ session
directory (subagents included, no empty shell left), and the session-env
companion are all gone, while the file-history rewind buffer is preserved.
Verified failing when the executor's removal is stubbed out. Runs on Linux CI.
* fix(ai-vault): address review findings on the session-delete flow
Three points raised in review:
- Disable Delete for a still-running session. resolveAiVaultSessionDeletability
now gates on liveState (working/blocked/waiting) last — an otherwise-deletable
session that is mid-run shows "wait for it to finish" instead of an enabled
Delete, so trashing a live agent's transcript can't drop writes it is still
appending. Unsupported/remote sessions keep their permanent reason.
- Realpath the roots, not just the target, in the executor's escape check. The
roots were only resolve()'d (text), so a session under a symlinked root
(~/.claude -> /Volumes/…) was falsely rejected; realpath each root (falling
back to its text form when it can't be resolved) before the membership check.
- Invalidate the parse cache with the raw filePath, not resolve(filePath). The
cache is keyed by the exact path the scanner discovered, so resolve() could
normalise it away from the stored key and miss. Drops the now-unused import.
Also moves AiVaultDeleteSessionArgs/Result out of ai-vault-types.ts (which the
upstream merge pushed over the max-lines limit) into the ai-vault-session-deletion
domain module they belong to, and updates importers.
Regression tests added for the live gate, the symlinked-root accept, and the
reason string; verified failing without each fix.
* fix(ai-vault): type the deleteSession preload bridge as its real result
The bridge declared Promise<unknown> while AiVaultApi.deleteSession promises
AiVaultDeleteSessionResult, so the preload object leaned on the api-types
declaration to stay honest instead of being checked against it.
Co-authored-by: Orca <help@stably.ai>
* refactor(ai-vault): tighten the session-delete code to house style
Comments across the delete flow explained HOW alongside WHY and ran to a dozen
lines; they now carry only the non-obvious reasoning. The excluded-agent
rationale, the caller contract on the validator, and the file-history carve-out
are kept — those are knowledge, not narration.
Also removes three duplications the feature introduced:
- AiVaultSessionDeleteExecutionResult was an alias for AiVaultDeleteSessionResult
whose comment pointed at a module the type no longer lives in.
- The synthetic-path predicate existed twice under near-identical names; the
renderer now re-exports the shared one it already had a sibling import of.
- The delete-failure toast was written out verbatim in both the rejected and
the thrown branch.
Co-authored-by: Orca <help@stably.ai>
* refactor(ai-vault): use a design-system dialog width and a stable row selector
The confirm dialog pinned an arbitrary sm:max-w-[440px]; every other dialog in
the right sidebar uses a scale token, and md (448px) covers the role.
The row-expand test selected the row by [draggable="true"], which stopped
naming the row when draggable moved to the title element upstream. It still
passed by bubbling, so the comment was the only thing wrong — now it selects
the title deliberately and says why the query is first-match (Radix's asChild
trigger repeats the subtree, so screen.get* sees duplicates).
Also types the e2e delete helper as AiVaultDeleteSessionResult instead of a
hand-written { outcome: string }, now that the preload bridge returns it.
Co-authored-by: Orca <help@stably.ai>
* refactor(ai-vault): consolidate agent sources and use system dialog
Discovery and deletion now share the same agent source definitions, eliminating the risk of them drifting apart. A single `AI_VAULT_AGENT_SOURCES` table declares each agent's root directories, file extensions, and acceptance predicates. Replaced the custom delete confirmation dialog with the system dialog, simplifying the delete action hook and removing boilerplate state management.
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Jinjing <6427696+AmethystLiang@users.noreply.github.com>
Co-authored-by: Orca <help@stably.ai>
* fix(terminal): defer metric option writes to unmeasurable panes
Writing fontSize/fontFamily/fontWeight/lineHeight makes xterm re-measure
cell size against the pane's current box. A hidden or mid-layout pane can
measure a wrong-but-nonzero size, which latches (hasValidSize) and mis-keys
the shared WebGL glyph atlas until a manual resize — the stuck variant of
the P0 bold/blurry-font reports.
Metric writes now land only on measurable panes; otherwise the latest
values park per-pane and flush on the next safe fit or reveal (with a refit
on the light tab-resume path, which otherwise skips fitting). Measurability
helpers move to pane-fit-measurability.ts to stay under the pane-fit.ts
line cap.
* fix(terminal): key metric deferral by terminal, not pane view
getPanes() returns a fresh toPublicPane() wrapper per call, so a
WeakMap keyed on ManagedPane never matched across call sites: deferred
metric options were dropped, not deferred. Key on pane.terminal, which
is carried by reference and dies with the pane.
Also from review:
- flushDeferredPaneMetricOptionsIfMeasurable checks the pending WeakMap
before the measurability probe, so the common no-deferral case costs
zero forced style/layout on every reveal.
- applyTerminalAppearance skips the apply (and the probe) when all five
values are already live and nothing is parked; any settings write
re-runs the pass over every mounted pane, and arming a no-op deferral
would trigger a refit on the next reveal.
- fitRevealedPane flushes first: its pixel/grid checks can both no-op
and return without fitting, stranding parked options.
- Font zoom folds its direct fontSize write into any pending deferral so
the flush inside safeFit cannot clobber the user's zoom.
Corrects comments that asserted a cell-size re-measure mechanism xterm
does not have: CharSizeService measures via OffscreenCanvas TextMetrics,
independent of the pane box, and only fontSize/fontFamily re-measure.
Test fixtures now allocate a fresh pane view per getPanes() call, which
is what production does and what hid the keying bug.
* fix(terminal): re-check the fit floor after a metric flush
performSafeFit evaluated the min cols/rows gate with the pre-flush cell
size, then flushed and fit unconditionally. A large font jump on a
narrow pane passes the gate at the old size and lands under it at the
new one, so fit() pinned the PTY to the tiny grid the floor exists to
reject. Re-check after a flush that actually landed.
The parked values still apply, so the pane is never stuck on stale
metrics; only the fit is skipped.
* fix(terminal): route a reveal metric flush through the stable fit
fitRevealedPane's new flush branch called safeFit directly, which is
exactly what the function's contract forbids on reveal: resumeRendering
has just re-attached WebGL, whose cell metrics transiently differ from
the DOM renderer's, so a raw fit can propose a one-column-off grid and
reflow — and xterm's wrap/unwrap is not a perfect inverse, leaving a
diff-painting inline TUI corrupted.
A landed flush leaves pixels unchanged with a diverged grid, the same
shape as a snapshot resize, so it takes the same steady-grid repair.
A real resize still fits synchronously, after the flush.
Reachable via window wake, which calls fitAllRevealedPanes with no
pre-flush loop.
* fix(terminal): gate metric writes on the pixel box, not the fit floor
canApplyPaneMetricOptions reused canMeasurePaneForFit, whose >=8 cols /
>=4 rows floor exists to stop a fit pinning the PTY to a sliver. But the
divider clamp is 50px, which clears the 48px pixel floor and proposes
~5 cols — so a pane dragged to the clamp deferred every font change and
never flushed: it never hides, and its box never changes, so no reveal
and no ResizeObserver entry ever arrives. It rendered a stale font until
widened, where pre-PR the write was unconditional.
Gate metric writes on display plus the pixel box only. Hidden panes and
the transient worktree-switch overlay are near-zero, so they still
defer — the deferral's purpose is unchanged. The cols/rows floor stays
on the fit, including the post-flush re-check in performSafeFit.
Apply and flush share the same predicate, so no "applies but never
flushes" state can open up.
* fix(terminal): flush heavy reveal metrics after WebGL resume
Wake idle Run coordinators with durable orchestration mail pointers while keeping message payloads in the store until check consumes them. Preserve waiter, Cursor, restart, real Codex title, and PTY replacement behavior.\n\nPart of #12953.
* fix(terminal): gate Ctrl+Enter CSI-u on a negotiated kitty pane
Ctrl+Enter emitted \x1b[13;5u unconditionally, so a pane that never
negotiated the kitty keyboard protocol (local Windows ConPTY, plain
shell) printed the escape verbatim into the prompt. Mirror the
Shift+Enter guard and fall back to the legacy CR every emulator sends
for this chord. Keeps the intercept, so IME commit ordering and the
single-send dedupe still apply.
Fixes#12329
Co-authored-by: Orca <help@stably.ai>
* test(e2e): negotiate kitty via PTY output in the Ctrl+Enter spec
The Ctrl+Enter gate reads the PTY-output kitty tracker, which
enableKittyKeyboardReporting never feeds (it writes straight into
xterm's parser), so the spec pressed the chord on a pane the policy
still saw as un-negotiated and got the CR fallback. Negotiate from the
application side like the neighbouring Shift+Enter spec, and reset the
flags afterwards for the serial suite.
Co-authored-by: Orca <help@stably.ai>
* fix(terminal): preserve trusted Ctrl+Enter routing
* fix(terminal): scope IME redispatch ownership
* fix(terminal): reject conflicting Ctrl+Enter evidence
---------
Co-authored-by: Orca <help@stably.ai>
Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com>
* fix: show full paths in quick open results
* refactor: use native file path tooltips
* fix: position file path tooltips
* refactor: share the cursor path tooltip with quick open
Co-authored-by: Orca <help@stably.ai>
* fix: let path tooltips run wider before wrapping
Co-authored-by: Orca <help@stably.ai>
---------
Co-authored-by: Orca <help@stably.ai>
* fix(terminal): clear stranded link hover tooltip
* fix(terminal): declare the tooltip reserve var where it resolves
--orca-terminal-link-tooltip-height was declared on .pane-manager-root, a
class no live element carries, so both .xterm-container height calc()s were
invalid at computed-value time and collapsed to height:auto — the element
FitAddon measures, making rows a fixed point.
Also isolate _clearCurrentLink() so a throwing provider leave() cannot skip
the cache invalidation, and bound the e2e gap assertion on both sides.
Co-authored-by: Orca <help@stably.ai>
---------
Co-authored-by: Orca <help@stably.ai>
Mouse events posted with CGEventPostToPid reach the target app with no
window association, so AppKit never routes the press to a view: hover
states fire but the control is never activated, and the mouseUp is
dropped outright when posted back-to-back. Post click events to the HID
event tap instead (as keyboard synthesis already does), pace them, and
stamp mouseEventClickState so multi-clicks register.
Synthetic clicks now also report verification unverified/synthetic_input
from the helper itself, matching the other synthetic actions.
Stacked on the #12793 revert. Widens the passive-identity set from
{inactive} back to {active, done, inactive}, so the PR/check glyph
returns to the left status lane for workspaces that are actively being
worked, not just idle ones.
Tradeoff, deliberate: #12658 was not purely a regression. It also fixed
#8813, where an active workspace with branch identity and no PR showed
the grey branch glyph instead of the emerald Active dot. This revert
reintroduces that, and removes its e2e guard.
The left lane holds one glyph, so activity, branch identity, and review
status cannot all be shown. This picks review status.
`orca serve` publishes a ready graph under HEADLESS_RUNTIME_WINDOW_ID with no
BrowserWindow behind it. `shouldCreateInBackground` only degraded when the
create was renderer-backed, so any focus-requested create fell through to
getAuthoritativeWindow() and threw "No renderer window available" — leaving
`terminal create --focus` with no workaround on a remote server (#10333).
With a worktree selector and no renderer window, a background spawn is the only
usable path, so collapse the renderer-backed window check into a plain
"no window" check. That is the existing rendererBacked clause plus exactly the
missing focus case, and it drops the confusing `rendererWindow === null`
indirection (rendererWindow is already gated on rendererBacked).
Focus is not lost by the degrade: the spawned pane is still published to the
session-tab model and revealed with `activate: true`, which is how a paired
client learns about it. Mirrors the in-tree precedent in
runCreateMobileSessionTerminal.
Headed hosts are unaffected — the clause only fires when no window exists.
Co-authored-by: Jinwoo-H <Jinwoo-H@users.noreply.github.com>
* Reorder source control to show staged changes first by default
Stages are closest to the commit action and most relevant to the
commit workflow. Merges untracked files into Changes visually while
preserving their Git area. Removes the untracked-first preset and
includes migration logic for existing user settings.
* Drop source control group order user preference
Remove the sourceControlGroupOrder setting and related UI, migrations, and persistence logic. The source control view now always displays sections in the order: staged changes, unstaged changes, untracked files.
* Reorder source control to show changes before staged
Aligns with the edit-stage-commit workflow by showing unstaged
changes (active edits) before staged changes (queued for commit).
- Replace 'Send answer' with 'Submit' for clarity and consistency
- Update all locale translations (en, es, ja, ko, zh)
- Remove fixed button width and add whitespace-nowrap for flexible sizing
- Update component and test references
PR 9501 shipped real-home routing for the host system default, and the
env override that could turn it back off was never a shipped control. The
managed-account half of the shared runtime mirror has been unreachable
since: every host account routes to its own self-contained CODEX_HOME
before that code runs.
Delete the flag module and its env plumbing plus the managed branch of
syncForCurrentSelection and the six helpers only it called. The three
lanes that still use the shared mirror -- Windows, a custom CODEX_HOME,
and a hook-lane gate that reports unusable -- are untouched, as are every
legacy migration and the WSL read-back helpers.
With a desktop client paired to a remote Orca runtime, terminal panes could report connected, writable, and `terminal.send` returning accepted — yet keystrokes never reached the agent. No error, no banner, no recovery; input silently vanished.
The ticket was really two bugs. The attach half was already fixed by #12589 (subscriber-driven daemon attach), confirmed by reproducing against current main. This fixes the remaining half: a write the host refuses had no way to tell anyone.
A capability-negotiated `WriteUnavailable` opcode carries that refusal back to the client, where it feeds the pane's pre-existing recovery hook. Capability gating matters because decoders reject unknown opcodes on desktop — and, worse, silently drop them on mobile — so the signal is negotiated in the subscribe handshake. Verified per direction: an old host strips the unknown Subscribe key, an old client omits it so the host never emits, and capability cannot be inherited across resubscribe.
Independent review then found the signal was being delivered and discarded: recovery demanded an authoritative liveness answer, and `pty:hasPty` had no `remote:` guard, so a paired pane's id fell through to the LOCAL provider, which returned false, and recovery bailed before remounting. Every test stopped at the transport boundary, so all of them passed while the pane stayed just as stuck. `pty:kill` already had exactly that guard.
The fix makes main answer LESS rather than claim more: `pty:hasPty` now returns unknown for a `remote:` id instead of a fabricated false, because main cannot speak for another host's PTY. The remount is then authorized by positive evidence — the process that owns the PTY stating it refused this specific write over a live negotiated connection — not by inference from silence. Local and app-SSH ids keep the probe, where a false genuinely means the shell died. Nothing is destroyed on this path; the remount rebuilds the renderer over the session it already had.
An end-to-end test now carries a rejected write from the host through to an actual remount, which no prior test did. A surviving mutant was also killed: the legacy-binary capability gate could previously be deleted with nothing turning red.
The reliability gate stays experimental — live paired journeys and mixed installed-release evidence remain uncollected. Fixes STA-2830.
Mixed versions are the normal state of the remote-server feature: users update clients and servers independently. Until now nothing tested that. Every cross-version claim was made by code reading plus unit tests with hand-written old/new shapes — enough to catch design problems, not enough to catch a real skew regression.
This runs the REAL protocol implementations from two builds against each other in one process: the actual host methods and RPC dispatcher on one side, the actual renderer multiplexer on the other, with a transport that reproduces the production asymmetry — each side decodes with its OWN codec and drops frames whose opcode it does not know. A frame survives only if the RECEIVING build understands it, which is what makes this level sufficient without launching two apps. The old side is a genuine checkout extracted from the release tag; the extracted client was confirmed to lack a symbol that exists only on main.
Journey: subscribe, first snapshot, input reaching the process, live output, hide/reveal snapshot, transport drop, resubscribe, input landing again — across old->new, new->old, and a current/current control. Every step ends on an observed-state barrier; no sleeps. The oracle asserts the recorded step list, the exact 16-frame named sequence, negotiated capabilities, the exact input the host wrote to the PTY, rendered content, and zero decoder-rejected frames. A host method the stub lacks is recorded by name and asserted empty, so a harness gap cannot masquerade as a wire break.
Detection is proven per violation shape, and it attributes each to the correct side: an unnegotiated opcode goes red only where a decoder would reject it, a removed published field goes red only where an old client consumes it, and a legal additive field stays green in all three pairings so the harness will not cry wolf on safe changes.
It also documents the three compatibility rules in docs/reference/remote-wire-compatibility.md, linked from AGENTS.md, since they previously existed only as folklore — notably that "decoders reject unknown opcodes" is true for the desktop decoder but NOT for mobile, which silently drops them.
Deliberately scoped: terminal stream only. The session-tab sync channel is not covered, nor agent-session publications, file/Git RPCs, mobile E2EE framing, or the relay transport. Two version points, so a regression introduced and reverted between them is invisible.
CI selection was verified rather than assumed — `vitest list` confirms 0 matches under the shard's exclude and 4 under the dedicated job — because a lane silently running zero tests is precisely how a host-side defect escaped CI earlier in this series. Closes STA-3469.
* Add host and project filtering to the worktree jump palette
Filter the search results by execution host (local, SSH, runtime) and project/repo, with a drill-down options menu, chips for active selections, and overflow hints for large result sets. Filters reset on open to prevent silently hiding results, and stale selections auto-prune. Caps rendered rows per section to prevent DOM bloat from single-character queries. Host badges appear when filtering to clarify which rows survived the cut.
* Add E2E test for worktree jump-palette host filtering
Tests filter interaction via keyboard, filter/project intersection, empty state when filters exclude all results, and ephemeral filter reset on modal close.
* Fix worktree jump-palette filter persistence and search UX
- Persist filter state when filter model changes to prevent dropped IDs from silently re-activating
- Fix result count to distinguish between query matches (all items) vs. empty list (capped sections)
- Improve filter field options: use state for scroller to handle unmount/remount, clamp highlight index to valid range, only reset on query/field change, not re-ranks
- Context-aware space key: allow toggle in listbox only, not in search input (preserves for typing)
- Replace generic "Clear field" translation with field-specific strings to preserve capitalization in non-English languages
* fix(cmd-j): reset filter highlight without prop-change effect
Derive the active option index from field/query identity instead of
resetting it in a useEffect so React Doctor and first-paint stay correct.
* fix static analysis issue
* fix(e2e): rename project entity in jump-palette filter seed
Filter options use project.displayName when a Project exists, so only
renaming the repo left the local option labeled with the path basename.
Reconnect could never recover a terminal pane whose host-side process was gone (host restarted, or the workspace was never opened there): recovery only polled the tab inventory, which can never create the surface it is waiting for, so Reconnect spun for ~60s and gave up permanently.
Verified with a deterministic reproduction: on main the recovery path issues 51 inventory polls and zero activations across both an automatic online trigger and a manual Reconnect click; with this change the pane re-materializes, rebinds and accepts input.
Review found and fixed three further defects beyond the original change:
- an activation answered with a stale ready handle left the loop polling forever instead of re-activating;
- a non-missing activation failure (e.g. an older host without the method) never fell back to inventory;
- host-side, activating a parked surface permanently deleted the host tab, because an already-absent persisted binding was read as a competing owner *after* the destructive retirement had already run.
Independent review confirmed by mutation testing that every production change is covered by a test that fails when it is reverted, that only an authoritative inventory can retire a pane, that the loop is bounded under every failure mode, and that the unknown-liveness guard (proven death required before retirement) is intact.
Fixes STA-3002.
* test(diff): repro for STA-3420 combined-diff invalidation freeze
Co-authored-by: Orca <help@stably.ai>
* Fix diff-view freeze when large diff invalidated by rebase writes
Staged-diff sections now reload in-place on external file changes instead of remounting every visible Monaco editor and bumping the virtualizer generation, which wedged the renderer during rebase bursts.
* test(diff): calibrate STA-3420 burst assertions against an idle baseline
The burst window's peak lag is dominated by a one-off stall from opening 8x15k-line
Monaco editors, which reproduces identically with invalidation disabled. Measure an
equal-length idle window first and assert p95, sample coverage, and lag relative to
that floor. Adds unit coverage for isUnchangedDiffSectionReload.
Co-authored-by: Orca <help@stably.ai>
* fix(diff): keep renderedIndicesRef pure during render
React Doctor blocks ref mutation during render; sync the on-screen
section set in a layout effect instead so static analysis can pass.
* Fix unchanged diff-section reload detection for truncated diffs
When a diff exceeds render limits, content is pruned to '' for memory.
The old check compared content equality, so limited reloads always
appeared changed, triggering unnecessary revalidation that froze the UI.
Compare render-limit metadata instead — it's the sole change signal
and full description of what the fallback banner displays.
Also calibrate STA-3420 e2e assertions relative to idle baseline for
machine independence instead of absolute thresholds.
* fix(diff): defer invalidation reloads for in-flight stale-token loads
When a diff section is invalidated while a large-diff load is in-flight:
- Don't delete the in-flight load from loadingIndicesRef, since a newer load may own it
- Bump the reload token but defer the reload if there's still an in-flight load
- Let the in-flight load settle first, then reschedule the reload at settle-time
- Prevents the freeze by avoiding race conditions that leave sections stuck loading
This fixes STA-3420 where rebase-driven invalidations could hang the diff view.
* test(diff): relax STA-3420 burst assertions to inclusive comparisons
Switch from strict inequality checks (toBeLessThan, toBeGreaterThan) to
inclusive variants (toBeLessThanOrEqual, toBeGreaterThanOrEqual) to allow
measurements landing exactly on the threshold boundaries.
---------
Co-authored-by: Orca <help@stably.ai>
* fix(status-bar): invalidate the CLI session count on kill and restart
`pty:management:killOne` / `killAll` / `restart` tear sessions down via `adapter.shutdown()` and broadcast nothing — unlike `pty:kill`, which ends in `sendPtyExitToRenderer`. The status-bar count is an event-sourced cache, so killing sessions from Manage Sessions or "Kill all terminals" left the `>_ N` chip frozen until the popover was opened, which itself triggers a refresh.
> [!NOTE]
> The dual-source split described in the issue text was already fixed by merged #9387. This closes a *different* remaining invalidation gap that produces the same reported symptom.
Broadcast the teardown so the chip updates without needing the popover opened.
Fixes#8372
Co-authored-by: Orca <help@stably.ai>
* test(e2e): add recordable proof for status-bar-cli-session-count
Fails on origin/main, passes on this branch.
Test: drops after Manage Sessions kills a foreign daemon session, popover never opened
Co-authored-by: Orca <help@stably.ai>
* fix(status-bar): avoid duplicate inventory refresh after kill all
---------
Co-authored-by: Orca <help@stably.ai>
* fix(setup): stop caching an unreadable orca.yaml as "no setup script"
`checkRepoHooks` returned `{hasHooks:false, hooks:null, mayNeedUpdate:false}` with no `status` field when the SSH filesystem provider was unavailable, and inside a blanket catch for any read error. The renderer only bails on `status === 'error'`, so that status-less false negative was cached as an authoritative "no setup script" and the prompt stayed on screen.
Mirror the `hooks:check` IPC twin exactly: `status:'error'` for a missing provider, ENOENT-aware in the catch, `status:'ok'` on the folder-repo, binary, SSH-success and local branches.
Fixes#8752
Co-authored-by: Orca <help@stably.ai>
* test(e2e): add recordable proof for setup-script-prompt-false-negative
Fails on origin/main, passes on this branch.
Test: recovers from an unreadable orca.yaml instead of pinning the failed verdict
Co-authored-by: Orca <help@stably.ai>
---------
Co-authored-by: Orca <help@stably.ai>
* perf(renderer): bail out of identity-equal terminal layout and cache-timer writes
setCacheTimerStartedAt and setTabLayout spread a fresh object and returned it
unconditionally, so every redundant call published a new AppState and ran every
zustand subscriber's selector across all mounted panes. Both have a real
redundant cadence: parked-terminal-byte-watcher writes a null cache timer on
each agent working/exit/stale-title transition, and TerminalPane re-persists an
identical layout on pane-title churn.
Extract the existing terminalLayoutEqual comparator out of web-session-tabs-sync
into a shared module and use it to gate the layout write, and dedupe the
remote-runtime layout IPC against the last snapshot pushed per tab.
* fix(renderer): retry failed remote pane layout pushes
* test(renderer): cover stale remote layout failures
* test(e2e): cover remote pane layout retry
* fix(native-chat): show Claude's AskUserQuestion card when the agent runs on a paired headless host
Three gaps kept the question card off the desktop when the agent ran on a
remote `orca serve` host:
- The `session.tabs` projection reduced HTTP agent-hook rows to identity only,
hard-coding `state: 'done'` and an empty prompt, so `toolName` and the full
`interactivePrompt` never left the host. It now publishes the newest fresh
hook row's status fields, bounded by the same staleness window `agentType`
uses, excluding `providerSessionOnly` resume rows, and yielding to live
title evidence unless a question is actually pending.
- Nothing republished `session.tabs` when only a hook row changed, and the
re-emit carried an unchanged `snapshotVersion` that clients drop on their
monotonic gate. Material hook transitions and pane/SSH status clears now
bump the version and schedule a coalesced emit.
- The desktop card resolved only from live status. It now falls back to the
pending ask in the transcript, matching mobile, so a relay gap can no longer
leave the composer mounted over a pane parked on a selector.
Closes#11761
Co-authored-by: Orca <help@stably.ai>
* fix(native-chat): date the hook-row recency guard against a real clock
`resolveHookLiveAgentRow` compared a hook `receivedAt` (epoch ms) against
title stamps that are title-observation sequence numbers, so the guard could
never fire — any fresh hook row overrode live title-derived state, and a manual
rename (the one epoch writer) inverted it. Stamp the live OSC title path with
wall-clock ms and compare against that alone.
The regression test fabricated epoch-valued title stamps production never
writes; it now drives the title through `onPtyData`, and a new case pins the
opposite direction (hook row newer than the title wins).
Co-authored-by: Orca <help@stably.ai>
* fix(native-chat): stop an orphaned tool call from pinning a dead question card
extractPendingAsk pairs tool results to calls by a global FIFO (tool_use_id
is dropped at decode time), so one call that never gets a result desyncs the
queue for the rest of the transcript and strands an answered ask as pending.
Real transcripts also hold asks the user escaped and typed past. On desktop
that card replaces the composer, so the pane became unsendable.
Drop in-flight calls at a turn boundary — a user turn or the decoders'
interrupt row — since the turn that owned them is over. Claude's tool-result
turns decode as role 'tool', so normal FIFO resolution is untouched.
Co-authored-by: Orca <help@stably.ai>
* refactor(native-chat): trim the headless AskUserQuestion projection
Reuse rather than restate: the invalidator now takes the shared
`AgentHookEventPayload` instead of a locally redeclared row shape, and the
hook live row is a `Pick<>` of the retained OSC snapshot so one projection
branch consumes either carrier. Fold the immediate/coalesced session-tabs
emit into one method (also drops a redundant re-emit on the
provider-session push). Drop card tests that re-route shared-parser
assertions through React. Isolate pane-status-clear subscribers and prove
the no-republish case by version arithmetic instead of a timed silence.
Co-authored-by: Orca <help@stably.ai>
* test(native-chat): pin the AskUserQuestion card render under real Electron
Why: the 13 parser unit tests pin extraction, but nothing proved a card
actually renders where an inert tool call used to. This spec reproduces the
paired-headless topology from the client side — live status carrying agent
identity and state 'working' but no interactivePrompt/toolName, with the
pending ask present only in the transcript — and fails on main.
Refs #11761
Co-authored-by: Orca <help@stably.ai>
* test(native-chat): drop the unused testInfo parameter
Why: oxlint no-unused-vars fails the lint gate on an unused test parameter.
Co-authored-by: Orca <help@stably.ai>
* test(native-chat): drop leftover proof scaffolding from the ask-card spec
The env-var screenshot label and the fixed 2s settle only existed to make
the pre-fix capture comparable; the card assertion already waits.
Co-authored-by: Orca <help@stably.ai>
* test(runtime): use a truly unresolvable pane key in the hook republish guard
#11203 taught pane lookup to recover a reminted tab id by leaf id, so the old
fixture (new tab id, live leaf id) resolved and bumped the snapshot a second
time once this branch merged with main.
Co-authored-by: Orca <help@stably.ai>
* fix(runtime): refuse a hydrated unconfirmed hook row as live pane status
#12346 landed on main after this branch was cut: a nonterminal row restored from
last-status.json is stamped `restoredUnconfirmed` because its transition may have
fired while no receiver was up, and every freshness gate treats it as never-fresh.
The new headless `live` projection here only checked `receivedAt`, so a restart
inside the 30-minute window would republish the hydrated row — resurrecting the
AskUserQuestion card with no agent left to answer it.
`agentType` still reads those rows: they prove identity, just not liveness.
Co-authored-by: Orca <help@stably.ai>
---------
Co-authored-by: Orca <help@stably.ai>
Co-authored-by: Neil <nwparker@users.noreply.github.com>