* Revert "test(ime): restore coverage the composition-ownership change removed (#13168)"
This reverts commit 25a8c517e1.
* Revert "refactor(terminal): return IME composition ownership to xterm (#13128)"
This reverts commit 17b3dff3c4.
* test(ime): keep the architecture-neutral Korean trace coverage
The recorded IBus/fcitx5 and Windows MS-Korean traces from #13168 assert PTY
byte order, not composition ownership, so they still hold once the terminal
composition layer is restored. The mobile accessory-order test pinned the new
handleLiveInputChange signature and does not.
Co-authored-by: Orca <help@stably.ai>
* fix(terminal): keep the macOS Backslash bypass through the revert
The restored native-text forwarder only claims keys for input sources in its
hardcoded CJK allowlist, so third-party IMEs off that list (Qingg, #10896) still
get a raw backslash. #13128 added this bypass as a partial replacement; keep it
rather than trade the open issue back.
Scoped to the bare backslash key. The rest of shouldBypassXtermForMacNativeText
bypassed all unmodified non-ASCII text, which would race the restored forwarder.
Co-authored-by: Orca <help@stably.ai>
* fix(mobile): move the mirror-step ref write out of render
The restored hook assigned runMirrorStepRef during render, which is not
replay-safe — React can discard render work, so the mutation can leak from UI
that never commits. Its only read is inside the held-commit timer, which fires
long after commit, and the ref has a safe default, so an effect is soon enough.
Surfaced by the changed-lines React Doctor gate: the rule postdates this code,
so restoring the file re-introduced it as a new violation.
Co-authored-by: Orca <help@stably.ai>
---------
Co-authored-by: Orca <help@stably.ai>
* test(terminal): pin the recorded Korean commit-before-newline order (STA-3132)
Recorded first-party on Windows 11 + Microsoft Korean (HKL 0412) against the
defect-era v1.4.164 build, with bytes read on the far side of the PTY: the
terminal received ea b0 80 0d, the syllable strictly before the CR.
The capture did not reproduce the suspected deferred-newline inversion. That
route needed a session end carrying dataPendingReconciliation, which plain
compose-then-Enter cannot produce because the IME finalizes first and the
newline is never held; back-to-back arms at 25/60/120 ms did not reach it
either. The test therefore pins the ordering rather than discriminating a fix.
Co-authored-by: Orca <help@stably.ai>
* test(terminal): restore Hangul back-to-back flush coverage deleted with the composition layer
#12278 fixed a Hangul syllable that was not flushed before the next
composition began — the force-end path, and the one that leaves stale glyphs
behind. Returning composition ownership to xterm deleted both that patch and
its test, so nothing guarded the behavior any more.
Replays the recorded back-to-back arms (25/60/120 ms, read as 가\r나 at the
PTY) against a real xterm Terminal. It passes on main: stock xterm flushes the
committed syllable natively, so the removal was safe rather than a silent
regression.
Co-authored-by: Orca <help@stably.ai>
* test(mobile): pin accessory-byte ordering behind a Hangul commit
Returning composition ownership to xterm deleted the accessory-input commit
tests along with the hook they targeted, but the guarantee they protected is
user-visible and still applies: an accessory-bar keystroke must not overtake
the syllable being committed, and must be suppressed when that commit fails.
Drives the current hook with an Android composing-region trace rather than
reconstructing the deleted coordinator.
Co-authored-by: Orca <help@stably.ai>
* test(terminal): replay recorded IBus and fcitx5 Hangul traces offline
Commits interleaved with ASCII (한abc글) are the Linux IME gesture users report
on, and its failure modes are a lost syllable and a doubled one. That gesture
was only covered by tests/e2e/terminal-linux-ime-native.spec.ts, which needs a
Linux host running a real input framework.
Fixtures are the recorded captures from the sealed linux-final evidence run,
replayed against a real xterm Terminal: exact onData, exactly-once counts across
five repetitions, and the PTY bytes the recorded run actually received.
Co-authored-by: Orca <help@stably.ai>
---------
Co-authored-by: Orca <help@stably.ai>
* fix(terminal): return IME composition ownership to xterm
* fix(mobile): derive terminal input from native replacement ranges
* test(mobile): record iOS Japanese IME traces
* fix(mobile): preserve native IME replacement ranges
* fix(xterm): flush queued application input after IME commit
* test(terminal): pin Korean intermediate commit
* test: pin Windows IME shortcut ownership
* test: replay IBus number candidate commit
* fix: preserve native macOS input-method punctuation
* refactor(terminal): remove stale mac focus override
* fix(mobile): preserve soft keyboard deletion ranges
* fix: keep IME-owned palette chords in renderer
* fix: stop carried IME shortcuts at renderer owner
* fix: preserve carried IME shortcut dispatch
* fix: narrow main-owned shortcut actions
* test(mobile): pin Japanese IME replacement traces
* test(terminal): retain paired native IME trace
* fix(chat): preserve browser IME composition ownership
* fix(chat): retain macOS IME confirm gesture
* fix(chat): expire unmatched IME confirm carry
* fix(chat): isolate IME confirmation expiry
* fix(chat): retain active IME confirmation
* refactor(terminal): remove dead composition handler
* feat(ime): add shared Enter-ownership seams for CJK composition
The confirming Enter of a CJK composition arrives as two keydowns and the
orderings differ by platform: Windows/Linux redispatch the unmarked Enter/13
before keyup, macOS delivers keyup first. A guard reading only isComposing or
keyCode 229 misses the redispatch, so surfaces submitted on a confirm.
Adds useImeEnterGestureOwnership (carry token, next-frame expiry), a shared
ImeEnterGuardedForm for native implicit submission, and the cmdk seam covering
18 CommandInput surfaces at one site.
A chorded Enter arms the carry but is never swallowed — the reverse would eat a
user's deliberate Cmd/Ctrl+Enter. Both failure modes are pinned by
ime-enter-gesture-ownership-contract.test.ts.
Co-authored-by: Orca <help@stably.ai>
* refactor(terminal): consolidate native input listeners and parked-screen owner
Extracts the shared native-input listener installer and renames the parked-screen
detector for what it actually does, replacing per-call-site duplication. The
listener installer keeps a forgetOptionKeyLocationOnBlur flag so per-window
semantics are preserved rather than flattened.
Net deletion; no behaviour change intended.
Co-authored-by: Orca <help@stably.ai>
* test(terminal): pin recorded IME shapes as regression tests
Nine regression tests built from hashed affected-platform captures, each with a
paired ordinary negative and a discriminating mutation verified to take the file
from all-passing to exactly one failure.
Covers the Windows MS-Korean Shift family (#12179, #11878, #12151, #11946,
#12152) and the Korean TUI line-break rows (STA-3237, STA-3222, STA-3129).
STA-3237 pins the empirical 3-Shift / 2-active-composition / 2-newline ratio the
device run established — the third Shift produces nothing because Space has
already committed. That ratio is not derivable from a static capture.
Co-authored-by: Orca <help@stably.ai>
* fix(ime): guard Enter-commit surfaces against CJK confirm
Applies the Enter-ownership guards across the surfaces whose Enter commits
something: publishes, clones, pairs, installs, posts, or persists.
Tiered deliberately rather than uniformly. Irreversible and remote-effect sites
take the carry token, which also blocks the unmarked redispatch. Locally
reversible sites take the oracle check with a one-line comment naming the
residual, because a spurious commit there costs one undo.
Three numeric fields are left unguarded with the reason in-code: Chromium blanks
number inputs at compositionstart, so a confirm-Enter only ever reaches an
empty-draft reset. Measured with a CDP probe rather than assumed — a guard that
cannot fire is noise.
Co-authored-by: Orca <help@stably.ai>
* test(ime): teeth-check the Enter guards on every guarded surface
One suite per guarded surface, each verified by deleting the guard and
confirming the test fails. A green guard test without that check is unverified,
not verified.
Two shapes pass vacuously in happy-dom and are avoided here: native implicit
form submission never fires, and blur() is inert on an unfocused element. Both
made "the commit did not happen" assertions pass with the guard removed, so the
suites assert the guard's contract directly instead.
Co-authored-by: Orca <help@stably.ai>
* fix(mobile): keep iOS Korean commits whole through the live-input path
iOS Korean reports isComposing: false on every event, so it bypasses the
composition guard entirely. The strict owner rejected UIKit's transformed
post-change field and sent only the leading jamo — the reported symptom.
Prefers the authoritative same-event field text over the predicted text when the
supplied operation cannot produce it. Generic: no Korean special-case, no locale
classifier, no normalization. Adds the RN-target-keyed submit carry alongside it.
Co-authored-by: Orca <help@stably.ai>
* test(e2e): make IME capture harnesses fail loudly instead of silently
Four instruments recorded silence as success, so a void run scored as a clean
one:
- readTerminalImeBoundaryTrace returned an empty trace when the probe never
installed, making every "nothing leaked" negative pass vacuously
- summarizeLatencies([]) returned a perfect zero distribution that passed all
three latency thresholds
- the macOS Vietnamese spec pinned an input-source ID that does not exist, and
failed as though the operator had chosen the wrong source
- the expectedLineCount=1 prefix property was undocumented and one edit from
silently downgrading a PTY assertion
Input sources now resolve by enumeration and name the near-matches on failure.
Co-authored-by: Orca <help@stably.ai>
* test(terminal): cover Cangjie cancellation and fix a cross-namespace assertion
Adds #11951's recorded Cangjie cancel shape to the existing cancellation suite,
which covered Pinyin and Sogou but not Cangjie. One keystroke then Backspace
arriving as deleteContentBackward with data: null, so the stale preedit is the
only thing a fallback could replay.
Verified against the historical pre-6cd944c62b3 bundle: the positive fails with
['尸'] where [] is expected, while the ordinary negative stays green.
Also fixes the Vietnamese spec, which asserted a TIS-space input-source ID
against getKeyboardInputSourceId(). Those two Orca APIs report the same source
in different namespaces — TIS nests it under VietnameseIM, the app API does not.
The resolver stays as an installation precondition; the assertion matches the
leaf.
Co-authored-by: Orca <help@stably.ai>
* test(e2e): add a real-IME macOS arm for the Korean chord commit
The existing korean-ime-terminal-shift-enter-commit spec synthesizes composition
over CDP: Input.imeSetComposition sets the preedit directly and Input.insertText
performs the commit. Asserting the IME produced events you injected yourself is
circular, so that spec cannot certify real-IME behaviour.
This arm selects 2-Set Korean via TIS, reads it back live, and injects through
System Events key codes, so the OS owns the preedit, the commit instant, and
isComposing. PTY byte expectations are preserved verbatim.
Covers 2 of the original 4 cases by design. The other two are the Windows/Linux
redispatch-before-keyup ordering, which macOS cannot produce and which cannot be
selected -- the OS decides it. Reintroducing synthesis to "restore coverage"
would reintroduce the circularity.
Co-authored-by: Orca <help@stably.ai>
* test(e2e): assert the macOS chord arm at the PTY boundary, not the renderer
The byte expectations were transcribed from korean-ime-terminal-shift-enter-commit
:364/:383, which assert against onData -- a renderer boundary where the terminator
is CR. This spec reads the PTY child, where the tty has already converted CR to LF.
Names both forms per row rather than swapping the constant, so the conversion reads
as evidence that the capture reached past the renderer, as #11936 and #11951 record.
Ctrl+Enter's CSI-u sequence is unaffected and is identical at both boundaries.
Co-authored-by: Orca <help@stably.ai>
* test(e2e): measure composer-to-onData latency and stop dropping IME keystrokes
Two defects in the echo latency probe.
It hooked onWriteParsed and onRender but never onData, so it measured
key->parse->render echo rather than the composer-vs-onData delta the latency rows
need. Adds a third hook feeding its own sample set.
And `event.key.length !== 1` silently dropped IME keystrokes: Pinyin and Cangjie
keydowns arrive as key:'Process' (length 7). Replayed over the captured corpus,
the old filter accepted 580 of 4137 Chinese IME keydowns -- it was discarding 80%
of them. The new filter matches the shape the owner itself branches on.
Attribution charges each onData to the latest keydown rather than a FIFO head,
because composing jamo emit no onData at all and a queue would credit a whole
composition to its first keystroke. The consumer now asserts sample count before
any percentile, so a zero-sample run cannot render as a flawless distribution.
Co-authored-by: Orca <help@stably.ai>
* test(terminal): pin the WSL shifted-jamo newline shape for #11919
In Korean 2-set, Shift types ordinary letters -- the double consonants and the
compound vowels. Each such keystroke reaches Chromium as key='Process',
keyCode=229, shiftKey=true.
The v1.4.163 classifier matched exactly that pattern with no code guard, so it
called those keystrokes Enter, rewrote them to a synthetic Shift+Enter, and
injected a newline into the middle of the word -- with no Enter key pressed.
That is why the reporters said "no modifier key pressed": they had not chorded
Shift+Enter, but they had pressed Shift, to type the double consonant.
Asserts the row's own recorded capture: 40 immediate keydowns, exactly 3 of them
Shift-carrying inside a single syllable, and an onData stream with one newline
per Enter press and none mid-word. Two ordinary negatives keep it from being a
blanket mute -- the same session's non-IME keydowns still reach shortcut policy,
and an ordinary Shift+Enter still resolves through the real policy.
Co-authored-by: Orca <help@stably.ai>
* test(terminal): pin the composition commit lag that made Korean type one behind
macOS Korean 2-Set commits syllable N only when the first jamo of N+1 arrives, so
compositionend and compositionstart land in the same task. A composition-start
handler cancelled the pending finalizer that was the only path to triggerDataEvent
and ended the session without emitting bytes, so every committed syllable reached
onData exactly one syllable late and the backlog cleared only at a Space or Enter.
Types continuously with no Enter and no Space -- either would flush the backlog and
hide it -- and samples onData at every syllable boundary. Paired with a
length-matched ASCII arm that stays green throughout, so the positive is a fact
about composition rather than about timing in general.
Bisected to a single call site across five builds: pristine, 1.4.155 and 1.4.162
pass, 1.4.163 fails, removing the one call repairs it, restoring it fails
identically. That window is exactly the reporter's "started immediately after
updating".
Co-authored-by: Orca <help@stably.ai>
* test(mobile): cover the send-queue abort that silently drops queued keystrokes
One failed send in use-terminal-live-input-commit aborts every keystroke
queued behind it, with the error swallowed by .catch(() => false). The
existing test resolves(true) on every send, so the failure branch was
uncovered.
Four arms: the abort itself, an ordinary negative on the healthy path, a
throwing sender, and a liveness control proving the queue recovers once
the chain settles. Deleting the abort takes 4 passed to 3 failed, with the
ordinary negative correctly surviving.
Scope is stated in the docblock: this is a transport send-queue abort,
reachable only via a real disconnect or RPC error. REQUEST_TIMEOUT_MS is
30s, so latency alone cannot reach the branch — consistent with #7094's
symptom class, not proven to be its cause.
* test(terminal): pin that daemon snapshot/restore cannot disturb a composition
Two independent reporters attributed broken Korean composition to the
always-on PTY daemon repainting terminal state over the preedit. The
attribution is wrong on ancestry — the daemon shipped three months before
the version both call good — but the boundary was never actually tested.
Runs the real applyMainBufferSnapshot choreography against a live
composition, including the full 2J/3J/H wipe plus the resize and
alt-screen branches. textarea.value, selectionStart/End,
compositionView.textContent and .active all survive byte-identical, and
interleaving a restore between every jamo of 문제 still commits 문제 at
onData. Also pins that the uncommitted preedit is absent from the captured
snapshot: it lives in the textarea, never the buffer, so a restore has
nothing stale to echo back.
Injecting one textarea.value = '' into the restore fails exactly the three
restore-boundary tests.
* test(terminal): pin that Cmd tears down a composition where Ctrl and Shift do not
xterm's composition keydown exempts only keyCode 16/17/18 (Shift/Ctrl/Alt)
plus 20/229. macOS Meta — 91/93/224 — is absent, so a Cmd press mid-composition
takes _finalizeComposition(false): the overlay goes dark and never recovers,
because compositionstart is not re-fired. The user composes the rest of the
word blind. Linux and Windows users press Ctrl and are exempt.
xterm already has a Meta-aware modifier predicate in wasModifierKeyOnlyEvent,
so this is an internal inconsistency rather than a deliberate choice.
Owns no reported row and is version-neutral: 5/5 on both 1.4.162 and 1.4.163.
The branch is unexercised in all 328 recorded traces, so this is a hazard pin,
not a regression guard. Only the teardown is asserted; the likely duplicated
commit needs a compositionend the IME kept alive across the Cmd, which no
capture contains.
Deleting the exemption fails exactly the three paired negatives; adding Meta
to it fails exactly the two Cmd arms.
* test(native-chat): characterize preedit loss when a question card replaces the composer
An AskUserQuestion card fully replaces the composer by design, but the
in-flight composition goes with it: the composer unmounts before
compositionend reaches it, so the preedit is never committed to the draft.
The committed text survives only because the draft is cached and restored
via defaultValue. Node identity changes, value 'abc' is preserved, the 가
is gone.
Drives the real NativeChatView -> SessionGate -> InteractiveCard ->
questionActive swap -> Composer -> ComposerField, flipped by writing the
same store field an AskUserQuestion hook event writes. Flipping
questionActive to false fails exactly this test and nothing else across
639 native-chat tests, so the path was entirely unguarded.
CHARACTERIZATION TEST: it asserts the loss. Fixing the defect — committing
the preedit before the swap, or keeping the composer mounted — will make
this file fail. Update the expectations to the new contract rather than
working around them.
Owns no reported row. #12118/STA-3219 flicker is keyed to token counters,
which provably do not remount, and a question card arrives once per
question.
* test(terminal): pin the duplicated commit when Meta interrupts a composition
_finalizeComposition(false) sends textarea.value.substring(start, end) but
cannot clear the IME-owned textarea, so a later compositionend re-sends the
same range. Meta reaches that path because CompositionHelper exempts only
Shift/Ctrl/Alt; xterm's own wasModifierKeyOnlyEvent covers Meta four ways,
so the omission is an internal inconsistency rather than a choice.
Companion to the modifier-exemption guard, which deliberately pins only the
overlay teardown. This pins the data consequence.
HAZARD PIN: owns no reported row. The trigger is unverified on hardware —
no capture in the corpus contains a Meta-during-composition gesture, and
whether macOS keeps the composition alive across it is unmeasured. The
duplication follows from the code given that sequence; whether users reach
the sequence is the open half.
An earlier premise that Space (keyCode 32) reaches this path was refuted by
a corpus scan: 0 of 731 evidence files carry a keyCode-32 Space while
composing, against 171 at 229, and 229 returns early.
* test(terminal): characterize the syllable lost when the textarea blurs mid-composition
CoreBrowserTerminal._handleTextAreaBlur clears the helper textarea
unconditionally — "Text can safely be removed on blur" — while
CompositionHelper._finalizeComposition reads the committed text back out of
that same value from a deferred timeout. By the time it runs the value is
empty, the substring is '', and triggerDataEvent never sees the syllable.
xterm checks composition state in _syncTextArea and omits the same check
here.
Six cases. Blurring mid-composition loses the syllable in every ordering,
including compositionend-before-blur, which is Chromium's real order — so
it is not an ordering artifact. A bare textarea.blur() with no Orca code
loses it too, which places the owner upstream: Orca's unguarded release on
outside pointerdown is one trigger, not the cause. Committing 한 then
blurring mid-가 yields ['한'] where ['한','가'] is correct: one syllable
gone, surrounding text intact.
Teeth checked by inverting — adding an Orca-side composition guard flips
exactly the three cases that route through the release path and leaves the
bare-blur and no-blur cases green, which is the scope split: a fix in
regular-terminal-focus-ownership alone would not close this.
HAZARD PIN, but unlike the others this one has a real production injector —
clicking outside the terminal mid-composition. Owns no reported row. The
shape matches #9738's report; the injector does not, and a shape match with
a mismatched injector is not an owner.
* test(terminal): say which arm the STA-3237 fixture came from
The recorded keydowns are wave 4's A-shift-unmarked-only — the arm that
emits no PTY bytes. Nothing in the file said so, so two readers concluded
the row's events fail the owner's predicate and that STA-3237 and STA-3222
were different defects. They share an owner; the arm that fires is
Process/229+Shift, absent from this bubble-phase trace because the owner
claims it in the capture phase.
Also corrects "code-blind": the v1.4.163 policy emits \x1b\r only for a
shift-only key:'Enter', and a jamo keydown reaches that branch solely via
the isTerminalImeProcessEnter rewrite. The mock is deliberately wider so
the ownership guard stays under test if that rewrite moves.
Comments only — no assertion, fixture value, or mock behaviour changed.
* test(e2e): track the input-source selector the macOS specs shell out to
Five tracked macOS IME specs ran `swift .tmp/select-input-source.swift`, a
file that is gitignored and existed only on one machine. Anyone else
checking out the repo — or the same machine after .tmp is cleaned — could
not run them, and they are the capture drivers for the macOS rows that are
blocked waiting for exactly those runs.
Moves it to tests/e2e/ beside its callers. The chord spec now resolves it
from __dirname rather than reaching two levels up into .tmp.
* test(terminal): pin the CJK repaint decision against the reporter's own output
#12164 comment 1 and #5921 report agent output with double-width glyphs
rendering duplicated character-by-character while ASCII in the same line
stays clean. No IME, no composition, no keystroke — the user never types
the CJK.
Segmenting all three verbatim samples into maximal same-risk-class runs
gives 33 runs and zero violations of "this run is corrupted iff the
production detector flags it": 17 wide runs all corrupted, 16 narrow runs
all byte-identical. The paired negative is co-located in the same line
rather than in a separate run — the reporter supplied it without knowing.
Doubling is asserted as present, not uniform: 자바스크립트 and 시스템 each
leave a jamo undoubled, which is a repaint-region boundary artifact rather
than a per-character transform.
The discriminating arm is in the test rather than a source mutation:
be3f30e2f8 (#6890) elects a repaint for all 17 corrupted runs when the
agent types nothing, and reverting its disjunct elects none. Both
predicates agree once the user has recently typed, which is the pre-#6890
condition.
Samples inlined with per-sample SHA-256 because .tmp is gitignored and
cannot back a landed test.
* test(terminal): pin macOS period substitution landing after the composition
#11504's reporter published a DOM trace showing insertText ". " arriving
149ms after compositionend, when two spaces are typed with a CJK input
source and NSAutomaticPeriodSubstitutionEnabled is on. This replays that
trace against a real Terminal and asserts what reaches onData — bytes to
the PTY, not anything visual.
The owner is stock upstream CoreBrowserTerminal._inputEvent, not an Orca
module, confirmed at the resolved install and in the shipped bundle Vite
loads rather than in the TypeScript source.
Three mutations against that install, predictions written before the runs,
each failing exactly the arms predicted: dropping the composed/keyDownSeen
guard fails two, dropping Orca's intercept fails the one arm where the
payload arrives before the send window drains, and flipping || to && —
the candidate-fix shape — fails the arm that pins the defect itself.
CHARACTERIZATION: arm 1 asserts the broken behaviour and will fail the
moment #11504 is fixed. Update it to the new contract rather than working
around it.
composed is absent from every recorded bundle, so composed: true is the
spec-required value rather than a captured one; the test asserts it before
dispatching so a harness that dropped the field fails loudly.
* test(terminal): replay the recorded Windows Shift sessions through the IME guard
STA-3179 reports a Shift release sending Enter; #12171 reports delayed
Hangul plus doubled newlines. Both replay their own recorded Windows
MS-Korean keydowns through resolveTerminalKeyboardShortcutAction with the
shortcut policy mocked, so the assertions are about which events reach the
policy and what reaches terminal input.
STA-3179's held-Shift gesture yields exactly one newline, from the unmarked
Enter alone; its release arms nothing for the next composition, asserted
after a precondition check that the release really is keyups with shiftKey
already dropped; and an ordinary Shift press-and-release still routes every
keydown, which is the paired non-IME negative.
Teeth, verified by mutation: bypassing the isImeOwnedKeyboardEvent guard in
keyboard-handlers takes STA-3179 from 3 passed to 2 failed / 1 passed — the
survivor being the ordinary-session negative, which is correct, since a
non-IME session should not depend on that guard — and #12171 from 2 passed
to 2 failed. Source restored byte-identical.
Recorded shapes are inlined and the bundles cited in comments; nothing is
imported from .tmp, which is gitignored.
* test(native-chat): correct 61977d45177 — the preedit survives the question card
61977d45177 claimed a composed syllable vanishes silently when an
AskUserQuestion card replaces the composer, and characterized that loss.
The claim was false. Its premise was an artifact of the harness: the test
simulated a preedit with a silent textarea.value assignment and no input
event, which no IME does.
Real composition fires input with insertCompositionText on every keystroke
— the shape this repo already records in its own observed-event capture —
and React's change handler returns on input/change with no composition
gate, so onChange runs for each frame. The draft cache is written
synchronously inside the updater, so the preedit is already committed
before the card can arrive. Driven that way, it survives.
Renamed to match the contract that actually holds, and extended: Hangul
jamo-per-frame, Japanese kana accumulation followed by per-segment
conversion asserting the candidate the user was looking at survives, and a
pin on the mechanism itself — the draft cache holds the preedit while the
card is up.
Teeth: there is no fix to revert, so the mutation is the plausible wrong
one — gating onChange on isComposing(). That takes 4 passed to 3 failed,
with the English negative correctly surviving, since it has no composition
to gate.
Two consequences remain, recorded rather than fixed: the OS aborts the
composition when the field disappears, so a lone jamo returns as a
compatibility jamo the user cannot compose onto, and the remounted
composer is unfocused because the card owned focus.
* docs(native-chat): name the corrected commit and the degraded-jamo consequence
Records in the file itself that 61977d45177 is pushed and wrong, quoting
the two claims that are false, so a reader who finds it in git log reaches
the correction from the file that replaced it.
Also states the residual as a consequence rather than a curiosity: a lone
leading jamo returns as a standalone compatibility jamo (U+3131), which is
not a composable state — the user cannot resume the syllable, only delete
and retype. Preserved, but degraded into something unusable. That is the
note to find if a reporter ever describes exactly that.
The invariant these tests pin is not "the composer commits on unmount" but
"composition input events must reach React" — which is what a future IME
change would break, and is not visible from the swap site at all.
* test(terminal): replay the recorded macOS Telex commit boundaries
#6905 reports Vietnamese composed characters breaking in the terminal.
Replays the retained macOS built-in Simple Telex capture — recorded
selection and value set before each dispatch, since that is what the
commit range reads — and asserts what reaches onData: the first commit
alone, then through the real Enter, then the ASCII tail of the same run.
A code-point count would catch NFD normalisation.
ENGINE CAVEAT, stated first in the docblock: this is macOS built-in Simple
Telex, Telex only. The reporter's three named engines cannot run on the
platform they declared, and which macOS Vietnamese engine they used is
unconfirmed. This file certifies no engine, and does not imply VNI.
The owner is upstream's — CompositionHelper._finalizeComposition's
waitForPropagation branch — so the arms are copies under .tmp aliased by a
scratch config, with node_modules verified unchanged by shasum after every
run. Collapsing the range end onto its start fails all three; collapsing
the start to zero re-emits the first word into the second commit, which is
the reporter's "duplicated" direction. Different failure sets, so the
mutants are distinguishable rather than merely detectable, and the ASCII
assertion passes under both.
Falsifiability here is by mutation, not by a defective build: #6905 does
not reproduce at HEAD, so this has never been watched going red on a real
reproduction.
* docs(terminal): lead the #6905 test with its engine caveat
Comment-only. Moves the caveat above the source line so a reader meets what
the file does NOT establish before what it does — the capture is macOS
built-in Simple Telex, the reporter's named engines cannot run on the
platform they declared, and which engine they used is what gates this row.
Co-authored-by: Orca <help@stably.ai>
* docs(terminal): record that the swallow eats a keystroke after Japanese conversion
This pin framed the swallowed insertText around Cmd interrupting a
composition. A differential through Japanese multi-segment conversion shows
it is broader: type a segment, convert, then press `a`, and the `a` is
lost. No modifier, no exotic gesture. Korean surfaced it first only because
2-Set composes on nearly every keystroke.
Also records why it cannot simply be fixed. The suppression de-duplicates
IMEs that deliver their commit a task after compositionend, which a sibling
test pins; this swallow is that dedup's false positive, and the two events
differ only in payload, so no flag-timing change separates them. Both a
smaller redesign and a content-aware variant were built and measured — the
first duplicates on IBus, the second costs a reported row's test and is
blocked while the patch cannot be regenerated.
The Japanese arrays behind this are authored, not observed: no Japanese DOM
composition trace exists in the corpus.
* docs(terminal): a Japanese capture does exist — correcting dbeecb11bee
That commit said no Japanese DOM composition trace exists in the corpus.
False. One does, filed under the Linux bundles rather than the bundle named
for Japanese: 30 DOM events, two にほんご->日本語 conversions, with full
selection state per event. It is retained byte-identically in three further
bundles — one capture copied four times, not four observations, checked by
hash rather than by counting files.
The claim came from checking the bundle named for Japanese, finding nothing,
and generalising to the corpus without querying the rest of it.
Replaying it emits 日本語日本語 under both sequencing extremes on all four
arms, matching its own recorded onData. So "repeated conversion is
undisturbed" is now captured rather than authored. It carries no
post-compositionend insertText, so it cannot speak to the swallow: the
a-after-conversion figure stays authored and unobserved.
Also rewords the paragraph opener. It claimed to broaden a Cmd framing, but
hazard 2 was never Cmd-framed — the lines above already say Cmd does not
reach it. The real gap was that hazard 2 named no trigger at all, which
reads as exotic when it is ordinary.
* build(xterm): land the patch regeneration harness
The five dependency patches under config/patches/ shipped with no tracked
way to regenerate any of them. The xterm one is the hard case: it is derived
from an upstream build, so no fix could be made without rebuilding, and the
tooling to rebuild lived only in one machine's scratch directory. That
blocked a measured fix for a live keystroke-loss bug, and the EditContext
reduction an OSS survey identified as the only real one available.
Adds the regenerator, the upstream pin, the hand-written source patch the
bundle hunks derive from, tests, docs, and a PR job that verifies the
shipped patches still match the pinned build. The job caches the shallow
clone keyed on the manifest, so a cold run is minutes and a warm one under
one. Round-trip verified: regenerating from a clean checkout reproduces the
shipped patch byte-for-byte.
Marks the emitted patch -diff -text. pnpm hashes it byte-for-byte, so a
CRLF checkout would break install on Windows, and its minified bundle lines
make a diff nobody can read — review the source patch instead.
Also rejects unknown flags. --check was the fallback for any unrecognised
argument, so a typo, or --help, silently triggered a full upstream build
instead of what the caller asked for.
* fix(xterm): stop swallowing a keystroke typed after an IME commit
Type a Japanese segment, convert it, then press a key one macrotask later
and that key was lost. No modifier, nothing exotic — every user who keeps
typing straight after converting. Korean surfaced it first only because
2-Set composes on nearly every keystroke.
handleCompositionInput discarded the payload unconditionally in the window
after the deferred send: _isSendingComposition stays true for one macrotask
after the timer cleared _pendingCompositionStart, and the branch substituted
'' for whatever arrived. The suppression is not itself wrong — it
de-duplicates IMEs that deliver their commit an event-loop turn late, which
terminal-stock-composition.test.ts pins. It just could not tell a duplicate
from new input, because the two events are identical apart from payload.
Now it compares against _sentComposition, the text the deferred send
actually emitted, and discards only a match. A flag-timing redesign was
measured first and rejected: it fixed this and duplicated on IBus, because
no timing change can separate events that differ only in content.
Edited in config/patches/xterm-src/ and regenerated through the harness, so
the emitted patch and the lockfile hash are derived, not hand-written.
The commit-overlap pin's swallow arm now asserts the repaired contract —
the value its own comment already named as correct and as what stock
beta.287 emits. #11504's arm at :184 flips too; it never covered that
report, as its own prior note recorded, and the reporter's +149ms arm is
untouched and still asserting the defect. Provenance hashes in three test
docblocks are updated, since regenerating changes the patch hash and with it
the resolved install directory.
* docs(terminal): re-measure the #6905 mutation citations against the new bundle
Regenerating the patch moved the resolved install, so this docblock's
patch_hash, line count, two line numbers and three mutation outcomes all
described a bundle that no longer exists. The deferred branch is one the
fix writes into, so the outcomes could not be re-pointed on reasoning.
Line numbers read off both files by diffing anchors rather than derived by
arithmetic: 201 to 205, 159 to 163. Outcomes re-run through the retained
rig, which re-resolves through the module loader and re-derives each arm
from a unique minified anchor: pristine 3 passed, m1 3 failed, m2 2 failed,
m3 3 passed — identical to the old bundle. Guard controls in both
directions exit 1, so the counts are falsifiable.
Comment-only; the assertions and expectations are unchanged.
* fix(xterm): size the preedit overlay to the cells its text will occupy
updateCompositionElements computed the overlay's left edge from the grid
but never its width, so the preedit rendered at the font's natural advance
while the committed text takes two cells per wide glyph. Measured in
Chromium 150: 가나다라 drew 48.45px as a preedit and 69.20px once committed
— the same characters, same font, 30% narrower, and drifting further with
each syllable. Every macOS mono font carrying Hangul measured 0.49–0.72 of
two cells; never 1.0.
Deriving the width from wcwidth and the cell measure moves Korean, Japanese
and Chinese to 1.000 and leaves ASCII at 1.000, which it already was:
한 12.125 -> 17.297 (17.30 expected)
가나다라 48.453 -> 69.188 (69.20)
안녕하세요 60.563 -> 86.500 (86.50)
日本語 42.000 -> 51.906 (51.90)
abcdefgh 69.234 -> 69.203 (69.20, unchanged)
Edited in config/patches/xterm-src/ and regenerated through the harness, so
the emitted patch and lockfile hash are derived rather than hand-written.
The unit test asserts the arithmetic, which is what CI can run. The pixel
consequence was measured on macOS with SF Mono in an Electron harness, not
on the Windows font stack STA-3232 reports from — so this demonstrates the
mechanism and does not stand as that row's platform evidence.
* test(e2e): pin the macOS Korean preedit as visible only while composing
#11914 reports the composing text invisible until Space. Its c3 was recorded
as unobtainable, and the reason on file was wrong: the boundary IS
assertable, but not in happy-dom, which reports display:block in BOTH the
active and inactive states and zeros for every rect. A test there passes
with the defect present.
Captured on real hardware instead: hidden and 0x0 before, .active with
display:block, a 15.84x16 rect and checkVisibility() true while composing
그, hidden again after. 39 DOM events, 2 composition starts, onData
["한","그","\r"].
Two mechanism findings are carried in the setup because both are invisible
in the result and fatal if removed. The input source must be selected AFTER
the app takes focus — focusing resets it to ABC. And the IME must be warmed
until an observed keyCode 229; typed cold it emits raw QWERTY (g k s r m)
with no composition at all, which is indistinguishable from an IME that is
not installed. Two runs were voided on exactly that signature before the
warm-up was found.
The has229 and compositionStarts assertions exist to make such a run fail
loudly rather than pass as a clean negative.
Gated on darwin plus ORCA_E2E_NATIVE_MACOS_KOREAN, like its siblings. The
final spec form has not itself been executed — the machine became
unavailable — so it carries the probe's measured values as literals rather
than a run of its own.
* docs(e2e): correct 19a8d133db7 — the Korean preedit spec has been executed
That commit said the landed form had never run and carried the probe's
values as literals. It has now run on real hardware: 1 passed, 9.1s, rc=0,
with the capture and log sealed under a verified hash manifest.
The teeth check was also run rather than reasoned about, and it changes
which assertion matters. Forcing the active overlay to max-width:0 with
overflow:hidden — invisible on screen — leaves the active class, the
textContent, display:block AND checkVisibility() all passing. Only
during.rect.width fails. checkVisibility() is not sufficient against this
defect; the bounding rect is the single load-bearing assertion, which the
docblock already said and this run confirms.
An earlier teeth attempt injected the CSS mid-run and tripped the
hasActiveClass poll instead, failing at the wrong assertion. It is
inconclusive and excluded from the seal rather than counted.
* test(terminal): add #12171's ordinary-English arm from a real Windows capture
c4 was recorded as unmet and the ledger sourced its control to
evidence/windows-current/, which holds 12 captures and not one English one.
The arm here comes from windows-9803-final instead — same probe, same host
geometry, same injector, en-US with no IME, replayed keydown for keydown.
Two limits are stated in the file rather than left for a reader to find. It
is a different bundle and a different run about 3.6 hours later, so it is
not a same-run arm. And it is #9803's range-active MUTANT arm: ordinary
English stays byte-exact even with that saved-range mutation live, which is
why it reads as a negative rather than as a baseline.
Bundle cited by directory with its file SHA-256; MANIFEST.sha256 verifies
21/21, rc=0. Nothing imported from .tmp.
* docs(terminal): correct #12164's grounds — the cited comments say no such thing
The rejection of #12164 from this file's family was recorded as resting on its
comment 1 (output doubling) and comment 2 (filed against 1.4.163). Checked
against the API: the issue has exactly two comments, neither of which says
that, and the string 1.4.163 appears nowhere in the thread.
The conclusion survives on better grounds. The issue BODY's repro is "Run any
CLI agent (Codex, AGY, Claude, etc.) that outputs Korean text into the Orca
terminal" — untyped output, no keystrokes, no composition — so excluding
CompositionHelper is right, and the input-path hunt was looking in the wrong
place. The body is also LLM-authored (it still contains a literal
"## 5. GitHub Submission Draft (Ready to Post)") and its Root Cause section
blames a CJK IME preedit buffer its own repro never engages, so it should not
be read as observation.
Comment-only; suite unchanged at 5/5.
* test(native-chat): make composition frames carry isComposing, not just inputType
This suite's comment claimed "Gating onChange on `isComposing` breaks here."
It did not. composeFrame() fired `input` with `inputType` but never set
`isComposing`, so a gate on `isComposing` passed all four tests untouched —
the suite asserted a discriminator it did not exercise.
Composition frames now carry both, so neither gate is exempt. Verified by
pointing the mutant at it: with an `isComposing` gate on the composer's
onChange, this suite now fails 3 of 4 (it passed 4 of 4 before), and the
ordinary-English arm correctly survives, since a composition gate should not
touch it. Production code is unchanged and stays gate-free; the mutation was
applied, measured, and reverted.
Found while excluding NativeChatView's question-card remount as the owner of
#12118 / STA-3219: the remount is real, but the preedit survives it precisely
because this write path has no composition gate.
* fix(mobile): ship the patched xterm build, matching desktop
mobile pinned @xterm/xterm 6.1.0-beta.285 while the patch is keyed to
6.1.0-beta.287, so mobile shipped stock xterm and neither IME defect fix
reached it: the swallowed keystroke after an IME commit (9506039de72) and
the preedit sized to the font rather than the grid (e04e0c88da5).
Bumps the three xterm packages to the desktop versions and adds the patch
to mobile's own pnpm.patchedDependencies. No copy of the patch: pnpm
accepts the parent-relative path and records it in the lockfile against
hash 8d63166272e9040a…, byte-identical to what desktop resolves, so the
two stay in step by construction rather than by a drift check.
The workspace separation is untouched — root pnpm-workspace.yaml still
declares `packages: []` and mobile keeps its own lockfile, which is what
keeps the root's patches from failing as ERR_PNPM_UNUSED_PATCH.
Verified in the generated webview bundle rather than at the install:
alignPreeditToGrid 0->2, sentComposition 0->3, pendingInput 0->11, and the
stock-only _handleAnyTextareaChanges 2->0 and dataAlreadySent 4->0. pnpm
applies patches during linking before postinstall regenerates the bundle,
confirmed by a revert/reinstall/re-apply cycle in both directions.
Mobile suite 2971 passed, 3 skipped — identical before and after. Bundle
+1,514 B (+0.24%). Lockfile churn is xterm-only; --frozen-lockfile passes.
mobile/src/ime/ime-submit-carry.ts is NOT made redundant and is untouched:
it handles iOS firing onSubmitEditing on a React Native native TextInput
after unmarking a composition, which is outside the WebView entirely.
Known divergence left alone: desktop also patches @xterm/addon-webgl and
mobile now runs that version unpatched. That patch is glyph/texture-atlas
rendering with nothing IME-related, so it affects neither fix.
* test(terminal): pin the preedit overlay against already-committed cells
STA-3132 (arm A), STA-3170 and STA-3232 report a Korean preedit painted on
top of text already on screen. Builds v1.4.163-v1.4.166 cancel the pending
finalizer in compositionstart, so a committed syllable reaches onData one
syllable late and buffer.x is stale — the overlay lands on the cell the
flushed syllable is about to occupy.
Replays a recorded hardware trace rather than an authored one: the ordered
DOM event stream captured on Windows + MS Korean (wave5-r2 evidence, 64
events), echoing onData back as PTY output.
The load-bearing assertion is deliberately not the obvious one. Comparing
overlay style.left against cursorX is tautological — left is computed from
buffer.x. This counts committed syllables from the compositionend events
the IME fired, so the two sides are independently derived.
Discriminated by a historical re-add across seven real bundles, since the
owner is deletion-shaped: pristine beta287, v1.4.155 and v1.4.162 pass;
v1.4.163 fails; v1.4.163 with that single call removed passes; the byte
identical baseline restored fails again; head passes. Every failing arm
fails only this case — the ordinary negative stays green in all seven.
The negative asserts its own category rather than claiming it: zero
composition events, zero isComposing, zero keyCode 229, exactly 16 events,
paired against the Korean arm's 4 starts / 3 ends / 11 updates / 64 events.
Scope: cell indices, not pixels. happy-dom has no layout, so the recorded
8x16 cell metrics are supplied to the render service. This makes no claim
about pixels visually overlapping; that is affected-OS confirmation and
stays open. Covers the overlap arm only — STA-3132's auto-line-break arm
and STA-3232's half-line-capacity and a11y arms are untouched.
* test(e2e): matrix macOS period substitution against the OS preference
#11504 reports macOS inserting ". " after a Hangul Space commit. This
sweeps six arms across both states of NSAutomaticPeriodSubstitutionEnabled,
reading the preference live per run rather than asserting a literal.
Two results worth having on record.
The reporter's stated trigger did not reproduce. Their words are "There is
no second press at all. One space is enough", but korean-single-space emits
zero insertText with the preference on or off. So does word-space-word-space.
Their timing does reproduce, with different content. korean-double-space and
korean-longer-word-double-space emit a delayed insertText at +122.5-122.7ms
after compositionend — squarely the reported +149ms — but the payload is a
space, never ". ". Consistent with the double-space rule seeing two slots
under ABC and only one under Korean, where the IME commit consumes the first.
The substitution itself is real and preference-bound: latin-double-space
gives "ab . " with the preference on and "ab " with it off, on one build
with the preference as the sole variable, reproduced across two runs.
That also refutes a claim in PR #11506, which states the substitution "is
enforced outside the renderer and never reproduces in dev builds, so changes
here must be verified against a packaged app". It reproduced in the dev build
twice and did not reproduce on the signed packaged app. That claim should not
be used as a verification gate.
Gated @headful behind ORCA_E2E_NATIVE_MACOS_PERIOD, same shape as the Korean
preedit spec, so it does not run in ordinary CI. Evidence is onData and DOM
only — the PTY-child reader aborted and no packaged-app arm was stable.
* test(terminal): actually enforce the recorded jamo progression
The preedit assertion compared sample.overlayText against sample.overlayText
— the same expression on both sides. A lane proved it by mutation: corrupting
seven of the eight recorded preedit values left the suite fully green. So the
docblock's ㄱ→가→간→나→낟→다→달→라, which the matrix also cites as this row's
recorded shape, was cited and unenforced.
The first attempt at a fix was insufficient and is worth recording. Threading
stroke.preedit through to the expectation still passed on a corrupted fixture,
because that value both drives the rig and was the expectation — corrupting it
moved both sides together. Same tautology, one level down.
The expectation is now an independent literal. Verified by mutation rather
than by reading: corrupting two recorded values fails one arm; restoring them
passes 3/3.
overlayCell was never affected — it is compared against a count derived from
the compositionend events, not from the buffer, and remains the load-bearing
assertion for the overlap.
* fix(e2e): select the selectable input source, not the first match
TISCreateInputSourceList can return several entries for one input source
id. A third-party IME publishes a non-selectable parent alongside the
selectable mode, and taking sources.first can return the parent — after
which TISSelectInputSource fails with paramErr (-50) while the caller
reports success from the enable step.
Found with Qingg (com.aodaren.inputmethod.Qingg), which exposes exactly
that pair under one id. Its mode id equals the bundle id, so filtering by
name would not have helped; selectability is the discriminator.
Now filters on kTISPropertyInputSourceIsSelectCapable and falls back to
the old behaviour when nothing advertises it, so single-entry sources are
unaffected. Also enables every entry for the id rather than only the one
being selected: selecting a mode whose parent is still disabled fails the
same way.
Compile-checked, and selecting com.apple.keylayout.ABC still exits 0.
Unrelated to the enable path: on macOS 26.5.2 third-party IMEs are gated
behind a consent sheet in System Settings. TISEnableInputSource returns
noErr immediately regardless, and the enable only lands if that sheet is
answered while the requesting process is still alive.
* test(terminal): discriminate #12171 against the real shortcut policy
The prior candidate mutation for this row was correctly refused: its suite
mocked shortcut policy so Process/229 became actionable, while the real
resolveTerminalShortcutAction has no Process branch — so the kill measured
the mock. This does not mock it.
Replays a capture of this row's own gesture (d, l, Shift+T, e, k, Space,
Enter under MS Korean, committing 있다) taken on Orca 1.4.164, through the
real useTerminalKeyboardShortcuts hook, capturing bytes at terminal.input.
The earlier capture could not discriminate at all because it recorded no
shiftKey; this one records it on 10 of 10 keydowns with code populated.
One physical Shift+T produces two shifted Process/229 keydowns. Under the
pre-#12265 classifier each synthesizes {key:'Enter', shiftKey:true}, which
the real policy resolves to sendInput '\x1b\r' — twice, giving 1b0d1b0d,
the two escapes the known-bad ed96881b0d1b0d contains.
Mutation is the retained pre-12265-process-shift.patch applied to HEAD, not
an authored one: patch -p1 applies clean and diffs identical to the mutant
copy. Arms are copies; shared source hashes the same before and after.
The English arm stays clean under both modules, so the mutation
discriminates by language rather than by harness — and a real Shift+Enter
through the same rig yields exactly ['\x1b\r'] in every arm, so a silent
pristine result means the code is quiet rather than the harness dead.
Scope: 1b0d1b0d is measured at the renderer boundary. The capture recorded
no PTY bytes — window.api.pty is frozen on shipped builds and the onData
channel needs a build-time flag — so this shows the renderer producing the
two escapes that payload contains, not a re-observation of the payload.
* docs(native-chat): narrow this file's disclaimer to what is now true
It said "THIS OWNS NO REPORTED ROW". Half of that is stale: the remount site
is now the attributed owner of #12118 and STA-3219. On real Windows TSF the
questionActive swap aborts a live composition — the old node gets only a
blur and no compositionend, the text returns as committed, and the next jamo
yields 아ㄴ rather than 안.
The other half holds. This file pins the opposite property, that the text
survives, which is the half those reporters already agree with. Mutation
shows the gap rather than asserting it: deleting the unmount entirely leaves
three of four tests green, because every substantive assertion is
after.value === … and a composer that never unmounts keeps its value.
Also records why the abort cannot be asserted here. The DOM exposes no
observable separating committed text from a live preedit — value is the same
string either way, there is no EditContext, and the only composing-ness
state is a per-instance ref discarded with the node. A test pinning "no
compositionend fires" would be an anti-guard: red the day it is fixed.
The cadence objection is kept, since it is now the open question rather than
the reason for exclusion.
* refactor(terminal): drop the unread isComposing field from XtermBypassEvent
Added by #6396 for terminal IME candidate handling that this branch has since
removed. No production or test code reads it, and the policy is safe without it:
during composition `key` is 'Process', so the non-ASCII printable checks that
would care never match.
Co-authored-by: Orca <help@stably.ai>
* fix(native-chat): keep the composer mounted through an in-flight IME composition
A question card replaced the composer outright
(`{questionActive ? null : <NativeChatComposer/>}`). Unmounting the field
mid-composition aborts the composition in the OS: the node is detached before
`compositionend` can fire, the preedit returns as committed text, and a resumed
Hangul syllable degrades — 아 then ㄴ yields `아ㄴ`, never `안`.
Confirmed in rasterised pixels on Windows with a real MS Korean IME, at both
v1.4.171 and the reporter-era v1.4.164 (the swap block is byte-identical
across them): the preedit underline present before the swap, the composer
visibly absent during it, and the same glyph back afterwards WITHOUT the
underline — committed, not composing.
The swap is now deferred while a composition is in flight, which is what
editors that survive IME do: ProseMirror gates DOM work on `view.composing`,
CodeMirror protects the composing subtree from redraws. Hiding instead of
unmounting does not work — `display:none` and `visibility:hidden` both blur the
focused element and abort the composition the same way.
The hold releases on `compositionend`, which browsers also fire on blur, so
clicking into the card's own answer input yields the input region immediately;
with nothing composing the card still replaces the composer at once, so no
stray "Send a message" appears beside a question.
The existing characterization test flips to a regression guard: it pinned the
node being destroyed, which was the defect. Node identity is the load-bearing
assertion — value-only checks are trivially satisfied by a composer that never
unmounts and cannot tell a held composition from a destroyed one.
The typing-redirect handler moves to its own hook. That is not cosmetic: both
touched files sat at the 400-line cap, and `max-lines` suppressions are
forbidden, so the room had to come from a real extraction.
* fix(macos): opt Orca out of AppKit automatic period substitution
macOS "Add period with double-space" (`NSAutomaticPeriodSubstitutionEnabled`,
on by default) is applied by AppKit's text input system. Native terminals never
join that system; Chromium text fields do, so xterm's helper textarea inherits
it and a double space arrives as `". "` — a period nobody typed, handed straight
to the PTY (#11504).
Chromium answers AppKit for quote and dash substitution and defaults both off,
but declares no period accessor at all, so AppKit applies that one without
asking. This user default is the only lever: there is no per-field or
per-webContents opt-out to prefer over it. Writing the key into Orca's own
defaults domain overrides the global value for this app alone and leaves the
user's system-wide setting untouched. It necessarily covers every Orca text
field, not only terminals — AppKit offers no narrower scope, and that tradeoff
is deliberate rather than accidental.
Measured on the reporter's own build v1.4.161: with the preference ON, typing
a,b,space,space yields `onData ["a","b"," ",". "]`; with it OFF the same arm
yields two spaces.
Note the issue's causal model is wrong and this fix does not follow it. It
claims the substitution only fires with a CJK input source and never with ABC.
The measurement is the inverse — every Korean arm is clean and the ABC arm is
the one that fires — so the fix is not conditioned on input source.
NOT YET VERIFIED ON HARDWARE. The unit tests inject the writer, so they prove
the call is made on darwin and skipped elsewhere; they do not prove AppKit
honours an app-domain override for this key. That check is outstanding.
* fix(xterm): keep a live composition across a lone Cmd press on macOS
CompositionHelper.keydown exempted keyCodes 16/17/18 from tearing a composition
down, which covers Shift/Ctrl/Alt but not macOS Meta — 91/93 in Chromium, 224 in
Firefox. A lone Cmd press mid-composition therefore reached
_finalizeComposition(false), which dropped the preedit overlay's `active` class
and committed the live syllable early. macOS keeps the marked text alive across
that press, so no later compositionstart re-arms the overlay and the rest of the
word composes invisibly.
Measured on hardware (m4air, macOS 26.5.2, Apple M4, 2-Set Korean) with the Cmd
posted as a CGEventType.flagsChanged, which is what a physical modifier emits.
AppleScript `key code 55` posts nothing a browser can see — a bare `key code 56`
for Shift is equally silent — which is why no capture in the corpus ever reached
this branch. Three arms, same build otherwise: overlay live throughout with the
exemption, dark and prematurely committed without it, live again with it
restored. Evidence under
.tmp/ime-handoff/swarm-scratch/wave31-cmd-preedit/evidence/.
The fix cannot widen past a lone modifier: only a standalone press reports these
keyCodes, and a Cmd chord during composition is reported by Chromium as 229,
which was already exempt. Cmd+A still ends the composition, via the IME's own
compositionend. Ghostty draws the same line, returning early from flagsChanged
under hasMarkedText() for every modifier including Super.
Orca's terminal pane was never affected — shouldSuppressTerminalModifierKeyboardEvent
drops a standalone Meta keydown before xterm sees it, and deleting only 'Meta'
from that set is what flipped the hardware arm to broken. The popout preview
terminal and mobile's webview install no such guard and did reach the teardown.
terminal-ime-xterm-composition-commit-overlap.test.ts asked its fixer to update
the two Cmd arms to the values it named as correct; both now emit a single ['한'].
* test(native-chat): drop two byte-identical duplicate cases
`4632b86919d` copy-pasted two cases twice into the same describe block:
`retains carry across a same-frame non-Enter keyup before redispatch` and
`expires carry before a deliberate Enter after the next frame`. Each pair is
byte-identical — same title, same body — so the copies asserted nothing the
originals did not.
This is what has been failing `static analysis` on this branch since 2026-08-06:
`oxlint vitest(no-identical-title)` reports both under `--deny-warnings`, and
`verify` fails solely because it requires static analysis to pass. Every other
gate in `verify` was already green, including typecheck, xterm patch sync, the
full test shard set, and both package jobs.
12 cases still pass in the file.
* test(e2e): skip the WebGL arm when no WebGL renderer exists
The #12164 probe runs two arms, webgl and dom, and closes by asserting the
active renderer is the requested one. That assertion is right for the dom arm —
it is what proves the pane actually left WebGL, without which the arm is
meaningless — but headless CI has no GPU, xterm falls back to DOM silently, and
the webgl arm then fails.
The failure reads as a Korean rendering defect and is not one, so the webgl arm
now skips with the active renderer named. The dom arm keeps the assertion
unchanged.
This is the third of three checks that have been red on this branch since
2026-08-06. `static analysis` and `verify` were fixed in c51c6b5837e; the CI log
shows this job as 1 failed / 1 passed, the pass being the dom arm.
* chore(lint): drop five unused no-console disable directives
`check-changed-code-quality` reports unused eslint-disable directives as errors,
and these five sat above diagnostic `console.log` calls in IME test and spec
files where `no-console` is not enabled — so each suppressed nothing.
This is the second of the two static-analysis steps. `c51c6b5837e` fixed
"Enforce focused code-quality plugins" (duplicate test titles); this fixes
"Enforce changed-code quality". Both had been red on this branch since
2026-08-06, and I mistook the first for the whole job.
The diagnostic logs themselves are kept — they are what a failing IME arm prints
for a reader to inspect.
* test(e2e): cover #12164 under fractional device scale factor
Fractional display scaling was #12164's last unexplored branch, and the reason
is worth recording: earlier attempts were BLOCKED, correctly, because they
proposed mutating the Windows display scale on a remote physical machine with no
console recovery. `--force-device-scale-factor` reaches the same renderer state
per process, so nothing outside the Electron instance changes and there is
nothing to restore.
The hypothesis was specific: `프프로로젝젝트트` is what a half-pixel cell boundary
could produce on a 2-column glyph, and nothing else in the suite varies dpr.
Measured at 1.25 and 1.5, both under WebGL: ink extents 25/21/16 with identical
ink groups, matching the scale-1 run. No doubling.
The arm self-certifies before asserting — if the flag does not take, the test
fails rather than silently measuring at dpr 1. That matters here: the sibling
spec's WebGL arm went two days reporting a missing GPU as a Korean rendering
defect precisely because a silent fallback looked like a result.
* fix(terminal): match Mod+letter shortcuts by physical key, not IME-rewritten key
With a CJK input source active, macOS and Windows report the physical key through
`code` but rewrite `key` to the layout's character: Korean 2-Set turns Cmd+C into
`{ key: "ㅊ", code: "KeyC", metaKey: true }`. Every `key.toLowerCase() === 'c'`
match misses it, so the shortcut is not recognised and xterm encodes the chord as
PTY input instead — issue #13033 reports `ESC[12618;9u` and a terminal that jumps
to the bottom, because user input scrolls the viewport.
This is the same key-vs-code confusion that owned #12171, where a `Shift+T`
typing ㅆ was read as Enter for want of a `code` guard, so the fix is the same
shape: trust `code` when it is present, fall back to `key` and then the legacy
`keyCode` when it is not (Chromium omits `code` on synthetic and some keypress
events, and `keyCode` keeps its US value even when `key` is rewritten).
Applied to the four terminal-side sites, including the dashboard pop-out, which
#13033 called out specifically as having its own key handler:
pty-connection.ts Cmd/Ctrl+C copy guard
keyboard-handlers.ts Cmd+G search navigation
agent-interrupt-inference.ts interrupt inference
preview-terminal-key-handler.ts pop-out paste
Nine further `key.toLowerCase()` letter matches exist outside the terminal
(TaskPage, editor, GitHub composer, browser markup). They have the same defect
and are deliberately left for a separate change rather than widening this one.
An existing case, `matchSearchNavigate > returns null for wrong key`, overrode
only `key` and left `code: 'KeyG'`, so it began passing for the wrong reason. It
now overrides both — which is what "wrong key" means once matching is physical —
and a companion case pins the Korean-rewritten chord still matching.
#13033 was closed NOT_PLANNED; the reporter's event shapes drive the new test.
* fix(renderer): match every Mod+letter shortcut by physical key, not IME-rewritten key
Completes the previous commit. A CJK input source rewrites `event.key` while
`event.code` keeps the physical key, so `key.toLowerCase() === 'z'` and friends
silently stop matching — the shortcut is not recognised and the keystroke falls
through to whatever handles unclaimed input.
The helper moves to `@/lib/ime-latin-shortcut-key` first: it now serves the
editor, GitHub composer and browser markup, and importing terminal-pane
internals into those would be the wrong direction. `lib/` already hosts
`ime-composition-keyboard-event` for the same reason.
Nine remaining sites, all previously unreachable under Korean/Japanese/Chinese/
Vietnamese input:
TaskPage, ActivityPrototypePage, ProjectViewWrapper Cmd+F search
useMarkupKeyboardShortcuts Cmd+Z undo
GitHubMarkdownComposer, RichMarkdownLinkBubble,
rich-markdown-link-shortcut Cmd+K link
native-chat-shortcut Cmd+J
rich-markdown-key-handler Cmd+Shift+X
Six of the nine test `!== 'letter'` as early-return guards and three test
`=== 'letter'`; the negation is applied per site, since a blind substitution
would have inverted six of them.
Full suite: 4249 files pass. Three files fail locally and none is caused by this
change — the branch touches no file under `src/main/` or `src/relay/`, all four
failures reproduce on an unmodified tree or pass in isolation (the worktree
poller passes 21/21 alone, so it is full-suite parallelism), and all 16 CI test
shards are green.
* docs(ime): scope the IME composition rules to the terminal-pane directory
#11893 proposed adding these to the root `AGENTS.md`, which every agent loads on
every task regardless of what it is doing. They only bind keyboard handling, the
composer and the terminal input path, so they belong next to that code —
`tests/e2e/AGENTS.md` already establishes the nested convention here.
Kept from #11893: range-derived commits, guarding above the key dispatch,
the `attachCustomKeyEventHandler` / `CompositionHelper` interaction, no
normalization at commit, and the recorded-trace evidence bar.
Added from defects found since it was written:
- match shortcuts on `event.code`, not `event.key` (#12171, #13033)
- `keyCode === 229` means an IME owns the press
- do not unmount a field mid-composition, and hiding is not a fix because
`display:none` blurs and aborts it too (#12118, STA-3219, #11332)
The evidence bar now also names the mutation check, since a test that survives
deleting the code it guards is guarding nothing — a failure this effort hit more
than once.
* fix(terminal): gate Ctrl+Enter CSI-u on a negotiated pane, porting #12462
Found while scoping the rebase onto `main`: #12462 landed on 2026-08-06 and
fixes a real defect this branch does not carry. Ctrl+Enter emitted
`\x1b[13;5u` unconditionally, so a pane that never negotiated the kitty
keyboard protocol — local Windows ConPTY, plain shell — printed the escape
verbatim into the prompt.
This branch deletes `terminal-ime-deferred-newline.ts`, which is one of the
files #12462 touched, so a rebase resolving those conflicts by taking our side
wholesale would silently reintroduce the defect. Porting it forward now means
the fix survives the rebase however the conflicts are resolved.
Mirrors the Shift+Enter guard already here: local ConPTY falls back to the
legacy CR every emulator sends, and a negotiated pane keeps the chord, so the
fallback is scoped to panes that cannot receive CSI-u rather than to Windows.
NARROWER THAN #12462 BY ONE CONDITION, deliberately. `main` also allows CSI-u
via `hasCtrlEnterCsiUAuthority()` (trusted consumer evidence, #12329); that
helper and its plumbing do not exist on this branch. Omitting it is the
conservative direction — an authorised pane gets `\r` instead of the chord,
rather than an unnegotiated pane printing an escape — but it should be restored
when the two histories are reconciled.
Test covers both directions and is mutation-checked: forcing the gate open
fails it, so it cannot pass by construction.
* fix: reconcile two more fixtures main moved while the stack waited
Both caught by CI, not locally, and the reason the local run missed one is
worth recording:
1. `browser-toolbar-profile-dialogs.ime-enter.test.tsx` did not pass
`useNativeUserAgent` / `onUseNativeUserAgentChange`, which `main` added to
`BrowserToolbarProfileDialogsProps`.
Local `pnpm typecheck` reported 0 errors on the same commit CI failed. The
cause was a stale `config/*.tsbuildinfo` — tsc reused an incremental cache
from before the merge. Deleting it reproduced CI's error exactly. Any
"typecheck clean" during this merge should be treated as unverified unless
the cache was cleared first.
2. Localization keys for `SshDisconnectedDialog` were absent from `en.json`:
the merge took this branch's component alongside `main`'s catalog.
Regenerated with `pnpm run sync:localization-catalog` rather than hand-added.
* fix(mobile): regenerate the lockfile the merge resolved by taking one side
CI's `verify` failed with `ERR_PNPM_OUTDATED_LOCKFILE` on `mermaid (lockfile:
11.16.0, manifest: 11.16.1)`. The mismatch was in `mobile/`, not the root — the
root lockfile was consistent throughout, which is why inspecting it (and even
GitHub's merge ref) found nothing wrong.
Cause: during the merge I resolved `mobile/pnpm-lock.yaml` by taking this
branch's side wholesale rather than merging it, so it kept `mermaid 11.16.0`
while `mobile/package.json` came from `main` at `11.16.1`. Taking one side of a
lockfile is only safe when the corresponding manifest comes from the same side.
Regenerated with `pnpm install --lockfile-only`; `--frozen-lockfile` now passes
in `mobile/`. Verified the xterm patch entry survives intact — same hash
`4f1b42d268f3964d…` and the parent-relative path into `config/patches/`, which
is the desktop/mobile coupling that would silently break the mobile build.
Two earlier diagnoses of this failure were wrong and are worth recording: it was
not the root lockfile, and it was not a stale merge ref (a rerun reproduced it
exactly).
---------
Co-authored-by: Orca <help@stably.ai>
mobile/ had no .npmrc, so unlike the root workspace it would resolve
packages published moments ago. #13113 surfaced this concretely: it
pulled nanoid 3.3.18 and postcss 8.5.26 at 0.4 and 1.7 days old, both
newer than anything the root gate would have allowed.
Copies the root's minimum-release-age=4320 (3 days). Deliberately not
shamefully-hoist -- that one is Electron-specific and would change how
mobile hoists.
Re-resolves nanoid to 3.3.17 and postcss to 8.5.25 in the same commit
because the gate is otherwise unusable: pnpm install fails with
ERR_PNPM_NO_MATCHING_VERSION on the locked nanoid 3.3.18. Both picks stay
above their advisory floors (CVE-2026-67213 needs >=3.3.17,
CVE-2026-69153 needs >=8.5.23), so this is not a security regression.
Co-authored-by: Orca <help@stably.ai>
Clears 47 of 49 open Dependabot alerts across the root and mobile lockfiles.
The 2 remaining (image-size) have no patched upstream release.
Direct bumps: pdfjs-dist 5.7.284 -> 6.2.108 (CVE-2026-16633), mermaid
11.16.0 -> 11.16.1 (root + mobile), dompurify 3.4.12 -> 3.4.13.
In-range re-resolves: brace-expansion, fast-uri, hono, ip-address,
js-yaml 4.3.1/3.15.1, nanoid, postcss, tar, undici 6.28.0/7.29.0.
Drops the @modelcontextprotocol/sdk>@hono/node-server override by bumping
shadcn's transitive SDK to 1.30.0, which widens its range to
^1.19.9 || ^2.0.5 so @hono/node-server resolves to a patched 2.1.0 on its
own. The other two overrides must stay: monaco-editor hard-pins dompurify
3.2.7 and xcode wants uuid ^7.0.3, both vulnerable.
pdf.js 6 removed PDFDocumentProxy.destroy(); PdfViewer now tears the
document down via the loading task it was already destroying.
Supersedes #13074, #13090, #12960, #12952.
Co-authored-by: mondaychen <monday.chen@gmail.com>
Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: Orca <help@stably.ai>
* feat(native-chat): add model and effort pickers for grok
Grok had no session-option catalog, so the native chat composer showed no
pills and every launch ran the CLI's own defaults with no way to change them.
Adds a `GROK_SESSION_OPTION_CATALOG` (model via `-m`/`/model`, reasoning
effort via `--reasoning-effort`/`/effort`) and the discovery plumbing behind
it. Grok's selectable ids depend on the signed-in account and on `[model.*]`
config, so the seed carries only `grok-4.5` and a runtime `grok models` probe
supplies the rest as authoritative — a retired id must be droppable, since
launching one is a fatal exit rather than a warning.
Because `grok models` publishes `Default model:` and marks the row
`(default)`, the picker can name the model a fresh session is actually
running: `defaultModelIsCliDefault` plus an untracked record means no `-m`
was ever emitted, so the CLI is on its own default. That default scopes the
effort row but is never written to persisted settings — that field is what
authorizes `-m` on every later launch, and adopting a model the user never
picked would pin today's default forever, fatally so on an account without
it. `grok --help` publishes no default for `--reasoning-effort`, so the
effort value stays unnamed until something sets it.
Known gap: that refusal to persist is also a limit. An option set while on
the CLI default is dispatched and honored in-session, but reaches no later
launch — it persists under the default's id with `model` left unset, and
both `resolveNativeChatSessionOptionDefaults` and
`resolveAgentSessionOptionLaunch` bail without that key. Picking a model
explicitly persists normally. Closing this means teaching both to resolve
options from the default model while still refusing to emit `-m`, which is
the launch-args path and wants its own review.
Known gap: the picker infers "no `-m` was emitted" from its own in-memory
record, so a model reaching argv from outside it — the user's own
`agentDefaultArgs`, or a renderer reload that drops the record while the
flagged PTY lives on — leaves the pill claiming the CLI default while
another model runs. No wrong model is persisted.
Extracts `hasFlag` and `labelFromModelId`, and splits the model-probe spec
out of the commit-message registry so discovery no longer implies an agent
can write commit messages.
Co-authored-by: Orca <help@stably.ai>
* docs(native-chat): note the invariant keeping modelIsCliDefault agent-safe
The flag is computed without checking the catalog, so it reads as unsafe for
the four agents with no CLI default. It is safe only because `persist` bails
unless `modelId` is truthy, which for those agents implies a tracked model.
Widening that guard would silently change persistence for every agent.
Co-authored-by: Orca <help@stably.ai>
* fix: retire persisted models on mount and handle -- terminator
- When a pane mounts after model discovery has already settled, it now checks
the cache and retires persisted models that are no longer available.
- CLI flag detection now respects the `--` option terminator, treating
everything after it as positional arguments rather than flags.
* Fix: persist grok session options under probe-confirmed defaults
Options set under the CLI default were silently lost on restart.
Distinguish seed guesses from probe-confirmed defaults by renaming
`modelIsCliDefault` to `modelIsUnverifiedDefault`. Once confirmed,
adopt the default as a persisted flag so options survive restarts.
* fix(native-chat): close the retired-model fatal-launch paths from counsel review
Counsel report C1/C2 (High), C3, P1, C4:
- Untrack a session model an authoritative discovery dropped and gate every
persist path, so option writes can never re-adopt a retired id (C1).
- Resolve launch defaults through the enrichment cache: a persisted model
missing from every settled probe no longer becomes a fatal `-m` (C2).
- Serialize retirement and picks on one settings write queue that re-reads
live state at apply time (C3).
- Stabilize onSwitchToTerminal so the session-option surface is not rebuilt
every TerminalPane render (P1), and cap the enrichment host map (C4).
Co-authored-by: Orca <help@stably.ai>
* Store agent in enrichment entry and extract token utilities
Refactor enrichment to store the agent field directly instead of
parsing it from a composite key, and extract CLI flag token filtering
into a shared utility. Use a dedicated function for tracked model ID
lookup. Improves code reuse and reduces parsing overhead.
* Rename modelIsUnverifiedDefault to adoptModelAsLaunchDefault
Move the model adoption gate into the core session-options module, where probe confirmation and discovered-model status are known. This ensures adoption decisions are gate-checked before persisting to avoid fatal launch flags, and simplifies the picker surface by moving the logic to where it belongs.
* Keep model probe evidence by agent, not host
Store probed model IDs in agent-keyed cache independent of host cache, so
evidence persists across host eviction. Prevents retired models from being
treated as valid when host cache entries are evicted.
* Store agent in enrichment entries instead of separate proof-evidence map
Model probe evidence is now tied to enrichment entries rather than maintained in a separate per-agent map, eliminating the need for eviction logic that could disconnect proof from entries.
---------
Co-authored-by: Orca <help@stably.ai>
* Bump mobile app.json to 0.0.42
* fix(mobile-ios): stop TestFlight CI from waiting on ASC processing
0.0.42 builds 1–2 uploaded successfully then hung for hours polling
processing with no Ready build and no Apple email. Exit after upload
and cap the job at 90m so the next cut does not repeat that hang.
* fix(mobile-ios): fully skip Pilot wait (no changelog)
Pilot only returns immediately after upload when changelog is nil;
passing notes re-enters the ASC build-list poll.
* fix(native-chat): stop diff colouring from misreading -- / ++ content lines as file headers
diffFromText skipped every line starting with --- / +++ as a file header, so a
deleted SQL/Lua '-- comment' (git emits '---<content>') or an added '++flag' fell
through to gray context with its marker still attached — and when it was the only
change, the two-marker gate dropped the coloured diff entirely.
Detect real headers structurally instead: an adjacent '--- <old>' / '+++ <new>'
pair outside any hunk. A hunk header or 'diff --git' line now also proves the text
is a diff, so a genuine single-line change renders while prose keeps the guard.
Co-authored-by: Orca <help@stably.ai>
* test(native-chat): adopt #12335 diff-collision vectors and add mobile parity
Pulls in @YuriNachos's test vectors from #12335 (header-less --- deletion, an
adjacent --x/++y content pair, mobile re-export parity) and adds the spaced
-- / ++ pair inside a hunk, which the pair-only rule in that PR misreads.
Co-authored-by: Orca <help@stably.ai>
* fix(native-chat): keep bare --- / +++ rules out of the diff marker count
Dropping the `---`/`+++` prefix exclusions made a bare `---` — a Markdown
thematic break or YAML document separator — classify as a deletion. Tool
results routinely carry those, so `---\na: 1\n---\nb: 2` went from correctly
rejected to rendering as a red diff.
A bare rule is never a file header (those need a path after the marker) and is
only content inside a hunk, so treat it as meta when outside one.
Fold the separate `isStructuredDiff` scan into the same pre-pass and skip
non-marker lines early, so the added guard costs no extra traversal: 5.1 -> 4.3
us per 120-line prose result, diff path unchanged.
---------
Co-authored-by: Orca <help@stably.ai>
* feat(native-chat): render omp transcripts
omp already ships as a first-class launchable agent with session_id resume, but
its transcripts had no decoder, so native chat could not render it — the agent
runs and the conversation stays a raw terminal. This adds the decoder and wires
it through the same path Claude, Codex and Grok use.
omp writes one envelope per line, `{ type, id, parentId, timestamp, … }`, where
conversation turns are `type: 'message'` and the rest is session bookkeeping.
Reasoning arrives as a `thinking` content block inside the assistant turn, so
the mapping follows Claude rather than Codex: thinking becomes a text block on
an assistant message, where Codex and Grok emit a separate reasoning role only
because their transcripts carry dedicated reasoning records.
- toolCall -> tool-call, arguments passed through as the object omp writes
- toolResult -> tool role, isError preserved
- developer -> system, matching the Codex non-user/non-assistant fallback
- blob-handle images drop, as the Claude mapper drops an image record with
neither path nor url
- bookkeeping and unrecognized types skip rather than throw
Session files are `<ISO timestamp>_<session id>.jsonl` under a per-cwd directory,
so the resolver matches the id as a base-name suffix the way Codex rollout files
are matched, and honors OMP_CODING_AGENT_DIR through normalizeAgentSessionsDir
so it stays consistent with the AI Vault scanner.
omp records no interruption or abort event, so unlike Claude and Codex there is
no NATIVE_CHAT_INTERRUPTED_STATUS_TEXT path.
Verified against 94,603 lines of real omp transcripts across four sessions:
50,546 records decoded, zero malformed, zero thrown.
* fix(native-chat): complete omp record coverage and gate remote transcripts
Review fixes on the omp transcript decoder.
omp writes several record types with no `content` field, so they decoded
to zero blocks and disappeared from the chat view entirely:
- `bashExecution` / `pythonExecution`: TUI `!command` runs, now a tool turn
- `fileMention`: `@path` attachments, listed by path (never `files[].content`,
which is an auto-read dump)
- `custom_message` and legacy `custom` / `hookMessage` rows, gated on
`display` the way omp's own renderer gates them
Also:
- `stopReason: 'aborted'` turns now surface as the interrupted row, matching
the Claude and Codex decoders. An abort carrying partial content keeps it.
- A cancelled command cell now reads as errored. Every omp cancel path emits
`exitCode: undefined`, which JSON drops, so an `exitCode !== 0` check read a
cancelled run as a clean success.
- omp joins Grok in requiring a locally readable transcript. Its hook reports
no transcript path, so under Model-A SSH the chat view opened against a disk
this process cannot read and never loaded. Applies on mobile too, which
shares the same allowlist.
- The session-file walk prunes omp's per-session subagent artifact
directories, matching the AI Vault scanner. It was returning a subagent
transcript instead of the parent session, and cost a full recursive readdir
on every resolve.
* style(native-chat): apply oxfmt to the omp review fixes
Mobile CI gates `oxfmt --check`; the two root files were unformatted too,
just ungated there. Line wrapping only, no behavior change.
---------
Co-authored-by: plotarmordev <299844489+plotarmordev@users.noreply.github.com>
Co-authored-by: Brennan Benson <79079362+brennanb2025@users.noreply.github.com>
* fix(github-project): index fork upstream slugs for project row matching
Project cards often reference the public upstream repo while the open
clone's origin is a personal fork. Map the parent slug to the same Repo
so selected-repo filters no longer hide every board row.
Preserves origin-based getRepoSlug identity for non-project callers.
Fixes#12647
* fix(github-project): match project rows against fork upstream slugs
Resolve the referenced call to a nonexistent `resolveRepoUpstreamSlug` and
match the persisted `repo.upstream` parent instead of issuing an extra
`github.repoUpstream` RPC per repo on every index build — that lookup shells
out to `gh repo view` for non-forks, so it would have gated the Projects tab
on N network calls. `repo.upstream` is already resolved at repo-add time and
backfilled at startup, so the fix costs no IPC.
Origin matches take precedence over upstream ones so an open clone of the
upstream repo itself is never made ambiguous by someone's fork of it.
Also covers the two surfaces the origin-only match broke alongside the desktop
table: mobile's project row matcher and the store-slice row-mutation routing.
* fix(github-project): scope fork upstream matching by host and selection
Round-1 review fixes on top of the upstream-slug index:
- Apply origin-over-upstream precedence among *selected* repos instead of
globally. An open-but-unselected clone of the upstream repo was shadowing the
selected fork, so #12647 still reproduced for anyone holding both — and repo
selection collapses to one repo per project key, which is exactly that case.
- Scope a fork's upstream identity key to the fork's own origin host.
Persistence strips upstream.host, so GHES forks never matched their own rows
and a GHES fork's parent could bind a same-named github.com row.
* fix(github-project): skip the fork alias when its own origin is unresolved
Round-2 review fix. `githubHostFromIdentityKey` cannot tell "origin resolved to
github.com" from "origin did not resolve" — both yield no host. A GHES fork
whose slug resolution had failed (auth lapse, unreachable runtime) therefore
landed in the github.com namespace, so an unrelated public Project row matched
it and Start work opened the wrong clone on the wrong server.
Require a resolved origin before indexing the upstream alias: it is the only
host evidence there is, and a repo with an unresolved origin was already absent
from the origin index, so nothing is lost that origin matching had.
* fix(repos): persist the fork upstream host instead of dropping it
`sanitizeRepoUpstream` kept only `{owner, repo}`, so a fork's parent lost the
server it lives on every time the record round-tripped through disk.
That forced the Project row matcher to re-infer the host from `origin`. The
inference is right for an API-resolved fork parent — `getRepoUpstream` stamps
`origin.host` there precisely because "a fork parent lives on the same server as
the fork". It is wrong for the other branch: a local `upstream` remote carries
its own host, so a github.com clone with a GHES `upstream` remote was indexed
into the github.com namespace, where an unrelated same-owner/name public repo
could claim it and Start work would open the wrong clone.
Keeping the host removes the guess. Absent stays absent, so records written
before this hydrate unchanged and the origin-derived fallback still covers them.
Also fixes the avatar for rehydrated GHES forks, which resolved against
github.com for the same reason.
* docs(github-project): correct upstream host fallback comment
Persistence now keeps non-empty upstream.host; originIdentityKey remains
the host fallback for older records without one (CodeRabbit nit).
* fix(github-project): own slug-index retry timer cleanup
Move the failure-retry setTimeout into its own effect so cleanup always
clears it. Scheduling from the async buildIndex then-handler failed the
react-doctor effect-needs-cleanup gate in static analysis.
* test(github-project): guard the slug-index retry timer, fix the mobile twin comment
Two follow-ups on 52298d82 and 2f89c20d:
- Cover the retry timer both ways: a failed resolution still re-resolves after
the TTL and recovers the match, and the pending timer is gone after unmount.
The second fails if the timer moves back into the async then-handler, so the
property is guarded by more than the lint rule.
- The mobile matcher's comment made the same stale "persistence strips
upstream.host" claim that 2f89c20d fixed on the renderer side.
* test(github-project): unmount slug-index hooks so React cannot flush after teardown
CI shard `tests node 24 6/16` failed with 10 unhandled
`ReferenceError: window is not defined` traced to this file. The tests mounted
hooks without unmounting, so React scheduler work flushed after the DOM
environment was disposed. All assertions passed; the shard failed on the
unhandled errors alone.
`cleanup()` after each test unmounts the trees. Does not reproduce locally in
isolation — it needs CI's worker pooling and file ordering.
---------
Co-authored-by: Jinwoo-H <Jinwoo-H@users.noreply.github.com>
* fix(mobile): keep the cached transcript visible while reconnecting
A manual retry closes the client and opens a fresh one, so the chat session
hook saw a new client under an unchanged identity, dropped its settled read,
and handed out an empty list — the transcript collapsed to a full-screen
spinner until the swapped client's snapshot landed.
Hold the last settled list per identity (captured post-commit) and keep
rendering it while the re-read is in flight. `transcriptLoading` still gates
consumers that decide from an empty transcript, so the launch-draft seed is
unaffected. The held list is keyed by a new `sourceIdentity` (host/workspace)
in addition to agent/session/transcript, so it can never serve another
source's messages.
Refs STA-3333.
* test(mobile): assert the whole reconnect window, not just its first frame
The re-subscribe lands a commit after the first render of the swap, so a
regression that cleared the held list there left frame 0 green and still
blanked the transcript. Verified: clearing the cache in the subscribe
cleanup now fails this test, where before only the view-toggle test caught it.
* fix(mobile): don't derive a tappable ask card from the held transcript
The cache this PR adds keeps the previous list rendered while a swapped
client re-reads. useMobileNativeChatPrompts was the one consumer reading
`messages` without honouring `transcriptLoading`, so an ask answered on
the terminal resurrected as a live, tappable card during that window.
Gating on `transcriptLoading` is exactly base behaviour: `setRead` only
ever stores 'ready'/'error', so status==='loading' implied an empty list
before this PR. The live `askFromStatus` path is untouched.
* chore: keep merge formatting scoped
* fix(file-explorer): sort numbered file names naturally
The File Explorer compared names with bare localeCompare, so numbered
files listed 100, 200 before 99. Hoist the numeric collator Source
Control file rows already use (#10850) into src/shared and apply it to
the local and runtime directory listings, the name-filtered view, and
Source Control directory nodes, which were inconsistent with the file
rows one line below (#11426).
* fix(file-explorer): natural sort on SSH funnels, relay, and pickers
Adversarial-review round 1 rework:
- Both readDir funnels short-circuited to the SSH filesystem provider
before the patched sort, so SSH workspaces kept lexicographic order;
re-sort locally after the provider returns (the remote relay may be an
older build), and fix the relay's own comparator for relay-native
consumers.
- sortDirEntries (shared, unit-tested) owns the directories-first +
natural-order listing contract used by every funnel.
- compareFileNames breaks numeric-collation ties ('2' vs '02') by code
units so sibling order stays total instead of readdir order, and pins
the collator locale to 'en' so every host produces one order.
- The SSH folder browser and runtime server dir picker now match the
Explorer they browse into.
- Ordering pinned by tests at the relay, source-control tree, and shared
helper.
* fix(mobile): natural sort in the mobile file explorer
Mobile re-sorted host readDir results with bare localeCompare, undoing
the host funnel's natural order (round-2 review). Reuse the shared
comparator and pin the order in the mobile suite.
* fix(file-explorer): natural sort at the renderer choke point and remaining ties
Round-3 review: the remote-runtime RPC and paired-web routes return the
host's order verbatim, so re-sort in readFileExplorerDirectory where
every desktop route converges; pin the SSH funnel with a handler-level
test; and route Source Control path compares through compareFileNames so
numeric-collation ties share one total order with the Explorer.
* docs(file-name-sort): state the real perf baseline in the hoist comment
* refactor(source-control): drop the dead collator export; pin the test oracle locale
* fix(file-listings): cover remaining natural-sort surfaces
* Reorder source control to show staged changes first by default
Stages are closest to the commit action and most relevant to the
commit workflow. Merges untracked files into Changes visually while
preserving their Git area. Removes the untracked-first preset and
includes migration logic for existing user settings.
* Drop source control group order user preference
Remove the sourceControlGroupOrder setting and related UI, migrations, and persistence logic. The source control view now always displays sections in the order: staged changes, unstaged changes, untracked files.
* Reorder source control to show changes before staged
Aligns with the edit-stage-commit workflow by showing unstaged
changes (active edits) before staged changes (queued for commit).
- Replace 'Send answer' with 'Submit' for clarity and consistency
- Update all locale translations (en, es, ja, ko, zh)
- Remove fixed button width and add whitespace-nowrap for flexible sizing
- Update component and test references
`activateMobileSessionTab` gated only on `publicTab.status !== 'ready'`. A deliberately slept pane publishes as `pending-handle` indefinitely — indistinguishable at that call site from a pane awaiting reconnect — so the reconnect probe added by #11542 respawned it with a re-resolved agent launch, waking something the user had deliberately put to sleep.
The first attempt refused activation for any pane with a `worktree-sleep` record, applied to every path. Independent review found that broke the documented wake gesture: opening the tab IS how those panes are meant to cold-restore (`wake-sleeping-agents-in-background.ts`: "Those panes cold-restore --resume when their own tab is opened"). A mobile tap sends the byte-identical call the reproduction test used, and in three of four topologies no wake clears the record first — so the tap became a permanent no-op with no feedback.
This carries intent explicitly instead of inferring it. A new shared `TabActivationIntent` ('user' | 'automatic') rides the existing ActivateTab schema as an optional additive field; `isAutomaticTabActivation` returns true only for an explicit 'automatic', so an absent value is permissive BY CONSTRUCTION in one place — an older client that does not send it keeps today's behavior rather than silently losing its wake gesture. The field is required on the mobile helper's params, so no call site can be added without declaring who asked.
Every user path (mobile tab switches, paired tab clicks, shortcuts, palette, the pane's own open) is labelled 'user'. The only automatic sender in the codebase is `waitForResubscribeHostSessionHandle`, the #11542 reconnect probe.
Verified per topology: user activation materializes a parked pane under headless serve, a paired runtime client, a completed agent with restoreOnTabOpenOnly, and a running agent whose wake cleared the record. The automatic probe is refused without retiring the surface, and #11542's reconnect tests stay green.
Also fixes a test fixture that made a real bug untestable: the store stub ignored the host id, so mutating the partition lookup to 'local' left the suite green. Correcting it exposed three existing SSH reattach tests that had been relying on that looseness — their workspace session sat in the local partition while their repo was SSH-hosted, a store production would never read. Production was always right; the tests described an impossible world.
Fixes STA-3465.
* fix(mobile): keep healthy relays green through focus and network nudges (F1+F2)
Focus/app-resume nudges probe the active relay instead of suspending it;
network-change nudges replace it make-before-break, suspending only after a
failed dial. Mount, Retry, and host-swap windows read 'connecting' instead of
'disconnected'; the host list keeps last-known worktrees for every
not-connected state and spins instead of rendering nothing.
* feat(mobile): surface the pairing relay path in the pairing log (F3)
The relay candidate was silent during pairing: dialing, E2EE handshake,
director recovery, and the winning path now emit redacted phase lines through
the same connectOptions.onLog the direct path already used.
* docs(mobile): relay UX investigation findings and F0-F10 fix plan
* feat(mobile): name and narrate relay dials while they happen (F5)
migrateTo forwards the dialing session's connecting/handshaking/reconnecting
phases whenever the client is suspended or disconnected — never downgrading a
live session — and exposes getPendingPath so the host card can say
'· Orca Relay' during the dial instead of only after it.
* feat(mobile): race a relay dial when the direct dial stalls (F6)
A 2.5s grace timer starts relay recovery while an unauthenticated direct dial
is still inside its 12s connect window; the race gets one attempt through the
existing mutex/cooldown machinery, cancels when direct authenticates, and
never arms for hosts without a relay endpoint.
* fix(mobile): overlay the protocol gate instead of unmounting the host stack (F9)
A pending status.get used to swap the mounted HostStack for a spinner at the
moment the socket connected, destroying in-flight nested navigation. Once
children have rendered for a host they stay mounted under an opaque
touch-blocking overlay; first visits and blocked verdicts keep the old
behavior.
* fix(mobile): keep loaded data through transient connection blips (F10)
Git history no longer blanks on reconnect (and commit files refetch instead
of caching an offline empty answer), the repo picker keeps its last-good list
when an in-flight repo.list rejects, the diff review's ready-state
preservation actually runs, and proven host capabilities survive a drop
flagged unverified instead of being wiped.
* feat(mobile): coordinate every home deep push and bounce dead resume targets (F4+F7+F8)
Notification taps, the Accounts card, and host-edit now use the shared
mount-then-replace transition (with a focused-route walker so root-layout
scope works); the Resume card renders from the snapshot in a disabled state
so its late arrival can't shift Tasks under the thumb; resume targets are
validated against proven catalog data, and a session route whose worktree
the host proves missing bounces to the host index with a notice banner
instead of stranding on a dead screen.
* test(mobile): cover the resume-target and notice policies (F7)
Key notice dismissal by code so closing one banner cannot swallow a later,
different one, and move the visibility rule into host-route-notice.ts where it
is testable without a screen.
Adds the missing units for F7's decision points: isResumeTargetConfirmedMissing
(unproven catalog is silence, synthetic routes exempt), the validating
last-visited reader, and the notice visibility rule.
* fix(mobile): review-pass hardening for the gate overlay and diff preservation
Adversarial review findings: the reader's hunk position now survives a
connection blip (reset only on item change), the covered stack is hidden from
TalkBack while the gate overlay is up, and the overlay's hit-test comment is
scoped honestly to in-tree views (native-Modal drawers present above it —
follow-up).
* fix(mobile): CI + CodeRabbit review fixes for #12609
Move the findings doc under docs/ (root directory guard), drop two unused
eslint-disable directives, and address review findings: an unproven snapshot
seed can no longer downgrade a proven worktree catalog; a locally-aborted
relay dial skips the director fallback; post-migration bookkeeping failures
log instead of masquerading as dial failures (which could suspend the healthy
session); the auth wait arms its timeout before subscribing; forwarded dial
phases stop at close(); the legacy selector_not_found fallback requires
runtime_error; the diff-loading effect depends on the fields it reads; and
host-edit auto-cancellation is now pinned by a test.
* fix(mobile): second review round — queued replacements, race fence, confirmed bounces
A network-change replacement now survives the recovery mutex and cooldowns as
a queued intent instead of being dropped or suspending a healthy session —
only a failed dial or a dead probe tears one down. The happy-eyeballs
migration withdraws when direct authenticated during the relay dial
(first-authenticated-wins). A worktree bounce requires two consecutive
host-proven misses, since a transient desktop repo-scan rejection answers
selector_not_found for a live worktree. Background network flaps no longer
wake a billed relay splice, the lifecycle foreground flag stays in sync, a
screen unmount cancels only its own pending host-stack transition, and diff
review keeps the loaded review when its reconnect refresh rejects.
Extracted mobile-endpoint-nudge-router.ts and the establisher's dialEligible
pass, and split the supervisor nudge tests, to stay under max-lines.
* fix(mobile): satisfy the React Doctor changed-code gate
Render-phase ref writes move into effects: the protocol gate's resolved/mounted
latches now record committed outcomes only (a discarded children render can no
longer count as mounted), and the bounce hook syncs its callback ref in an
effect. Array<T> annotations become T[] in the extracted modules.
* fix(mobile): keep the loaded diff when the reconnect refetch rejects (F10)
The diff-loading hook's catch was the one path still erasing a ready diff —
the same keepLoadedDiff guard its disconnect and loading branches already use,
now pinned by a reject-after-ready test.
* fix(mobile): process foreground revival nudges
---------
Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com>
* fix(mobile): label tool rows with a clean summary, expand full input (STA-3333)
Mobile tool rows showed the raw input JSON (`{"file_path":…}`) as the row
label, and the expanded detail just repeated that same truncated string.
- `describeToolInput` labels a row with the target file path, else the
primary argument (command/cmd/query/pattern/url/description), else the
bounded JSON preview.
- Codex delivers tool arguments as a JSON string; normalize those into the
object shape the helpers already understand, so labels, file links,
run summaries and the expanded detail all work for Codex calls too.
- The expanded detail now renders the fully formatted input, capped at
MAX_TOOL_RESULT_CHARS like desktop's tool detail (and like the result
body), and a structured input makes the row expandable.
* fix(mobile): name search rows by their term and keep the filename in path labels (STA-3333)
Review follow-ups to the tool-row summary, all in the shared helper:
- A Grep/Glob row labelled itself with the directory it scanned and dropped
the pattern entirely, because `toolFilePath` treats `path` as a file target.
That path is a scan root, so it also rendered a tap-to-open link that asked
the app to open a folder. `toolFilePath` now ignores the generic `path` key
for search-shaped input, which lets the pattern win the label and drops the
bogus link; an explicit `file_path` still wins.
- An overlong path was truncated from the head, cutting off the basename —
the one part that tells two rows apart. Trim from the front instead, so
the label reads `…/session/MobileNativeChatMessage.tsx`.
- The primary-argument chain used `??`, so a present-but-blank key selected
itself and swallowed the keys ranked after it, dropping the label all the
way back to raw JSON. Take the first key that actually yields a label.
Refs STA-3333.
* fix(mobile): don't offer an expander whose detail repeats the row (STA-3333)
An empty tool input formats back to the row label verbatim, so `{}` and `[]`
advertised an expander and then re-showed the label — the same repeat-the-JSON
problem this change set out to remove. Gate `isStructuredToolInput` on the
collection actually having contents; the lazy detail path is untouched.
Also pins the overlong-path test to the path itself: asserting only length<=80
plus a `…` passed just as well with path labelling deleted.
* fix(mobile): gate the tool detail panel on having detail (STA-3333)
The Tools toggle opens every row at once, bypassing the row's tap guard,
so a row with nothing to expand rendered its own label again underneath
itself — and the tap that would dismiss it is a no-op. Matches desktop.
* fix(mobile): keep a blank tool argument out of the run header (STA-3333)
Skipping a present-but-blank primary key let `briefToolArg` fall through
to the raw JSON preview, so a run header read `Bash {"command":""}` where
it used to read `Bash`. Also state the search-path trade-off honestly:
suppressing the link costs a file-scoped search its tap target.
* fix(mobile): only treat a blank primary key as a missing argument (STA-3333)
The previous guard tested key presence, so a populated but non-string
argument — a mixed argv like ['kill','-9',pid], or a structured query —
dropped out of the run header instead of falling back to the preview.
* test(mobile): pin the tool-row chevron to the detail panel (STA-3333)
The panel gate was covered but the chevron beside it was not: swapping
`showDetail` back to `expanded` on the icon alone left all 909 mobile
tests green, so the affordance lie this branch fixes could return
unnoticed — a down-chevron over no panel, on a row whose tap is guarded
off.
Asserts both icon counts on the fixture that test already renders. The
two halves now die for distinct reasons: the panel gate on the duplicate
label text, the chevron on the icon count.
* test(shared): pin the blank-search-key guard in the tool label (STA-3333)
Dropping `.trim()` from summarizePrimaryToolArg left all 32 tests green,
yet it leaks through isSearchToolInput: a whitespace-only `query` starts
counting as a search term, which suppresses `path`. One character takes
the row's label, its tap-to-open link and its run-header argument at
once, and puts the raw JSON label back — the bug this branch removes.
Asserts all three outputs on that shape. Kills only that mutant; the
isSearchToolInput mutant still dies on the existing search test.
* fix(native-chat): share tool input display semantics (STA-3333)
Build the tool row label, file target, detail eligibility and bounded detail from one normalized input model. Mobile no longer reparses JSON-string input across independent helpers or repeats an already-complete plain label, and desktop now uses the same clean row summary instead of retaining raw JSON.\n\nKeep full detail formatting lazy for collapsed rows and share the 4000-character detail cap across both renderers. Tests pin desktop adoption, mobile disclosure parity, one-pass JSON parsing and the shared bound.
* fix(mobile): keep native chat ask dismissals tab-scoped and gated
Dismissal state lived in the chat view subtree, which unmounts on a
chat<->terminal toggle, so an answered ask card came back on return. It
also had no tab scope and no waiting/blocked gate.
- move dismissal into the controller, keyed per session tab
- gate ask cards on waiting/blocked like the permission path already is,
and retire a dismissal off the ungated detected prompt so a working/done
status can't be mistaken for the prompt clearing
- ignore a dismissal that settles after its prompt cleared or was replaced
Refs STA-3333.
* fix(mobile): keep an ask dismissal through the transcript re-subscribe
A view toggle or tab switch re-subscribes the native-chat transcript, and
useMobileNativeChatSession withholds `messages` until that read settles. A
transcript-derived ask therefore reads as null while the chat surface is
already visible, so the reset effect took it as "the agent moved on" and
retired a live dismissal — the answered card came back, which is the bug
the off-chat guard was meant to close.
Treat an unobserved null as unobserved: `observing` now also requires the
read to have settled. A prompt that is already detected stays observable on
its own, so a status-derived ask still registers on first paint and an
answer taken during that first load is still accepted.
* fix(mobile): keep the transcript-derived ask outside the paused gate
A hook row idle past AGENT_STATUS_STALE_AFTER_MS (30m) projects to `done`
with no interactivePrompt, so the transcript fallback is the only source
left for a still-pending question. Gating it behind waiting/blocked made
that question unanswerable from mobile. Only the sticky status payload
needs the gate; `extractPendingAsk` clears itself on the tool result.
Also pins the load-window clause in the ask-observability guard, which
was behaviourally load-bearing but killed no test.
* fix(mobile): treat a never-read transcript as unobserved, not as "no ask"
The ask-observability guard only excused `transcriptLoading`, which is true
for an in-flight read alone. useMobileNativeChatSession also withholds
`messages` when the client is gone ('idle') or the tab has not reported a
provider session yet ('waiting-session') — both leave the flag false over an
empty list that was never read. The derived prompt then read as null, the
reset effect took that as "the agent moved on", and a live dismissal was
retired; when the read landed with the question still pending the answered
card came back — the resurfacing bug this guard exists to close.
Gate on the read having actually settled instead. 'error' still counts: it
keeps the last successful read in `messages`, so a prompt that clears under
it is real evidence, unlike a list that was never populated.
Also locks three guards that killed no test: the sticky-status suppression
of the transcript fallback (which is what makes the new paused gate hold in
the post-answer window), the reset effect's identity bail-out, and showAsk's
empty-prompt case. The transcript stand-in now derives `transcriptLoading`
from `status` the way the real hook couples them, so these tests can only
express states the session hook can reach.
Refs STA-3333.
* test(mobile): pin the ask dismissal's tab scope and ungated retirement input
Both wirings were unpinned: swapping `scopeKey` to a constant or feeding the
gated `ask` in as `detectedAsk` left the whole mobile suite green.
* fix(mobile): require a landed read before an errored transcript retires a dismissal
`status === 'error'` was treated as settled on the claim that an error keeps
the last successful read in `messages`. That only holds for an error that lands
on top of an earlier read. The host forwards an initial-drain failure as an
error frame carrying an EMPTY list (transcript-watch-error.test.ts), the mobile
frame applier checks `frame.error` before the messages array so those rows are
discarded, and the session hook's error path never calls `setMessages` — so a
first-read error leaves `messages` at the `[]` the identity-change effect wrote.
That frame is also not terminal: the watcher keeps `initialDrain` true and a
real snapshot follows once the read recovers. So a re-subscribe whose first
read errors made the never-populated list read as "no ask", retired the live
dismissal, and the recovered snapshot brought the answered card back over the
composer — the exact resurfacing this guard exists to close, and most likely on
remote/SSH transcript reads.
Require rows for the error case. Rows can only be present once a read landed,
so the predicate is never wrong in the resurfacing direction; it only declines
to retire a dismissal when the transcript was never observed.
Also drop the dismiss hook's `detectedAsk = ask` default and make both prompts
required. That default silently fed the gated prompt in as the detected one,
which is the pre-fix behavior: a paused-out card would read as "prompt gone"
and retire the dismissal. tsc now enforces the ungated payload at every call
site instead of leaving a trap for the next caller.
* fix(mobile): scope the ask dismissal to the provider session, not the tab
A restart, /clear, or resume swaps the provider session inside one tab. The
next session's first question is often byte-identical, so a tab-keyed dismissal
hid the live card and left the turn blocked with nothing to act on.
* chore: restore upstream formatting
* feat(mobile): native-chat model/session-option picker + shared slash catalog (STA-3332)
Piece A — shared slash catalog + send classification:
- Mobile composer now serves getVerifiedNativeChatCommands from the shared
catalog (agent-aware, with description rows) instead of a hardcoded
provider-agnostic list that advertised commands Claude does not have.
- classifyNativeChatSend moves to src/shared/native-chat-slash-commands.ts
(renderer re-exports keep desktop import paths stable); mobile's send seam
now gates optimistic echoes on it, so slash sends no longer create a
'Queued' bubble that no transcript echo can ever retire, and the
ack-lost hold only arms for chat sends.
Piece B — mobile model/session-option pickers:
- New per-tab session-option tracking (state/commands/labels modules) ported
from the desktop live flow, reading the shared agent-session-option
catalog for Claude AND Codex.
- Composer pill row (model + options) opening an inline choice card in the
proven Ask-card pattern; applies use catalog modelApply semantics
(/model <value> via the existing send path), Codex-style agent-picker
entries dispatch the picker command and flip the tab to the terminal view.
- Current model seeds from the hook-reported provider model when derivable;
typed /model-style commands update tracked state (recordOutgoingCommand
parity); dispatched values render as sent-not-confirmed.
* fix(mobile): keep session option sends scoped
* fix(mobile): synchronize native chat refs after commit
* refactor: share native chat session option logic
* fix(mobile): keep the live tab's session-option record from eviction
`getScopedRecord` returned an existing record without re-inserting it, so the
per-tab record map evicted by insertion order rather than recency. A long-lived
active tab is the oldest key, so crossing the 32-scope cap silently dropped its
tracked model and reset the pill to "Model". Desktop's scope cache does
delete-then-set for exactly this reason.
Also moves the shared session-option tests to src/shared so the root suite runs
them (they only exercised src/shared logic the Electron renderer consumes, but
sat under mobile/ where only mobile's vitest project sees them), and restores
two "why" comments dropped while extracting the shared modules.
* fix(mobile): stop a stale session-start report reverting a model pick
Re-entering a chat tab re-delivers the same `agentStatus.model`, and the
reported-model effect re-applied it unconditionally — so picking a model, moving
to another tab, and coming back reverted the pill to the model the agent reported
at session start, which cannot have observed the `/model` sent after it. The
status stream reconnecting had the same effect.
A report is now only treated as evidence when the matched catalog id CHANGES for
that scope; a genuinely new report still supersedes a local pick. Mobile has no
screen read to confirm a switch against, so the repeat is all we can key off.
* fix(mobile): close four session-option picker defects found in review
D1 — a picker apply could interleave with a composer send. The composer already
blocks a text send while an apply is dispatching, but not the reverse: the host
spaces a send's body and its Enter ~500ms apart, so an apply tapped inside that
window was submitted as part of the user's prompt, and the pill then claimed a
model change that never ran as a command. The pickers render inside the composer,
so they now take its in-flight state directly — the same guard, mirrored.
D2 — an option was filed under the wrong model. `setTrackedSessionOption` resolves
the owning model when it commits, not when the command was built, and the report
effect mutates the same record off-queue. A report landing mid-dispatch therefore
recorded `/effort low` against the model it switched TO. Ports desktop's
supersession guard, which skips the commit when the baseline moved.
D3 — a command template's prefix also matches prose that starts with it, so
"/model is a weird word" tracked that prose as the current model, rendered it as
the pill label, and matched no catalog model, dropping every per-model option.
Parsed values are now canonicalized against the catalog; a typed value containing
whitespace is treated as a prompt rather than a command.
Perf — `/` on a Codex tab returned all 45 commands into a non-virtualized
ScrollView showing ~5, re-reconciled on every streaming tick above the transcript.
Capped at 12.
Also splits the row primitives out of MobileNativeChatSessionOptionPickers.tsx,
which the D1 guard pushed to 402 effective lines against a 400 cap.
* refactor: share the session-option display ordering
CATEGORY_ORDER and the non-model sort were byte-identical in
NativeChatSessionOptionPickers.tsx and mobile's labels module — pure logic with
no i18n in it, so there was no reason for two copies that can drift. Both now
call sortNativeChatSessionOptions from the shared snapshot module.
* refactor(mobile): align model picker layout
* style(mobile): round native chat composer
* fix(mobile): inset rounded chat composer
files.resolveTerminalPath began returning a foreign worktree id + relativePath
for absolute paths owned by a sibling workspace, with no protocol or capability
gate. Mobile 0.0.36 in the field ignores resolved.worktree and reuses its own
worktree id for the follow-up files.open, so a tap on a sibling-worktree path
opened the WRONG worktree's copy of that file (on 1.4.168 the tap was a safe
no-op).
Gate the sibling-workspace lookup behind a new optional crossWorkspace request
field: clients that honor resolved.worktree opt in; everything else keeps the
pre-sibling-resolution contract. Old servers strip the unknown field (zod), so
every version pairing degrades to the safe legacy behavior. Optional-field
addition, so no RUNTIME_PROTOCOL_VERSION bump per protocol-version.ts rules.
The terminal-path RPC tests move to files-terminal-path-resolution.test.ts
because files.test.ts sits at the max-lines cap.
Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com>
* perf(runtime): gate terminal.list visual layouts and stop the false writable claim
visualLayouts is ~31% of a large terminal.list payload (44,208 B of 137,412 B on a live 134-terminal remote runtime) and has exactly one consumer: the human-readable CLI formatter. Gate it behind an includeVisualLayouts request param that defaults to included, so pre-flag clients are unaffected, and have every --json/internal caller opt out.
Also drop the record-backed builder's writable, which was a verbatim copy of connected. terminal.show now states writability explicitly as exactly what terminal.send's PTY gate enforces.
* test(runtime): type the payload-size fixture arrays for tsc
* fix(runtime): preserve terminal list compatibility
* test(runtime): guard terminal list optimization
* fix(cli): preserve agent access to terminal layouts
* fix(mobile): open the Resume workspace through a mounted host stack
Tapping Resume on Home landed on a blank host screen instead of the
session. A cold push straight into the nested /h/[hostId] navigator
resolves to the host index route without the dynamic id, so
HostProtocolGate mounts with hostId undefined and never connects.
Home already worked around this for the host editor (#11635) and Tasks
(#11853) by mounting /h/[hostId] first and replacing it once the stack
is committed. Extract that mechanism into host-stack-navigation so
Resume uses the same transition instead of a direct push.
The previous Resume fix (#11876) swapped the manual href for a typed
dynamic href, but expo-router's encodeParam already applies
encodeURIComponent to dynamic segments, so it resolved to the same URL
the manual string produced and left the cold-navigator path unchanged.
Claude-Session: https://claude.ai/code/session_01RMoaxp7MLg2ydP28KFLX7B
* fix(mobile): harden the host-stack transition after bot review
- match a host route committed as the encoded segment it was pushed as,
so an id containing `/`, `#`, or `%` still triggers the REPLACE
- share one pending transition across the Home entry points; per-hook
refs let a Tasks tap and a Resume tap arm two pushes that could not
cancel each other
- assert the source markers before slicing in the Resume wiring test
Claude-Session: https://claude.ai/code/session_01RMoaxp7MLg2ydP28KFLX7B
* test(mobile): lock the host-stack transition state machine
Assert the replace waits for the host mount (zero dispatches before it,
exactly one after), and cover cancel/retarget — the paths the shared
pending transition relies on.
* test(mobile): model listener removal in the navigation harness
A no-op unsubscribe let setState keep calling a canceled listener, so the
teardown assertions only exercised the active guard. Dropping the
unsubscribe call from dispose now fails the suite.
---------
Co-authored-by: kaynan <kaynan.camargo@terceiro-sky.com.br>
* fix(mobile): keep paged chat history coherent across reconnect replays
The transport replays nativeChat.subscribe with its original params after an
in-place reconnect, and the session hook treated every snapshot as a fresh
base — so a socket blip truncated paged-in history back to the initial 40.
A replay snapshot that extends a contiguous retained tail now merges in by id;
a disjoint replay (long outage, compaction while away) still replaces, since
stitching would leave a silent gap. Only a genuinely replaced window resets the
grown read limit and paging cursor, and any snapshot invalidates an in-flight
older-page request so a stale cursor result cannot land on the new window.
Refs STA-3333.
* test(mobile): pin the replay contiguity-scan rejection branches
The scan's three rejection rules were unpinned: deleting the
`sawNewMessage` guard, the ordering check, or the retained-tail anchor
each left the whole suite green. Cover the interleaved-new-row,
reordered-id, and short-of-tail cases, and the replaced-window hasMore
fallback. Each new test is mutation-proven to kill exactly one mutant.
* test(mobile): pin replay paging-metadata and base-snapshot authority
Two more branches of the replay logic were unpinned. Adopting a replay's
`beforeOffset` when it starts partway into paged-in history would make the
next loadEarlier re-fetch on-screen rows and prepend duplicates; treating a
post-replacement snapshot as a replay would retain a row the authoritative
window dropped. Both mutants now fail exactly one test.
* test(mobile): pin the replay removal boundary and base-snapshot bookkeeping
Two branches introduced by this PR survived the suite unpinned:
- `firstIndex > 0` was only pinned one-directionally. Weakening it to
`firstIndex > 1` kept all 28 tests green, so an off-by-one would silently
retain one row the host had already dropped.
- `snapshotSeenRef` is set only for snapshot frames. Setting it
unconditionally is invisible in normal flows, where the first frame is the
snapshot, but demotes the real base snapshot to a replay when a live append
lands first.
Each new test kills exactly one of those mutants and nothing else. No source
change.
* test(mobile): pin the older-page fence against a cursor-re-cutting replay
The snapshot arm of `if (applied.windowReplaced || frame.type === 'snapshot')`
was unpinned: deleting it kept the suite green.
It is load-bearing. A replay that merges cleanly can still carry a new
`beforeOffset`, which the hook adopts via `replayStillStartsAtOldest`. The page
already in flight was addressed with the old offset, so without the fence it
lands and writes its own stale cursor back over the fresh one, leaving the next
`loadEarlier` addressed from a byte offset that no longer describes the file.
Row order alone stays correct, which is why the ordering-only reasoning missed
this.
Sole failure under the mutation. No source change.
* fix(mobile): make native-chat file links and path citations tappable (STA-3331)
- Linkify POSIX absolute paths in chat prose (leading-/ regex alternative;
URL guard now keys off the char before the matched slash)
- Parse agent-style path:line(:col) citations in prose, code spans, and the
open flow; line/column ride into the mobile file preview route
- Route non-web markdown hrefs (file: URIs, relative/absolute paths) to the
file opener instead of silently dropping them; unknown schemes stay dead
- Resolve chat paths against the worktree root, not the terminal's live cwd
- Reuse the terminal tap-to-open flow for chat taps (haptic, preview route,
tab activation with retries) via a shared identity-stable hook, and toast
on misses instead of silent no-ops
- Keep snake_case paths whole (intraword underscores are literal text),
scan bold/italic/strike spans for paths, split trailing punctuation off
autolinks, and let taps land while the composer keyboard is up
* fix(mobile): harden chat file tap handling
* refactor(chat): share native chat href routing
* fix(mobile): detect files directly under path roots
* fix(mobile): keep inline tokens and dunder paths intact around emphasis
Review follow-ups on the chat file-link work:
- A rejected intraword `_` token left the scan index past its closing
underscore, so every inline token between two snake_case words was
swallowed and rendered as literal source — including markdown links,
which became untappable. Rescan from just past the opening delimiter.
- Treat a path separator as an intraword flank so `src/__init__.py` and
`a/__tests__/x.ts` stay whole; previously they rendered as bold plus a
remnant that the new absolute-root pattern turned into a tap on `/x.ts`.
- Bound the `:line(:col)` tail so `src/app.ts:1e3` and `:80%` no longer
parse a line number, while a cited range still opens its first line.
- Route chat tap failures through the composer banner (toast fallback):
chat taps happen with the keyboard up, which covers the toast.
- Drop the tap-handler mirror's dep list; the call site rebuilds its
accessors every render, so it could never skip on a route that
rerenders per keystroke.
* Revert "fix(mobile): keep inline tokens and dunder paths intact around emphasis"
This reverts commit 308bfaf22b88dafc5c43c6a2b8fb8b73c40e2972.
* fix(mobile): preserve chat file-link parsing and feedback
* fix(mobile): serialize native chat PTY writes
Two composed native-chat write sequences into one PTY interleaved their
bytes: the per-terminal send-in-flight guard lived inside the image
attachments hook, so ask answers, permission choices, and question
answers wrote straight past it. Move the guard to a shared module-scope
write lock (the terminal outlives any one screen) and take it on all
four paths.
Ask answers additionally queue behind the prior chain's RPC rather than
racing it, and a superseding answer is fenced once a key has actually
landed: an accepted or ambiguously-delivered keystroke already moved the
remote selector, so a replacement's from-scratch key plan would answer
the wrong question. A superseding answer inherits the cancelled chain's
hold (refcounted, last chain out releases) so changing your mind
mid-answer still works, and Stop/cancel stay unguarded so an interrupt
can never deadlock against the send it cancels.
Refs STA-3333.
* fix(mobile): report a fenced native-chat answer instead of dropping it
The fence added for superseding Ask answers returned false with no
onSendError, and the card re-enables on a false result — so a queued
answer vanished with no banner and no toast, looking exactly like a dead
button. Every other bail in answerAsk reports. Keep the generation-
mismatch branch silent: a newer answer owns the error surface.
Also covers the hold-count release, which had no test at all: replacing
it with an unconditional release left all 58 tests green while silently
reopening the terminal mid-sequence — the exact interleave this PR fixes.
* fix(mobile): stop a superseded answer from clearing the fence banner
A chain that finished its key plan reported `true` even after a newer answer
superseded it. The route sends through useNativeChatAcceptedAction, whose
accepted callback retires the send-error banner — and that callback runs after
the successor's fence report, because finishTurn() fires in `finally`, before
the chain's own promise settles. So the successful predecessor deterministically
wiped the fence message the successor had just raised: every healthy write took
that branch, which made the previous commit's report vacuous exactly where it
mattered.
A superseded chain now reports no success, matching every other supersession
checkpoint in this hook.
* fix(mobile): fence on the turn slot, not the generation counter
The supersession guards added in c36f2ce14e read `generationRef`, but
`cancelPending()` bumps that counter from three callers with no successor
chain: Stop, ask-cancel, and the lease effect that fires on every
disconnect. An answer whose key had already landed then reported false —
`fail()` never runs on an accepted write, so the card stayed up silent
with Submit re-enabled, and the retry wrote the same key into a selector
that key had already advanced.
Test the turn slot instead: a successor takes it synchronously before
its first await and cannot resolve ahead of this chain, so it means
"a successor took over" exactly, without catching bare cancellations.
* fix(mobile): report a fenced answer when a dropped lease bumps the generation
The fence-report guard tested generationRef, which answers "was I
cancelled", not "did a successor replace me". Stop and ask-cancel both
write an Escape that clears the ask card, so their silence is harmless.
A dropped input lease bumps the same counter and writes nothing: the
card stays up with Submit re-enabled while the predecessor's option key
has already advanced the live selector, so the retry double-steps it.
Gate on the turn slot, matching the two success returns.
* test(mobile): cover the turn-slot release after a landed answer
The slot delete in the finally was the one line in this hook no test
killed: without it a landed answer's resolved-false turn stays parked,
so every later answer on that handle reads a fenced predecessor and the
ask card dies after its first use.
* test(mobile): keep PTY write locks terminal-scoped
* fix(mobile): keep repeated-prefix chat replies streaming
Text alone can't tell "the transcript caught up with this stream" from
"a new reply repeats the previous turn's prefix", so the old suppress-on-
prefix rule swallowed genuine repeated replies. A stateful gate remembers
which transcript tail predates the current stream segment and hides the
bubble only when that tail moved during the segment, scoped to the active
host/workspace/tab/session so a swapped chat can't inherit a baseline.
Refs STA-3333.
* fix(mobile): keep the streaming gate alive across chat/terminal toggles
The gate lived in MobileNativeChatView, but MobileNativeChatOverlay returns
null whenever the user peeks at the terminal — that unmounts the view and
throws the baseline away, so the repeated-prefix reply was swallowed again on
the way back. Move the gate (and the fold memo it reads) up to the overlay,
which stays mounted across those toggles.
While hidden the transcript is empty and the throttled stream reports no text,
which the gate would have read as "idle" and re-anchored on. Pass the agent's
working state so a textless tick inside a live segment holds the baseline
instead. The scope key is now keyed off the tab rather than the view-gated
chat resolution, so it survives the toggle too; streamIdentity keeps its exact
previous value because the delayed-send guards compare against it.
Also drops a dead disjunct in the caught-up test: a null baseline is already
unequal to every real tail id.
* test(mobile): model the real re-show ordering in the streaming-gate tests
The overlay regression test replayed the transcript before the stream text on
the way back from the terminal view. That ordering is backwards: the session
withholds `messages` until a fresh read settles (an RPC round trip) while the
throttled stream text returns in ~50ms — and with the transcript already back,
a gate that got discarded on the toggle still passes. Replay the real order,
which pins the gate's lifetime as intended.
Swaps the hidden-gap duplicate case for the in-view one (a tool frame clears
the assistant text mid-turn), which is where the hold actually earns its keep;
the hidden-gap direction stays covered at the gate level.
* fix(mobile): stop the streaming gate adopting a reply as its own history
A textless status tick was re-anchoring the gate's pre-stream baseline, so
two paths still rendered wrong:
- The reply's transcript push beats its throttled status text whenever the
pane stays `working` past the turn (a live subagent or background task).
The tick in between adopted the just-landed reply as history, and the
status text that followed rendered it a second time — a duplicate bubble,
and a regression against main's suppress-on-prefix rule.
- Peeking at the terminal between turns empties the transcript. That empty
tail was adopted as the baseline, so the next repeated-prefix reply was
swallowed again — the bug this PR exists to fix.
Only a tick that carries a real tail and sits outside a live turn anchors
now, with an exception for a gate that has never anchored: mounted mid-turn,
the first real tail it sees is the best history it will ever get.
Also drop `buildMobileNativeChatData`, a test-only builder this PR had wired
the new gate into; its green test asserted the exact suppression this PR
removes. Its fold/pending/image coverage moves to the builder the view calls.
* test(mobile): pin the textless anchor's text reset
Mutation testing found the `prevText` reset on an anchoring textless tick
unpinned: keeping the previous turn's text there reads the next turn's
opener as a new segment, re-anchors onto the reply that just landed, and
renders it a second time — the same duplicate-bubble class already fixed
twice on this branch.
* fix(repos): remove a paired computer's deleted projects from every connected device
A project deleted on a paired Orca host stayed in every connected client's
sidebar and could not be removed there.
Two independent defects:
1. Host-local repo IPC mutations only sent `repos:changed` to the host's own
renderer (src/main/ipc/repos.ts:2711). The runtime client-event stream was
fed only by mutations arriving over runtime RPC, and clients refetch a remote
catalog only on a `reposChanged` event -- there is no polling on desktop -- so
the deleted rows persisted indefinitely. The shared `notifyReposChanged`
helper now also calls the new
`OrcaRuntimeService.notifyReposChangedForRemoteClients()`
(src/main/runtime/orca-runtime.ts:5175), mirroring the existing
`notifyWorktreesChangedForRemoteClients` precedent. This covers every repo,
project-group and folder-workspace IPC mutation, so renames, colors, reorders
and adds propagate too.
2. Deleting the ghost row on the client routed `repo.rm` to the owner, which
answered `repo_not_found`. `removeProject` wrapped its whole body in one
try/catch, so the rejection aborted the local purge before the `set()`
(src/renderer/src/store/slices/repos.ts:3466) and the delete button appeared
to do nothing. Only `repo_not_found` is now tolerated; any other failure still
keeps the row, and an opt-in `errorFeedback: 'toast'` makes it visible at the
three single-project user-initiated entry points. Bulk and background callers
keep today's silence plus their own aggregate reporting.
Closes#11994
Co-authored-by: Orca <help@stably.ai>
* fix(repos): revert inert RepositoryPane removeProject arg
The settings pane's only render site drops the argument; the toast is
already delivered by removeSettingsProjectFromAllHosts.
Co-authored-by: Orca <help@stably.ai>
* fix(repos): scope duplicate-repo-id deletes to the owning execution host
Cover the cross-host collisions #11994's broadcast now fans out to every paired
device. Same-name projects on different hosts were already isolated (per-host
UUIDs, host-scoped catalog merge and purge) and are pinned by regression tests.
Two same-repo-id paths were not: `repo.rm` with a `path:`/`name:` selector and
`deleteProjectHostSetup` both resolved one row and then deleted by bare id,
taking the sibling host's registration with it.
Co-authored-by: Orca <help@stably.ai>
* test(mobile): align the poll-interval rationale with the new reposChanged emission
Co-authored-by: Orca <help@stably.ai>
* fix(repos): resolve deleteProjectHostSetup's repo row only on the setup's own host
The sibling-host fallback could only ever pick a row on a host the caller
did not name; with no exact match the setup is stale and the existing path
already drops just the setup.
Co-authored-by: Orca <help@stably.ai>
---------
Co-authored-by: Orca <help@stably.ai>
* fix(mobile): keep cached workspace counts across a transient RPC failure
The Home host card showed "12 worktrees · 2 active" until any worktree.ps
failed — a backgrounded app, a Wi-Fi→cellular handoff, or a sleep/resume
that kills the socket mid-request. Two things then went wrong:
- render dropped the counts: `markHomeWorktreeCatalogUnavailable` kept the
proven numbers in state, but the card only rendered them when
`catalogUnavailable` was unset, so the line collapsed to "Worktree list
unavailable" even though the last successful counts were right there.
- nothing re-drove the fetch: the per-host wiring latched a `statsFetched`
boolean on the first connect, and the logical client survives socket
drops, so its reconnect never re-read the catalog. The card stayed wrong
until the user navigated away and back.
Keep the proven counts and flag them stale (`staleCounts`), rendered as
"Last known: 12 worktrees · 2 active"; a host whose catalog never loaded
still reads "Worktree list unavailable" (STA-3123). Replace the one-shot
latch with createHostConnectRefetchGate, which fires on each transition
INTO 'connected' — one refetch per reconnect, no polling timer — mirroring
useWorktreeResync on the host screen. fetchHomeHostWorktreeInfo moves out
of app/index.tsx so its rejection path is covered by tests.
* fix(mobile): bound "Last known" counts and survive a path cutover
Review found two ways the home host card's stale-count fix misbehaves.
1. A migrateTo cutover (relay->direct probe, forced replacement) rejects
in-flight requests with LogicalClientCutoverError and republishes
'connected' from 'connected', so the connect gate never re-arms and the
card latched on "Last known: ..." with nothing left to clear it.
worktree.ps now re-issues on the authenticated replacement, bounded,
like runtime-capability-probe and worktree-create-retry already do.
2. "Last known: N worktrees" had no age bound. The home snapshot is
persisted, so a cold start whose first worktree.ps failed rendered
counts proven days ago exactly like counts proven seconds ago - the case
STA-3123 deliberately rendered as "Worktree list unavailable". Counts now
carry countsProvenAt and expire out of the "last known" wording after
10 minutes; counts persisted by an older build count as expired.
Also, per review: the card derives its own worktree line from
HostWorktreeInfo, so a caller can no longer re-gate the counts away (that
was the original defect), and the derivation is covered by a render test -
mobile/vitest.config.ts never collected *.test.tsx, so component tests
were silently dead. Home stats are keyed by host and summed instead of
letting whichever desktop replied last overwrite the shared header row,
which the per-reconnect refetch made churn on flaky links.
* fix(mobile): age bounds liveness, not the counts; scope the header total to paired hosts
Round-2 review follow-up.
Age bound was anchored on proof time inside the failure branch only, so a
session connected past the window that then hit one failed refresh rendered
the pre-fix "Worktree list unavailable" — the exact case this PR exists for —
while identically aged counts still rendered unlabeled as live whenever the
refresh was merely pending. Age now decides live vs "Last known" and the
failure branch keeps whatever the host last proved; "Worktree list unavailable"
is reserved for a catalog that never loaded.
Header stats summed every entry ever cached, so removing a desktop left its
lifetime numbers in the total for the rest of the session. totalHomeStats now
sums the hosts still paired, which also covers removal from the host screen.
wireHostSubscriptions is the effect body moved verbatim out of useEffect;
react-doctor's effect-needs-cleanup false-positives on `subscribe` inside one
and the changed-code gate has no working suppression path (an inline directive
reads as unused to the plugin-less scan).
---------
Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com>
Completes the 0.0.37 release attempted on 2026-08-03 (run 30791649691
failed on the version assertion). Ships the post-0.0.36 transport fixes:
relay session recovery when the LAN endpoint is unreachable (#12344,
#11368, #11465, #11690) and honest worktree-catalog failure states
(#12235) — the released-app defect class verified live tonight.
Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com>
- skip proactive rotation when the resume confirmation reports renewed=false
(a re-resume provably returns the same unchanged deadline; rotating churned
one session replacement per clamp floor, ~60/hour, until a fresh credential)
- armCredentialReprobe under a held gate mints the tick's pass token so the
effective reprobe cadence stays 60s..15min instead of doubling to ~30min
- registerFailure honors scheduleRetry=false in gate branches: no reprobe
timer is armed while backgrounded/stopped; foreground resume re-arms
- extract RelayRetryDelays and supervisor test fakes into their own modules
(max-lines)
Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com>
"Hide sleeping" swept each project's main workspace out of the sidebar as soon as
it had no live PTY, browser tab or agent — even with "Hide default branch" off.
For a project whose only row is that workspace (a folder workspace, a fresh
clone, a detached-HEAD main), the entire project vanished with no in-place way
back.
Adds a shared `isSleepingSweepExemptWorkspace` predicate keyed on
`isMainWorktree` rather than the branch name, so folder workspaces (no branch),
detached-HEAD mains, and SSH rows whose head/branch are blanked while a provider
is disconnected all stay put. Wired into `computeVisibleWorktreeIds` (sidebar,
Cmd+1-9, workspace board), the jump palette's duplicate inline pass, and mobile's
`filterWorktrees`.
Ships default-on with an escape hatch: a persisted
`alwaysShowDefaultBranchWorkspace` setting surfaced as "Except default branch"
under "Hide sleeping". Explicit "Hide default branch" still wins, since it
filters before the sleeping sweep.
Mobile reads the setting but never writes it back, so a desktop opt-out can't be
clobbered by a filter tap before the ui.get roundtrip lands.
Combines the two PRs open against #8873. #8966's exempt set is a strict subset of
this one, so its production diff was subsumed rather than ported; its jump-palette
render harness and e2e spec were carried over, and are the only such coverage here.
Fixes#8873Closes#8966
Co-authored-by: Rod Boev <rod.boev@gmail.com>
Co-authored-by: Orca <help@stably.ai>
* fix(mobile): keep relay runtime recovery alive without direct connectivity
A phone paired over the relay whose direct LAN endpoint is unreachable
(e.g. a Tailscale IP with Tailscale off) could lose the runtime channel
permanently: the reconnect controller's recovery gates parked with no
timer and no logs, the supervisor snapshotted relay credentials once at
start (dying silently if the read failed and dialing stale tokens after
rotation), and the only path that cleared a rejected-credential gate
required a working direct connection. Field symptom: home card shows
"Connected - Orca Relay" (or "Can't connect - check Tailscale") while
the host page sits at zero worktrees forever.
- gates (fresh-credential, external-signal) now arm a slow 60s reprobe
instead of parking; each gated attempt re-reads the durable credential
bundle and adopts it when its version is fresher than the rejected one
- supervisor start no longer dies for the process lifetime when the
initial Keychain read fails or the bundle is expired
- every recovery decision now reaches logcat and the in-app connection
log ([relay] lines); previously the whole relay dial path was silent
- direct-return probing extracted to mobile-direct-return-probe.ts,
credential selection to mobile-relay-credential-selection.ts
Regression suite mirrors the field failure (rejected outer credential,
unreadable bundle at start, expired bundle, E2EE rejection without a UI
nudge) plus real-rpc-client failover integration tests; the four
deterministic scenarios fail on the previous code.
* fix(mobile): adopt durable relay credentials by outcome, not version
Adversarial review caught two blockers in the version-comparison rule:
renewals extend expiresAt without bumping current.version, and a re-pair
restarts the version counter — both left the durable bundle unadopted
and reproduced the original outage. Selection now adopts the disk bundle
exactly when it yields a dialable (unexpired, non-rejected) credential
while memory does not, which also keeps revoked versions unresurrectable.
Also from review: the gate reprobe cadence now escalates 60s -> 15min
ceiling with 0.75-1.25x jitter (no fleet phase-alignment, no permanent
one-minute beacon); clearing a gate drops its timer, pending tick, and
cadence so an orphaned reprobe cannot swallow the next fast backoff; the
reprobe tick token is only minted while its gate still holds; and a
merely missing/expired bundle uses a plain cooldown instead of the
fresh-credential gate so it cannot force rotations on direct reconnects.
New regression tests (all red on the previous code): renewal without a
version bump, re-pair with a restarted counter, orphaned-timer backoff
swallowing, escalating gated cadence, and background/foreground recovery
after an E2EE rejection.
* fix(mobile): reset gated relay cadence on app resume
Review round 2: an escalated fresh-credential gate kept its cadence
across background/foreground, so reopening the app could wait out a
15-minute tick (measured 11.25min to first attempt after a 2h
background) — indistinguishable from the outage itself. A resume now
resets the streak even when it cannot lift the credential gate, and a
successful direct connection does the same in resetForDirectConnection.
Also: the streak now advances once per fired tick instead of once per
armed-delay computation (three arms per cycle escalated 60s -> ceiling
in ~7 minutes instead of the documented eight steps); delay computation
is a pure read.
* fix(mobile): rotate relay sessions on resume expiry, not attach deadline
Live phone verification of the failover fix exposed a second defect the
old latch had been masking: the relay-hello's leaseExpiresAt is the
cell's attach-reservation deadline (now + 10s for resumes,
credential-store.ts:213 server-side), but the supervisor scheduled
proactive rotation from it with a 30s margin clamped to 1s — so every
relay runtime session force-replaced itself ~1s after connecting
(measured every ~2.5s on device, 253 dials per 5 simulated minutes in
the red test). Any RPC slower than the cycle could never complete,
which is the "Worktree list unavailable" symptom.
The session now captures resumeExpiresAt from the hello (updated by the
resume confirmation) and rotation keys off it. Test fakes previously
used a 120s lease, which is why no suite ever reproduced the loop; they
now mirror the production 10s attach deadline, and a churn regression
holds one session across 5 minutes with direct unreachable.
* fix(mobile): clamp lease rotation delay on both ends
Adversarial review of the resume-expiry rotation fix caught an int32
setTimeout overflow: production resumeTtlMs is 30 days, and
30d - 30s = 2,591,970,000ms exceeds INT32_MAX, so Node (and vitest's
fake timers) clamp the timer to 1ms — 3001 relay dials and credential
writes in 3 simulated seconds, ~2500x worse than the churn being fixed.
The delay is now clamped to [60s, 6h]: the ceiling makes overflow
unreachable regardless of server TTL (a harmless re-resume every 6h on
long sessions), and the floor bounds any bad deadline to one forced
rotation per minute instead of a sub-second loop — which also disarms
the Math.max(1000, ...) landmine for return-unchanged-grace resumes
whose stored expiry can be arbitrarily near.
Also from review: getLeaseExpiresAt is renamed getAttachDeadlineAt (it
had zero production callers left; the plausible name is how the churn
bug happened), the expired-vs-missing bundle cases now log distinct
strings, and both test fakes use production constants (10s attach
deadline, 30-day resume TTL) — fictional fake values hid all three
defects in this subsystem. The four forced-rotation lease tests are
retimed to the 60s floor with direct pinned unreachable so return
probes cannot race their windows.
* style(mobile): merge duplicate imports in relay failover test
* style(mobile): use T[] array syntax in credential selection
---------
Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com>