Commit Graph

510 Commits

Author SHA1 Message Date
Jinjing 8331fe5995
docs: update GitHub star history chart (#13380)
Refresh README star history graphic to the latest ~40K stars curve (through August).
2026-08-09 15:17:11 -07:00
Neil 17cfc968cf
Revert the terminal IME composition-ownership change (#13282)
* Revert "test(ime): restore coverage the composition-ownership change removed (#13168)"

This reverts commit 25a8c517e1.

* Revert "refactor(terminal): return IME composition ownership to xterm (#13128)"

This reverts commit 17b3dff3c4.

* test(ime): keep the architecture-neutral Korean trace coverage

The recorded IBus/fcitx5 and Windows MS-Korean traces from #13168 assert PTY
byte order, not composition ownership, so they still hold once the terminal
composition layer is restored. The mobile accessory-order test pinned the new
handleLiveInputChange signature and does not.

Co-authored-by: Orca <help@stably.ai>

* fix(terminal): keep the macOS Backslash bypass through the revert

The restored native-text forwarder only claims keys for input sources in its
hardcoded CJK allowlist, so third-party IMEs off that list (Qingg, #10896) still
get a raw backslash. #13128 added this bypass as a partial replacement; keep it
rather than trade the open issue back.

Scoped to the bare backslash key. The rest of shouldBypassXtermForMacNativeText
bypassed all unmodified non-ASCII text, which would race the restored forwarder.

Co-authored-by: Orca <help@stably.ai>

* fix(mobile): move the mirror-step ref write out of render

The restored hook assigned runMirrorStepRef during render, which is not
replay-safe — React can discard render work, so the mutation can leak from UI
that never commits. Its only read is inside the held-commit timer, which fires
long after commit, and the ref has a safe default, so an effect is soon enough.

Surfaced by the changed-lines React Doctor gate: the rule postdates this code,
so restoring the file re-introduced it as a new violation.

Co-authored-by: Orca <help@stably.ai>

---------

Co-authored-by: Orca <help@stably.ai>
2026-08-08 19:06:50 -07:00
github-actions[bot] 54a9b5840b Update README downloads badge 2026-08-08 12:28:05 +00:00
Neil 17b3dff3c4
refactor(terminal): return IME composition ownership to xterm (#13128)
* fix(terminal): return IME composition ownership to xterm

* fix(mobile): derive terminal input from native replacement ranges

* test(mobile): record iOS Japanese IME traces

* fix(mobile): preserve native IME replacement ranges

* fix(xterm): flush queued application input after IME commit

* test(terminal): pin Korean intermediate commit

* test: pin Windows IME shortcut ownership

* test: replay IBus number candidate commit

* fix: preserve native macOS input-method punctuation

* refactor(terminal): remove stale mac focus override

* fix(mobile): preserve soft keyboard deletion ranges

* fix: keep IME-owned palette chords in renderer

* fix: stop carried IME shortcuts at renderer owner

* fix: preserve carried IME shortcut dispatch

* fix: narrow main-owned shortcut actions

* test(mobile): pin Japanese IME replacement traces

* test(terminal): retain paired native IME trace

* fix(chat): preserve browser IME composition ownership

* fix(chat): retain macOS IME confirm gesture

* fix(chat): expire unmatched IME confirm carry

* fix(chat): isolate IME confirmation expiry

* fix(chat): retain active IME confirmation

* refactor(terminal): remove dead composition handler

* feat(ime): add shared Enter-ownership seams for CJK composition

The confirming Enter of a CJK composition arrives as two keydowns and the
orderings differ by platform: Windows/Linux redispatch the unmarked Enter/13
before keyup, macOS delivers keyup first. A guard reading only isComposing or
keyCode 229 misses the redispatch, so surfaces submitted on a confirm.

Adds useImeEnterGestureOwnership (carry token, next-frame expiry), a shared
ImeEnterGuardedForm for native implicit submission, and the cmdk seam covering
18 CommandInput surfaces at one site.

A chorded Enter arms the carry but is never swallowed — the reverse would eat a
user's deliberate Cmd/Ctrl+Enter. Both failure modes are pinned by
ime-enter-gesture-ownership-contract.test.ts.

Co-authored-by: Orca <help@stably.ai>

* refactor(terminal): consolidate native input listeners and parked-screen owner

Extracts the shared native-input listener installer and renames the parked-screen
detector for what it actually does, replacing per-call-site duplication. The
listener installer keeps a forgetOptionKeyLocationOnBlur flag so per-window
semantics are preserved rather than flattened.

Net deletion; no behaviour change intended.

Co-authored-by: Orca <help@stably.ai>

* test(terminal): pin recorded IME shapes as regression tests

Nine regression tests built from hashed affected-platform captures, each with a
paired ordinary negative and a discriminating mutation verified to take the file
from all-passing to exactly one failure.

Covers the Windows MS-Korean Shift family (#12179, #11878, #12151, #11946,
#12152) and the Korean TUI line-break rows (STA-3237, STA-3222, STA-3129).

STA-3237 pins the empirical 3-Shift / 2-active-composition / 2-newline ratio the
device run established — the third Shift produces nothing because Space has
already committed. That ratio is not derivable from a static capture.

Co-authored-by: Orca <help@stably.ai>

* fix(ime): guard Enter-commit surfaces against CJK confirm

Applies the Enter-ownership guards across the surfaces whose Enter commits
something: publishes, clones, pairs, installs, posts, or persists.

Tiered deliberately rather than uniformly. Irreversible and remote-effect sites
take the carry token, which also blocks the unmarked redispatch. Locally
reversible sites take the oracle check with a one-line comment naming the
residual, because a spurious commit there costs one undo.

Three numeric fields are left unguarded with the reason in-code: Chromium blanks
number inputs at compositionstart, so a confirm-Enter only ever reaches an
empty-draft reset. Measured with a CDP probe rather than assumed — a guard that
cannot fire is noise.

Co-authored-by: Orca <help@stably.ai>

* test(ime): teeth-check the Enter guards on every guarded surface

One suite per guarded surface, each verified by deleting the guard and
confirming the test fails. A green guard test without that check is unverified,
not verified.

Two shapes pass vacuously in happy-dom and are avoided here: native implicit
form submission never fires, and blur() is inert on an unfocused element. Both
made "the commit did not happen" assertions pass with the guard removed, so the
suites assert the guard's contract directly instead.

Co-authored-by: Orca <help@stably.ai>

* fix(mobile): keep iOS Korean commits whole through the live-input path

iOS Korean reports isComposing: false on every event, so it bypasses the
composition guard entirely. The strict owner rejected UIKit's transformed
post-change field and sent only the leading jamo — the reported symptom.

Prefers the authoritative same-event field text over the predicted text when the
supplied operation cannot produce it. Generic: no Korean special-case, no locale
classifier, no normalization. Adds the RN-target-keyed submit carry alongside it.

Co-authored-by: Orca <help@stably.ai>

* test(e2e): make IME capture harnesses fail loudly instead of silently

Four instruments recorded silence as success, so a void run scored as a clean
one:

- readTerminalImeBoundaryTrace returned an empty trace when the probe never
  installed, making every "nothing leaked" negative pass vacuously
- summarizeLatencies([]) returned a perfect zero distribution that passed all
  three latency thresholds
- the macOS Vietnamese spec pinned an input-source ID that does not exist, and
  failed as though the operator had chosen the wrong source
- the expectedLineCount=1 prefix property was undocumented and one edit from
  silently downgrading a PTY assertion

Input sources now resolve by enumeration and name the near-matches on failure.

Co-authored-by: Orca <help@stably.ai>

* test(terminal): cover Cangjie cancellation and fix a cross-namespace assertion

Adds #11951's recorded Cangjie cancel shape to the existing cancellation suite,
which covered Pinyin and Sogou but not Cangjie. One keystroke then Backspace
arriving as deleteContentBackward with data: null, so the stale preedit is the
only thing a fallback could replay.

Verified against the historical pre-6cd944c62b3 bundle: the positive fails with
['尸'] where [] is expected, while the ordinary negative stays green.

Also fixes the Vietnamese spec, which asserted a TIS-space input-source ID
against getKeyboardInputSourceId(). Those two Orca APIs report the same source
in different namespaces — TIS nests it under VietnameseIM, the app API does not.
The resolver stays as an installation precondition; the assertion matches the
leaf.

Co-authored-by: Orca <help@stably.ai>

* test(e2e): add a real-IME macOS arm for the Korean chord commit

The existing korean-ime-terminal-shift-enter-commit spec synthesizes composition
over CDP: Input.imeSetComposition sets the preedit directly and Input.insertText
performs the commit. Asserting the IME produced events you injected yourself is
circular, so that spec cannot certify real-IME behaviour.

This arm selects 2-Set Korean via TIS, reads it back live, and injects through
System Events key codes, so the OS owns the preedit, the commit instant, and
isComposing. PTY byte expectations are preserved verbatim.

Covers 2 of the original 4 cases by design. The other two are the Windows/Linux
redispatch-before-keyup ordering, which macOS cannot produce and which cannot be
selected -- the OS decides it. Reintroducing synthesis to "restore coverage"
would reintroduce the circularity.

Co-authored-by: Orca <help@stably.ai>

* test(e2e): assert the macOS chord arm at the PTY boundary, not the renderer

The byte expectations were transcribed from korean-ime-terminal-shift-enter-commit
:364/:383, which assert against onData -- a renderer boundary where the terminator
is CR. This spec reads the PTY child, where the tty has already converted CR to LF.

Names both forms per row rather than swapping the constant, so the conversion reads
as evidence that the capture reached past the renderer, as #11936 and #11951 record.
Ctrl+Enter's CSI-u sequence is unaffected and is identical at both boundaries.

Co-authored-by: Orca <help@stably.ai>

* test(e2e): measure composer-to-onData latency and stop dropping IME keystrokes

Two defects in the echo latency probe.

It hooked onWriteParsed and onRender but never onData, so it measured
key->parse->render echo rather than the composer-vs-onData delta the latency rows
need. Adds a third hook feeding its own sample set.

And `event.key.length !== 1` silently dropped IME keystrokes: Pinyin and Cangjie
keydowns arrive as key:'Process' (length 7). Replayed over the captured corpus,
the old filter accepted 580 of 4137 Chinese IME keydowns -- it was discarding 80%
of them. The new filter matches the shape the owner itself branches on.

Attribution charges each onData to the latest keydown rather than a FIFO head,
because composing jamo emit no onData at all and a queue would credit a whole
composition to its first keystroke. The consumer now asserts sample count before
any percentile, so a zero-sample run cannot render as a flawless distribution.

Co-authored-by: Orca <help@stably.ai>

* test(terminal): pin the WSL shifted-jamo newline shape for #11919

In Korean 2-set, Shift types ordinary letters -- the double consonants and the
compound vowels. Each such keystroke reaches Chromium as key='Process',
keyCode=229, shiftKey=true.

The v1.4.163 classifier matched exactly that pattern with no code guard, so it
called those keystrokes Enter, rewrote them to a synthetic Shift+Enter, and
injected a newline into the middle of the word -- with no Enter key pressed.
That is why the reporters said "no modifier key pressed": they had not chorded
Shift+Enter, but they had pressed Shift, to type the double consonant.

Asserts the row's own recorded capture: 40 immediate keydowns, exactly 3 of them
Shift-carrying inside a single syllable, and an onData stream with one newline
per Enter press and none mid-word. Two ordinary negatives keep it from being a
blanket mute -- the same session's non-IME keydowns still reach shortcut policy,
and an ordinary Shift+Enter still resolves through the real policy.

Co-authored-by: Orca <help@stably.ai>

* test(terminal): pin the composition commit lag that made Korean type one behind

macOS Korean 2-Set commits syllable N only when the first jamo of N+1 arrives, so
compositionend and compositionstart land in the same task. A composition-start
handler cancelled the pending finalizer that was the only path to triggerDataEvent
and ended the session without emitting bytes, so every committed syllable reached
onData exactly one syllable late and the backlog cleared only at a Space or Enter.

Types continuously with no Enter and no Space -- either would flush the backlog and
hide it -- and samples onData at every syllable boundary. Paired with a
length-matched ASCII arm that stays green throughout, so the positive is a fact
about composition rather than about timing in general.

Bisected to a single call site across five builds: pristine, 1.4.155 and 1.4.162
pass, 1.4.163 fails, removing the one call repairs it, restoring it fails
identically. That window is exactly the reporter's "started immediately after
updating".

Co-authored-by: Orca <help@stably.ai>

* test(mobile): cover the send-queue abort that silently drops queued keystrokes

One failed send in use-terminal-live-input-commit aborts every keystroke
queued behind it, with the error swallowed by .catch(() => false). The
existing test resolves(true) on every send, so the failure branch was
uncovered.

Four arms: the abort itself, an ordinary negative on the healthy path, a
throwing sender, and a liveness control proving the queue recovers once
the chain settles. Deleting the abort takes 4 passed to 3 failed, with the
ordinary negative correctly surviving.

Scope is stated in the docblock: this is a transport send-queue abort,
reachable only via a real disconnect or RPC error. REQUEST_TIMEOUT_MS is
30s, so latency alone cannot reach the branch — consistent with #7094's
symptom class, not proven to be its cause.

* test(terminal): pin that daemon snapshot/restore cannot disturb a composition

Two independent reporters attributed broken Korean composition to the
always-on PTY daemon repainting terminal state over the preedit. The
attribution is wrong on ancestry — the daemon shipped three months before
the version both call good — but the boundary was never actually tested.

Runs the real applyMainBufferSnapshot choreography against a live
composition, including the full 2J/3J/H wipe plus the resize and
alt-screen branches. textarea.value, selectionStart/End,
compositionView.textContent and .active all survive byte-identical, and
interleaving a restore between every jamo of 문제 still commits 문제 at
onData. Also pins that the uncommitted preedit is absent from the captured
snapshot: it lives in the textarea, never the buffer, so a restore has
nothing stale to echo back.

Injecting one textarea.value = '' into the restore fails exactly the three
restore-boundary tests.

* test(terminal): pin that Cmd tears down a composition where Ctrl and Shift do not

xterm's composition keydown exempts only keyCode 16/17/18 (Shift/Ctrl/Alt)
plus 20/229. macOS Meta — 91/93/224 — is absent, so a Cmd press mid-composition
takes _finalizeComposition(false): the overlay goes dark and never recovers,
because compositionstart is not re-fired. The user composes the rest of the
word blind. Linux and Windows users press Ctrl and are exempt.

xterm already has a Meta-aware modifier predicate in wasModifierKeyOnlyEvent,
so this is an internal inconsistency rather than a deliberate choice.

Owns no reported row and is version-neutral: 5/5 on both 1.4.162 and 1.4.163.
The branch is unexercised in all 328 recorded traces, so this is a hazard pin,
not a regression guard. Only the teardown is asserted; the likely duplicated
commit needs a compositionend the IME kept alive across the Cmd, which no
capture contains.

Deleting the exemption fails exactly the three paired negatives; adding Meta
to it fails exactly the two Cmd arms.

* test(native-chat): characterize preedit loss when a question card replaces the composer

An AskUserQuestion card fully replaces the composer by design, but the
in-flight composition goes with it: the composer unmounts before
compositionend reaches it, so the preedit is never committed to the draft.
The committed text survives only because the draft is cached and restored
via defaultValue. Node identity changes, value 'abc' is preserved, the 가
is gone.

Drives the real NativeChatView -> SessionGate -> InteractiveCard ->
questionActive swap -> Composer -> ComposerField, flipped by writing the
same store field an AskUserQuestion hook event writes. Flipping
questionActive to false fails exactly this test and nothing else across
639 native-chat tests, so the path was entirely unguarded.

CHARACTERIZATION TEST: it asserts the loss. Fixing the defect — committing
the preedit before the swap, or keeping the composer mounted — will make
this file fail. Update the expectations to the new contract rather than
working around them.

Owns no reported row. #12118/STA-3219 flicker is keyed to token counters,
which provably do not remount, and a question card arrives once per
question.

* test(terminal): pin the duplicated commit when Meta interrupts a composition

_finalizeComposition(false) sends textarea.value.substring(start, end) but
cannot clear the IME-owned textarea, so a later compositionend re-sends the
same range. Meta reaches that path because CompositionHelper exempts only
Shift/Ctrl/Alt; xterm's own wasModifierKeyOnlyEvent covers Meta four ways,
so the omission is an internal inconsistency rather than a choice.

Companion to the modifier-exemption guard, which deliberately pins only the
overlay teardown. This pins the data consequence.

HAZARD PIN: owns no reported row. The trigger is unverified on hardware —
no capture in the corpus contains a Meta-during-composition gesture, and
whether macOS keeps the composition alive across it is unmeasured. The
duplication follows from the code given that sequence; whether users reach
the sequence is the open half.

An earlier premise that Space (keyCode 32) reaches this path was refuted by
a corpus scan: 0 of 731 evidence files carry a keyCode-32 Space while
composing, against 171 at 229, and 229 returns early.

* test(terminal): characterize the syllable lost when the textarea blurs mid-composition

CoreBrowserTerminal._handleTextAreaBlur clears the helper textarea
unconditionally — "Text can safely be removed on blur" — while
CompositionHelper._finalizeComposition reads the committed text back out of
that same value from a deferred timeout. By the time it runs the value is
empty, the substring is '', and triggerDataEvent never sees the syllable.
xterm checks composition state in _syncTextArea and omits the same check
here.

Six cases. Blurring mid-composition loses the syllable in every ordering,
including compositionend-before-blur, which is Chromium's real order — so
it is not an ordering artifact. A bare textarea.blur() with no Orca code
loses it too, which places the owner upstream: Orca's unguarded release on
outside pointerdown is one trigger, not the cause. Committing 한 then
blurring mid-가 yields ['한'] where ['한','가'] is correct: one syllable
gone, surrounding text intact.

Teeth checked by inverting — adding an Orca-side composition guard flips
exactly the three cases that route through the release path and leaves the
bare-blur and no-blur cases green, which is the scope split: a fix in
regular-terminal-focus-ownership alone would not close this.

HAZARD PIN, but unlike the others this one has a real production injector —
clicking outside the terminal mid-composition. Owns no reported row. The
shape matches #9738's report; the injector does not, and a shape match with
a mismatched injector is not an owner.

* test(terminal): say which arm the STA-3237 fixture came from

The recorded keydowns are wave 4's A-shift-unmarked-only — the arm that
emits no PTY bytes. Nothing in the file said so, so two readers concluded
the row's events fail the owner's predicate and that STA-3237 and STA-3222
were different defects. They share an owner; the arm that fires is
Process/229+Shift, absent from this bubble-phase trace because the owner
claims it in the capture phase.

Also corrects "code-blind": the v1.4.163 policy emits \x1b\r only for a
shift-only key:'Enter', and a jamo keydown reaches that branch solely via
the isTerminalImeProcessEnter rewrite. The mock is deliberately wider so
the ownership guard stays under test if that rewrite moves.

Comments only — no assertion, fixture value, or mock behaviour changed.

* test(e2e): track the input-source selector the macOS specs shell out to

Five tracked macOS IME specs ran `swift .tmp/select-input-source.swift`, a
file that is gitignored and existed only on one machine. Anyone else
checking out the repo — or the same machine after .tmp is cleaned — could
not run them, and they are the capture drivers for the macOS rows that are
blocked waiting for exactly those runs.

Moves it to tests/e2e/ beside its callers. The chord spec now resolves it
from __dirname rather than reaching two levels up into .tmp.

* test(terminal): pin the CJK repaint decision against the reporter's own output

#12164 comment 1 and #5921 report agent output with double-width glyphs
rendering duplicated character-by-character while ASCII in the same line
stays clean. No IME, no composition, no keystroke — the user never types
the CJK.

Segmenting all three verbatim samples into maximal same-risk-class runs
gives 33 runs and zero violations of "this run is corrupted iff the
production detector flags it": 17 wide runs all corrupted, 16 narrow runs
all byte-identical. The paired negative is co-located in the same line
rather than in a separate run — the reporter supplied it without knowing.

Doubling is asserted as present, not uniform: 자바스크립트 and 시스템 each
leave a jamo undoubled, which is a repaint-region boundary artifact rather
than a per-character transform.

The discriminating arm is in the test rather than a source mutation:
be3f30e2f8 (#6890) elects a repaint for all 17 corrupted runs when the
agent types nothing, and reverting its disjunct elects none. Both
predicates agree once the user has recently typed, which is the pre-#6890
condition.

Samples inlined with per-sample SHA-256 because .tmp is gitignored and
cannot back a landed test.

* test(terminal): pin macOS period substitution landing after the composition

#11504's reporter published a DOM trace showing insertText ". " arriving
149ms after compositionend, when two spaces are typed with a CJK input
source and NSAutomaticPeriodSubstitutionEnabled is on. This replays that
trace against a real Terminal and asserts what reaches onData — bytes to
the PTY, not anything visual.

The owner is stock upstream CoreBrowserTerminal._inputEvent, not an Orca
module, confirmed at the resolved install and in the shipped bundle Vite
loads rather than in the TypeScript source.

Three mutations against that install, predictions written before the runs,
each failing exactly the arms predicted: dropping the composed/keyDownSeen
guard fails two, dropping Orca's intercept fails the one arm where the
payload arrives before the send window drains, and flipping || to && —
the candidate-fix shape — fails the arm that pins the defect itself.

CHARACTERIZATION: arm 1 asserts the broken behaviour and will fail the
moment #11504 is fixed. Update it to the new contract rather than working
around it.

composed is absent from every recorded bundle, so composed: true is the
spec-required value rather than a captured one; the test asserts it before
dispatching so a harness that dropped the field fails loudly.

* test(terminal): replay the recorded Windows Shift sessions through the IME guard

STA-3179 reports a Shift release sending Enter; #12171 reports delayed
Hangul plus doubled newlines. Both replay their own recorded Windows
MS-Korean keydowns through resolveTerminalKeyboardShortcutAction with the
shortcut policy mocked, so the assertions are about which events reach the
policy and what reaches terminal input.

STA-3179's held-Shift gesture yields exactly one newline, from the unmarked
Enter alone; its release arms nothing for the next composition, asserted
after a precondition check that the release really is keyups with shiftKey
already dropped; and an ordinary Shift press-and-release still routes every
keydown, which is the paired non-IME negative.

Teeth, verified by mutation: bypassing the isImeOwnedKeyboardEvent guard in
keyboard-handlers takes STA-3179 from 3 passed to 2 failed / 1 passed — the
survivor being the ordinary-session negative, which is correct, since a
non-IME session should not depend on that guard — and #12171 from 2 passed
to 2 failed. Source restored byte-identical.

Recorded shapes are inlined and the bundles cited in comments; nothing is
imported from .tmp, which is gitignored.

* test(native-chat): correct 61977d45177 — the preedit survives the question card

61977d45177 claimed a composed syllable vanishes silently when an
AskUserQuestion card replaces the composer, and characterized that loss.
The claim was false. Its premise was an artifact of the harness: the test
simulated a preedit with a silent textarea.value assignment and no input
event, which no IME does.

Real composition fires input with insertCompositionText on every keystroke
— the shape this repo already records in its own observed-event capture —
and React's change handler returns on input/change with no composition
gate, so onChange runs for each frame. The draft cache is written
synchronously inside the updater, so the preedit is already committed
before the card can arrive. Driven that way, it survives.

Renamed to match the contract that actually holds, and extended: Hangul
jamo-per-frame, Japanese kana accumulation followed by per-segment
conversion asserting the candidate the user was looking at survives, and a
pin on the mechanism itself — the draft cache holds the preedit while the
card is up.

Teeth: there is no fix to revert, so the mutation is the plausible wrong
one — gating onChange on isComposing(). That takes 4 passed to 3 failed,
with the English negative correctly surviving, since it has no composition
to gate.

Two consequences remain, recorded rather than fixed: the OS aborts the
composition when the field disappears, so a lone jamo returns as a
compatibility jamo the user cannot compose onto, and the remounted
composer is unfocused because the card owned focus.

* docs(native-chat): name the corrected commit and the degraded-jamo consequence

Records in the file itself that 61977d45177 is pushed and wrong, quoting
the two claims that are false, so a reader who finds it in git log reaches
the correction from the file that replaced it.

Also states the residual as a consequence rather than a curiosity: a lone
leading jamo returns as a standalone compatibility jamo (U+3131), which is
not a composable state — the user cannot resume the syllable, only delete
and retype. Preserved, but degraded into something unusable. That is the
note to find if a reporter ever describes exactly that.

The invariant these tests pin is not "the composer commits on unmount" but
"composition input events must reach React" — which is what a future IME
change would break, and is not visible from the swap site at all.

* test(terminal): replay the recorded macOS Telex commit boundaries

#6905 reports Vietnamese composed characters breaking in the terminal.
Replays the retained macOS built-in Simple Telex capture — recorded
selection and value set before each dispatch, since that is what the
commit range reads — and asserts what reaches onData: the first commit
alone, then through the real Enter, then the ASCII tail of the same run.
A code-point count would catch NFD normalisation.

ENGINE CAVEAT, stated first in the docblock: this is macOS built-in Simple
Telex, Telex only. The reporter's three named engines cannot run on the
platform they declared, and which macOS Vietnamese engine they used is
unconfirmed. This file certifies no engine, and does not imply VNI.

The owner is upstream's — CompositionHelper._finalizeComposition's
waitForPropagation branch — so the arms are copies under .tmp aliased by a
scratch config, with node_modules verified unchanged by shasum after every
run. Collapsing the range end onto its start fails all three; collapsing
the start to zero re-emits the first word into the second commit, which is
the reporter's "duplicated" direction. Different failure sets, so the
mutants are distinguishable rather than merely detectable, and the ASCII
assertion passes under both.

Falsifiability here is by mutation, not by a defective build: #6905 does
not reproduce at HEAD, so this has never been watched going red on a real
reproduction.

* docs(terminal): lead the #6905 test with its engine caveat

Comment-only. Moves the caveat above the source line so a reader meets what
the file does NOT establish before what it does — the capture is macOS
built-in Simple Telex, the reporter's named engines cannot run on the
platform they declared, and which engine they used is what gates this row.

Co-authored-by: Orca <help@stably.ai>

* docs(terminal): record that the swallow eats a keystroke after Japanese conversion

This pin framed the swallowed insertText around Cmd interrupting a
composition. A differential through Japanese multi-segment conversion shows
it is broader: type a segment, convert, then press `a`, and the `a` is
lost. No modifier, no exotic gesture. Korean surfaced it first only because
2-Set composes on nearly every keystroke.

Also records why it cannot simply be fixed. The suppression de-duplicates
IMEs that deliver their commit a task after compositionend, which a sibling
test pins; this swallow is that dedup's false positive, and the two events
differ only in payload, so no flag-timing change separates them. Both a
smaller redesign and a content-aware variant were built and measured — the
first duplicates on IBus, the second costs a reported row's test and is
blocked while the patch cannot be regenerated.

The Japanese arrays behind this are authored, not observed: no Japanese DOM
composition trace exists in the corpus.

* docs(terminal): a Japanese capture does exist — correcting dbeecb11bee

That commit said no Japanese DOM composition trace exists in the corpus.
False. One does, filed under the Linux bundles rather than the bundle named
for Japanese: 30 DOM events, two にほんご->日本語 conversions, with full
selection state per event. It is retained byte-identically in three further
bundles — one capture copied four times, not four observations, checked by
hash rather than by counting files.

The claim came from checking the bundle named for Japanese, finding nothing,
and generalising to the corpus without querying the rest of it.

Replaying it emits 日本語日本語 under both sequencing extremes on all four
arms, matching its own recorded onData. So "repeated conversion is
undisturbed" is now captured rather than authored. It carries no
post-compositionend insertText, so it cannot speak to the swallow: the
a-after-conversion figure stays authored and unobserved.

Also rewords the paragraph opener. It claimed to broaden a Cmd framing, but
hazard 2 was never Cmd-framed — the lines above already say Cmd does not
reach it. The real gap was that hazard 2 named no trigger at all, which
reads as exotic when it is ordinary.

* build(xterm): land the patch regeneration harness

The five dependency patches under config/patches/ shipped with no tracked
way to regenerate any of them. The xterm one is the hard case: it is derived
from an upstream build, so no fix could be made without rebuilding, and the
tooling to rebuild lived only in one machine's scratch directory. That
blocked a measured fix for a live keystroke-loss bug, and the EditContext
reduction an OSS survey identified as the only real one available.

Adds the regenerator, the upstream pin, the hand-written source patch the
bundle hunks derive from, tests, docs, and a PR job that verifies the
shipped patches still match the pinned build. The job caches the shallow
clone keyed on the manifest, so a cold run is minutes and a warm one under
one. Round-trip verified: regenerating from a clean checkout reproduces the
shipped patch byte-for-byte.

Marks the emitted patch -diff -text. pnpm hashes it byte-for-byte, so a
CRLF checkout would break install on Windows, and its minified bundle lines
make a diff nobody can read — review the source patch instead.

Also rejects unknown flags. --check was the fallback for any unrecognised
argument, so a typo, or --help, silently triggered a full upstream build
instead of what the caller asked for.

* fix(xterm): stop swallowing a keystroke typed after an IME commit

Type a Japanese segment, convert it, then press a key one macrotask later
and that key was lost. No modifier, nothing exotic — every user who keeps
typing straight after converting. Korean surfaced it first only because
2-Set composes on nearly every keystroke.

handleCompositionInput discarded the payload unconditionally in the window
after the deferred send: _isSendingComposition stays true for one macrotask
after the timer cleared _pendingCompositionStart, and the branch substituted
'' for whatever arrived. The suppression is not itself wrong — it
de-duplicates IMEs that deliver their commit an event-loop turn late, which
terminal-stock-composition.test.ts pins. It just could not tell a duplicate
from new input, because the two events are identical apart from payload.

Now it compares against _sentComposition, the text the deferred send
actually emitted, and discards only a match. A flag-timing redesign was
measured first and rejected: it fixed this and duplicated on IBus, because
no timing change can separate events that differ only in content.

Edited in config/patches/xterm-src/ and regenerated through the harness, so
the emitted patch and the lockfile hash are derived, not hand-written.

The commit-overlap pin's swallow arm now asserts the repaired contract —
the value its own comment already named as correct and as what stock
beta.287 emits. #11504's arm at :184 flips too; it never covered that
report, as its own prior note recorded, and the reporter's +149ms arm is
untouched and still asserting the defect. Provenance hashes in three test
docblocks are updated, since regenerating changes the patch hash and with it
the resolved install directory.

* docs(terminal): re-measure the #6905 mutation citations against the new bundle

Regenerating the patch moved the resolved install, so this docblock's
patch_hash, line count, two line numbers and three mutation outcomes all
described a bundle that no longer exists. The deferred branch is one the
fix writes into, so the outcomes could not be re-pointed on reasoning.

Line numbers read off both files by diffing anchors rather than derived by
arithmetic: 201 to 205, 159 to 163. Outcomes re-run through the retained
rig, which re-resolves through the module loader and re-derives each arm
from a unique minified anchor: pristine 3 passed, m1 3 failed, m2 2 failed,
m3 3 passed — identical to the old bundle. Guard controls in both
directions exit 1, so the counts are falsifiable.

Comment-only; the assertions and expectations are unchanged.

* fix(xterm): size the preedit overlay to the cells its text will occupy

updateCompositionElements computed the overlay's left edge from the grid
but never its width, so the preedit rendered at the font's natural advance
while the committed text takes two cells per wide glyph. Measured in
Chromium 150: 가나다라 drew 48.45px as a preedit and 69.20px once committed
— the same characters, same font, 30% narrower, and drifting further with
each syllable. Every macOS mono font carrying Hangul measured 0.49–0.72 of
two cells; never 1.0.

Deriving the width from wcwidth and the cell measure moves Korean, Japanese
and Chinese to 1.000 and leaves ASCII at 1.000, which it already was:

  한        12.125 -> 17.297   (17.30 expected)
  가나다라   48.453 -> 69.188   (69.20)
  안녕하세요 60.563 -> 86.500   (86.50)
  日本語     42.000 -> 51.906   (51.90)
  abcdefgh  69.234 -> 69.203   (69.20, unchanged)

Edited in config/patches/xterm-src/ and regenerated through the harness, so
the emitted patch and lockfile hash are derived rather than hand-written.

The unit test asserts the arithmetic, which is what CI can run. The pixel
consequence was measured on macOS with SF Mono in an Electron harness, not
on the Windows font stack STA-3232 reports from — so this demonstrates the
mechanism and does not stand as that row's platform evidence.

* test(e2e): pin the macOS Korean preedit as visible only while composing

#11914 reports the composing text invisible until Space. Its c3 was recorded
as unobtainable, and the reason on file was wrong: the boundary IS
assertable, but not in happy-dom, which reports display:block in BOTH the
active and inactive states and zeros for every rect. A test there passes
with the defect present.

Captured on real hardware instead: hidden and 0x0 before, .active with
display:block, a 15.84x16 rect and checkVisibility() true while composing
그, hidden again after. 39 DOM events, 2 composition starts, onData
["한","그","\r"].

Two mechanism findings are carried in the setup because both are invisible
in the result and fatal if removed. The input source must be selected AFTER
the app takes focus — focusing resets it to ABC. And the IME must be warmed
until an observed keyCode 229; typed cold it emits raw QWERTY (g k s r m)
with no composition at all, which is indistinguishable from an IME that is
not installed. Two runs were voided on exactly that signature before the
warm-up was found.

The has229 and compositionStarts assertions exist to make such a run fail
loudly rather than pass as a clean negative.

Gated on darwin plus ORCA_E2E_NATIVE_MACOS_KOREAN, like its siblings. The
final spec form has not itself been executed — the machine became
unavailable — so it carries the probe's measured values as literals rather
than a run of its own.

* docs(e2e): correct 19a8d133db7 — the Korean preedit spec has been executed

That commit said the landed form had never run and carried the probe's
values as literals. It has now run on real hardware: 1 passed, 9.1s, rc=0,
with the capture and log sealed under a verified hash manifest.

The teeth check was also run rather than reasoned about, and it changes
which assertion matters. Forcing the active overlay to max-width:0 with
overflow:hidden — invisible on screen — leaves the active class, the
textContent, display:block AND checkVisibility() all passing. Only
during.rect.width fails. checkVisibility() is not sufficient against this
defect; the bounding rect is the single load-bearing assertion, which the
docblock already said and this run confirms.

An earlier teeth attempt injected the CSS mid-run and tripped the
hasActiveClass poll instead, failing at the wrong assertion. It is
inconclusive and excluded from the seal rather than counted.

* test(terminal): add #12171's ordinary-English arm from a real Windows capture

c4 was recorded as unmet and the ledger sourced its control to
evidence/windows-current/, which holds 12 captures and not one English one.
The arm here comes from windows-9803-final instead — same probe, same host
geometry, same injector, en-US with no IME, replayed keydown for keydown.

Two limits are stated in the file rather than left for a reader to find. It
is a different bundle and a different run about 3.6 hours later, so it is
not a same-run arm. And it is #9803's range-active MUTANT arm: ordinary
English stays byte-exact even with that saved-range mutation live, which is
why it reads as a negative rather than as a baseline.

Bundle cited by directory with its file SHA-256; MANIFEST.sha256 verifies
21/21, rc=0. Nothing imported from .tmp.

* docs(terminal): correct #12164's grounds — the cited comments say no such thing

The rejection of #12164 from this file's family was recorded as resting on its
comment 1 (output doubling) and comment 2 (filed against 1.4.163). Checked
against the API: the issue has exactly two comments, neither of which says
that, and the string 1.4.163 appears nowhere in the thread.

The conclusion survives on better grounds. The issue BODY's repro is "Run any
CLI agent (Codex, AGY, Claude, etc.) that outputs Korean text into the Orca
terminal" — untyped output, no keystrokes, no composition — so excluding
CompositionHelper is right, and the input-path hunt was looking in the wrong
place. The body is also LLM-authored (it still contains a literal
"## 5. GitHub Submission Draft (Ready to Post)") and its Root Cause section
blames a CJK IME preedit buffer its own repro never engages, so it should not
be read as observation.

Comment-only; suite unchanged at 5/5.

* test(native-chat): make composition frames carry isComposing, not just inputType

This suite's comment claimed "Gating onChange on `isComposing` breaks here."
It did not. composeFrame() fired `input` with `inputType` but never set
`isComposing`, so a gate on `isComposing` passed all four tests untouched —
the suite asserted a discriminator it did not exercise.

Composition frames now carry both, so neither gate is exempt. Verified by
pointing the mutant at it: with an `isComposing` gate on the composer's
onChange, this suite now fails 3 of 4 (it passed 4 of 4 before), and the
ordinary-English arm correctly survives, since a composition gate should not
touch it. Production code is unchanged and stays gate-free; the mutation was
applied, measured, and reverted.

Found while excluding NativeChatView's question-card remount as the owner of
#12118 / STA-3219: the remount is real, but the preedit survives it precisely
because this write path has no composition gate.

* fix(mobile): ship the patched xterm build, matching desktop

mobile pinned @xterm/xterm 6.1.0-beta.285 while the patch is keyed to
6.1.0-beta.287, so mobile shipped stock xterm and neither IME defect fix
reached it: the swallowed keystroke after an IME commit (9506039de72) and
the preedit sized to the font rather than the grid (e04e0c88da5).

Bumps the three xterm packages to the desktop versions and adds the patch
to mobile's own pnpm.patchedDependencies. No copy of the patch: pnpm
accepts the parent-relative path and records it in the lockfile against
hash 8d63166272e9040a…, byte-identical to what desktop resolves, so the
two stay in step by construction rather than by a drift check.

The workspace separation is untouched — root pnpm-workspace.yaml still
declares `packages: []` and mobile keeps its own lockfile, which is what
keeps the root's patches from failing as ERR_PNPM_UNUSED_PATCH.

Verified in the generated webview bundle rather than at the install:
alignPreeditToGrid 0->2, sentComposition 0->3, pendingInput 0->11, and the
stock-only _handleAnyTextareaChanges 2->0 and dataAlreadySent 4->0. pnpm
applies patches during linking before postinstall regenerates the bundle,
confirmed by a revert/reinstall/re-apply cycle in both directions.

Mobile suite 2971 passed, 3 skipped — identical before and after. Bundle
+1,514 B (+0.24%). Lockfile churn is xterm-only; --frozen-lockfile passes.

mobile/src/ime/ime-submit-carry.ts is NOT made redundant and is untouched:
it handles iOS firing onSubmitEditing on a React Native native TextInput
after unmarking a composition, which is outside the WebView entirely.

Known divergence left alone: desktop also patches @xterm/addon-webgl and
mobile now runs that version unpatched. That patch is glyph/texture-atlas
rendering with nothing IME-related, so it affects neither fix.

* test(terminal): pin the preedit overlay against already-committed cells

STA-3132 (arm A), STA-3170 and STA-3232 report a Korean preedit painted on
top of text already on screen. Builds v1.4.163-v1.4.166 cancel the pending
finalizer in compositionstart, so a committed syllable reaches onData one
syllable late and buffer.x is stale — the overlay lands on the cell the
flushed syllable is about to occupy.

Replays a recorded hardware trace rather than an authored one: the ordered
DOM event stream captured on Windows + MS Korean (wave5-r2 evidence, 64
events), echoing onData back as PTY output.

The load-bearing assertion is deliberately not the obvious one. Comparing
overlay style.left against cursorX is tautological — left is computed from
buffer.x. This counts committed syllables from the compositionend events
the IME fired, so the two sides are independently derived.

Discriminated by a historical re-add across seven real bundles, since the
owner is deletion-shaped: pristine beta287, v1.4.155 and v1.4.162 pass;
v1.4.163 fails; v1.4.163 with that single call removed passes; the byte
identical baseline restored fails again; head passes. Every failing arm
fails only this case — the ordinary negative stays green in all seven.

The negative asserts its own category rather than claiming it: zero
composition events, zero isComposing, zero keyCode 229, exactly 16 events,
paired against the Korean arm's 4 starts / 3 ends / 11 updates / 64 events.

Scope: cell indices, not pixels. happy-dom has no layout, so the recorded
8x16 cell metrics are supplied to the render service. This makes no claim
about pixels visually overlapping; that is affected-OS confirmation and
stays open. Covers the overlap arm only — STA-3132's auto-line-break arm
and STA-3232's half-line-capacity and a11y arms are untouched.

* test(e2e): matrix macOS period substitution against the OS preference

#11504 reports macOS inserting ". " after a Hangul Space commit. This
sweeps six arms across both states of NSAutomaticPeriodSubstitutionEnabled,
reading the preference live per run rather than asserting a literal.

Two results worth having on record.

The reporter's stated trigger did not reproduce. Their words are "There is
no second press at all. One space is enough", but korean-single-space emits
zero insertText with the preference on or off. So does word-space-word-space.

Their timing does reproduce, with different content. korean-double-space and
korean-longer-word-double-space emit a delayed insertText at +122.5-122.7ms
after compositionend — squarely the reported +149ms — but the payload is a
space, never ". ". Consistent with the double-space rule seeing two slots
under ABC and only one under Korean, where the IME commit consumes the first.

The substitution itself is real and preference-bound: latin-double-space
gives "ab . " with the preference on and "ab  " with it off, on one build
with the preference as the sole variable, reproduced across two runs.

That also refutes a claim in PR #11506, which states the substitution "is
enforced outside the renderer and never reproduces in dev builds, so changes
here must be verified against a packaged app". It reproduced in the dev build
twice and did not reproduce on the signed packaged app. That claim should not
be used as a verification gate.

Gated @headful behind ORCA_E2E_NATIVE_MACOS_PERIOD, same shape as the Korean
preedit spec, so it does not run in ordinary CI. Evidence is onData and DOM
only — the PTY-child reader aborted and no packaged-app arm was stable.

* test(terminal): actually enforce the recorded jamo progression

The preedit assertion compared sample.overlayText against sample.overlayText
— the same expression on both sides. A lane proved it by mutation: corrupting
seven of the eight recorded preedit values left the suite fully green. So the
docblock's ㄱ→가→간→나→낟→다→달→라, which the matrix also cites as this row's
recorded shape, was cited and unenforced.

The first attempt at a fix was insufficient and is worth recording. Threading
stroke.preedit through to the expectation still passed on a corrupted fixture,
because that value both drives the rig and was the expectation — corrupting it
moved both sides together. Same tautology, one level down.

The expectation is now an independent literal. Verified by mutation rather
than by reading: corrupting two recorded values fails one arm; restoring them
passes 3/3.

overlayCell was never affected — it is compared against a count derived from
the compositionend events, not from the buffer, and remains the load-bearing
assertion for the overlap.

* fix(e2e): select the selectable input source, not the first match

TISCreateInputSourceList can return several entries for one input source
id. A third-party IME publishes a non-selectable parent alongside the
selectable mode, and taking sources.first can return the parent — after
which TISSelectInputSource fails with paramErr (-50) while the caller
reports success from the enable step.

Found with Qingg (com.aodaren.inputmethod.Qingg), which exposes exactly
that pair under one id. Its mode id equals the bundle id, so filtering by
name would not have helped; selectability is the discriminator.

Now filters on kTISPropertyInputSourceIsSelectCapable and falls back to
the old behaviour when nothing advertises it, so single-entry sources are
unaffected. Also enables every entry for the id rather than only the one
being selected: selecting a mode whose parent is still disabled fails the
same way.

Compile-checked, and selecting com.apple.keylayout.ABC still exits 0.

Unrelated to the enable path: on macOS 26.5.2 third-party IMEs are gated
behind a consent sheet in System Settings. TISEnableInputSource returns
noErr immediately regardless, and the enable only lands if that sheet is
answered while the requesting process is still alive.

* test(terminal): discriminate #12171 against the real shortcut policy

The prior candidate mutation for this row was correctly refused: its suite
mocked shortcut policy so Process/229 became actionable, while the real
resolveTerminalShortcutAction has no Process branch — so the kill measured
the mock. This does not mock it.

Replays a capture of this row's own gesture (d, l, Shift+T, e, k, Space,
Enter under MS Korean, committing 있다) taken on Orca 1.4.164, through the
real useTerminalKeyboardShortcuts hook, capturing bytes at terminal.input.
The earlier capture could not discriminate at all because it recorded no
shiftKey; this one records it on 10 of 10 keydowns with code populated.

One physical Shift+T produces two shifted Process/229 keydowns. Under the
pre-#12265 classifier each synthesizes {key:'Enter', shiftKey:true}, which
the real policy resolves to sendInput '\x1b\r' — twice, giving 1b0d1b0d,
the two escapes the known-bad ed96881b0d1b0d contains.

Mutation is the retained pre-12265-process-shift.patch applied to HEAD, not
an authored one: patch -p1 applies clean and diffs identical to the mutant
copy. Arms are copies; shared source hashes the same before and after.

The English arm stays clean under both modules, so the mutation
discriminates by language rather than by harness — and a real Shift+Enter
through the same rig yields exactly ['\x1b\r'] in every arm, so a silent
pristine result means the code is quiet rather than the harness dead.

Scope: 1b0d1b0d is measured at the renderer boundary. The capture recorded
no PTY bytes — window.api.pty is frozen on shipped builds and the onData
channel needs a build-time flag — so this shows the renderer producing the
two escapes that payload contains, not a re-observation of the payload.

* docs(native-chat): narrow this file's disclaimer to what is now true

It said "THIS OWNS NO REPORTED ROW". Half of that is stale: the remount site
is now the attributed owner of #12118 and STA-3219. On real Windows TSF the
questionActive swap aborts a live composition — the old node gets only a
blur and no compositionend, the text returns as committed, and the next jamo
yields 아ㄴ rather than 안.

The other half holds. This file pins the opposite property, that the text
survives, which is the half those reporters already agree with. Mutation
shows the gap rather than asserting it: deleting the unmount entirely leaves
three of four tests green, because every substantive assertion is
after.value === … and a composer that never unmounts keeps its value.

Also records why the abort cannot be asserted here. The DOM exposes no
observable separating committed text from a live preedit — value is the same
string either way, there is no EditContext, and the only composing-ness
state is a per-instance ref discarded with the node. A test pinning "no
compositionend fires" would be an anti-guard: red the day it is fixed.

The cadence objection is kept, since it is now the open question rather than
the reason for exclusion.

* refactor(terminal): drop the unread isComposing field from XtermBypassEvent

Added by #6396 for terminal IME candidate handling that this branch has since
removed. No production or test code reads it, and the policy is safe without it:
during composition `key` is 'Process', so the non-ASCII printable checks that
would care never match.

Co-authored-by: Orca <help@stably.ai>

* fix(native-chat): keep the composer mounted through an in-flight IME composition

A question card replaced the composer outright
(`{questionActive ? null : <NativeChatComposer/>}`). Unmounting the field
mid-composition aborts the composition in the OS: the node is detached before
`compositionend` can fire, the preedit returns as committed text, and a resumed
Hangul syllable degrades — 아 then ㄴ yields `아ㄴ`, never `안`.

Confirmed in rasterised pixels on Windows with a real MS Korean IME, at both
v1.4.171 and the reporter-era v1.4.164 (the swap block is byte-identical
across them): the preedit underline present before the swap, the composer
visibly absent during it, and the same glyph back afterwards WITHOUT the
underline — committed, not composing.

The swap is now deferred while a composition is in flight, which is what
editors that survive IME do: ProseMirror gates DOM work on `view.composing`,
CodeMirror protects the composing subtree from redraws. Hiding instead of
unmounting does not work — `display:none` and `visibility:hidden` both blur the
focused element and abort the composition the same way.

The hold releases on `compositionend`, which browsers also fire on blur, so
clicking into the card's own answer input yields the input region immediately;
with nothing composing the card still replaces the composer at once, so no
stray "Send a message" appears beside a question.

The existing characterization test flips to a regression guard: it pinned the
node being destroyed, which was the defect. Node identity is the load-bearing
assertion — value-only checks are trivially satisfied by a composer that never
unmounts and cannot tell a held composition from a destroyed one.

The typing-redirect handler moves to its own hook. That is not cosmetic: both
touched files sat at the 400-line cap, and `max-lines` suppressions are
forbidden, so the room had to come from a real extraction.

* fix(macos): opt Orca out of AppKit automatic period substitution

macOS "Add period with double-space" (`NSAutomaticPeriodSubstitutionEnabled`,
on by default) is applied by AppKit's text input system. Native terminals never
join that system; Chromium text fields do, so xterm's helper textarea inherits
it and a double space arrives as `". "` — a period nobody typed, handed straight
to the PTY (#11504).

Chromium answers AppKit for quote and dash substitution and defaults both off,
but declares no period accessor at all, so AppKit applies that one without
asking. This user default is the only lever: there is no per-field or
per-webContents opt-out to prefer over it. Writing the key into Orca's own
defaults domain overrides the global value for this app alone and leaves the
user's system-wide setting untouched. It necessarily covers every Orca text
field, not only terminals — AppKit offers no narrower scope, and that tradeoff
is deliberate rather than accidental.

Measured on the reporter's own build v1.4.161: with the preference ON, typing
a,b,space,space yields `onData ["a","b"," ",". "]`; with it OFF the same arm
yields two spaces.

Note the issue's causal model is wrong and this fix does not follow it. It
claims the substitution only fires with a CJK input source and never with ABC.
The measurement is the inverse — every Korean arm is clean and the ABC arm is
the one that fires — so the fix is not conditioned on input source.

NOT YET VERIFIED ON HARDWARE. The unit tests inject the writer, so they prove
the call is made on darwin and skipped elsewhere; they do not prove AppKit
honours an app-domain override for this key. That check is outstanding.

* fix(xterm): keep a live composition across a lone Cmd press on macOS

CompositionHelper.keydown exempted keyCodes 16/17/18 from tearing a composition
down, which covers Shift/Ctrl/Alt but not macOS Meta — 91/93 in Chromium, 224 in
Firefox. A lone Cmd press mid-composition therefore reached
_finalizeComposition(false), which dropped the preedit overlay's `active` class
and committed the live syllable early. macOS keeps the marked text alive across
that press, so no later compositionstart re-arms the overlay and the rest of the
word composes invisibly.

Measured on hardware (m4air, macOS 26.5.2, Apple M4, 2-Set Korean) with the Cmd
posted as a CGEventType.flagsChanged, which is what a physical modifier emits.
AppleScript `key code 55` posts nothing a browser can see — a bare `key code 56`
for Shift is equally silent — which is why no capture in the corpus ever reached
this branch. Three arms, same build otherwise: overlay live throughout with the
exemption, dark and prematurely committed without it, live again with it
restored. Evidence under
.tmp/ime-handoff/swarm-scratch/wave31-cmd-preedit/evidence/.

The fix cannot widen past a lone modifier: only a standalone press reports these
keyCodes, and a Cmd chord during composition is reported by Chromium as 229,
which was already exempt. Cmd+A still ends the composition, via the IME's own
compositionend. Ghostty draws the same line, returning early from flagsChanged
under hasMarkedText() for every modifier including Super.

Orca's terminal pane was never affected — shouldSuppressTerminalModifierKeyboardEvent
drops a standalone Meta keydown before xterm sees it, and deleting only 'Meta'
from that set is what flipped the hardware arm to broken. The popout preview
terminal and mobile's webview install no such guard and did reach the teardown.

terminal-ime-xterm-composition-commit-overlap.test.ts asked its fixer to update
the two Cmd arms to the values it named as correct; both now emit a single ['한'].

* test(native-chat): drop two byte-identical duplicate cases

`4632b86919d` copy-pasted two cases twice into the same describe block:
`retains carry across a same-frame non-Enter keyup before redispatch` and
`expires carry before a deliberate Enter after the next frame`. Each pair is
byte-identical — same title, same body — so the copies asserted nothing the
originals did not.

This is what has been failing `static analysis` on this branch since 2026-08-06:
`oxlint vitest(no-identical-title)` reports both under `--deny-warnings`, and
`verify` fails solely because it requires static analysis to pass. Every other
gate in `verify` was already green, including typecheck, xterm patch sync, the
full test shard set, and both package jobs.

12 cases still pass in the file.

* test(e2e): skip the WebGL arm when no WebGL renderer exists

The #12164 probe runs two arms, webgl and dom, and closes by asserting the
active renderer is the requested one. That assertion is right for the dom arm —
it is what proves the pane actually left WebGL, without which the arm is
meaningless — but headless CI has no GPU, xterm falls back to DOM silently, and
the webgl arm then fails.

The failure reads as a Korean rendering defect and is not one, so the webgl arm
now skips with the active renderer named. The dom arm keeps the assertion
unchanged.

This is the third of three checks that have been red on this branch since
2026-08-06. `static analysis` and `verify` were fixed in c51c6b5837e; the CI log
shows this job as 1 failed / 1 passed, the pass being the dom arm.

* chore(lint): drop five unused no-console disable directives

`check-changed-code-quality` reports unused eslint-disable directives as errors,
and these five sat above diagnostic `console.log` calls in IME test and spec
files where `no-console` is not enabled — so each suppressed nothing.

This is the second of the two static-analysis steps. `c51c6b5837e` fixed
"Enforce focused code-quality plugins" (duplicate test titles); this fixes
"Enforce changed-code quality". Both had been red on this branch since
2026-08-06, and I mistook the first for the whole job.

The diagnostic logs themselves are kept — they are what a failing IME arm prints
for a reader to inspect.

* test(e2e): cover #12164 under fractional device scale factor

Fractional display scaling was #12164's last unexplored branch, and the reason
is worth recording: earlier attempts were BLOCKED, correctly, because they
proposed mutating the Windows display scale on a remote physical machine with no
console recovery. `--force-device-scale-factor` reaches the same renderer state
per process, so nothing outside the Electron instance changes and there is
nothing to restore.

The hypothesis was specific: `프프로로젝젝트트` is what a half-pixel cell boundary
could produce on a 2-column glyph, and nothing else in the suite varies dpr.

Measured at 1.25 and 1.5, both under WebGL: ink extents 25/21/16 with identical
ink groups, matching the scale-1 run. No doubling.

The arm self-certifies before asserting — if the flag does not take, the test
fails rather than silently measuring at dpr 1. That matters here: the sibling
spec's WebGL arm went two days reporting a missing GPU as a Korean rendering
defect precisely because a silent fallback looked like a result.

* fix(terminal): match Mod+letter shortcuts by physical key, not IME-rewritten key

With a CJK input source active, macOS and Windows report the physical key through
`code` but rewrite `key` to the layout's character: Korean 2-Set turns Cmd+C into
`{ key: "ㅊ", code: "KeyC", metaKey: true }`. Every `key.toLowerCase() === 'c'`
match misses it, so the shortcut is not recognised and xterm encodes the chord as
PTY input instead — issue #13033 reports `ESC[12618;9u` and a terminal that jumps
to the bottom, because user input scrolls the viewport.

This is the same key-vs-code confusion that owned #12171, where a `Shift+T`
typing ㅆ was read as Enter for want of a `code` guard, so the fix is the same
shape: trust `code` when it is present, fall back to `key` and then the legacy
`keyCode` when it is not (Chromium omits `code` on synthetic and some keypress
events, and `keyCode` keeps its US value even when `key` is rewritten).

Applied to the four terminal-side sites, including the dashboard pop-out, which
#13033 called out specifically as having its own key handler:
  pty-connection.ts            Cmd/Ctrl+C copy guard
  keyboard-handlers.ts         Cmd+G search navigation
  agent-interrupt-inference.ts interrupt inference
  preview-terminal-key-handler.ts  pop-out paste

Nine further `key.toLowerCase()` letter matches exist outside the terminal
(TaskPage, editor, GitHub composer, browser markup). They have the same defect
and are deliberately left for a separate change rather than widening this one.

An existing case, `matchSearchNavigate > returns null for wrong key`, overrode
only `key` and left `code: 'KeyG'`, so it began passing for the wrong reason. It
now overrides both — which is what "wrong key" means once matching is physical —
and a companion case pins the Korean-rewritten chord still matching.

#13033 was closed NOT_PLANNED; the reporter's event shapes drive the new test.

* fix(renderer): match every Mod+letter shortcut by physical key, not IME-rewritten key

Completes the previous commit. A CJK input source rewrites `event.key` while
`event.code` keeps the physical key, so `key.toLowerCase() === 'z'` and friends
silently stop matching — the shortcut is not recognised and the keystroke falls
through to whatever handles unclaimed input.

The helper moves to `@/lib/ime-latin-shortcut-key` first: it now serves the
editor, GitHub composer and browser markup, and importing terminal-pane
internals into those would be the wrong direction. `lib/` already hosts
`ime-composition-keyboard-event` for the same reason.

Nine remaining sites, all previously unreachable under Korean/Japanese/Chinese/
Vietnamese input:
  TaskPage, ActivityPrototypePage, ProjectViewWrapper   Cmd+F search
  useMarkupKeyboardShortcuts                            Cmd+Z undo
  GitHubMarkdownComposer, RichMarkdownLinkBubble,
    rich-markdown-link-shortcut                         Cmd+K link
  native-chat-shortcut                                  Cmd+J
  rich-markdown-key-handler                             Cmd+Shift+X

Six of the nine test `!== 'letter'` as early-return guards and three test
`=== 'letter'`; the negation is applied per site, since a blind substitution
would have inverted six of them.

Full suite: 4249 files pass. Three files fail locally and none is caused by this
change — the branch touches no file under `src/main/` or `src/relay/`, all four
failures reproduce on an unmodified tree or pass in isolation (the worktree
poller passes 21/21 alone, so it is full-suite parallelism), and all 16 CI test
shards are green.

* docs(ime): scope the IME composition rules to the terminal-pane directory

#11893 proposed adding these to the root `AGENTS.md`, which every agent loads on
every task regardless of what it is doing. They only bind keyboard handling, the
composer and the terminal input path, so they belong next to that code —
`tests/e2e/AGENTS.md` already establishes the nested convention here.

Kept from #11893: range-derived commits, guarding above the key dispatch,
the `attachCustomKeyEventHandler` / `CompositionHelper` interaction, no
normalization at commit, and the recorded-trace evidence bar.

Added from defects found since it was written:
- match shortcuts on `event.code`, not `event.key` (#12171, #13033)
- `keyCode === 229` means an IME owns the press
- do not unmount a field mid-composition, and hiding is not a fix because
  `display:none` blurs and aborts it too (#12118, STA-3219, #11332)

The evidence bar now also names the mutation check, since a test that survives
deleting the code it guards is guarding nothing — a failure this effort hit more
than once.

* fix(terminal): gate Ctrl+Enter CSI-u on a negotiated pane, porting #12462

Found while scoping the rebase onto `main`: #12462 landed on 2026-08-06 and
fixes a real defect this branch does not carry. Ctrl+Enter emitted
`\x1b[13;5u` unconditionally, so a pane that never negotiated the kitty
keyboard protocol — local Windows ConPTY, plain shell — printed the escape
verbatim into the prompt.

This branch deletes `terminal-ime-deferred-newline.ts`, which is one of the
files #12462 touched, so a rebase resolving those conflicts by taking our side
wholesale would silently reintroduce the defect. Porting it forward now means
the fix survives the rebase however the conflicts are resolved.

Mirrors the Shift+Enter guard already here: local ConPTY falls back to the
legacy CR every emulator sends, and a negotiated pane keeps the chord, so the
fallback is scoped to panes that cannot receive CSI-u rather than to Windows.

NARROWER THAN #12462 BY ONE CONDITION, deliberately. `main` also allows CSI-u
via `hasCtrlEnterCsiUAuthority()` (trusted consumer evidence, #12329); that
helper and its plumbing do not exist on this branch. Omitting it is the
conservative direction — an authorised pane gets `\r` instead of the chord,
rather than an unnegotiated pane printing an escape — but it should be restored
when the two histories are reconciled.

Test covers both directions and is mutation-checked: forcing the gate open
fails it, so it cannot pass by construction.

* fix: reconcile two more fixtures main moved while the stack waited

Both caught by CI, not locally, and the reason the local run missed one is
worth recording:

1. `browser-toolbar-profile-dialogs.ime-enter.test.tsx` did not pass
   `useNativeUserAgent` / `onUseNativeUserAgentChange`, which `main` added to
   `BrowserToolbarProfileDialogsProps`.

   Local `pnpm typecheck` reported 0 errors on the same commit CI failed. The
   cause was a stale `config/*.tsbuildinfo` — tsc reused an incremental cache
   from before the merge. Deleting it reproduced CI's error exactly. Any
   "typecheck clean" during this merge should be treated as unverified unless
   the cache was cleared first.

2. Localization keys for `SshDisconnectedDialog` were absent from `en.json`:
   the merge took this branch's component alongside `main`'s catalog.
   Regenerated with `pnpm run sync:localization-catalog` rather than hand-added.

* fix(mobile): regenerate the lockfile the merge resolved by taking one side

CI's `verify` failed with `ERR_PNPM_OUTDATED_LOCKFILE` on `mermaid (lockfile:
11.16.0, manifest: 11.16.1)`. The mismatch was in `mobile/`, not the root — the
root lockfile was consistent throughout, which is why inspecting it (and even
GitHub's merge ref) found nothing wrong.

Cause: during the merge I resolved `mobile/pnpm-lock.yaml` by taking this
branch's side wholesale rather than merging it, so it kept `mermaid 11.16.0`
while `mobile/package.json` came from `main` at `11.16.1`. Taking one side of a
lockfile is only safe when the corresponding manifest comes from the same side.

Regenerated with `pnpm install --lockfile-only`; `--frozen-lockfile` now passes
in `mobile/`. Verified the xterm patch entry survives intact — same hash
`4f1b42d268f3964d…` and the parent-relative path into `config/patches/`, which
is the desktop/mobile coupling that would silently break the mobile build.

Two earlier diagnoses of this failure were wrong and are worth recording: it was
not the root lockfile, and it was not a stale merge ref (a rerun reproduced it
exactly).

---------

Co-authored-by: Orca <help@stably.ai>
2026-08-07 22:58:40 -07:00
github-actions[bot] 7df7aa5fe9 Update README downloads badge 2026-08-07 01:27:54 +00:00
github-actions[bot] 2539889197 Update README downloads badge 2026-08-05 18:28:05 +00:00
Jinwoo Hong 06780260c0
test(remote-runtime): run an old client and an old server against current code (#12682)
Mixed versions are the normal state of the remote-server feature: users update clients and servers independently. Until now nothing tested that. Every cross-version claim was made by code reading plus unit tests with hand-written old/new shapes — enough to catch design problems, not enough to catch a real skew regression.

This runs the REAL protocol implementations from two builds against each other in one process: the actual host methods and RPC dispatcher on one side, the actual renderer multiplexer on the other, with a transport that reproduces the production asymmetry — each side decodes with its OWN codec and drops frames whose opcode it does not know. A frame survives only if the RECEIVING build understands it, which is what makes this level sufficient without launching two apps. The old side is a genuine checkout extracted from the release tag; the extracted client was confirmed to lack a symbol that exists only on main.

Journey: subscribe, first snapshot, input reaching the process, live output, hide/reveal snapshot, transport drop, resubscribe, input landing again — across old->new, new->old, and a current/current control. Every step ends on an observed-state barrier; no sleeps. The oracle asserts the recorded step list, the exact 16-frame named sequence, negotiated capabilities, the exact input the host wrote to the PTY, rendered content, and zero decoder-rejected frames. A host method the stub lacks is recorded by name and asserted empty, so a harness gap cannot masquerade as a wire break.

Detection is proven per violation shape, and it attributes each to the correct side: an unnegotiated opcode goes red only where a decoder would reject it, a removed published field goes red only where an old client consumes it, and a legal additive field stays green in all three pairings so the harness will not cry wolf on safe changes.

It also documents the three compatibility rules in docs/reference/remote-wire-compatibility.md, linked from AGENTS.md, since they previously existed only as folklore — notably that "decoders reject unknown opcodes" is true for the desktop decoder but NOT for mobile, which silently drops them.

Deliberately scoped: terminal stream only. The session-tab sync channel is not covered, nor agent-session publications, file/Git RPCs, mobile E2EE framing, or the relay transport. Two version points, so a regression introduced and reverted between them is invisible.

CI selection was verified rather than assumed — `vitest list` confirms 0 matches under the shard's exclude and 4 under the dedicated job — because a lane silently running zero tests is precisely how a host-side defect escaped CI earlier in this series. Closes STA-3469.
2026-08-05 01:31:29 -07:00
Jinwoo Hong 738f640428
fix(mobile): relay UX overhaul — steady status colors, visible relay dials, coordinated deep links (F1-F10)
* fix(mobile): keep healthy relays green through focus and network nudges (F1+F2)

Focus/app-resume nudges probe the active relay instead of suspending it;
network-change nudges replace it make-before-break, suspending only after a
failed dial. Mount, Retry, and host-swap windows read 'connecting' instead of
'disconnected'; the host list keeps last-known worktrees for every
not-connected state and spins instead of rendering nothing.

* feat(mobile): surface the pairing relay path in the pairing log (F3)

The relay candidate was silent during pairing: dialing, E2EE handshake,
director recovery, and the winning path now emit redacted phase lines through
the same connectOptions.onLog the direct path already used.

* docs(mobile): relay UX investigation findings and F0-F10 fix plan

* feat(mobile): name and narrate relay dials while they happen (F5)

migrateTo forwards the dialing session's connecting/handshaking/reconnecting
phases whenever the client is suspended or disconnected — never downgrading a
live session — and exposes getPendingPath so the host card can say
'· Orca Relay' during the dial instead of only after it.

* feat(mobile): race a relay dial when the direct dial stalls (F6)

A 2.5s grace timer starts relay recovery while an unauthenticated direct dial
is still inside its 12s connect window; the race gets one attempt through the
existing mutex/cooldown machinery, cancels when direct authenticates, and
never arms for hosts without a relay endpoint.

* fix(mobile): overlay the protocol gate instead of unmounting the host stack (F9)

A pending status.get used to swap the mounted HostStack for a spinner at the
moment the socket connected, destroying in-flight nested navigation. Once
children have rendered for a host they stay mounted under an opaque
touch-blocking overlay; first visits and blocked verdicts keep the old
behavior.

* fix(mobile): keep loaded data through transient connection blips (F10)

Git history no longer blanks on reconnect (and commit files refetch instead
of caching an offline empty answer), the repo picker keeps its last-good list
when an in-flight repo.list rejects, the diff review's ready-state
preservation actually runs, and proven host capabilities survive a drop
flagged unverified instead of being wiped.

* feat(mobile): coordinate every home deep push and bounce dead resume targets (F4+F7+F8)

Notification taps, the Accounts card, and host-edit now use the shared
mount-then-replace transition (with a focused-route walker so root-layout
scope works); the Resume card renders from the snapshot in a disabled state
so its late arrival can't shift Tasks under the thumb; resume targets are
validated against proven catalog data, and a session route whose worktree
the host proves missing bounces to the host index with a notice banner
instead of stranding on a dead screen.

* test(mobile): cover the resume-target and notice policies (F7)

Key notice dismissal by code so closing one banner cannot swallow a later,
different one, and move the visibility rule into host-route-notice.ts where it
is testable without a screen.

Adds the missing units for F7's decision points: isResumeTargetConfirmedMissing
(unproven catalog is silence, synthetic routes exempt), the validating
last-visited reader, and the notice visibility rule.

* fix(mobile): review-pass hardening for the gate overlay and diff preservation

Adversarial review findings: the reader's hunk position now survives a
connection blip (reset only on item change), the covered stack is hidden from
TalkBack while the gate overlay is up, and the overlay's hit-test comment is
scoped honestly to in-tree views (native-Modal drawers present above it —
follow-up).

* fix(mobile): CI + CodeRabbit review fixes for #12609

Move the findings doc under docs/ (root directory guard), drop two unused
eslint-disable directives, and address review findings: an unproven snapshot
seed can no longer downgrade a proven worktree catalog; a locally-aborted
relay dial skips the director fallback; post-migration bookkeeping failures
log instead of masquerading as dial failures (which could suspend the healthy
session); the auth wait arms its timeout before subscribing; forwarded dial
phases stop at close(); the legacy selector_not_found fallback requires
runtime_error; the diff-loading effect depends on the fields it reads; and
host-edit auto-cancellation is now pinned by a test.

* fix(mobile): second review round — queued replacements, race fence, confirmed bounces

A network-change replacement now survives the recovery mutex and cooldowns as
a queued intent instead of being dropped or suspending a healthy session —
only a failed dial or a dead probe tears one down. The happy-eyeballs
migration withdraws when direct authenticated during the relay dial
(first-authenticated-wins). A worktree bounce requires two consecutive
host-proven misses, since a transient desktop repo-scan rejection answers
selector_not_found for a live worktree. Background network flaps no longer
wake a billed relay splice, the lifecycle foreground flag stays in sync, a
screen unmount cancels only its own pending host-stack transition, and diff
review keeps the loaded review when its reconnect refresh rejects.

Extracted mobile-endpoint-nudge-router.ts and the establisher's dialEligible
pass, and split the supervisor nudge tests, to stay under max-lines.

* fix(mobile): satisfy the React Doctor changed-code gate

Render-phase ref writes move into effects: the protocol gate's resolved/mounted
latches now record committed outcomes only (a discarded children render can no
longer count as mounted), and the bounce hook syncs its callback ref in an
effect. Array<T> annotations become T[] in the extracted modules.

* fix(mobile): keep the loaded diff when the reconnect refetch rejects (F10)

The diff-loading hook's catch was the one path still erasing a ready diff —
the same keepLoadedDiff guard its disconnect and loading branches already use,
now pinned by a reject-after-ready test.

* fix(mobile): process foreground revival nudges

---------

Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com>
2026-08-04 19:47:39 -07:00
github-actions[bot] 9058059154 Update README downloads badge 2026-08-04 09:51:50 +00:00
Neil d8e5944b60
Stop a duplicate headless orca serve from crash-looping and exhausting AppImage FUSE mounts (#12212)
* fix(startup): stop a duplicate headless serve from crash-looping and leaking AppImage mounts

A second Orca launch that loses the single-instance lock called app.quit()
before `ready`. That quit is deferred, so the doomed process kept booting into
Chromium's Linux display initialization, failed with "Missing X server or
$DISPLAY", and died with SIGSEGV. systemd read that as a crash and restarted it
forever; each restart re-mounted the AppImage and left the squashfuse mount
behind, until the host hit the 1000-mount FUSE ceiling and every later launch
failed.

The lock-losing launch now calls app.exit(3), which terminates synchronously
before any display init. Exit code 3 is a stable "another process already owns
this userData profile" contract, and the documented systemd unit uses
RestartPreventExitStatus=3 plus a real StartLimitIntervalSec/StartLimitBurst
window so a permanently failing launch can no longer retry unbounded.

Second-instance argv is now forwarded to the owner, and a duplicate `orca serve`
no longer asks the live headless server to open a desktop window. Desktop
activation for ordinary launches and macOS dock re-activation is unchanged.

Closes #11935

* docs(headless): clear the start limit before the scripted service starts

StartLimitIntervalSec=300/StartLimitBurst=5 rate-limits operator starts too, so
after a crash-loop trips the burst systemd refuses a plain `systemctl start` for
the rest of the window. The Upgrade and Roll back scripts run under
`set -euo pipefail`, so that refusal aborted the rollback mid-flight and left the
server down on the exact recovery path the doc prescribes.

Both scripts (and their EXIT-trap recoveries) now run `systemctl reset-failed`
first, the unit reference explains the interaction, and the crash-loop bullet
points at it for manual starts.

Co-authored-by: Orca <help@stably.ai>

* test(startup): reproduce the #11935 duplicate-serve crash loop under real Electron

The committed coverage for #11935 was source-text greps, so nothing gated the
mechanism the fix rests on: pre-`ready` `app.quit()` is deferred, which is why
the lock-losing headless `orca serve` kept booting into Linux display init.

This runs two real Electron processes against one disposable profile. The
duplicate executes the lock-loss gate's own `app.*` statement, lifted out of
`src/main/index.ts`, so reverting to `app.quit()` fails the test. It also feeds
the owner's real forwarded argv through `shouldActivateDesktopForSecondInstance`.

Also record why the activation predicate matches `--serve` and not the `serve`
subcommand: an AppImage launched as `orca serve` exits at the CLI redirect
before requesting the lock.

* test(startup): wait for the owner process to exit before removing its profile

Windows holds the profile's handles for a beat after SIGKILL, so an immediate
rmSync can fail with EBUSY/EPERM.

Co-authored-by: Orca <help@stably.ai>

* test(startup): pass the fixture marker path by env, not argv

Chromium reorders argv and the duplicate's argv is itself under test, so a
trailing positional was the wrong channel for it.

Co-authored-by: Orca <help@stably.ai>

* test(startup): only the activation case waits on the owner notification

The exit-contract cases assert on the duplicate's own already-terminated
process, so they should not block on cross-process delivery.

Co-authored-by: Orca <help@stably.ai>

* test(startup): drop the staged lock race, keep the real-Electron gate contract

CI proved the two-process form cannot work on a display-less Linux runner:
Chromium's ProcessSingleton needs the browser IO thread, which needs `ready`,
which needs a display. The pre-`ready` owner looked stale and the duplicate took
the lock (`expected [ 'DUPLICATE_WON_LOCK' ] to include 'DUPLICATE_LOST_LOCK'`).

Lock acquisition and argv forwarding are already covered in
single-instance-lock.test.ts. What only a real process can settle is what the
loser does next, so that is all this file now runs -- display-independent.

Co-authored-by: Orca <help@stably.ai>

---------

Co-authored-by: Orca <help@stably.ai>
2026-08-03 23:34:22 -07:00
OrcaWin ed7849eb7b
fix(worktrees): stop silently switching existing Windows setup scripts to Git Bash (#12406)
* fix(worktrees): stop silently switching existing Windows setup scripts to Git Bash

#6967 derived the Windows setup-runner shell from `terminalWindowsShell`. On
upgrade, any Windows user whose terminal preference resolved to Git Bash had
their existing `orca.yaml` setup script (and issue command) handed to bash
instead of cmd.exe. Scripts authored against the cmd runner — `copy`, `xcopy`,
`set VAR=value`, `if errorlevel 1`, `%VAR%`, backslash paths — broke with no
migration and no warning, and the failure looked like Orca broke the project.

The conflation is also wrong in the steady state: a terminal preference is
per-user, so two people on the same repo got different interpreters for the
same orca.yaml and no project could write a setup script that worked for all
of its Windows contributors.

The interpreter is now a property of the script, declared the standard way:
a leading `#!` line. Native Windows keeps the historical `.cmd` runner unless
the script declares a POSIX shell, so no existing script changes behavior.
`resolveSetupRunnerShell` keeps its role as the feasibility gate — a bash
runner still requires the terminal to resolve to Git Bash, since the launch
command is typed into that shell and uses MSYS `/c/...` paths.

`buildWindowsRunnerScript` now drops a leading `#!` line rather than `call`ing
it, so a declared-bash script that falls back to cmd (Git Bash missing) fails
on a real setup line instead of aborting on errorlevel at line one.

WSL worktrees, POSIX platforms, and SSH hosts are untouched.

* fix(worktrees): keep the cmd setup runner launchable from a Git Bash pane

Adversarial review of this PR found that pinning the runner format per script
reopened issue #6896 one layer down.

- `WorktreeSetupLaunch.shell` had been redefined to mean "the format the runner
  file was written in". `resolveSetupRunnerCommand` consumes it as "the shell
  that types the launch command", so a Git Bash terminal with a batch setup
  script produced `cmd.exe /c "C:\...\setup-runner.cmd"` typed into a bash pane,
  where MSYS rewrites the `/c` switch into a drive path: cmd opens interactively
  and setup never runs. `shell` is the terminal's family again; the runner file's
  .cmd/.sh extension carries the format, and a batch runner launched from a POSIX
  pane reuses the existing PowerShell ProcessStartInfo launcher.
- The cmd runner dropped a leading `#!` line and ran the rest as batch, so a bash
  script reaching cmd (PowerShell/cmd terminal, or any SSH-to-Windows host) got
  its interpreter-agnostic prefix executed before failing mid-way. It now prints
  why and exits 1 without running anything.
- A `#!` line's option flags were discarded: `#!/usr/bin/env -S bash -euo
  pipefail` lost pipefail because the runner is launched as `bash <path>`. The
  generated posix runner now replays declared flags via `set` and drops the
  duplicate interpreter line.
- Docs cover the per-user setup command in repository hook settings, which goes
  through the same `#!` rule, and describe what the `#!` line does and does not
  select.

Tests: composed launch command for a POSIX pane + cmd runner (hooks, shared
runner command, setup sequencing gate, observed-setup signal), the cmd runner's
shebang refusal, and shebang flag replay. Each fails with the source reverted.

* fix(worktrees): replay only real `set` flags and keep the gate in the pane's shell

Two round-2 review findings:

- `#!/bin/bash -l` replayed `set -l`, which exits 2 and aborted the runner under
  its own `set -e` before a single setup line ran (all platforms). Only the flags
  `set` documents are replayed now; a bare `-o` with no option name is dropped
  instead of dumping the shell-option table.
- The wait-for-setup gate picked its language from the runner file, so a batch
  runner launched from a Git Bash pane got the PowerShell gate while the agent
  startup command was already POSIX-quoted — `Invoke-Expression` cannot parse
  `'\''`. The gate now follows the pane; the runner still launches through the
  ProcessStartInfo launcher, never through bash.

---------

Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com>
2026-08-03 23:25:50 -07:00
Jinjing 6f28c118bb docs: drop WeChat group 5/6 QRs; keep group 7 only
Older groups are full; community join path is group 7 only.
2026-08-02 19:48:51 -07:00
Jinjing 616d22406e docs: add WeChat group 7 QR with overflow guidance
Groups 5/6 filled; point community join flow at group 6 + group 7.
2026-08-02 19:47:46 -07:00
github-actions[bot] e0f597a351 Update README downloads badge 2026-08-02 00:58:08 +00:00
Jinjing 96c954f3be
chore: remove force-added design docs from docs/ (#11891)
Keep only the durable docs already allowlisted for tracking
(STYLEGUIDE, assets, localized readme, and reference compatibility
guides). Drop feature design notes, plans, and repro artifacts that
were force-added past the existing docs ignore rules.
2026-08-01 00:33:20 -07:00
OrcaWin 3b7ea59c5b
fix(windows): make the GPU fallback actually remove the GPU child, and stop WSL latching absent (#11295)
* fix(windows): make the GPU fallback actually remove the GPU child, and stop WSL latching absent

Three Windows crash/regression fixes from shipped 1.4.156/1.4.158/1.4.159 crash reports.

GPU fallback (cluster D, 14 reports, exit 0x80000003 STATUS_BREAKPOINT):
the software-rendering fallback called disableHardwareAcceleration() plus
--disable-gpu, neither of which removes the GPU child process — Chromium still
spawns it to host Viz and merely drops the backend to software GL. Measured on
Windows 11 / Electron 43.1.0: gpuProcessCount stays 1. So a GPU process being
killed by a bad driver or an injected DLL kept dying after the fallback engaged,
on every launch, for the life of that build (the marker is sticky per version).
The crash tails show exactly this: gpu_fallback_applied followed by another GPU
crash 1.3s later. --in-process-gpu is the only switch that drops the child count
to 0; --disable-software-rasterizer is deliberately excluded because it also
kills SwiftShader, which would drop every terminal to the DOM renderer.

WSL distro list: a successful-but-empty `wsl --list --quiet` was cached for the
process lifetime. `wsl --install` reports zero distros while one is still
provisioning, so an early probe latched "no WSL" until restart — WSL appeared
during setup and then vanished from the terminal picker. Empty results now
re-probe on an exponential window (15s doubling to a 5min cap) while staying
readable, so a missing distro is still visible to isKnownMissingDistro.

WSL availability: isWslAvailable() latched false on any failure via a bare catch,
so one slow wsl.exe activation disabled WSL for the whole session. Failures are
now classified — a numeric exit status or ENOENT is answer-shaped and holds for
10min, anything else (timeout, spawn failure) retries after 45s — and both back
off per consecutive failure, mirroring isPwshAvailable.

Windows-only: every changed path is behind an existing process.platform check,
so macOS and Linux behaviour is unchanged.

* fix(windows): drop a stale WSL availability failure once a distro list succeeds

The distro-list and availability caches expire independently, and
getWslRepairReason checks availability first. So a definitive availability
failure (numeric exit status or ENOENT) held for 10-30min would keep reporting
`wsl-unavailable` even after `wsl --list --quiet` successfully returned a
distro — i.e. over a WSL that demonstrably just answered. That is the same
latch class this branch fixes, surviving in the gap between the two caches.

A non-empty distro list proves wsl.exe ran, so drop the negative availability
cache and let the next call re-probe. Scoped to non-empty lists only: those are
cached for the process lifetime, so this cannot re-spawn the blocking 5s probe
more than once. An empty list keeps its failure cache, since it re-probes on a
15s-to-5min schedule and would otherwise pay the blocking probe far too often.

* fix(windows): harden GPU safe mode and WSL recovery

* fix(wsl): make capability refresh cleanup explicit

---------

Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com>
2026-07-30 17:49:32 -07:00
github-actions[bot] 9f5aa41a7a Update README downloads badge 2026-07-30 18:48:03 +00:00
Sebastián Castaño 676ef7fab8
feat(cli): add orca skills install and orca skills update for headless skill setup (#9201)
Adds `orca skills install` and `orca skills update` so skills can be set up without the GUI — SSH hosts, containers, CI. Previously `orca skills` had only `list` and `get`, so there was no headless path.

**Agent targeting is scoped explicitly rather than delegated to detection.** The `skills` CLI decides which agents to install into, and with `-y` and zero detected agents it takes `targetAgents = validAgents` — all ~75. That is not a corner case for a headless CLI: a fresh SSH box or container with no agent installed is the normal starting state. Measured on a bare host, the unscoped command created **52 top-level agent directories and 54 junctions** (one real payload in `~/.agents/skills`, the rest links) on Windows, and 52/53 on macOS.

The CLI now passes `--agent` derived from Orca's own detection, mapped to the `skills` key namespace, plus `universal`. Supplying `--agent` makes `runAdd` use it directly and never call `detectInstalledAgents()`, so the fan-out branch is unreachable. On a bare host it now refuses with `No coding agent detected on this host` and exit 1, creating nothing. Same command with scoping: **1 directory, 0 junctions.**

`universal` alone would under-install — Claude Code is not in that set, and 19 of 28 mapped keys write agent-private homes `universal` never touches. `--agent '*'` is the bug itself. The mapping is hedged three ways: `null` for any agent whose key could not be confirmed, `satisfies Record<TuiAgent, …>` so a new Orca agent is a compile error, and a test pinning every mapped key against the CLI's own valid list.

Fixed during review — two holes that each restored the full fan-out through a different door:
- `--agent ','` trimmed to nothing, which skipped the refusal *and* emitted no `--agent`.
- `--agent -y` passed an emptiness check, and the vendor CLI silently drops `-`-leading values, re-emptying its list.

The real invariant is argument *shape*, not emptiness, and it is now enforced at the choke point in `buildAgentFeatureSkillInstallArgs`, so no caller can emit `-y` without a usable target. `*` remains allowed — asking for every agent explicitly is a choice, not an accident. Verified with 51 hostile inputs through the built binary, each recorded argv replayed through the vendor's own parser.

Also fixed: the `ORCA_CLI_CWD` refusal now runs before target resolution (it was quoting the wrong host's agent list), and `--dry-run` is refused in a forwarded shell rather than printing a command naming the wrong machine.

Validated on a real Windows host across PowerShell 7, PowerShell 5.1, cmd.exe and Git Bash: `.cmd` shims route through `cmd.exe` and `.exe` shims spawn directly (proved with instrumented shims, not inferred), the ENOENT path produces an actionable error rather than a silent failure, and `skills update` genuinely restores a corrupted skill byte-for-byte.

Known, not addressed here — both upstream behaviours this only forwards: a partial install failure exits 0, and "no installed skills found" exits 0. Both are invisible to the headless callers this feature exists for.

Co-authored-by: scastanoh21 <scastanoh21@gmail.com>
2026-07-30 11:20:29 -07:00
Jinjing 5f7807497e
feat(ssh): bound relay PTY output end to end (#11005)
* docs: design SSH relay PTY backpressure

* fix(ssh): bound relay frame decoding

* fix(relay): bound PTY output publication

* fix(ssh): bound PTY model admission

* fix(ssh): settle closed model admissions

* feat(ssh): negotiate bounded PTY consumer sessions

* fix(ssh): fence exit on renderer settlement

* feat(ssh): track PTY source credit end to end

* fix(ssh): recover bounded PTY output across reconnect

* feat(ssh): complete relay PTY output backpressure

* fix(ssh): close final PTY source credit races

* docs(ssh): record final backpressure validation

* feat(ssh): complete relay PTY source-credit lifecycle

* test(ssh): complete provider notification fixture

* fix(ssh): preserve terminal source credit across rotation

* fix(ssh): fail closed on recovery cancellation

* fix(ssh): prioritize mux control writes after drain

* fix(ssh): retire canceled relay restore deliveries

* fix(ssh): order exit cancellation cleanup

* fix(ssh): gate provisional source activation

* test(ssh): register mux drain-priority coverage

* fix(ssh): type stale owner recovery mismatches

* fix(ssh): close projection replacement races

* fix(relay): contain streaming edge failures

* fix(ssh): secure relay endpoint credentials

* docs(ssh): reconcile final backpressure lifecycle

* fix(ssh): bound main IPC output lifecycle

* fix(ssh): close recovery ownership gaps

* docs(ssh): record exact artifact validation

* fix(ssh): reject reclaimed snapshot replacements

* fix(ssh): fence model admission across reconnect

* fix(ssh): contain migration failure per PTY

* docs(ssh): record final exact-head validation

* test(ssh): align deploy fixtures with credential publication

* feat(ssh): add per-target bounded output setting

* fix(ssh): close source recovery review gaps

* fix(ssh): latch source credit environment override

* feat(ssh): make PTY source credit the default

* docs(ssh): record always-on relay validation

* docs(ssh): bind validation to current main

* test(ssh): grant source credit in IPC fixture

* test(ssh): grant source credit in fake relay

---------

Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com>
2026-07-29 17:03:15 -07:00
OrcaWin dde72f85de
fix(windows): separate updater from orchestration migration (#11405)
* fix(windows): separate updater from orchestration migration

* fix(terminal): attest adopted reveal identity

---------

Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com>
2026-07-29 15:23:06 -07:00
OrcaWin 363e478909
fix(orchestration): preserve active workers across updates (#11271)
* fix(orchestration): preserve active workers across updates

* test(ssh): model absent legacy adoption

* test(orchestration): align compatibility contracts

* fix(windows): escape updater PowerShell booleans

* fix(windows): restore stock uninstall process check

* fix(orchestration): keep recovery off renderer startup barrier

* fix(orchestration): harden legacy recovery migration

* fix(orchestration): close recovery review gaps

* fix(orchestration): complete legacy worker cutover recovery

* fix(orchestration): preserve legacy workers across updates

---------

Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com>
2026-07-29 11:31:35 -07:00
Neil 3a80fbe162
Revert terminal rendering changes from #10692, #10794, #10871, and #10907 (#11338)
* Revert "fix(terminal): avoid flash while restoring parked terminals (#10871)"

This reverts commit 5a6a9e0b28.

Reverted for terminal rendering regressions (flashing, lost content).
Conflict resolution preserves the forwardRef signature from #10433 and
drops the parked-presentation gating #11016 fed with its effective set.

Co-authored-by: Orca <help@stably.ai>

* Revert "fix(terminal): limit pre-paint WebGL resume to macOS (#10794)" and "fix(terminal): stop switch bold flash and Windows lag (#10692)"

This reverts commits 4681edb520 and
8f5a45401f.

#10794 was itself a partial revert of #10692, so both are reverted
together: the Windows retained-WebGL LRU and the macOS pre-paint
(layout-phase) visibility transition that survived it. Terminal
visibility resume returns to passive disposal and recreation on every
platform, and the WebGL context ceiling returns to a flat 128.

Co-authored-by: Orca <help@stably.ai>

* Revert "fix(terminal): release an abandoned synchronized-output frame on reveal (STA-2694) (#10907)"

This reverts commit 97cb32c1cc.

---------

Co-authored-by: Orca <help@stably.ai>
2026-07-29 01:47:20 -07:00
Neil a721125d06
fix(perf): correct three 07-27 perf regressions (#11234)
* fix(perf): correct three 07-27 perf regressions

Traversal capacity cap no longer scales with worker concurrency
(#11026). retainWorkspaceSpaceScanEntry charged a traversal-wide entry
counter, so N workers each holding a listing multiplied the live charge.
At concurrency 48 a 48x2,100 tree (100,848 entries) hit the 100,000 cap
while 100x1,500 (150,100 entries, 50% more) passed, and scanLocalWorktree
treats the capacity error as terminal, reporting an intact worktree as
"Unavailable" with sizeBytes 0. The cap is now per directory listing --
the only quantity fixed by directory shape -- restoring the invariant
docs/workspace-space-scan-resource-bounds.md already states. Aggregate
live retention stays bounded by the unchanged 64 MiB byte cap.

Note: releasing each entry's charge at dispatch (the originally suggested
fix) was measured and does not help; the peak is set at admission, before
any entry is dispatched.

Repo image icons are no longer fully base64-decoded on every snapshot
publish (#11012). sanitizeRepoIcon reached decodeBase64Prefix, which
sized its buffer to the whole payload to read a 24-byte header, running
synchronously inside ipcMain.handle at a 250 ms throttle. Validation is
now memoized on source+src in a BoundedMap. Measured for 10 icons x
256 KB: 37.34 ms -> 0.67 ms per publish.

One over-long card label no longer discards the entire snapshot (#11012).
isDashboardSnapshot was all-or-nothing and dashboard-popout returned
early with no log while replaying lastSnapshot, so `orca terminal rename
--title "<1025+ chars>"` froze the pop-out board on its last good paint
with nothing surfaced. Labels are truncated at the producer, the
validator drops only the offending card, and both the rejection and the
drop are logged. The bound now lives in the shared snapshot contract so
producer and validator cannot drift.

Co-authored-by: Orca <help@stably.ai>

* fix(perf): charge a scan listing's parent path once, not per entry

The 4.1 fix made the entry cap per-listing but left the 64 MiB byte cap
charging parentPath.length for every entry in a listing. Because a
listing's entries all share one parent-path string, that multiplied the
path by the directory's width, so the byte cap measured checkout depth
rather than live heap.

The reported symptom therefore still reproduced at the production default
limits: 48 x 2,100 @ concurrency 48 raised a capacity error once the
worktree path passed ~58 characters, while the same layout at concurrency
1 succeeded. The shipped regression test could not see this because it
passes maxRetainedBytes: Number.MAX_SAFE_INTEGER, disabling the only cap
still in play. Measured at a real 65-char worktree root, 3 of the report's
4 documented layouts still failed.

The parent path is now charged once per listing, with its first entry, so
an empty listing strands no charge. Per-entry overhead is unchanged at
512 B + name, which still dominates the estimate, so the OOM protection
the original PR added is preserved.

Adds a production-default-limits case covering the report's layouts under
a deep root, plus an assertion that a short and a deep root reach the same
verdict -- the path independence docs/workspace-space-scan-resource-bounds.md
requires and which no existing test enforced.

Co-authored-by: Orca <help@stably.ai>

* fix(perf): prove the icon cache by decode count, not wall clock

The caching test asserted a per-publish millisecond budget, which failed
on CI at 5.64 ms against a 5 ms ceiling. Any threshold flakes on a loaded
box, so count real sanitizeRepoIcon entries instead: 10 repos x 20
publishes is 200 icon checks against exactly 1 decode. Added cases pin
the cache key (payload and source both re-decode; a cached image verdict
never answers for an emoji) and that a rejection is cached too.

Also drops budget.entries, which the per-listing cap left as a
traversal-wide counter no check reads -- exactly the shape a future
guard could reintroduce the concurrency bug from.

Co-authored-by: Orca <help@stably.ai>

* fix(dashboard): bound the project filter label the whole board rides on

#11042 added snapshot-level filterOptions whose project labels are
repo.displayName -- the same unbounded source this PR already bounds for
card.repoName, but one level up where dropping a card cannot recover it.
An over-long project name would fail isDashboardFilterOptions and take
the entire snapshot with it, which is the exact frozen-board failure the
per-card drop was added to end. Workspace-status labels are already
capped at 32 by workspace-statuses.ts, so only projects needed this.

Co-authored-by: Orca <help@stably.ai>

* fix(dashboard): disambiguate the repo icon cache key

The memoization key joined `source` and `src` with a space, but the
sanitizer's base64 pattern admits whitespace inside a valid `src`. A
rejected icon can therefore split the same concatenation differently and
inherit an accepted icon's cached verdict, reaching the pop-out's
`<img src>` without ever being sanitized. Length-prefix the source.

Co-authored-by: Orca <help@stably.ai>

---------

Co-authored-by: Orca <help@stably.ai>
2026-07-28 17:08:14 -07:00
Jinjing a40183389b
feat: bound direct SSH reconnect fan-out and recovery (#11003)
* docs: design for direct SSH reconnect fan-out

Capture the implementation-ready plan for host-qualified, epoch-fenced
SSH reconnect recovery after two rounds of multi-model LLM counsel review.

* docs: reconcile SSH reconnect fan-out design

* docs: close reconnect design consistency gaps

* feat: implement bounded direct SSH reconnect recovery

* fix: bound direct SSH retry settlement

* fix: harden direct SSH reconnect authority

* fix: preserve split SSH retry ownership

* fix: preserve SSH split continuation authority

* docs: record final SSH reconnect validation

* fix: preserve SSH authority through retained and detached state

* fix: retain SSH authority across delayed split mounts

* fix: close SSH authority recovery gaps

* fix: fence stale SSH transport replacement

* fix: serialize SSH target teardown

* fix: settle SSH teardown failures before reconnect

* fix: retire failed SSH reset sessions

* test: reconcile current main E2E contracts

* fix: close direct SSH reconnect review gaps

* fix: fence stale SSH reconnect side effects

* fix: close final SSH reconnect lifecycle gaps

* test: stabilize current-main reliability gates

* test: prove plugin navigation containment

* test: make plugin navigation oracle authoritative

* test: make plugin navigation oracle deterministic

* ci: allow sharded e2e suite to finish

* test: wait for runtime pane publication

* test: classify pane readiness by error code

* test: select close persistence terminal by tab identity

* docs: mark reconnect implementation validated

---------

Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com>
2026-07-28 12:33:17 -07:00
Jinjing efcc015d69 docs: add WeChat group 6 QR with overflow guidance
Show group 5 and group 6 QR codes side by side so people can join group 6 if group 5 is full.
2026-07-28 01:36:06 -07:00
Jinjing 9a8e21a47e
fix(workspace-space): bound traversal memory and serialize local disk scans (#11026)
* fix(workspace-space): serialize local disk and cap traversal memory

Prevent resource exhaustion during large workspace scans by limiting
local disk access to one concurrent `du` call and capping portable
traversal memory to 100k entries or 64 MiB per worktree. Fixes July 27
incident with 298 worktrees causing host stalls and renderer OOM.

Portable traversals now use fixed-worker iterative frames instead of
recursive promises. Capacity failures become unavailable rows. Behavior
below limits is unchanged.

* fix(workspace-space): bound concurrent SSH fallback traversals

Desktop-side SSH fallback traversals run in the main process with independent admission
budgets. Without limiting, up to six concurrent traversals could stack six 64 MiB budgets.
Cap remote fallback traversals to 2 concurrent, keeping aggregate admission at 2 × 64 MiB.

Also make capacity error messages reflect configured limits instead of hardcoded defaults.
2026-07-27 18:50:52 -07:00
Neil 97cb32c1cc
fix(terminal): release an abandoned synchronized-output frame on reveal (STA-2694) (#10907)
* fix(terminal): release an abandoned synchronized-output frame on reveal

Alt-screen agent TUIs (OpenCode/OpenTUI, Codex, grok) bracket every repaint
in `?2026h … ?2026l`. Hiding a pane mid-bracket — which a worktree switch or
cold-park lands on routinely, since these brackets are written many times a
second — leaves xterm's `decPrivateModes.synchronizedOutput` latched.

RenderService.refreshRows checks that latch *before* rendering, so while it
holds, every repaint Orca owns is a no-op: the forced render-pause repaint,
the plain `refresh()` fallback, and the shared glyph-atlas rebuild all render
zero rows while the xterm buffer is perfectly correct. Release the latch at
the two reveal repaint entry points so those repaints actually paint.

Also adds an OpenCode-shaped alt-screen e2e fixture and spec. The existing
inline-TUI convergence spec covers the normal-buffer shape (live block glued
to the bottom, history scrolling into scrollback); this covers the
full-screen alternate-buffer shape, where nothing scrolls and so no row ever
self-heals through the scroll path.

Scope note: xterm arms a 1s watchdog that clears this latch on its own, so
this closes a bounded window rather than the whole STA-2694 report. The e2e
spec passes with and without the production change for that reason; the unit
tests are what pin the behavior. Refs STA-2694.

* fix(terminal): clear the render model on the plain-refocus repaint path

`schedulePaneRevealPresent` — the atlas-preserving path a plain window
refocus takes — only called `terminal.refresh()`. xterm's renderers are
diff-based: `_updateModel` early-continues on any cell whose code/fg/bg/ext
still match the cached model, so a refresh repaints nothing for a pane whose
buffer never changed. When an occluded window loses its canvas contents while
that model stays populated, the refresh skips exactly the cells that went
stale and the pane keeps compositing pre-hide pixels — until a window resize
reallocates the model, which is the repair users find by hand.

Clear the model first (`RenderService.clear()` → renderer `clear()` →
`_clearModel(true)`) so the refresh becomes a guaranteed full repaint. That
drops cached cells and glyph vertices but NOT the texture atlas, which is
shared by every same-config terminal and whose mid-stream wipe re-arms xterm's
page-merge garble race (xterm.js #4480) — the reason this path is
atlas-preserving in the first place.

Also covers the DOM-renderer fallback in `resetWebglTextureAtlas`:
`clearTextureAtlas()` is what invalidated the model on the WebGL path, so a
pane without an addon had nothing invalidate it and hit the same skip.

Scope note: the e2e spec guards buffer/geometry convergence across the
hide/reveal boundaries and adds idle-agent and headful desktop-hide cases, but
it cannot observe a stale canvas — both oracles built for that (canvas-vs-buffer
ink sampling, screenshot-vs-forced-repaint) were proven blind by injecting the
defect, and the spec header documents why. The unit tests pin the ordering and
the atlas-preservation invariant. Refs STA-2694.

Co-authored-by: Orca <help@stably.ai>

* docs(terminal): hand off the STA-2694 reveal-artifact investigation

Records both fixed defects with their xterm mechanisms, the reveal/wake call
graph, why every e2e oracle for a stale canvas was proven blind, how to arm the
in-app render-desync sentinel on real hardware, and the one unverified lead
(dimension staleness) that would explain why a window resize specifically is
the repair users find. Refs STA-2694.

Co-authored-by: Orca <help@stably.ai>

* Revert "fix(terminal): clear the render model on the plain-refocus repaint path"

This reverts commit 0f7ec4458d37010338f16e70ff06957cb335e074.

* test(terminal): add a draw-command oracle for reveal repaints, and correct the STA-2694 scope

Every pixel oracle tried for STA-2694 was blind: `drawImage` on a
non-preserveDrawingBuffer WebGL canvas returns a re-rendered copy, and
Playwright's screenshot drives a fresh compositor frame that heals a stale paint
before capture. Reading pixels is self-defeating here — the read triggers the
repaint that hides the bug.

Count the WebGL draw commands instead, by wrapping GlyphRenderer.updateCell and
gl.drawElementsInstanced on the live pane. A draw command cannot be healed after
the fact, so "did the reveal actually repaint?" becomes directly observable.
Teeth-verified: removing releaseAbandonedSynchronizedOutput from
schedulePaneRevealPresent fails the stranded-latch test.

Two findings, both of which change previously-committed claims:

1. The 1s watchdog does NOT bound the synchronized-output defect. It is armed
   only inside `bufferRows`, and `refreshRows` returns at its `_isPaused` check
   first — so while a pane is occluded nothing reaches `bufferRows` and no timer
   is ever pending. A pane hidden mid-`?2026h` holds the latch with no watchdog
   behind it, indefinitely. ed1eaf55f1's "closes a bounded window" scope note was
   wrong; this is the unbounded garble the report describes, and the fix closes
   it. Corrected in the module doc comment.

2. It refutes the diff-based-staleness hypothesis behind 0f7ec4458d (reverted in
   8d5eacecb4). `_updateModel` does early-continue per unchanged cell, but
   `GlyphRenderer.render` then copies vertices for EVERY row up to
   `lineLengths[y]` and issues ONE full-viewport draw — measured identical
   instance counts (562) for a diff-skipped and a model-cleared refresh, with
   updateCell at 0 vs 561. The DOM renderer likewise replaceChildren()s every
   row unconditionally. Clearing the model could not change what reached the
   screen, and `_clearModel(true)` zeroes every glyph vertex while
   `RenderService.clear()` fires no repaint of its own — so it opened a
   blank-viewport window (also asserted here) for no benefit.

Also keeps the idle-agent and headful desktop-hide cases from the reverted
commit, since those were independent of the refuted production change, and
rewrites the alt-screen spec header to point paint questions at this oracle.
Refs STA-2694.

Co-authored-by: Orca <help@stably.ai>

* docs(terminal): rewrite the STA-2694 handoff after the refutation

Records that the garble window is unbounded (the 1s watchdog never arms for an
occluded pane), that the diff-based-staleness hypothesis was refuted by
measurement and reverted, why pixel oracles are structurally blind here, and the
two leads now closed by measurement (dimension staleness, lazy atlas bindings).
Refs STA-2694.

Co-authored-by: Orca <help@stably.ai>

* test(terminal): capture visual proof of the STA-2694 stale paint

The earlier screenshot oracles were blind because they compared a revealed pane
against a repaired one and both ran the same repaint code. Capturing the defect
directly works instead, because the mechanism is self-preserving: while
synchronizedOutput is latched, refreshRows returns before reaching the renderer,
so a compositor frame just re-composites the existing canvas texture and the
stale pixels survive the screenshot rather than being healed by it.

Latch a frame, write a full new frame the pane cannot paint, and capture. The
screenshot comes back byte-identical to the pre-hide one while the buffer holds
the new frame — the buffer/screen divergence users report — and differs after
the reveal repaint runs. Asserts both halves, so it fails if either the defect
stops reproducing or the fix stops repairing it.

Refs STA-2694.

Co-authored-by: Orca <help@stably.ai>

* test(terminal): note where the xterm gate-order double is pinned for real

The unit double encodes RenderService's paused-then-latch gate order, which can
drift on an xterm upgrade. Point at the e2e oracle that pins the same order
against the real renderer, so a future upgrade has a trail to the authoritative
check. Refs STA-2694.

Co-authored-by: Orca <help@stably.ai>

* test(terminal): add a perf budget for the synchronized-output release

releaseAbandonedSynchronizedOutput runs inside resetWebglTextureAtlas, which a
streaming alt-screen TUI can reach through the terminal-output atlas recovery
path — not only on reveal. Measure rather than assert that this costs nothing.

Steady state (a TUI that closes every frame it opens): 200 bracketed frames
produce zero releases, zero extra draw calls, and an unmeasurable early-out
cost. Worst case (every reveal finds a latched frame): 50 latched atlas resets
at 0.08ms each. Both are asserted with headroom, so the guard catches a future
change that makes this scan the buffer per pane rather than flaking on machine
speed. Refs STA-2694.

Co-authored-by: Orca <help@stably.ai>

* test(terminal): address review — drive real code paths, close vacuity gaps

CodeRabbit caught a genuine tautology in the perf budget: it timed a
hand-copied mirror of the early-out rather than the shipped function, so the
assertion would have held even if the real code grew a buffer scan. Driving
resetWebglTextureAtlases instead moved the measured cost from ~0 to ~0.03ms per
call, which is the honest number for the whole recovery; bound re-set to 0.4ms
(10x measured).

Other review fixes:
- Assert the draw counts both perf tests were measuring and logging but never
  checking, so the 'no extra draws' titles now mean something.
- Fail fast when decPrivateModes is unavailable; previously the latched test
  would pass without ever exercising the fix.
- Re-check the latch right after the worktree switch in the mid-frame test: the
  pane is visible until then, so the 1s watchdog can arm and clear it before the
  hide, making the run vacuous.
- Count scheduleRevealPresent invocations instead of returning a literal true,
  so a missing test hook no longer masquerades as a production failure.
- Assert the latch clears on every reveal iteration, not just the last.
- Make the fixture heartbeat write atomic (tmp + rename); writeFileSync
  truncates first, so a reader could see '' and read it as frame 0.
- Relabel assertRevealPixelsNeedNoRepair as the weak secondary check it is; it
  contradicted the file header by calling itself 'the decisive paint assertion'.

Refs STA-2694.

Co-authored-by: Orca <help@stably.ai>

---------

Co-authored-by: Orca <help@stably.ai>
2026-07-27 16:32:24 -07:00
OrcaWin cd05f2ff93
Implement robust orchestration primitives and connected-server workers (#9925) 2026-07-27 12:31:37 -07:00
github-actions[bot] 94009f63ff Update README downloads badge 2026-07-27 13:20:45 +00:00
Neil 6b16c20796
fix(memory): clarify Resource Manager accounting (#10821) 2026-07-26 19:46:29 -07:00
github-actions[bot] 1c17405431 Update README downloads badge 2026-07-26 07:11:08 +00:00
github-actions[bot] c4b7aeddd7 Update README downloads badge 2026-07-26 00:59:08 +00:00
github-actions[bot] 9eff3728a3 Update README downloads badge 2026-07-25 18:35:56 +00:00
github-actions[bot] be0212e278 Update README downloads badge 2026-07-25 07:03:47 +00:00
github-actions[bot] a903304ad4 Update README downloads badge 2026-07-25 00:55:11 +00:00
github-actions[bot] 506b75c768 Update README downloads badge 2026-07-24 21:44:43 +00:00
github-actions[bot] 3dfbb10775 Update README downloads badge 2026-07-24 18:47:45 +00:00
github-actions[bot] 17959a22b8 Update README downloads badge 2026-07-24 12:47:46 +00:00
Neil 662d23f9b3
fix(readme): point French star badge to repo root (#10335)
Co-authored-by: Orca <help@stably.ai>
2026-07-24 00:39:12 -07:00
github-actions[bot] cda97cec41 Update README downloads badge 2026-07-24 07:10:19 +00:00
Jinjing c88cf8413f docs: finish Android APK 0.0.32 link rollup
Point localized READMEs and in-app Android download CTAs at
mobile-android-v0.0.32 (English README and orca-site were already updated).
2026-07-23 22:20:43 -07:00
github-actions[bot] 43f626574b Update README downloads badge 2026-07-24 00:34:31 +00:00
github-actions[bot] e55e15d44b Update README downloads badge 2026-07-23 18:43:10 +00:00
github-actions[bot] 434f433287 Update README downloads badge 2026-07-23 12:48:59 +00:00
OrcaWin 569e9d88f8
docs: update WeChat QR to group 5 (#10126)
All earlier WeChat groups are full; point the README QR and copy at group 5.
2026-07-23 00:19:06 -07:00
github-actions[bot] 2b7aa0aead Update README downloads badge 2026-07-23 07:09:53 +00:00
Neil a0944cc129
fix(linux): restore Ubuntu 20.04 launch — pin node-pty glibc symbols + add glibc/libstdc++ packaging gate (#9902) (#10019)
* fix(linux): restore Ubuntu 20.04 launch by pinning node-pty glibc symbols (#9902)

The bundled node-pty pty.node is compiled from source in release CI on
ubuntu-latest (glibc 2.39). glibc's 2.32-2.34 libpthread/libutil merge
relocated openpty/forkpty (GLIBC_2.34) and pthread_sigmask (GLIBC_2.32)
into libc under new symbol versions, so the from-source build bound to
versions absent on Ubuntu 20.04 (glibc 2.31). The main process imports
node-pty at startup, so the app crashed on launch. pty.node is the sole
blocker (Electron needs GLIBC_2.25; other native modules <= 2.17).

- Patch node-pty: a .symver shim pins the 3 symbols to their pre-merge
  version (GLIBC_2.2.5 x64 / GLIBC_2.17 arm64), and Linux-only ldflags
  force libutil.so.1/libpthread.so.0 back into DT_NEEDED. Guarded to
  Linux; macOS/Windows untouched.
- Add a packaging gate (verify-linux-glibc-floor.cjs, afterPack): reads
  each bundled native binary's objdump -p version needs and fails the
  Linux build if any strong GLIBC_/GLIBCXX_/CXXABI_ node exceeds stock
  Ubuntu 20.04 (glibc 2.31 / GLIBCXX_3.4.28 / CXXABI_1.3.12). Catches
  GLIBC_ABI_DT_RELR, rejects GLIBC_PRIVATE, skips weak needs, fail-closed.
- Docs + tests; the lazy sherpa-onnx speech prebuilt (GLIBCXX_3.4.29,
  never loaded at launch) is a documented libstdc++-floor exemption.

* fix(linux): assert DT_NEEDED provider deps in the glibc-floor gate

Harden the packaging gate (flagged in adversarial re-eval): the version-floor
check alone can false-pass if the patch's forced `-l:libutil.so.1` ever silently
drops — the pinned openpty@GLIBC_2.2.5 still resolves from libc's compat alias at
build time, but fails to load on Ubuntu 20.04 where openpty/forkpty live only in
libutil. The gate now also asserts that any binary importing openpty/forkpty
keeps libutil.so.1 in DT_NEEDED. Validated on a real symver-pinned .so with
libutil dropped (now fails) vs. present (passes). Documents the recommended
real-host smoke-test follow-up.
2026-07-22 19:11:44 -07:00
github-actions[bot] 6a8e992c32 Update README downloads badge 2026-07-23 00:55:31 +00:00
github-actions[bot] 05ec4b749b Update README downloads badge 2026-07-22 18:41:48 +00:00
github-actions[bot] 334027cfd5 Update README downloads badge 2026-07-22 12:47:33 +00:00