orca/tests
Neil c06bf64b48
test(e2e): make Codex typing-latency harness measure real echo latency (#10660)
* test(e2e): make Codex typing-latency harness measure real echo latency

The local Codex typing-latency spec produced meaningless numbers. Four
defects, all fixed here:

1. False-positive readiness. `/Ask Codex|OpenAI/i` matched "OpenAI's
   command-line coding agent" on the *sign-in* screen, so the test went
   "ready" against a login prompt and measured typing into a non-composer.
   Now gated on the composer status bar (`/Context \d+% used/i`), which
   only the live composer draws. Banner text is unusable: the serialized
   buffer interleaves ANSI escapes through those glyphs.

2. Missing auth. The E2E profile runs an isolated HOME with a managed
   CODEX_HOME that has no auth.json, guaranteeing the sign-in screen. The
   launch now pins the real ~/.codex, and skips with a clear message when
   auth.json is absent instead of silently measuring a login screen.

3. Measurement overhead swamped the signal. Per-key latency was measured
   by polling getTerminalContent() every 5ms, so each sample was real echo
   latency + full buffer serialize + CDP round-trip + poll granularity.
   Measurement now happens entirely in-renderer: an in-page hook stamps
   performance.now() on keydown (window capture phase, before xterm
   forwards to the PTY) and again in xterm's onWriteParsed once the glyph
   is in the viewport, with onRender giving a separate time-to-paint.
   Samples are drained in one page.evaluate after typing ends — zero CDP
   round-trips inside the measured window.

4. Thresholds were meaningless (median<150ms / worst<500ms). Replaced with
   p50<35 / p95<60 / max<120, based on 10 local runs.

Also: 60 keystrokes instead of 24 with the first 10 discarded as warmup,
p50/p95/max instead of a lone median, lowercase-only input so the slash
and file-mention popups can't perturb later keys, an assertion that no
keystroke went unechoed, and a terminal dump on readiness failure.

Measured (10 local runs, headless, real Codex 0.145.0):
  echo (key->parse)   p50 21.6-22.6ms, p95 23.2-41.5ms, max 23.4-58.7ms
  paint (key->render) p50 25.5-32.9ms, p95 34.3-49.7ms
A plain-shell control on the same probe reads p50 2.0ms / p95 3.0ms,
confirming the ~22ms is Codex composer redraw cost rather than a harness
floor — the old harness reported ~29-30ms for everything.

Co-authored-by: Orca <help@stably.ai>

* test(e2e): widen Codex latency tail budgets and assert terminal focus

Follow-up calibration over ~20 local runs: the per-key distribution is
unimodal at p50 21.3-22.7ms with rare isolated spikes to ~90-125ms that
are not a steady-state shift. Tail budgets move to p95<80 / max<150 so
only a sustained regression fails; p50<35 still gates the steady state.

Also assert the xterm helper textarea actually took focus. One run typed
all 60 keys with only 5 parse events because focus was lost, which
previously surfaced as an opaque sample-count mismatch.

Co-authored-by: Orca <help@stably.ai>

---------

Co-authored-by: Orca <help@stably.ai>
2026-07-25 20:06:19 -07:00
..
e2e test(e2e): make Codex typing-latency harness measure real echo latency (#10660) 2026-07-25 20:06:19 -07:00
.gitignore feat: add idempotent E2E test suite with headless Electron support (#671) 2026-04-19 12:09:32 -07:00
playwright.config.ts Stabilize sharded e2e CI concurrency (#6312) 2026-06-24 18:39:02 -07:00