* fix(pty): release ptmx fd on natural exit + defuse SIGHUP-to-recycled-pid
Daemons accumulated ptmx fds over time because node-pty's UnixTerminal
only releases the master fd when destroy() runs. On the natural-exit
path (the common case — user closes a tab, shell runs `exit`) nothing
ever calls destroy(), so the fd leaks until GC. On macOS this
eventually hits kern.tty.ptmx_max=511 and all new terminals fail to
spawn.
Fix: release the fd synchronously on every teardown path (natural
exit, explicit kill, stale SSH spawn, daemon shutdown) and close the
concurrent SIGHUP-to-recycled-pid hazard inside node-pty's
UnixTerminal.destroy().
- src/main/daemon/pty-subprocess.ts: synchronous POSIX proc.kill
neutralization inside proc.onExit; dead guards on forceKill/signal
so they never target a reaped-and-possibly-recycled pid
- src/main/daemon/session.ts: new disposeSubprocess() for already-
exited sessions (fd release only, no SIGKILL) — avoids sending
SIGKILL to a recycled pid during daemon shutdown
- src/main/daemon/terminal-host.ts: dispose loop routes on isAlive —
live sessions get forceKillAndDisposeSubprocess (SIGKILL + fd
release), exited sessions get disposeSubprocess (fd release only)
- src/main/providers/local-pty-provider.ts: same POSIX kill
neutralization at top of onExit for the legacy local path
- src/relay/pty-handler.ts: same neutralization in wireAndStore;
disposed flag guards all public entry points; dispose() uses
SIGKILL (not SIGTERM) before destroy since the relay is exiting;
killTimer fallback + immediate-shutdown + stale-spawn cleanup all
call disposeManagedPty + ptys.delete so wedged children (D-state,
bad NFS) can't leak map entries against the 50-PTY cap
Windows is exempt everywhere — WindowsTerminal.destroy IS a kill()
call internally (closes the ConPTY agent), so neutralizing would
turn destroy into a no-op and leak the agent.
See docs/fix-pty-fd-leak.md for the full design.
Co-authored-by: Orca <help@stably.ai>
* fix(pty): patch node-pty native off-by-one leaking /dev/ptmx per spawn
node-pty 1.1.0's pty_posix_spawn on macOS walks low_fds[0..2] in an
allocation loop that breaks at the first fd >= STDERR_FILENO, then
cleans up via `for (; count > 0; count--) close(low_fds[count])`. In
the typical case (break at count=0) the cleanup body never runs and
low_fds[0] — a /dev/ptmx handle — leaks per spawn. Fixed upstream in
microsoft/node-pty af053f2 (PR #882), not in any 1.1.0 release.
Backport the 3-line cleanup-loop fix as a pnpm patch. E2E validated
against a dev daemon: 200 spawn/kill cycles kept the daemon's ptmx
fd count flat at baseline; prior runs reproduced linear 1-per-spawn
growth. Also documents the native root cause as a status addendum in
docs/fix-pty-fd-leak.md — the JS-side destroy() discipline previously
landed is still load-bearing for the SIGHUP-to-recycled-pid hazard and
for synchronous fd release on daemon shutdown.
Co-authored-by: Orca <help@stably.ai>
* fix(pty): capture stable kill spy ref in pty.test.ts
destroyPtyProcess reassigns proc.kill = () => {} on POSIX to defuse
the SIGHUP-to-recycled-pid hazard (see docs/fix-pty-fd-leak.md). After
that reassignment, proc.kill.mock is undefined and the assertions
crashed in CI. Capture a stable reference to the vi.fn() before it
gets reassigned.
Co-authored-by: Orca <help@stably.ai>
---------
Co-authored-by: Orca <help@stably.ai>
* fix(daemon): defer startup-command flush past shell raw-mode switch
When launching an agent (e.g. `claude`) through the quick-launch menu,
the command name appeared twice in the terminal — once from kernel
echo, once from readline's prompt redraw. The daemon session was
flushing its pre-ready stdin queue the moment the OSC 777 shell-ready
marker arrived, but that marker fires from precmd_functions /
PROMPT_COMMAND — before the shell draws its prompt and before
zle/readline flips the PTY into raw mode. Writing while ECHO was still
on produced the visible duplicate.
Mirror the gating already used by the non-daemon path
(local-pty-shell-ready.ts::writeStartupCommandWhenShellReady): wait
for the next data chunk after the marker (the prompt draw) plus a
short 30ms delay, with a 50ms wall-clock fallback for the case where
the prompt arrives in the same chunk as the marker.
This regression became visible after #1025 made the shell-ready
wrappers persist reliably — previously the marker was often missed
and the 15s fallback path fired long after the shell was already in
raw mode, masking the race.
* refactor(daemon): extract PostReadyFlushGate out of Session
Factor the post-ready flush gating out of Session into its own class in
post-ready-flush-gate.ts. Session's only responsibility is now to
arm() the gate on shell-ready, notifyData() on subsequent PTY data,
and clear() on teardown. Removes the max-lines oxlint-disable added in
the previous commit and puts the timing behavior behind focused unit
tests.
No behavior change — kept as a separate refactor commit from the
behavior fix for reviewability.
* fix(daemon): keep queueing writes while post-ready flush gate is pending
Codex review caught an ordering regression: once transitionToReady()
sets _shellState to 'ready', any Session.write() that arrives during
the 30–50ms gate window was being written directly to the subprocess,
bypassing the still-unflushed preReadyStdinQueue. That let late input
race ahead of the buffered startup command.
Expose PostReadyFlushGate.isPending and continue queuing while the gate
is armed so queued writes drain in their original order before any
fresh input reaches the subprocess.