* fix(cascade): per-kind prune staleness + rebuild safety
Adversarial review of the review-response fixes found five real defects,
all in code this PR introduced.
Health signal (P1): prune staleness reported the time since the NEWEST
successful prune across kinds, so on a multi-kind deployment (every real
one) a single kind whose cleanup died was masked by the others pruning
on schedule — /health stayed green while that table's index dir grew
unbounded, the exact incident the signal exists to catch. Report the
WORST kind instead and name it in the reason. The failure streak could
not cover this either: an intervening light beat resets it, so it never
reaches the threshold for a prune-only failure. Documented that split of
duties.
Spurious fallback rebuilds (P1): the benign-conflict carve-out excluded
the heavy beat, justified by "runs under the write lock, so it can't hit
this benignly" — but that lock is in-process only, so a second process
(a long `cascade backfill`, a `cascade sync`) preempts prune's Rewrite
commit. Those counted as real failures, and ~25min of cross-process
churn reached the threshold and fired a fallback rebuild, which drops
every index before recreating it; a rebuild that also lost the race was
swallowed as a warning, leaving the table with no FTS index (every
/search on that kind 500s) until the next 12h sweep. Treat commit
conflicts as benign on both beats and let prune-staleness detect a prune
that genuinely stops succeeding.
cascade rebuild (P1 ×2 + P2): it drops and recreates tables with no
guard while --help/docstrings advertised it as safe, so `rebuild --yes`
against a live daemon corrupted the rebuild (the daemon keeps writing
through cached handles). Refuse when the OME jobstore lock is held,
reusing backfill's detection and its exit code 3. It also ran the
pre-drop migration pass (`ensure_business_indexes`) against the damaged
table, so on the corruption classes it exists to repair (missing column,
un-alterable type) the recovery path died on the damage itself — skip it
via `_runtime(ensure=False)`. Reset the queue BEFORE dropping so every
crash window converges on "queue pending → re-index" instead of empty
tables with a fully-done queue (a silently empty deployment), and handle
Ctrl-C with exit 130 plus a resume hint.
Recovery guidance (P1): the nullable-vector migration error still told
users to wipe the index directory — which this PR's own runbook documents
as the wrong recovery (queue stays done, index comes back empty). Point
it at `everos cascade rebuild`. Dropped the schema-drift error's
"restart first" step too: the startup migrations only alter nullability,
never a name or type, so a name/type drift never self-heals.
Also: backfill's post-write prune passed a zero retention window from a
separate process, able to delete files under a daemon /search still
holding that version — pass the daemon's window instead. Runbook gains
the /health cascade block (thresholds, what flips healthy, why
failed_permanent does not) and its quoted schema-drift error now matches
the code.
Tests: the three safety mechanisms this PR adds were unpinned — a
one-line revert of any of them passed the suite. Added per-kind
staleness, heavy-beat benign conflict, benign-filter negative case
(an error whose message merely contains "retryable" must still count),
prune recurrence across light beats (mutation-verified: hoisting the
attempt-clock advance out of the heavy branch turns it red), the prune
timeout releasing the write lock, timeout-below-cadence, the rebuild
server guard, and a tier3 assertion that the /health cascade block is
actually wired. Froze the last fabricated-monotonic test.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Also: treat ``last_prune_attempt_at == 0.0`` as "never attempted" instead of
comparing clocks. ``monotonic()`` is boot-relative, so ``now - 0 >= cadence``
is false for the first ~cadence of container uptime — the catch-up prune was
skipped exactly when a fresh process most needs it (and made a test depend on
the runner uptime, which CI caught).
* fix(lancedb): bound every write-lock critical section
run7 (1h at 2.5x rate, concurrent CLI maintenance, doubled fuzz) reproduced a
table whose version cleanup stopped permanently: 150 versions retained, disk
11x live size, while the other two tables sat at 1 version each — and with no
error logged anywhere, because nothing failed. It simply never returned.
Three things combined. The maintenance scheduler allows one task per table (a
LanceDB table takes one writer), so it skips a kind whose task is still in
flight. The prune timeout sat *inside* the lock and covered only the cleanup
call. And the other six critical sections on that lock — add, upsert, update,
delete, delete_by_md_path, rebuild_indexes — had no deadline at all. So one
operation stuck anywhere outside that narrow window wedged the table for good:
every writer blocked on acquire, and every later heartbeat was turned away
because the stuck task never finished.
Make it structurally impossible instead of patching prune: all seven sections
now go through `LanceRepoBase._locked(budget, op)`, where the deadline covers
**acquisition and the body**. No path can wait for this lock, or hold it,
indefinitely. Budgets are hang-catchers, not throughput limits: 120s for row
writes, 600s for an index rebuild, the existing 60s for prune.
Expiry raises `VectorStoreBusyError`, deliberately under `ExternalServiceError`
so the cascade worker retries the row; under `VectorStoreError` a transient
lock contention would be marked permanently failed and need a manual
`cascade fix`.
Tests: a stuck holder now makes a waiter fail its deadline and release (the
lock is reusable afterwards), and the prune timeout is pinned as retryable.
Verified by mutation — moving the timeout back inside the lock makes a waiter
block until the enclosing observation window expires (1001ms vs 51ms), i.e.
wait forever in production.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(lancedb): size write-lock budgets from measurements
120s for a row write was a guess, and a bad one: the budget doubles as the
detection latency for a wedged table, so an over-slack value means minutes of
blocked writers before anything surfaces — the failure this change exists to
prevent.
Measured the four locked write ops on a local SSD across table sizes and batch
sizes (10k-100k rows, 50-500 rows per call): add 3-22ms, upsert (merge_insert,
the read-modify-write one) 6-25ms, update 2-4ms, delete 2-3ms; worst observation
63ms, and flat in both dimensions since these are append-and-commit, not scans.
So: writes 120s -> 15s (~240x the worst observation, enough for a contended disk
and several waiters queued ahead — the deadline includes acquisition and
asyncio.Lock is FIFO), rebuild 600s -> 300s (still the one genuinely slow
section at ~0.3s per 50k rows per indexed column). Prune stays 60s.
Test pins the sizing intent: writes stay in the tens of seconds, and
rebuild > prune > write so the slowest section is not the most eagerly killed.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* feat(lancedb): record wait/hold time on write-lock critical sections
A soak run stalled one table's writes for ~16s and the logs could not say why:
maintenance beats only log at `debug`, so a section that is slow but still
inside its deadline is invisible, and the timeout warning did not distinguish
"never acquired the lock" from "acquired it and overran".
`_locked` now carries that apart. The deadline warning gains `acquired`,
`waited_seconds` and `held_seconds` — `acquired` alone answers whether a holder
was slow or this operation was — and a completed section that held the lock for
at least a second logs `lancedb_write_lock_slow_hold` at info, so a stall that
never reaches a deadline still leaves a trace.
Uses `time.monotonic` (elapsed measurement, not wall clock — the datetime
discipline bans `time.time`).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(search): reject a mismatched query vector before it reaches LanceDB
A soak run showed every slow search was a failing search: `search:vector` p50
251ms / p99 1.6s, but its 8 requests over 10s were exactly its 8 failures
(13-14s each). The cause was a query vector whose width disagreed with the
index. LanceDB only notices after the query is built and reports it as an
opaque `ValueError: Invalid input, No vector column found to match…`, which
escaped as an unhandled 500.
Validate at `_embed_query` — the single point every query vector passes
through — against the provider's declared `dim`. Microseconds instead of 13s,
and a named `ConfigurationError` (500 + CONFIGURATION_ERROR) instead of an
unhandled crash. Deliberately not `InvalidInputError`/422: callers only send
query *text*, so a bad width is our provider's fault, not the caller's.
Also cap traceback rendering. structlog's default is
`RichTracebackFormatter(show_locals=True, max_frames=100, extra_lines=3)`,
which on an async stack rendered 82 frames into 6423 log lines per exception —
85MB of server.log across 11 of them — at ~290ms of synchronous CPU each, and
risks printing request payloads into logs. With locals off and 15 frames the
same traceback is 103 lines and 10ms.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* docs(changelog): record the storage-reliability work under Unreleased
#379 merged without changelog entries, so this covers both it and the
follow-up work in this branch: the maintenance split (compaction vs
reclamation) that fixes unbounded index growth, bounded write-lock critical
sections, the /health cascade readiness block and its alert contract,
`cascade rebuild`, schema type-drift detection, the query-vector width check,
and the traceback-rendering cap.
Each entry states the operator-visible consequence, not just the change —
`cascade rebuild` now refusing to run against a live server, benign-conflict
warnings dropping in volume, and `/health` being able to report a stalled kind
that was previously invisible are all behaviour changes someone will notice.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: zhanghui <zhanghui@shanda.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* refactor(config): make [embedding] and [rerank] soft dependencies
Make [embedding] and [rerank] soft runtime dependencies so a
freshly-onboarded user can run EverOS end-to-end with only [llm]
configured. Previously the server refused to start without embedding,
locking out anyone who just wanted keyword-only search.
## Capability tiers
- Tier 1 ([llm] only) KEYWORD search, add/flush, md
writes, cascade sync
- Tier 2 ([llm] + [embedding]) + VECTOR / HYBRID search,
reflection, skill extraction,
backfill
- Tier 3 ([llm] + [embedding] + rerank) + AGENTIC search, knowledge
Tier upgrades require a server restart (capability accessors cache
for the process lifetime). Tier downgrades are read-safe: a Tier-3
user who drops [rerank] can still read/rename/delete existing
knowledge documents; only write/search endpoints return 422.
## What changed
- Component accessors — component/{embedding,rerank,llm}/accessor.py
are the single process-wide provider singletons. service/* never
maintains parallel singletons; it consumes get_embedding_capability()
/ get_rerank_capability() / get_llm_client() directly. Build-time
ValueError from the factory is logged as capability_build_failed
(was silently swallowed).
- Error mapping — ProviderNotConfiguredError -> 422 with everos.toml
section hints (never EVEROS_* env-var strings).
LanceDBMigrationError fails loud with escalating recovery guidance
(restart -> wipe index). LLMNotConfiguredError in search maps to
None for KEYWORD degradation.
- Nullable-vector LanceDB migration — schema v2 makes the vector
column nullable so Tier-1 rows can land without embeddings.
Migration is guarded by a cross-process memory_root_lock
(fcntl.flock + anyio.to_thread) and runs optimize() per table
after Phase-1 backfill to reclaim manifest bloat.
- Cascade — knowledge handlers register unconditionally (Tier-3 ->
Tier-2/1 downgrade no longer strands DELETE); embed-requiring
strategies use body-guards that check capability.available at
execution time. _TABLE_SPECS has an import-time drift assertion
against BUSINESS_SCHEMAS_WITH_VECTOR.
- `everos cascade backfill` CLI — Phase-1 (embed missing vectors) /
Phase-2 (emit synthetic events for cascaded processing) / Phase-3
(sync new skill files). Exit codes: 0 / 1 / 2 / 3 (server running
preflight) / 4 (COMPLETED_WITH_FAILURES — per-row failures rolled
up) / 130 (SIGINT). OMEConfig.crash_recovery_enabled=False in
backfill engines prevents stale-RUNNING rows re-enqueuing into a
smaller strategy registry.
- /health — reports capabilities + disabled_features per tier so ops
can distinguish "boots but degraded" from "boots and full".
- Presentation split — memory / service / infra never import typer /
click. TyperPresenter Protocol + run_backfill() live in
entrypoints/cli/commands/_backfill_cmd.py. Enforced by
import-linter.
- Startup hint — unconditional count_rows(filter="vector IS NULL")
sweep emits unbackfilled_memory_rows (event name + hint text
pinned) when Tier-1 rows exist. ParserLifespanProvider warms the
everalgo.parser import at boot so /health doesn't block on first
call.
- Knowledge upload UTF-8 short-circuit — _looks_like_utf8_text()
routes text/* mime and known plaintext extensions (md/txt/rst)
straight to UTF-8 decode instead of the parser. Prevents 503
Multimodal-not-configured when Tier 3 sans [multimodal] uploads a
markdown doc.
## Sync history with main (2 merges collapsed into this squash)
Merged origin/main at 6dcd3eb (v1.1.4 -> v1.2.0 adds OTel tracing,
/api/v2 alias, TracingLifespanProvider, per-cascade-embedding span
fix, memory-op instrumentation) and later at 42629df (PR #366
backfills v1.1.4 CWE-22 knowledge path traversal fix + cascade
retry-budget rework + errors.py -> core.errors.ExternalServiceError).
Key merge decisions:
- service/search.py adopts single wrap site — component.llm accessor
already applies UsageRecordingClient when observability is on;
service layer never keeps a parallel LLM singleton (Round-1 CR
rule: "service layer never maintains parallel singletons").
- Knowledge router prefix moved to /knowledge; create_app() mounts
it under both /api/v1 and /api/v2.
- Cascade retry classification uses ExternalServiceError from
core.errors (cascade/errors.py deleted). _MAX_TOTAL_RETRIES=12
cross-cycle budget preserved.
- Fixed backport typo: MemoryRoot.default() -> MemoryRoot.resolve()
(no .default() classmethod exists — main PR #366 shipped a broken
call).
## Verified layering
$ git grep -l "^import typer\|^from typer" src/everos/{memory,service,infra}
# empty
$ git grep -l "^import click\|^from click" src/everos/{memory,service,infra}
# empty
Memory / service / infra layers clean of CLI presentation libraries.
## Review history
Three rounds of Fable 5 (opus) code review across the pre-squash
commit history closed 38 findings total:
- Round 1: 10 findings (fail-loud migration, backfill hardening,
knowledge router gate scoping, SearchManager guards, profile
throttle lift)
- Round 2: 13 findings (hermetic test env, hot-reload doc drift,
knowledge handler registration, Phase 3 sync guarantee, Phase 2
idempotency, profile event-first path, OMEConfig crash-recovery
gate, cross-process migration lock, batch embed per-row fallback,
LanceDB optimize, typer/click layer split)
- Round 3: 15 Minor cleanups (accessor unification, marker revert,
episode query hygiene, --verbose subcommand, parser lifespan warm,
task-number scrub, temporal-overlap test, real-SIGINT slow mark,
4 design-note back-references)
Full per-round context lives in the PR description on GitHub.
## Test plan
- make lint (ruff + import-linter 3 contracts + assets +
deprecated-names + github-docs + datetime + OpenAPI drift)
- Hermetic env full pytest — 2027 passed / 7 deselected (7 = slow +
live_llm markers)
- Manual e2e across Tier 1/2/3 (21/21 assertions across v1/v2
double-mount and Tier 3 -> Tier 2 downgrade)
- /health reports correct capabilities + disabled_features per tier
## Known follow-ups
- .superpowers/sdd/followup-http-bridge.md (gitignored) — Path A for
spec §10's "backfill 期间 EverOS 完全可用" promise
- _TABLE_SCHEMA_VERSION docstring — v3+ migrations need a version
dispatch table
- extract_user_profile.py throttle-counter block — replace LanceDB
count_by_owner with a sqlite memcell count
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(review): close 3 blockers surfaced by round-4 review
N1. cluster_repo.find_cluster_id_for_member was cross-owner-unsafe.
Its reverse index (member_type, member_id) alone cannot disambiguate
two owners whose entry_id happens to collide — entry_id is
deliberately only per-owner unique (see entries.py:47:
'Cross-user uniqueness is handled at the database layer via a
composite <user_id>_<entry_id> field; it is not encoded into the
EntryId string itself'). Phase 2's _scan_all_rows crosses all
owners, so on any multi-owner root, same-day seq=1 episodes under
different owners would either false-hit each other's cluster or
be silently skipped from clustering. Add required (app_id,
project_id, owner_id) keyword args + JOIN Cluster (which already
carries scope) to filter by parent scope. Prior signature had zero
production callers except two the same PR just added, so the
API break is contained. Regression test: two owners persist a
cluster each around the same entry_id, each lookup resolves to
its own owner's cluster, a third owner's lookup returns None.
N2. Ctrl-C / EOF at the y/N prompt was landing on the generic
except Exception branch (exit 2 with rich traceback) instead of
the exit-130 interrupt path. Root cause: typer 0.15+ vendored
click under typer._click, so typer.Abort and the standalone
click.exceptions.Abort are distinct classes. The interrupt-branch
catch only listed the standalone one; every existing 'abort'
test was manually raising click.exceptions.Abort so the miss
was a false-positive guard rail. Widen the catch to
(typer.Abort, click.exceptions.Abort) and declare click as a
first-class dependency (it was only pulled in via uvicorn).
Regression test: raise real typer.Abort() at the confirm step
and assert exit 130 + INTERRUPTED banner.
N3. _looks_like_utf8_text used mime.startswith('text/'), which
caught text/html as well. HTML uploads then bypassed everalgo's
_aparse_html — losing clean_html_for_llm (strips <script>/<style>
/<nav>/<iframe> + HTML comments) and the 1 MiB output cap. A
40 MiB .html with <script> bodies and <!-- prompt injection -->
comments would flow straight into the extraction LLM. Replace
with explicit allowlist {text/plain, text/markdown, text/x-rst,
text/x-markdown}; text/html and any future text/* mime now
default to the parser path. Test matrix asserts text/html →
False (was regressed as True by the earlier commit).
Hermetic env full pytest: 2033 passed / 7 deselected (+6 tests
from these regressions).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(review): close round-4 major + minor + PR body errata
Round-4 review-driven cleanup. Blocker fixes (N1/N2/N3) landed in
561b5fe. This commit closes the remaining CONFIRMED items:
Major:
- J3 MemoryRoot.default() -> resolve() was a breaking public API
rename this branch introduced (main still has default()). Adds
default() as a backward-compat alias forwarding to resolve()
with DeprecationWarning; CHANGELOG entry under Unreleased.
- J4 episode_repo.list_by_owner_after_ts(limit=N) truncates in
fragment order (== insertion order), NOT newest-first. Docstring
now spells out the trap so a future caller passing limit for a
'newest N' window doesn't silently get the oldest N.
- J5.2 TyperPresenter.nothing_to_backfill picked colour via
'could not be read' in message — a domain wording change
would silently flip yellow -> green. Signature gains explicit
scan_failed: bool kwarg; CLI colour-picks off the flag.
- J5.5 phase_header was Protocol-declared but never dispatched
(run_backfill calls _print_phase_header directly). Removed the
dead Protocol method + both no-op implementations.
- J6.3 3 inline from everos.core.errors import ... inside
Phase 1/2/3 preflights promoted to a single top-level import.
- J7 subject-side embed failure was silently exit-0 because
rows_processed advanced whenever any side wrote. Now: a row with
a needed side still NULL counts as rows_failed (exit 4 =
COMPLETED_WITH_FAILURES). Gated on spec.subject_of + row.needs_*
+ row.subject_text so non-Episode tables and subject-empty rows
don't false-positive.
- J9 test_migration_cross_process.py did NOT actually test cross-
process (all 5 tests mock memory_root_lock). Renamed to
test_migration_lock_wiring.py; docstring now scopes it to
'lock-invocation wiring' and points at test_core/…/test_locking.py
for real flock coverage.
- J10 Phase 1 lacked the server-running preflight Phase 2/3 have.
--phase all against a live server would burn Phase 1 embed API
calls (real cost) before Phase 2 halted with exit 3. Phase 1 now
probes _probe_ome_lock_available first; regression test in
test_backfill_preflight.py; upgrade_path integration patches the
probe so its in-process 'server + backfill' scenario stays valid.
Minor:
- M1 knowledge upload with NUL byte or filename > 255 bytes UTF-8
used to raise ValueError/OSError at write_bytes → 500 with a
half-written md left on disk. _safe_original_filename now
rejects both up front with InvalidInputError (→ 400).
- M2 backfill optimize() now passes cleanup_older_than=timedelta(0)
so older manifest versions are physically pruned (previous call
compacted fragments but left the manifest chain on disk).
- M3 verify_business_schemas remediation text used to jump straight
to 'rm -rf ~/.everos/.index/lancedb'; now walks restart → wipe.
- M5 multimodal/accessor.py capability_build_failed warning added
so all four provider accessors log symmetrically (was silent).
- M7 test_knowledge_api parser-absence tests call
parser_available.cache_clear() around the sys.modules patch so
the lru_cache doesn't strand a stale True/False.
- M8 cascade_handler_embed_skipped (6 handlers) demoted INFO → DEBUG:
Tier 1 imports were generating N × 6 handler-info lines per md.
- M10.1 test_drift_scenario_would_raise was a tautology (compared
two hardcoded string sets, never touched the guard). Now
monkey-patches BUSINESS_SCHEMAS_WITH_VECTOR to a superset and
reloads _backfill, proving the import-time RuntimeError fires.
- M11 test_cascade_verbose_position subprocess.run calls gain
env= — scrubs EVEROS_* from the developer environment so the
same footgun as round-2 B1 doesn't re-appear inside subprocesses.
- M14 health() -> dict degraded the OpenAPI schema to
additionalProperties: true. Introduce HealthResponse +
HealthCapabilities Pydantic models so clients get real field
shape; docs/openapi.json regenerated.
- M16 routes/knowledge.py:_require_knowledge_capabilities docstring
claimed cascade.registry.build_handlers still gates
knowledge_topic/knowledge_document, contradicting
registry.py:177-194 (gate removed there, moved to HTTP layer).
Rewritten to describe the current design accurately.
Hermetic env full pytest: 2037 passed / 7 deselected.
Explicitly deferred to followup:
- J1 lazy multimodal client (needs everalgo signature change)
- J2 tier definition (knowledge = Tier 3 whole)
- J5.1/3/4 broader presentation-split refactor
- J6.1/2 backfill dispatch and _backfill_table refactor
- J8 Phase 1 keyset pagination for bulk migration OOM
- M4 flock timeout/waiting log/re-entry (core/persistence refactor)
- M6 UTF-8 codec strategy (BOM/UTF-16/GBK)
- M9 count_by_owner monotonicity (latent, INTERVAL=1 short-circuits)
- M13 --phase all multi-phase combined-outcome test coverage
- M15 PR-marker rationale comments (16 occurrences, all inert)
M12 was refuted (both event names exist).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>