Commit Graph

9 Commits

Author SHA1 Message Date
Kendrick-Song 9d48544280
fix(memory): rescue skill extraction, disable foresight, tighten APIs (#393)
* feat(ome): exponential backoff + jitter between retry attempts

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(events): carry case body + vector on agent skill chain

* feat(strategies): populate extended agent-skill chain events

* feat(md): AgentSkillReader.list_by_cluster for md-first skill enum

* fix(strategies): rescue extract_agent_skill from cascade-lag dead-letter

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(agentic): honor radius, use kind-shaped rerank, non-empty passage

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(api/ome): expose dispatched + runs, distinguish not_dispatched

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* chore(release): v1.2.3 — agent-skill rescue + OME/agentic contracts

* fix(agentic): drop inert radius plumbing; correct release docs

* fix(md): sanitize LLM-generated skill names against path traversal

AgentSkillFrontmatter.name comes straight from LLM output
(memory.strategies.extract_agent_skill) and was concatenated
unsanitized into the skills/skill_<name>/ directory segment on both
the write path (agent_skill_writer._skill_dir) and the read path
(agent_skill_reader._skill_dir). Given a sufficiently long ../ prefix,
the write target could escape the memory root (CWE-22). A live run
also produced a skill name containing spaces and CJK characters,
proving the "keep snake_case" docstring convention on
AgentSkillFrontmatter.name is not enforced at runtime.

This is the same class of defect already fixed for knowledge-upload
titles/categories. Promote that fix's sanitizer
(knowledge_writer._sanitize_dirname) to a shared primitive,
everos.core.persistence.markdown.sanitize_dirname, so there is one
CWE-22 defense for md directory names instead of two independently
maintained copies:

- New core/persistence/markdown/path_safety.py holds sanitize_dirname
  (idempotent: sanitize(sanitize(x)) == sanitize(x)), exported through
  the markdown + persistence facades.
- SkillPathMixin gains skill_dir_name(), the single sanitization point
  both AgentSkillWriter._skill_dir and AgentSkillReader._skill_dir now
  derive from, replacing their previous independent string
  concatenation.
- KnowledgeWriter now imports the shared sanitize_dirname instead of
  keeping its own private copy.
- AgentSkillFrontmatter.name gains a field_validator rejecting path
  separators / ".." as defence in depth, so a hand-edited SKILL.md is
  caught on parse rather than silently relocating the skill on the
  next write.

Idempotency is what keeps the reader and writer in agreement even
though they recover a skill_name from different sources:
list_by_cluster derives it from the on-disk (already-sanitized)
directory name, while extract_agent_skill._hydrate_algo_skills
re-reads using the frontmatter's raw name field. A regression test
(test_agent_skill_reader.py) seeds a skill whose frontmatter name
contains CJK + a space and asserts both routes resolve to the same
file.

No data migration: agent-skill extraction has never once succeeded
before this branch (the cascade-lag defect this branch fixes meant
.skills/ was never created), so there is no legacy skill corpus whose
directory names would change under the new sanitizer.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(e2e): make the agent-skill chain assertion real

tests/conftest.py's autouse fixture pins embedding + rerank capability
to unavailable for hermeticity, and this test never opted back in.
trigger_skill_clustering and extract_agent_skill both body-guard on
get_embedding_capability().available and return early, so the module
docstring's "real embedder" claim was false and the skill chain never
ran: measured log counts were skill_cluster_updated=0,
agent_skills_extracted=0, strategy_gated_off_embedding_unavailable=10.
The three skill assertions were assert len(...) >= 0 — always true —
with a comment blaming "LLM-dependent" flakiness for a count that was
in fact deterministically zero. This is why a defect that made
agent-skill extraction fail 4/4 in production reached a release: there
was no working e2e coverage of the chain.

- New _opt_in_real_embedding_and_rerank autouse fixture, scoped to this
  file only, resets everos.component.embedding.accessor._capability and
  everos.component.rerank.accessor._capability to None (the mechanism
  the global fixture's own docstring prescribes) so both capabilities
  rebuild from the real .env credentials tests/e2e/conftest.py already
  loads. Restores to None on teardown; every other test keeps its
  hermetic default.
- Replaced the three vacuous per-agent assertions with one aggregate
  floor across all three agents (>= 1 total skill). A per-agent floor
  would be flaky: extract_agent_skill has no cluster-size gate, only
  everalgo's per-case skip_quality_threshold, so a single low-quality
  trajectory can legitimately yield 0 skills for one agent.
- Added a sharper, defect-specific check: assert no dead-lettered
  extract_agent_skill run in OME's run_record (via
  OfflineEngine.list_runs), since a dead-letter (retries exhausted)
  is unambiguously a failure, unlike a quality-gated 0-skill outcome.
- Corrected the module docstring's "real embedder" claim and the old
  "# 4.5" comment's reasoning: extract_agent_skill has no cluster-size
  gate, only everalgo's per-case quality threshold.

Unexecuted: this test is slow + live_llm and this machine has no
provider credentials (the verification .env was deleted), so make ci
does not run it and it could not be run here. Verified by inspection
instead: ran the test file with -m "" to override the marker
deselection and confirmed it proceeds past the new fixture and
through app lifespan startup without error, failing only at the
expected point — LLMNotConfiguredError from the missing API key —
which confirms the fixture and imports are wired correctly and the
only blocker is the missing credentials, not a bug in this change.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(strategies): sanitize skill name before frontmatter construction

Follow-up to afe1609 (path-traversal sanitization). That commit added a
field_validator on AgentSkillFrontmatter.name rejecting path separators
or "..", intended as a read-side defence for hand-edited SKILL.md
files. But _persist_skill (memory/strategies/extract_agent_skill.py)
constructs AgentSkillFrontmatter(name=skill.name) with skill.name
straight from LLM output — the raw, unsanitized string — so the
validator actually fires on the write path too. Sanitization already
makes the on-disk path safe (SkillPathMixin.skill_dir_name), so an
LLM emitting a traversal-shaped name (reachable via prompt injection,
since the LLM's input is user conversation content) gained nothing
from the validator except a new failure mode: ValidationError ->
strategy raises -> OME retries with backoff -> dead-letter -> that
case's skill is permanently lost. A DoS vector introduced by a
security fix, and it also contradicted the validator's own docstring
("catches a hand-edited SKILL.md").

Also fixes a latent second bug in the validator itself, caught by the
new tests below: it rejected any name containing the substring "..",
but sanitize_dirname keeps "." as a safe character, so
"../" * 8 + "tmp/pwned" sanitizes to "................tmppwned" —
still containing ".." many times over. The validator would have
rejected the sanitizer's own safe output. Narrowed the check to actual
path separators or the name being exactly ".." (the only case where
".." functions as a real traversal component when there's no separator
left to combine it with).

Fix:

- SkillPathMixin gains sanitize_skill_name(skill_name) — the bare
  sanitized name (no skill_ prefix), factored out of skill_dir_name so
  both share one sanitizer call.
- _persist_skill now sanitizes skill.name via sanitize_skill_name
  once, up front, and uses that same sanitized string for
  AgentSkillFrontmatter.id, .name, and the writer.write_main() call.
  A traversal-shaped LLM name is now made filesystem-safe before it
  ever reaches the frontmatter constructor, instead of tripping the
  validator.
- The validator's docstring now describes actual behaviour: the write
  path pre-sanitizes, so the validator only fires for a name that
  bypassed the writer (e.g. a hand-edited file, or any other direct
  AgentSkillFrontmatter construction that skips pre-sanitization).

Bonus: with the write path pre-sanitizing, frontmatter.name becomes
byte-identical to the directory-derived name for LLM-written skills —
an identity, not merely an idempotency argument. This also closes the
gap in the previous commit's reader/writer test, which proved
idempotency generically but never drove an adversarial name through
the actual production write path end-to-end.

Tests:
- test_agent_skill.py: constructing AgentSkillFrontmatter with a
  pre-sanitized adversarial name (mirroring _persist_skill's own call
  shape) succeeds and yields a separator-free name; the read-side
  rejection test for bypassed/hand-edited names is unchanged and still
  passes with the narrowed check.
- test_agent_skill_writer.py: new parametrized identity test — for
  both an adversarial and a CJK/space raw name, sanitize once, write
  via that sanitized name, and assert frontmatter.name equals the
  directory-derived name exactly.
- test_agent_skill_reader.py: docstring updated to clarify its
  existing round-trip test now covers the bypass case (a caller that
  writes via a raw, unsanitized name directly through the writer,
  skipping _persist_skill's pre-sanitization) rather than the normal
  production path, which is proven as an identity by the writer test
  above.
- Existing test_extract_agent_skill.py strategy tests (snake_case
  fixture names) are unaffected — sanitize_dirname is the identity
  function for already-safe names.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(md): reject degenerate sanitizer results, read globbed paths

Security review of 0c6820f found it did not close the DoS it was meant
to close: sanitize_dirname("../") returns ".." verbatim, because "."
is a safe character and is not stripped by the character-class filter
— only the leading "/" is removed. sanitize_skill_name("../") ->
sanitize_dirname("../") -> "..", and AgentSkillFrontmatter(name="..")
still raises ValidationError (name == ".." is exactly the case the
narrowed validator rejects). Same dead-letter DoS as 0c6820f, just a
shorter payload; the previous commit's tests only exercised the long
"../" * 8 + "tmp/pwned" payload, which happens to sanitize past the
fixpoint. The same fixpoint is a real one-level directory escape on
the knowledge path, which has no skill_ prefix to protect it:
Path(root) / sanitize_dirname("../", "Others") / "doc_123" resolved to
root/doc_123, skipping the category directory entirely.

Fix (path_safety.py): sanitize_dirname now falls back on "", ".", or
".." instead of only "". This one change closes both the skill
dead-letter DoS and the knowledge one-level escape, since both callers
already route through this single primitive. Also NFC-normalizes the
input before the character filter (unicodedata.normalize("NFC", raw)),
so an NFD-decomposed accented character (base letter + combining
mark, which is not \w) no longer silently loses its accent. Rewrote
the docstring, which previously claimed "..  sequences are always
stripped" (false — "." is explicitly a safe character) and "cannot
escape the directory it is concatenated into" (false for an unprefixed
caller before this fix); it now states what actually holds: no
separator survives, so the result is always exactly one path
component, and it is never "", ".", or "..".

Also fixes (per review, cheap and worth doing alongside):

- AgentSkillReader.list_by_cluster previously globbed skill_*/SKILL.md,
  stripped the prefix to recover a name, then called read_main(name),
  which re-derives (and re-sanitizes) the path from that name. Any
  on-disk directory whose suffix was not already a sanitizer fixpoint
  (e.g. "skill_My Skill", a raw space) re-derived to a path that
  doesn't exist and was silently dropped. Since list_by_cluster is the
  documented strong-consistency existence check, a dropped skill would
  make the LLM emit add() for a skill that already exists, duplicating
  it at the sanitized path and orphaning the original. Fixed by having
  list_by_cluster read each globbed path directly (new _read_path
  helper, shared with read_main) instead of round-tripping through a
  recovered name — the reader never derives a path at all on this
  route, which is a stronger guarantee than the idempotency argument
  the docstrings previously leaned on.
- e2e test: made fixture ordering explicit — the embedding opt-in
  fixture now takes _reset_embedding_capability_singleton and
  _reset_rerank_capability_singleton as parameters so pytest's
  dependency graph guarantees correct ordering, rather than relying on
  collection order between conftest files. Added a positive
  "extract_agent_skill actually ran" assertion (any status) before the
  dead-letter check — without it, the dead-letter assertion alone is
  vacuously satisfied by a strategy that never executed at all; it was
  only meaningful before because the skill-count floor happened to run
  first. Dropped the rerank capability opt-in and the module
  docstring's "real reranker credentials" claim: nothing on the
  agent-skill write path touches rerank, so opting it in only widened
  the credential surface with no coverage benefit.
- CHANGELOG: corrected the validator description (rejects a path
  separator or being exactly "..", not any string containing ".."),
  and added the previously-missing user-visible fact that
  AgentSkillFrontmatter.name and the agent_skill LanceDB primary key
  now hold the sanitized name, not the raw LLM output.

Tests: parametrized the sanitize -> construct -> (write, for the
writer-level test) tests over a boundary family instead of one long
payload: "..", "../", "/../", ".", "./", "!!!" (empty), "a" * 200
(truncation), a CJK+space name, and the original "../" * 8 +
"tmp/pwned". Each case asserts the sanitized name is a single
component, is never "" / "." / "..", frontmatter construction
succeeds, and (writer-level) frontmatter.name is byte-identical to the
directory-derived suffix. New test_path_safety.py cases pin the
degenerate-fixpoint fallback directly, the knowledge-style unprefixed
one-level-escape repro, and NFC normalization. New
test_list_by_cluster_finds_skill_whose_directory_suffix_has_a_space
reproduces the exact list_by_cluster drop bug against a directory
written outside the writer entirely.

Explicitly not in scope (per review): the collision behaviour where
"fix django" and "fix_django" now map to the same directory is a real
product-decision question the reviewer is raising separately, not
touched here.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(md): record the skill-name collision trade-off

No behaviour change. Documents a product decision the coordinator made
explicit: sanitize_dirname is lossy, so distinct raw skill names can
collapse onto the same directory ("fix django" and "fix_django" both
become "fix_django"; "fix!django" and "fixdjango" both become
"fixdjango"; names differing only past the 50-char cap also collide).
Because AgentSkillWriter.write_main is a full-file replace and the
LanceDB primary key is f"{agent_id}_{sanitized_name}", a collision
means the later skill silently overwrites the earlier one, losing its
accumulated source_case_ids, maturity_score, and body.

This is accepted rather than mitigated: the LLM's add/update decision
for a skill is keyed on the name it sees, so a collision usually reads
as an intended update anyway; and adding a disambiguating suffix would
break the frontmatter.name == directory-suffix identity the
reader/writer seam (from 0c6820f) relies on.

- SkillPathMixin.sanitize_skill_name docstring now states the
  collision consequence and the two reasons it is accepted, so a
  reader does not have to derive them.
- sanitize_dirname's docstring gains one line: the function is lossy
  and not injective; callers that need distinct outputs for distinct
  inputs must disambiguate themselves. General primitive — the
  knowledge path calls it too.
- CHANGELOG: added the collision consequence to the existing
  path-traversal entry, next to the already-documented fact that name
  / the LanceDB key hold the sanitized value.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(md): return skill bodies from list_by_cluster; correct docs

Security review of 1b8cf11 + add842d found the list_by_cluster fix was
incomplete: it stopped its own enumeration from dropping a skill whose
directory suffix wasn't a sanitizer fixpoint, but the caller
(_hydrate_algo_skills) still re-read each selected skill by
fm.name via read_main, which re-derives (and re-sanitizes) a path from
that name and drops it there instead. Reproduced on a real
filesystem: skill_My Skill/ enumerates fine, but
read_main("My Skill") re-derives to skill_My_Skill/ and misses. The
drop moved one layer downstream; existing_relevant_skills was empty
before and after the prior fix.

Fix: list_by_cluster now returns (frontmatter, body) pairs instead of
frontmatter alone, so the caller never needs a second, name-based
read. _select_existing_skills / _rank_skills_by_relevance updated to
carry (fm, body) tuples through selection; _hydrate_algo_skills is
deleted — the body is already in hand, so there's nothing left for it
to do. This closes the drop for real, removes the second disk read
(the 2n-read concern carried since Task 4), and makes "the reader
never derives a path" true end-to-end rather than true only for
list_by_cluster's own enumeration step.

New end-to-end regression test
(test_select_existing_skills... / test_existing_skills_reaches_llm_for_skill_whose_directory_has_a_space)
seeds a skill_My Skill/ directory directly on disk (bypassing the
writer) and runs the real extract_agent_skill strategy against it,
asserting the skill reaches existing_relevant_skills with non-empty
content — the property the previous commit's test docstring claimed
but the code didn't yet deliver. The reader-level regression test
gained the same body assertion.

Also, per review:

- path_safety.py: corrected the NFC docstring claim, which was wrong
  for the ~1,082 Unicode composition-exclusion codepoints (e.g.
  Devanagari क़/ख़, U+0958/U+0959) — NFC decomposes an
  already-precomposed exclusion character instead of preserving it, so
  the combining mark is stripped either way. Scoped the claim to
  "best-effort for the common case", not a guarantee for every script.
  New test pins this directly.
- Dropped the e2e test's unused _reset_rerank_capability_singleton
  fixture parameter: the reviewer adjudicated the earlier instruction
  conflict the other way — ordering is only meaningful between
  fixtures that touch the same state, and this fixture never reads or
  writes the rerank capability at all.
- Widened the skill-name collision documentation (SkillPathMixin.sanitize_skill_name,
  CHANGELOG) beyond dropped-punctuation / space-collapse / truncation
  to the larger case: every combining mark is non-\w and is stripped
  regardless of script, so e.g. Devanagari "किताब" and "कताब" both
  collapse to "कतब" (same for Thai tone marks, Hebrew niqqud, Arabic
  harakat).
- Corrected the collision justification: because _persist_skill
  sanitizes before frontmatter construction, the LLM sees the
  already-sanitized name in existing_relevant_skills, so a colliding
  raw name is an *affirmative* decision that two skills are different,
  not a probable intended update. The decision to accept collisions
  still stands, but on its real grounds: a disambiguating suffix would
  break the fm.name == directory-suffix identity the reader/writer
  seam relies on, and detecting-and-raising would reintroduce the
  dead-letter DoS.
- Merged two consecutive "# -- Internals --" banners in
  agent_skill_reader.py into one.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(md): skip unparseable SKILL.md instead of failing the cluster

`list_by_cluster` only handled a missing file, so a ValidationError from
`_read_path` aborted the whole enumeration. That enumeration is what
feeds `extract_agent_skill` its existing skills, so one bad SKILL.md
starved every skill in the cluster and dead-lettered that cluster's
extraction on every subsequent run -- the permanent-failure mode this
md-first read path was introduced to eliminate. The write path already
pre-sanitizes to avoid exactly this; the read path was left open.

The trigger surface is the whole schema, not just the traversal
validator this PR added: `_read_path` validates the full
`AgentSkillFrontmatter`, so a field a later revision makes required
would take out every existing file at once. Reproduced both ways.

`read_main` still propagates -- a caller naming one specific skill needs
an error, since `None` already means "not created yet" and reusing it
for "exists but is corrupt" would let an upsert overwrite the damage.

Also moves this route's `OfflineEngine` import under TYPE_CHECKING. It
is used only to annotate `_summarize_runs`, and an eager import
contradicted the deferred `_get_engine` import a few lines below. Note
this saves nothing at startup today: `service.memorize` imports the
engine eagerly to construct it, so any app process pays the ~750ms
apscheduler cost regardless.

* docs(md): correct two docstrings, add the case-collision dimension

Three fixes to claims that did not match behavior:

- `_rank_skills_by_relevance` claimed no skill is silently dropped from
  the prompt. The backfill loop is capped at MAX_SKILLS_IN_PROMPT, and
  the function only runs when the cluster already exceeds that budget, so
  skills beyond K are dropped by design. Reworded to what the backfill
  actually guarantees: a lagging index cannot under-fill the prompt.

- `sanitize_skill_name` enumerated collision causes in detail but omitted
  case, the dimension an LLM varies most freely. "Fix Django" and "fix
  django" sanitize to two distinct names -- two LanceDB rows, but one
  directory on a case-insensitive filesystem (macOS APFS, Windows NTFS
  defaults), so the index advertises a name whose content was overwritten.

- The same docstring justified accepting collisions partly on a
  disambiguating suffix breaking the `frontmatter.name` = directory-suffix
  identity. It would not: writing "fix_django_2" into both keeps that
  intact. Replaced with the real reason it is deferred rather than
  dismissed -- it needs a collision probe and a case-folding rule.

Adds the knowledge-writer sanitization tests that were missing entirely:
swapping in the shared primitive changed NFD input ("Résumé" no longer
degrades to "Resume") and made a "." / ".." topic fall back. Knowledge
upload predates this PR, so unlike skills it has a corpus whose
directory names those first cases affect. Both tests verified red against
the 1.2.2 sanitizer.

* docs(changelog): scope the migration claim, record run_record growth

Three corrections to the 1.2.3 entry:

- "No data migration" was asserted for the whole sanitization change but
  only holds for agent skills, which have no corpus because extraction
  never succeeded. Knowledge upload predates this release and does have
  one: NFD topics and `.`/`..` topics resolve to a different directory
  now. Scoped the claim and spelled out both cases.

- Added the case dimension to the collision list, and replaced the
  "disambiguating suffix breaks the name = directory identity" reason
  with the accurate one -- it does not break it, it just needs a probe
  and a case-folding rule, so it is deferred rather than rejected.

- Recorded that `SkillClusterUpdated` now persists a 1024-dim vector in
  `run_record.event_payload`: ~0.8 KB to ~14 KB per record, ~14 MB per
  strategy at the default 1000-record ring buffer. Operators sizing
  ome.db need this number, and it was not stated anywhere.

* fix(strategies): ship extract_foresight disabled by default

The sender scan reads `m.role` off every memcell item, but only
ChatMessage carries it: ToolCallRequest has `sender_id` and no `role`,
ToolCallResult has neither. So any memcell holding a tool call raises
AttributeError before the first sender resolves -- correct on plain user
chat, guaranteed to fail on agent trajectories, where it burns its
max_retries budget and dead-letters on output nothing consumes today.

Flipped the decorator rather than `default_ome.toml`, because `everos
init` skips an existing `~/.everos/ome.toml` (init_cmd.py:85), so a
template edit would reach new installs only. The toml opt-in is left
working on purpose -- a chat-only deployment does get correct
foresights -- and documented in both the module docstring and the
template comment.

This is a stop-gap. The fix is per-episode extraction, like
atomic_fact, which needs an everalgo entry point that does not exist
yet.

The one test that used foresight as its UserPipelineStarted subscriber
now opts back in through that same toml key, so the opt-in path is
covered rather than worked around. It has to wait for the override to
reach the registry first: ConfigReloader.start() fires its initial load
as a task, so engine.start() returns before ome.toml is applied, and an
emit inside that window is judged against the coded defaults and
dropped by the enabled gate with no redelivery.

* fix(ome): hold engine_sem per attempt, not across the retry chain

The backoff this PR added slept inside the semaphore block, so a run
waiting to retry kept its concurrency slot. That turns a partial outage
into a total stall: with max_concurrent_runs slots and a 1s/2s/4s
backoff, enough simultaneously-failing runs park every slot in
asyncio.sleep and starve strategies that would have succeeded.

The cap exists to bound concurrent strategy work -- LLM calls,
embeddings, storage IO -- and a sleeping coroutine consumes none of it.
Backpressure on the failing work is intended; backpressure on everything
else is not. Semantics change is deliberate and stated in the docstring:
the cap still applies to execution, no longer to waiting.

The guard uses a single-permit semaphore so locked() is unambiguous, and
asserts a second waiter actually acquires -- locked() alone would pass on
an implementation that freed the slot but left waiters unable to take it.
Verified red against the previous structure.

* fix(strategies): reap the directory a renamed skill leaves behind

everalgo treats a name change as a first-class update: _apply_update
preserves prior.id while swapping the name, so _persist_skill wrote the
skill to a new skill_<new_name>/ and the old directory survived with the
same cluster_id.

That is not a cosmetic leak now. Since existing skills are read from
markdown rather than LanceDB, the orphan returns in the next run's
existing_relevant_skills as a duplicate of a skill the LLM already
renamed -- feeding exactly the add-instead-of-update full-replace clobber
this PR set out to close, once more per rename. Reconciliation keys off
skill.id, the only field that survives a rename (a fresh add mints a
uuid4 and can never match), and never deletes a name another emitted
skill just claimed.

Also in this pass:

- AgentSkillWriter.delete_skill, the one destructive operation here. It
  fails closed: a directory it cannot resolve by the writer's own path
  rule is left alone rather than targeted by anything looser.
- reference_name / script_filename now sanitized on both reader and
  writer. They are appended after the skill_<name> segment, so
  skill_dir_name never covered them. Zero callers in src/ today; closing
  it before progressive disclosure wires them up.
- Agentic case rows with every passage field empty fall back to a
  placeholder instead of raising ValueError in everalgo's _format_docs
  and 500ing a whole search the row merely appears in.
- The retire op is documented as unimplemented rather than left implied.
  aextract returns a flat list with no discriminator, so a retirement
  arrives as an ordinary low-confidence skill and is written back like
  any other. Honouring it means either giving an LLM confidence score
  authority to delete the source of truth, or a retired flag that the
  enumeration, cascade, and search all learn to filter on -- a design
  decision, deferred.
- The embedding body-guard comment no longer claims to protect a local
  embed call; this strategy stopped embedding when the vector moved onto
  the event.

* fix(strategies): stop extract_foresight crashing on tool-call memcells

The sender scan read m.role off every memcell item, but only ChatMessage
carries it: ToolCallRequest has sender_id without it, ToolCallResult has
neither. The first tool call raised AttributeError before any sender was
resolved, so the strategy was correct on plain user chat and
dead-lettered every time on agent trajectories.

everalgo contracts for exactly this input -- user_memory/_render
.chat_messages says "the caller need not pre-filter; an
AgentMemCell-shaped MemCell is acceptable input" -- and every other
user-memory extractor gets that for free by delegating. This strategy was
the one place the filter was hand-rolled, and it was hand-rolled wrong.

It stays disabled by default, but for the correct reason: nothing in
EverOS reads foresights yet, so running it spends one LLM call per sender
per memcell on write-only data. The earlier justification (per-episode
extraction needs an everalgo entry point that does not exist) confused
extraction granularity with the crash; granularity is still open, the
crash was one line. Fixing it is what makes the documented ome.toml
opt-in actually usable.

The guard pins both directions: a pure agent trajectory extracts nothing
and never reaches the LLM, and a mixed memcell extracts for human senders
only -- an implementation that stopped raising but scanned tool-call
sender_id values would invent "agent" as a user.

CHANGELOG also records that the stale-index clobber is fully closed only
for clusters at or below MAX_SKILLS_IN_PROMPT; above it LanceDB orders
the markdown candidates, and the skill a lagging index omits is the one
written most recently.

* docs(changelog): fold the merged cascade work into the 1.2.3 entry

#392 landed on main with its entries under [Unreleased] and no version
bump. Since 1.2.3 ships that code, leaving them there would have the
release notes disclaim work the release contains. Merged section by
section into [1.2.3] and dated it to the actual release day.

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-08-07 13:07:42 +08:00
zhanghui e5118c52a8
fix(lancedb): reclaim stale versions via write-locked prune (#379)
* fix(lancedb): reclaim stale versions via write-locked prune

The storage soak (48h, sustained churn + fuzz) proved the bundled
lock-free `optimize(cleanup_older_than=...)` loses its commit-conflict
race against concurrent writes — cleanup ran only 16 of ~250 times, so
old dataset versions / FTS orphans piled up and the index dir grew to
the 40G disk guardrail and never reclaimed under load. main still had
that bundled call.

Split the maintenance path:
- `LanceRepoBase.optimize()` is now compact-only and lock-free (a commit
  conflict here is benign — the next beat retries, so it must not stall
  writers).
- `LanceRepoBase.prune(older_than)` runs `cleanup_older_than +
  delete_unverified=True` **under the per-table write lock**, so no
  writer is in flight: the Rewrite has the manifest to itself (cleanup
  completes every beat) and aggressive deletion is safe. It also removes
  the empty `_indices/<uuid>/` husks cleanup leaves behind (soak: 13061
  dirs, 98% empty), offloaded to a thread.
- The cascade worker's heavy beat calls `prune()`; the light beat calls
  `optimize()`. A benign light-beat commit conflict is logged at debug
  and does not count toward the failure streak or trigger a rebuild.
- Prune's retention window (`cleanup_older_than`) is decoupled from the
  prune cadence and defaulted short (60s). It runs under the write lock,
  so the window only needs to outlive an in-flight read; keeping it =
  cadence (300s) left ~2 cadences of superseded full-table copies on
  disk between beats (soak: transient ~15G/table peaks). 60s reclaims
  all but the last minute each beat — same live floor (~625MB/table),
  far lower transient footprint.

Result on the re-run soak: disk sawtooths and reclaims to live-data size
(~1.3G total) under active load — vs run1 stuck at 40G until writes
stopped — with 0 crashes / 0 OOM / 0 stuck cleanups over 48h.

Cascade projection health is now observable:
- `CascadeOrchestrator.health()` -> `CascadeHealth`, combining the
  worker's in-memory signals (drain-loop failures, unrecoverable count,
  optimize streak, prune staleness) with the SQLite queue summary.
- `GET /health` gains a typed `cascade` readiness block. `healthy`
  reflects operational health only (drain / optimize / prune);
  `failed_permanent` (files awaiting `cascade fix`) is a data-quality
  backlog reported as an informational count that does NOT flip
  `healthy` — otherwise the signal would sit red permanently.

The scanner-side retry cap and the `_MAX_TOTAL_RETRIES` budget already
on main handle re-enqueue storms, so no duplicate is added here.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(deps): pin lancedb to >=0.34.0,<0.35.0

The previous open-ended `>=0.13.0` let any environment float to an
untested release, including 0.32-0.34 which carry a compaction
offset-overflow regression (lance-format/lance#7653) that stalls
version cleanup and grows the index dir without bound.

- Floor 0.34.0: the current resolved version; runs safely thanks to
  the with_position=False FTS workaround shipped in #336. Verified that
  data written by lancedb 0.32.0 (lance v6) reads correctly under
  0.34.0 (lance v8), so existing deployments upgrade cleanly. Never
  widen the floor below 0.34 -- older lance cannot read v8-format data.
- Ceiling <0.35.0: 0.35 embeds lance-rust v9 (large encoding jump, not
  yet stable-released); it must pass the soak harness before we allow
  it.

Resolved version is unchanged (still 0.34.0); this only tightens the
declared constraint.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
(cherry picked from commit 94f9aa67d11a07c93c7d59d6b134e80858b60316)

* fix(search): bridge agent-kind metadata into the agentic doc contract

The agent AGENTIC path (`agent_case` / `agent_skill`) fed recall
candidates straight into `aagentic_retrieve`, whose `_format_docs`
(LLM sufficiency / multi-query prompt) reads `metadata["episode"]` as a
`{subject, content}` dict plus a ms-epoch `timestamp`. Agent-kind rows
carry their body in the recaller's `text_field` and time as a datetime,
so `_format_docs` raised `TypeError: Candidate ... has no episode dict`
and `POST /api/v*/memory/search` returned 500 for any
`owner_type=agent` + `method=agentic` request.

Mirror the episode path's bridge: reshape agent candidate metadata into
the everalgo doc contract before `aagentic_retrieve`, and revert it
before DTO shaping so the agent shapers still see a datetime timestamp.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
(cherry picked from commit bb9a3a585e044643d3852b9d2810b42c05a7a47c)

* test(search): regenerate search seed to current linkage + migrate e2e

The committed `search_seed` fixture and several search e2e tests were
written for the pre-1.5 memcell fact-linkage model. Current extraction
links atomic_facts to episodes via `parent_id == episode.entry_id`
(parent_type="episode"), and user_memory clusters store episode
entry_id members — so the stale fixture made VECTOR/AGENTIC recall and
the cluster-narrowing path find nothing, and stale assertions checked
an old error code.

- Regenerate `search_seed/*` from a fresh corpus in the current
  entry_id format; facts now bridge across multiple episodes (richer
  agentic / hierarchical-eviction coverage).
- Fix `_dump_search_seed.py` sampling: pick episodes that host facts
  first and keep facts by episode entry_id, so re-dumps stay coherent.
- Migrate e2e tests to the entry_id model (hierarchical-eviction,
  session/timestamp filters, cluster seeding helper) and update the
  filter-error assertion to the current `INVALID_INPUT` code.
- Provision `ome.toml` in the full-app pipeline fixture (the OME config
  reloader requires it; strategies are code-registered so the packaged
  default suffices), unblocking corpus regeneration.

Full search e2e suite now green (49/49).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
(cherry picked from commit 44cd9cecc07775cb9270104a5804f36e32cd788d)

* fix(lancedb): detect schema type drift and add cascade rebuild recovery

verify_business_schemas only compared column names, so a column whose
on-disk Arrow type had drifted (name unchanged) slipped through and
detonated later inside merge_insert as an opaque LanceError(IO)
"Spill has sent an error" (#337). Now compare each shared column's Arrow
type against schema.to_arrow_schema() — the exact schema get_table
builds tables from, so a healthy table never false-positives.

Reproduced #337 byte-identically: an episode.subject_vector column left
as string or null by an older build, plus a real 1024-d vector on
upsert, yields the exact crash. No lancedb version (0.13-0.34) renders
Optional[Vector] as a non-vector type, so the startup guard is what
should catch it — not the runtime.

Add `everos cascade rebuild` as the safe recovery: it drops the business
LanceDB tables and re-indexes from markdown, skipping the verify guard
(which the drift would otherwise trip on startup). Unlike removing only
.index/lancedb it re-populates already-done entries (reset_all clears
the cascade queue); unlike removing all of .index it preserves
unprocessed_buffer (messages not yet extracted).

Fixes #337.

* docs(cascade): document cascade rebuild and correct recovery guidance

Add the `everos cascade rebuild` command to the runbook, CLI, and
how-memory-works docs. Correct the old recovery guidance: a bare
`rm -rf .index/lancedb` leaves md_change_state marked `done`, so the
scanner skips those files and the index comes back empty — the runbook
previously claimed a full repopulation that does not happen. `cascade
rebuild` is the safe path (re-populates done entries, preserves
unprocessed_buffer). Also document that verify now checks column types.

* chore(rebase): adapt #354 integration test to soft-embedding main

CascadeOrchestrator dropped the embedder param when embedding became a
soft dependency (main); fold the schema-drift integration test onto it.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(cascade): freeze monotonic clock in prune-staleness health tests

Both prune-staleness health tests fabricated
`_started_at = time.monotonic() - (ALERT + 100)`, assuming monotonic()
is a large value. On a fresh CI runner monotonic() is only ~100-180s, so
the subtraction went negative, the source clamped the baseline to 0, and
staleness read back as ~130s < 900s — failing on CI while passing on
long-lived dev boxes where monotonic() is huge.

Freeze the monotonic clock via monkeypatch so staleness is deterministic
regardless of the runner's boot uptime. Source logic is unchanged; only
the tests are made hermetic.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(cascade): wait for stable terminal state in rename scenarios

`_wait_path_done` reached a terminal status, slept a 0.1s settle window,
then asserted the status was still terminal — which contradicted its own
docstring ("absorb any last-second re-enqueue"). A rename's delete event
or an atomic-replace echo can flip a done row back to `processing` inside
that window, so on a slow CI runner the assert fired
("flipped back to processing after reaching done"), failing
test_rename_cross_owner_keeps_frontmatter_owner intermittently (seen on
the 3.13 job). `make integration` runs without `--reruns`, so a single
flake fails the whole job.

Wait for a terminal state that *survives* the settle window instead:
absorb a transient re-enqueue by waiting for terminal again, still bounded
by `deadline` so a row that never settles surfaces as a timeout. Pre-
existing flake on main, unrelated to the prune change; the scenario's real
assertions (row counts, frontmatter owner) are untouched.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(cascade): cross-process prune safety + backfill reclaim fix

Review of #379 found two P0s plus P1/P2s, all verified against the code:

- P0-1: cascade backfill still called the removed
  optimize(cleanup_older_than=…) kwarg → TypeError swallowed by the
  best-effort try/except → backfill silently skipped all compaction +
  reclaim (the exact bloat this PR fixes). CI stayed green because the test
  double kept the stale signature. Fix: call optimize() + prune(0) at the
  call site; make the fake mirror the real signature so the drift can't hide
  again; pin prune in the backfill tests.

- P0-2: prune ran delete_unverified=True guarded only by an in-process
  asyncio lock, but the runbook promises `cascade sync` is safe alongside a
  live server — and the CLI's first optimize beat does prune, in a separate
  process. It could delete files the daemon is mid-commit on. Fix: switch
  prune to delete_unverified=False. Measured to reclaim identically on
  churned tables (both collapse superseded versions ~97%); True only
  additionally deletes in-flight/dangling files — exactly the corruption
  vector. No cross-process lock needed; the write-lock/commit fix (the real
  reclaim win) is unchanged.

- P1-3: /health called orch.health() (6 SQLite aggregates) with no guard →
  a locked/full/migrating DB would 500 the liveness probe and restart the
  container. Wrap it: unhealthy readiness + reason, HTTP stays 200.

- P1-4: rebuild drops + recreates tables; a live daemon holds cached handles
  pointing at the dropped dataset. Runbook now says stop the server first —
  the one cascade command unsafe alongside a live server.

- P2: narrow _is_benign_commit_conflict to the "commit conflict" phrase (a
  bare "retryable" swallowed unrelated recoverable errors); add a timeout
  around the prune cleanup so a hung lance call can't wedge the write lock;
  correct two stale docstrings.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs: fix schema-recovery guidance + drop dead internal doc refs

Follow-up to the review; two doc issues neither the review nor the fix
commit caught:

- The schema-drift startup error's docstrings (verify_business_schemas
  and LanceDBLifespanProvider) still described the recovery as
  `rm -rf ~/.everos/.index/lancedb` — which the runbook explicitly calls
  the WRONG recovery (it leaves the cascade queue `done`, so nothing
  re-indexes and the index comes back empty). The raised error already
  points to `everos cascade rebuild`; align the docstrings to match.

- 15 dangling references to an internal numbered design-doc set
  (12_/13_/16_/17_*.md) that was never shipped to this repo. Point the
  schema-recovery ones at docs/cascade_runbook.md; drop the rest (pure
  provenance in table/component docstrings) while keeping the substance.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs: align maintenance docstrings with split optimize/prune API

Two docstring residuals from the #379 review's P2 list:

- _run_optimize_once still described the pre-split bundled heavy beat
  ("same work plus cleanup_older_than ... older than one cadence");
  the heavy beat now calls prune() under the write lock and the
  retention window is decoupled from the cadence
  (DEFAULT_OPTIMIZE_PRUNE_RETENTION_SECONDS).

- _restore_shaper_metadata converts any numeric timestamp, wider than
  an exact inverse of the bridge; document that this is deliberate
  (the shaper contract requires a datetime either way).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* style(e2e): hoist fixture-body imports to conftest module top

shutil / importlib.resources.files were imported inside the
core_pipeline_runtime fixture body; move them to the module top to
match the repo import style (#379 review P2).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(cascade): back off a hung prune via a separate attempt clock (N1)

The write-lock timeout on prune (review P2) didn't achieve its goal: it
was 300s — equal to the prune cadence — and `last_prune_at` only advanced
on success. So a hung lance cleanup timed out after 300s, `should_prune`
was still true (clock never moved), and the next beat re-pruned ~10s
later — pinning the per-table write lock ~97% of the time, the exact
write-starvation the timeout was meant to prevent.

A real cleanup is milliseconds even on a heavily churned table (measured
~40ms at 320k writes / 100 versions), so the timeout is a pure hang-catcher:
lower it to 60s (~1500x headroom, never fires normally, well below the 300s
cadence).

Split the prune clock so a failed prune backs off without masking the
health signal:
- last_prune_attempt_at (new) gates scheduling, advanced before the call
  whether it succeeds or times out → a hung prune waits a full cadence
  before retrying (light lock-free compaction runs meanwhile), so the lock
  is held at most ~timeout/cadence ≈ 17% in the worst case.
- last_prune_at advances only on success and still drives the
  prune-staleness health signal, so a persistently failing prune surfaces
  as degraded instead of being hidden by an advanced schedule clock.

Regression test: a raising prune advances the attempt clock (next beat is
light, no immediate re-prune) but not the success clock.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: zhanghui <zhanghui@shanda.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-08-03 15:41:13 +08:00
Kendrick-Song b1441da607
refactor(config): make [embedding] and [rerank] soft dependencies (#361)
* refactor(config): make [embedding] and [rerank] soft dependencies

Make [embedding] and [rerank] soft runtime dependencies so a
freshly-onboarded user can run EverOS end-to-end with only [llm]
configured. Previously the server refused to start without embedding,
locking out anyone who just wanted keyword-only search.

## Capability tiers

- Tier 1 ([llm] only)                   KEYWORD search, add/flush, md
                                        writes, cascade sync
- Tier 2 ([llm] + [embedding])          + VECTOR / HYBRID search,
                                        reflection, skill extraction,
                                        backfill
- Tier 3 ([llm] + [embedding] + rerank) + AGENTIC search, knowledge

Tier upgrades require a server restart (capability accessors cache
for the process lifetime). Tier downgrades are read-safe: a Tier-3
user who drops [rerank] can still read/rename/delete existing
knowledge documents; only write/search endpoints return 422.

## What changed

- Component accessors — component/{embedding,rerank,llm}/accessor.py
  are the single process-wide provider singletons. service/* never
  maintains parallel singletons; it consumes get_embedding_capability()
  / get_rerank_capability() / get_llm_client() directly. Build-time
  ValueError from the factory is logged as capability_build_failed
  (was silently swallowed).
- Error mapping — ProviderNotConfiguredError -> 422 with everos.toml
  section hints (never EVEROS_* env-var strings).
  LanceDBMigrationError fails loud with escalating recovery guidance
  (restart -> wipe index). LLMNotConfiguredError in search maps to
  None for KEYWORD degradation.
- Nullable-vector LanceDB migration — schema v2 makes the vector
  column nullable so Tier-1 rows can land without embeddings.
  Migration is guarded by a cross-process memory_root_lock
  (fcntl.flock + anyio.to_thread) and runs optimize() per table
  after Phase-1 backfill to reclaim manifest bloat.
- Cascade — knowledge handlers register unconditionally (Tier-3 ->
  Tier-2/1 downgrade no longer strands DELETE); embed-requiring
  strategies use body-guards that check capability.available at
  execution time. _TABLE_SPECS has an import-time drift assertion
  against BUSINESS_SCHEMAS_WITH_VECTOR.
- `everos cascade backfill` CLI — Phase-1 (embed missing vectors) /
  Phase-2 (emit synthetic events for cascaded processing) / Phase-3
  (sync new skill files). Exit codes: 0 / 1 / 2 / 3 (server running
  preflight) / 4 (COMPLETED_WITH_FAILURES — per-row failures rolled
  up) / 130 (SIGINT). OMEConfig.crash_recovery_enabled=False in
  backfill engines prevents stale-RUNNING rows re-enqueuing into a
  smaller strategy registry.
- /health — reports capabilities + disabled_features per tier so ops
  can distinguish "boots but degraded" from "boots and full".
- Presentation split — memory / service / infra never import typer /
  click. TyperPresenter Protocol + run_backfill() live in
  entrypoints/cli/commands/_backfill_cmd.py. Enforced by
  import-linter.
- Startup hint — unconditional count_rows(filter="vector IS NULL")
  sweep emits unbackfilled_memory_rows (event name + hint text
  pinned) when Tier-1 rows exist. ParserLifespanProvider warms the
  everalgo.parser import at boot so /health doesn't block on first
  call.
- Knowledge upload UTF-8 short-circuit — _looks_like_utf8_text()
  routes text/* mime and known plaintext extensions (md/txt/rst)
  straight to UTF-8 decode instead of the parser. Prevents 503
  Multimodal-not-configured when Tier 3 sans [multimodal] uploads a
  markdown doc.

## Sync history with main (2 merges collapsed into this squash)

Merged origin/main at 6dcd3eb (v1.1.4 -> v1.2.0 adds OTel tracing,
/api/v2 alias, TracingLifespanProvider, per-cascade-embedding span
fix, memory-op instrumentation) and later at 42629df (PR #366
backfills v1.1.4 CWE-22 knowledge path traversal fix + cascade
retry-budget rework + errors.py -> core.errors.ExternalServiceError).
Key merge decisions:
- service/search.py adopts single wrap site — component.llm accessor
  already applies UsageRecordingClient when observability is on;
  service layer never keeps a parallel LLM singleton (Round-1 CR
  rule: "service layer never maintains parallel singletons").
- Knowledge router prefix moved to /knowledge; create_app() mounts
  it under both /api/v1 and /api/v2.
- Cascade retry classification uses ExternalServiceError from
  core.errors (cascade/errors.py deleted). _MAX_TOTAL_RETRIES=12
  cross-cycle budget preserved.
- Fixed backport typo: MemoryRoot.default() -> MemoryRoot.resolve()
  (no .default() classmethod exists — main PR #366 shipped a broken
  call).

## Verified layering

    $ git grep -l "^import typer\|^from typer" src/everos/{memory,service,infra}
    # empty
    $ git grep -l "^import click\|^from click" src/everos/{memory,service,infra}
    # empty

Memory / service / infra layers clean of CLI presentation libraries.

## Review history

Three rounds of Fable 5 (opus) code review across the pre-squash
commit history closed 38 findings total:
- Round 1: 10 findings (fail-loud migration, backfill hardening,
  knowledge router gate scoping, SearchManager guards, profile
  throttle lift)
- Round 2: 13 findings (hermetic test env, hot-reload doc drift,
  knowledge handler registration, Phase 3 sync guarantee, Phase 2
  idempotency, profile event-first path, OMEConfig crash-recovery
  gate, cross-process migration lock, batch embed per-row fallback,
  LanceDB optimize, typer/click layer split)
- Round 3: 15 Minor cleanups (accessor unification, marker revert,
  episode query hygiene, --verbose subcommand, parser lifespan warm,
  task-number scrub, temporal-overlap test, real-SIGINT slow mark,
  4 design-note back-references)

Full per-round context lives in the PR description on GitHub.

## Test plan

- make lint (ruff + import-linter 3 contracts + assets +
  deprecated-names + github-docs + datetime + OpenAPI drift)
- Hermetic env full pytest — 2027 passed / 7 deselected (7 = slow +
  live_llm markers)
- Manual e2e across Tier 1/2/3 (21/21 assertions across v1/v2
  double-mount and Tier 3 -> Tier 2 downgrade)
- /health reports correct capabilities + disabled_features per tier

## Known follow-ups

- .superpowers/sdd/followup-http-bridge.md (gitignored) — Path A for
  spec §10's "backfill 期间 EverOS 完全可用" promise
- _TABLE_SCHEMA_VERSION docstring — v3+ migrations need a version
  dispatch table
- extract_user_profile.py throttle-counter block — replace LanceDB
  count_by_owner with a sqlite memcell count

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(review): close 3 blockers surfaced by round-4 review

N1. cluster_repo.find_cluster_id_for_member was cross-owner-unsafe.
Its reverse index (member_type, member_id) alone cannot disambiguate
two owners whose entry_id happens to collide — entry_id is
deliberately only per-owner unique (see entries.py:47:
'Cross-user uniqueness is handled at the database layer via a
composite <user_id>_<entry_id> field; it is not encoded into the
EntryId string itself'). Phase 2's _scan_all_rows crosses all
owners, so on any multi-owner root, same-day seq=1 episodes under
different owners would either false-hit each other's cluster or
be silently skipped from clustering. Add required (app_id,
project_id, owner_id) keyword args + JOIN Cluster (which already
carries scope) to filter by parent scope. Prior signature had zero
production callers except two the same PR just added, so the
API break is contained. Regression test: two owners persist a
cluster each around the same entry_id, each lookup resolves to
its own owner's cluster, a third owner's lookup returns None.

N2. Ctrl-C / EOF at the y/N prompt was landing on the generic
except Exception branch (exit 2 with rich traceback) instead of
the exit-130 interrupt path. Root cause: typer 0.15+ vendored
click under typer._click, so typer.Abort and the standalone
click.exceptions.Abort are distinct classes. The interrupt-branch
catch only listed the standalone one; every existing 'abort'
test was manually raising click.exceptions.Abort so the miss
was a false-positive guard rail. Widen the catch to
(typer.Abort, click.exceptions.Abort) and declare click as a
first-class dependency (it was only pulled in via uvicorn).
Regression test: raise real typer.Abort() at the confirm step
and assert exit 130 + INTERRUPTED banner.

N3. _looks_like_utf8_text used mime.startswith('text/'), which
caught text/html as well. HTML uploads then bypassed everalgo's
_aparse_html — losing clean_html_for_llm (strips <script>/<style>
/<nav>/<iframe> + HTML comments) and the 1 MiB output cap. A
40 MiB .html with <script> bodies and <!-- prompt injection -->
comments would flow straight into the extraction LLM. Replace
with explicit allowlist {text/plain, text/markdown, text/x-rst,
text/x-markdown}; text/html and any future text/* mime now
default to the parser path. Test matrix asserts text/html →
False (was regressed as True by the earlier commit).

Hermetic env full pytest: 2033 passed / 7 deselected (+6 tests
from these regressions).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(review): close round-4 major + minor + PR body errata

Round-4 review-driven cleanup. Blocker fixes (N1/N2/N3) landed in
561b5fe. This commit closes the remaining CONFIRMED items:

Major:
- J3 MemoryRoot.default() -> resolve() was a breaking public API
  rename this branch introduced (main still has default()). Adds
  default() as a backward-compat alias forwarding to resolve()
  with DeprecationWarning; CHANGELOG entry under Unreleased.
- J4 episode_repo.list_by_owner_after_ts(limit=N) truncates in
  fragment order (== insertion order), NOT newest-first. Docstring
  now spells out the trap so a future caller passing limit for a
  'newest N' window doesn't silently get the oldest N.
- J5.2 TyperPresenter.nothing_to_backfill picked colour via
  'could not be read' in message — a domain wording change
  would silently flip yellow -> green. Signature gains explicit
  scan_failed: bool kwarg; CLI colour-picks off the flag.
- J5.5 phase_header was Protocol-declared but never dispatched
  (run_backfill calls _print_phase_header directly). Removed the
  dead Protocol method + both no-op implementations.
- J6.3 3 inline from everos.core.errors import ... inside
  Phase 1/2/3 preflights promoted to a single top-level import.
- J7 subject-side embed failure was silently exit-0 because
  rows_processed advanced whenever any side wrote. Now: a row with
  a needed side still NULL counts as rows_failed (exit 4 =
  COMPLETED_WITH_FAILURES). Gated on spec.subject_of + row.needs_*
  + row.subject_text so non-Episode tables and subject-empty rows
  don't false-positive.
- J9 test_migration_cross_process.py did NOT actually test cross-
  process (all 5 tests mock memory_root_lock). Renamed to
  test_migration_lock_wiring.py; docstring now scopes it to
  'lock-invocation wiring' and points at test_core/…/test_locking.py
  for real flock coverage.
- J10 Phase 1 lacked the server-running preflight Phase 2/3 have.
  --phase all against a live server would burn Phase 1 embed API
  calls (real cost) before Phase 2 halted with exit 3. Phase 1 now
  probes _probe_ome_lock_available first; regression test in
  test_backfill_preflight.py; upgrade_path integration patches the
  probe so its in-process 'server + backfill' scenario stays valid.

Minor:
- M1 knowledge upload with NUL byte or filename > 255 bytes UTF-8
  used to raise ValueError/OSError at write_bytes → 500 with a
  half-written md left on disk. _safe_original_filename now
  rejects both up front with InvalidInputError (→ 400).
- M2 backfill optimize() now passes cleanup_older_than=timedelta(0)
  so older manifest versions are physically pruned (previous call
  compacted fragments but left the manifest chain on disk).
- M3 verify_business_schemas remediation text used to jump straight
  to 'rm -rf ~/.everos/.index/lancedb'; now walks restart → wipe.
- M5 multimodal/accessor.py capability_build_failed warning added
  so all four provider accessors log symmetrically (was silent).
- M7 test_knowledge_api parser-absence tests call
  parser_available.cache_clear() around the sys.modules patch so
  the lru_cache doesn't strand a stale True/False.
- M8 cascade_handler_embed_skipped (6 handlers) demoted INFO → DEBUG:
  Tier 1 imports were generating N × 6 handler-info lines per md.
- M10.1 test_drift_scenario_would_raise was a tautology (compared
  two hardcoded string sets, never touched the guard). Now
  monkey-patches BUSINESS_SCHEMAS_WITH_VECTOR to a superset and
  reloads _backfill, proving the import-time RuntimeError fires.
- M11 test_cascade_verbose_position subprocess.run calls gain
  env= — scrubs EVEROS_* from the developer environment so the
  same footgun as round-2 B1 doesn't re-appear inside subprocesses.
- M14 health() -> dict degraded the OpenAPI schema to
  additionalProperties: true. Introduce HealthResponse +
  HealthCapabilities Pydantic models so clients get real field
  shape; docs/openapi.json regenerated.
- M16 routes/knowledge.py:_require_knowledge_capabilities docstring
  claimed cascade.registry.build_handlers still gates
  knowledge_topic/knowledge_document, contradicting
  registry.py:177-194 (gate removed there, moved to HTTP layer).
  Rewritten to describe the current design accurately.

Hermetic env full pytest: 2037 passed / 7 deselected.

Explicitly deferred to followup:
- J1 lazy multimodal client (needs everalgo signature change)
- J2 tier definition (knowledge = Tier 3 whole)
- J5.1/3/4 broader presentation-split refactor
- J6.1/2 backfill dispatch and _backfill_table refactor
- J8 Phase 1 keyset pagination for bulk migration OOM
- M4 flock timeout/waiting log/re-entry (core/persistence refactor)
- M6 UTF-8 codec strategy (BOM/UTF-16/GBK)
- M9 count_by_owner monotonicity (latent, INTERVAL=1 short-circuits)
- M13 --phase all multi-phase combined-outcome test coverage
- M15 PR-marker rationale comments (16 occurrences, all inert)

M12 was refuted (both event names exist).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-07-29 11:05:23 +08:00
Dani 649046b0df
docs(api): use /api/v2 in docs and examples, demote v1 to legacy (#370)
1.2.0 introduced /api/v2 as the canonical, cloud-aligned prefix and mounted
every business router twice, but the user-facing entry points (README,
README.zh-CN, QUICKSTART, the docs/ set, the Langfuse example) still taught
/api/v1 — so new users were pointed at the compatibility alias while
docs/api.md already declared v2 canonical.

- Switch every EverOS endpoint reference in docs, examples, and
  `everos demo --live` to /api/v2, plus the matching CLI test expectations.
- Describe /api/v1 as a legacy compatibility alias that may be removed in a
  future major release, rather than a permanent one. Nothing changes at
  runtime: both prefixes still resolve to the same handlers and the
  v1/v2 parity test is untouched.
- Add a short note in README / README.zh-CN / QUICKSTART so existing v1
  integrations know they keep working.
- Fix the five dead endpoint anchors in the docs/api.md table of contents,
  which still pointed at the pre-1.2.0 #post-apiv1... slugs.

Left on v1 deliberately: docs/migration-to-1.0.0.md (historical record),
CHANGELOG history, tests/** (v1 must stay covered), and the
use-cases/claude-code-plugin + openher READMEs, which document a different
cloud API.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-29 10:52:29 +08:00
zhanghui 21b845bf88 feat(api): serve endpoints under /api/v2, retain /api/v1 as alias
Every business endpoint (memory/*, ome/*, knowledge/*) is now served
under /api/v2, aligning the open-source API with the EverOS Cloud
contract. /api/v1 is retained as a permanent, backward-compatible alias:
the same router objects are mounted under both prefixes, so both resolve
to identical handlers and request/response contracts. Existing /api/v1
integrations keep working unchanged. Infra endpoints (/health, /metrics)
stay unversioned.

Fix the Prometheus request-metric label to build the path from the full
request URL (with path params folded) rather than the route's
router-relative path, so the version prefix is preserved and v1/v2
traffic stays distinguishable.

Docs (docs/api.md, docs/openapi.json), CHANGELOG, and route docstrings
updated to lead with /api/v2. Add test_api_versioning as the parity
guard: every v2 route has an identical v1 twin and vice versa.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-24 13:44:58 +08:00
Elliot Chen 15efd1198c
chore(release): update EverOS to 1.1.1 (#327) 2026-07-07 18:30:03 +08:00
Elliot Chen 0df88f5603
chore(release): update EverOS to 1.1.0 (#307) 2026-06-24 23:17:23 +08:00
Elliot Chen a10cdcd197
chore(release): prepare EverOS 1.0.1 (#290) 2026-06-16 21:46:17 +08:00
Elliot Chen 518b8eca85 chore: initialize EverOS 1.0.0
md-first memory extraction framework for AI agents.

Markdown is the single source of truth; SQLite holds state and LanceDB
provides the rebuildable vector + BM25 + scalar index. The codebase follows
a single-direction DDD layering (entrypoints -> service -> memory -> infra,
with component / core / config cross-cutting) enforced by import-linter.

Engineering surface:
- Coding conventions in .claude/rules/ (path-scoped) and workflows in
  .claude/skills/ (/commit, /new-branch, /pr).
- GitHub Actions CI runs make lint + test + integration; pre-commit mirrors
  the gates locally (ruff, hygiene hooks, gitlint commit-msg).
- Commit messages follow Conventional Commits, enforced by gitlint.
- make lint also enforces datetime two-zone discipline and OpenAPI drift.
2026-06-06 07:33:17 +08:00