The wheel shipped ~94 public headers under `include/` and 7 static/shared
libraries under `lib/` (libzvec*.a, libzvec*.dylib). These come from the
`cc_library(PACKED)` install rules, which are meant for a standalone C++
`make install`, not the Python wheel. `_zvec.so` links the zvec libraries
statically (verified via otool: it depends only on system libs), so none of
those `lib/` / `include/` files are used at runtime — they are pure bloat.
Tag the three runtime install rules (_zvec, the DiskANN plugin, the jieba
dict) with `COMPONENT python`, and set `install.components = ["python"]` in
pyproject.toml so scikit-build-core installs only that component into the
wheel. The SDK install rules (default component) are excluded.
The standalone `make install` is unaffected (it still installs every
component).
Verified on macOS arm64: wheel no longer contains any `include/`, `lib/`,
`.a`, or `.h` entries; `import zvec` and `from zvec import Query` still work.
* feat(query): add VectorViewClause zero-copy path and unify validate
- Add VectorViewClause (string_view-based) as zero-copy counterpart to
VectorClause; variant now holds VectorClause | VectorViewClause | FtsClause
- Add QueryTarget::get_vector_view() unified accessor via std::visit,
returns optional<VectorViewClause> regardless of which variant is held
- Split validate_and_sanitize into QueryTarget::validate (read-only) +
sanitize_sparse_vector (mutate); validate handles both VectorClause and
VectorViewClause via get_vector_view()
- Collection::Query passes original request directly to sqlengine when no
sparse sanitization is needed; only copies when sort is required
- Change build_query_info/BuildSQLInfoFromSearchQuery to take const
SearchQuery& so VectorMatrixNode string_views point to caller's data
This PR exposes a copy-on-write mmap option through the public StorageOptions API and fixes a bug where the MAP_POPULATE flag was applied to the wrong mmap() argument.
convert_c_index_params_to_cpp used a whitelist switch over IndexType,
prone to miss for newly index types.
Replace the entire switch with a single clone() call (pure-virtual on
IndexParams base, implemented by all subclasses).
* refactor: make Reranker stateless with std::variant value semantics (#461)
Replace class hierarchy (Reranker/ScoreBasedReranker/RrfReranker/
WeightedReranker/CallbackReranker) with std::variant<RrfParams,
WeightedParams, CallbackParams> value type and a stateless free function
reranker::rerank().
Key changes:
- reranker.h: define RerankParams variant + reranker::rerank() API
- query.h: MultiQuery::reranker (shared_ptr) -> MultiQuery::rerank (value)
- schema.h: add CollectionSchema::get_field_ptr() returning FieldSchema::Ptr
- collection.cc: push field lookup to caller, pass vector<FieldSchema::Ptr>
- c_api: remove opaque zvec_reranker_t, add zvec_multi_query_set_rerank_*
- python binding: expose _RrfParams/_WeightedParams/_CallbackParams + setters
- python layer: WeightedReRanker(list[float]), remove Python rerank logic
- all tests updated to new interface
Benefits:
- Thread-safe by design: no mutable state, safe to share across threads
- Collection-decoupled: no bind_schema(), field info passed as parameter
- Simpler lifecycle: value semantics, no shared_ptr management
Closes#461
* chore: remove nightly_build.yml unrelated to reranker refactor
* chore: remove uv.lock unrelated to reranker refactor
* fix: raise ValueError when multi-query has no reranker
After the reranker stateless refactor the C++ MultiQuery rerank
strategy uses a std::variant with a default value, so the implicit
'reranker required' validation no longer triggered. Restore the
check in QueryExecutor._execute_multi_query so that a hybrid
(multi-query) request without a reranker raises ValueError.
* fix(reranker): use index_type FTS check for non-vector normalization
Replace dynamic_cast nullptr check with explicit IndexType::FTS check
and map FTS/BM25 positive scores to (0.0, 1.0) via 2*atan(score)/pi.
* refactor(reranker): move Params types into reranker namespace and qualify usages
Move RrfParams, WeightedParams, CallbackParams and RerankParams into the
zvec::reranker namespace, and add explicit reranker:: qualification at all
usage sites outside the reranker module (query.h, python/c bindings, tests).
* refactor(query): drop unused PendingQuery wrapper, use std::vector<SearchQuery> directly
* refactor(reranker): make _to_cpp_params non-abstract with default NotImplementedError
Remove @abstractmethod from RerankFunction._to_cpp_params and provide a
default implementation raising NotImplementedError. Drop the redundant
_to_cpp_params overrides from Qwen and Sentence rerankers since they use
the Python rerank path and don't need the C++ conversion.
* feat(entity/search): add LinearPool/BlockHeap, split contiguous entity layout and add direct-pointer fast search path
Introduce single-heap candidate structures (LinearPool/BlockHeap) and
refactor greedy search to dispatch between pool and dual-heap paths.
Split node layout into separate vector/graph arrays for better cache
locality, add zero-copy get_vector_ptr() on the hot path, and extract
huge-page allocation into MemoryHelper. Also adds pyglass attribution.
* refactor: change rerank interface from map-based to vector-based (#452)
- Define QueryResult = list[Doc] type alias in doc.py
- Change C++ Reranker::rerank() signature from map<string, DocPtrList> to vector<DocPtrList>
- Extend bind_schema() to accept field_names for index-based field lookup
- Update ScoreBasedReranker/WeightedReranker/CallbackReranker implementations
- Adapt collection.cc MultiQuery path to use vector<DocPtrList>
- Update Python binding to expose rerank() and use vector<double> weights
- Refactor Python RerankFunction interface to list[QueryResult] -> QueryResult
- Remove Python-layer rerank logic from RrfReRanker/WeightedReRanker (delegate to C++)
- Update query_executor to return list[list[Doc]] instead of dict
- Update all related unit tests (C++ and Python)
* refactor: replace list[Doc] with QueryResult type alias in executor and rerank functions
* refactor: replace list[list[Doc]] with list[QueryResult] in query_executor
* fix: remove unused Doc import in rerank_function.py (ruff F401)
* refactor(query_executor): merge duplicate rerank return paths
* refactor: RrfReRanker/WeightedReRanker.rerank() directly call C++ reranker
* refactor: simplify QueryExecutor into unified class, remove Factory/subclasses/validation/concurrency
* refactor: rename _VectorQuery to _SearchQuery, from_vector_query to from_search_query
* refactor(query_executor): split execute into single/multi paths, rename core_vector to search_query, drop unused core_vectors
* style: apply ruff formatter to test_reranker.py and query_executor.py
* refactor: make rescore() private in ScoreBasedReranker hierarchy
* style: apply clang-format to reranker.h
* style: apply clang-format to all modified C++ files
* refactor: rename private methods in QueryExecutor for clearer semantics
* refactor: rename mvq to multi_query for clarity
* fix: make BasicRRF test order-independent for equal scores
* fix: update collection_test to use vector-based reranker interface
* fix: update reranker tests to expect TypeError instead of NotImplementedError
* refactor: remove PendingQuery wrapper, use SearchQuery directly in MultiQuery path
* refactor: simplify MultiQuery path - remove seen_fields, merge field_names into main loop
* fix: address review comments - defensive checks and remove fields param from C API
- ScoreBasedReranker::rerank(): early return empty list when topn <= 0
- WeightedReranker::rescore(): null-check schema_ before use
- CallbackReranker::rerank(): check callback_ is not empty before invoke
- C API zvec_reranker_create_weighted(): remove unused fields parameter
* fix: remove duplicate field name test (check was intentionally removed)
* fix: address egolearner review comments
- Rename QueryResult to DocList for clarity (见名知义)
- Change docstring to #: comment for type alias
- Fix output_fields check: use 'is not None' instead of truthy check
(None means unset, [] means explicit empty list - different semantics)
- Raise ValueError when search-by-id finds no document
* refactor: remove redundant output_fields assignment in _build_search_query
* refactor: address egolearner review comments (C++ refactoring)
- c_api.cc: simplify weighted reranker creation with inline vector ctor
- python_reranker.cc: refactor unwrap_rerank_result - take by value,
early error return, move semantics
- Rename C API functions for consistent naming:
zvec_reranker_create_rrf -> zvec_create_rrf_reranker
zvec_reranker_create_weighted -> zvec_create_weighted_reranker
zvec_reranker_destroy -> zvec_destroy_reranker
zvec_reranker_get_rank_constant -> zvec_get_reranker_rank_constant
- reranker.h/cc: bind_schema returns Result<void>, caches
vector<const FieldSchema*> to avoid repeated schema lookups in rescore
- python_param.cc: rename py::arg vector_query to search_query
* revert: rollback bind_schema refactoring due to thread-safety concern
The field_schemas_ caching approach introduces a data race when the same
WeightedReranker instance is shared across concurrent queries: bind_schema()
writes field_schemas_ while rerank() reads it concurrently.
Revert to storing schema_ + field_names_ and looking up fields in rescore().
Add @note thread-safety warning to WeightedReranker class documentation.
* fix: unify error message format in collection.cc
Change 'Vector field not found: X' to 'Invalid query: field X not found'
for consistent error formatting as suggested by zhourrr.
* fix: sort __all__ and remove duplicates in __init__.pyi
Fix RUF022 lint error: sort __all__ alphabetically and remove duplicate
entries (DenseEmbeddingFunction, ReRanker).
* style: format query_executor.py with ruff formatter
* fix: resolve Python test failures after FTS rebase integration
- test_query_executor.py: update method names to match refactored API
(_do_build -> _build_queries, _do_merge_rerank_results -> _merge_and_rerank)
- test_reranker.py: fix expected exception type (TypeError from pybind11)
- test_collection_fts.py: update error message match patterns
- test_collection_fts_vector_hybrid.py: remove obsolete 'metrics' param,
update weights from dict to positional list, adapt validation tests
for multi-vector queries (now supported with reranker)
- test_collection_dql.py: remove 'metrics' param, update weights format
- collection.cc: distinguish FTS vs vector fields in MultiQuery path
using get_fts_clause() to route field lookup correctly
- reranker.cc: use get_field() instead of get_vector_field() in rescore
to support FTS+vector hybrid weighted reranking
* refactor: pass topn as rerank() parameter, move rerank_field to model rerankers
* fix: address review comments - rename test functions and restore duplicate field check
* refactor: simplify MultiQuery field lookup, let validate_and_sanitize handle type check
Previously MultiQuery only accepted vector sub-queries by using
get_vector_field() to look up each sub-query's field. FTS sub-queries
(whose field_name points to an FTS-indexed string column) would fail
with "Vector field not found".
Changes:
- collection.cc: use get_field() uniformly in MultiQuery path; let
validate_and_sanitize() check type compatibility internally, which
is consistent with the single-query path.
- query_executor.py: allow SingleVectorQueryExecutor to accept
multi-query when it contains an FTS query (with reranker), and route
to C++ MultiQuery fast path.
- Add test_collection_fts_vector_hybrid.py covering hybrid retrieval
ranking, scoring, filter, validation, and edge cases.
* feat: add FTS support for Collection::CreateIndex/DropIndex
Enable dynamic creation and removal of FTS indexes on existing STRING
columns through the standard CreateIndex/DropIndex API, matching the
lifecycle model already used by vector and scalar (invert) indexes.
Key changes:
- New FtsIndexer class (fts_indexer.h/cc) encapsulating per-segment FTS
RocksDB management: multi-field lifecycle, snapshot, insert, seal
- New BlockType::FTS_INDEX with block_id-based directory naming
(fts.<block_id>.rocksdb) for crash-safe snapshot-and-swap
- Segment::create_fts_index builds FTS index on a snapshot copy by
scanning forward store, then outputs new SegmentMeta + FtsIndexer
for atomic reload (same pattern as create_scalar_index)
- Segment::drop_fts_index snapshots, removes field CFs, outputs updated
meta (or nullptr when last FTS field is removed)
- Collection layer wires FTS into the existing task dispatch, version
update, and reload loops alongside vector/invert paths
- CreateIndex/DropIndex reject unsupported index types explicitly
instead of falling through to the wrong branch
* fixup! feat: add FTS support for Collection::CreateIndex/DropIndex
fix: address review comments for FTS CreateIndex/DropIndex
- Reject CreateIndex when column already has a different index type
(e.g. FTS on an INVERT-indexed column) at both Collection and
Segment layers
- Allow same-type different-params CreateIndex to rebuild the index
(remove old + create new + replay data), aligned with INVERT behavior
- Return OK when CreateIndex is called with identical params
- Rename operator[] to get() in both FtsIndexer and InvertedIndexer
- Add test cases: create→drop→create→drop cycle, and params-change
rebuild with case-sensitivity verification
* android skip Feature_CreateOrDropFtsIndex
* fixup! feat: add FTS support for Collection::CreateIndex/DropIndex
Add a fast path that copies the first segment's index file as the merge base and only merges the tail segments into it. Limited to streaming indexes (HNSW, HNSW_RABITQ, FLAT) with matching index_type + quantize_type and no filter; IVF/VAMANA always rebuild (their Merge is dump-then-reopen and would drop the base docs).
Also, fix the incorrect concurrency of the compaction task.
- Upgrade ilammy/msvc-dev-cmd from v1 to v1.13.0 to resolve Node.js 20
deprecation warning (affects 05-windows-build and build_wheel_on_windows)
- Remove /MP from antlr4 and protobuf patches to eliminate 332
non-cacheable sccache calls caused by "multiple input files"
- Add sccache statistics steps after pip install and after C++ Tests to
capture build and test compilation cache metrics separately
Previously execute() and execute_group_by() filled the query_infos
vector with copies of the same shared_ptr, so all segments pointed to
the identical QueryInfo object. The optimizer mutates QueryInfo
in-place (set_invert_cond, set_filter_cond, set_forward, set_text,
etc.), which meant optimizations applied for segment 0 silently
corrupted the input state for subsequent segments.
Build a separate QueryInfo via build_query_info() for each segment so
they can be optimized independently. The filter-satisfiability check
is query-global and done once before the loop.
feat(ci): enable sccache on Windows and propagate compiler launcher to ExternalProject
Windows CI builds were ~6-9x slower than other platforms (35-40 min vs
4-7 min) due to missing compiler cache. This commit enables sccache for
Windows and ensures all ExternalProject_Add dependencies (Arrow, lz4)
receive the compiler launcher on every platform.
Changes:
MSVC /Zi → /Z7 compatibility:
- Replace /Zi with /Z7 in root CMakeLists.txt for DEBUG and
RELWITHDEBINFO configurations (MSVC cache variable level)
- Patch rocksdb, antlr4, and googletest source CMakeLists to replace
hardcoded /Zi with /Z7 at the source, eliminating D9025 warnings
- Clear COMPILE_PDB_NAME on googletest targets to prevent CMake from
adding /Fd<path>, which causes sccache to look for nonexistent .pdb
- Remove /FS (forced synchronous PDB writes) from Arrow ExternalProject
MSVC flags — unnecessary with /Z7
Windows CI workflow:
- Add mozilla-actions/sccache-action@v0.0.10
- Set SCCACHE_GHA_ENABLED=true to use GitHub Actions Cache backend
- Pass CMAKE_C/CXX_COMPILER_LAUNCHER=sccache via --config-settings
Compiler launcher propagation:
- Propagate CMAKE_C/CXX_COMPILER_LAUNCHER to Arrow ExternalProject on
all platforms (Windows, Linux, macOS, Android, iOS)
- Propagate to lz4 ExternalProject on Windows via CMAKE_ARGS, using
CMP0141=NEW + CMAKE_MSVC_DEBUG_INFORMATION_FORMAT=Embedded for /Z7
* fix(vamana): medoid entry point, concurrent build crash and prune quality
- Add DiskANN-standard medoid (centroid-closest point) as the persisted
entry point, computed at dump time, replacing the fixed node-0 start.
- Fix data races on node_chunks_ / dist_chunks_ / node_chunk_bases_ during
concurrent build: add mutexes with double-checked locking and pre-reserve
capacity so the lock-free read fast path never hits reallocation.
- Fix RobustPrune for metrics with signed internal distance (quantized int8
cosine): introduce IndexMetric::build_distance_offset() to shift distances
into a non-negative range, restoring a geometrically meaningful
occlude_factor and improving graph quality / low-ef recall.
- Remove redundant `virtual` keyword from methods already marked `override`
- Replace `virtual ~Derived()` with `~Derived() override` for derived class destructors
- Add missing `override` on methods that override base class virtuals
* fix(cpu_features): refactor architecture detection to explicit x86 whitelist
Currently, `cpu_features.cc` assumes any non-ARM architecture is x86/x64, which leads to a fatal missing `<cpuid.h>` error on architectures like RISC-V.
This commit refactors the preprocessor macros to explicitly whitelist x86 architectures (`__x86_64__`, `__i386__`, `_M_X64`, `_M_IX86`). All other architectures (RISC-V, ARM, etc.) will now safely fall back to the default zero-initialization, allowing cross-compilation to succeed.
Signed-off-by: ihb2032 <hebome@foxmail.com>
* ci: Add RISE RISC-V runner
Introduce the RISC-V CI runner provided by the RISE project.
This enables automated testing and building for the RISC-V architecture.
Signed-off-by: ihb2032 <hebome@foxmail.com>
* ci: use python 3.12 for RISC-V64
Signed-off-by: ihb2032 <hebome@foxmail.com>
* ci: add RISC-V numpy dependencies
Signed-off-by: ihb2032 <hebome@foxmail.com>
* ci: add RISC-V wheel cache
Signed-off-by: ihb2032 <hebome@foxmail.com>
* ci: use pre-built RISE numpy wheel to speed up riscv builds
Signed-off-by: ihb2032 <hebome@foxmail.com>
* ci: use pre-built RISE cmake wheel to speed up riscv builds
Signed-off-by: ihb2032 <hebome@foxmail.com>
* ci: split RISC-V build and test into separate jobs
Signed-off-by: ihb2032 <hebome@foxmail.com>
* ci: fix
Signed-off-by: ihb2032 <hebome@foxmail.com>
* ci: fix
Signed-off-by: ihb2032 <hebome@foxmail.com>
* Update hnsw_streamer_test.cc
* ci: add cache
Signed-off-by: ihb2032 <hebome@foxmail.com>
* Update hnsw_streamer_test.cc
* Update test_gil_release.py
* ci: schedule workflow to run overnight
---------
Signed-off-by: ihb2032 <hebome@foxmail.com>
Co-authored-by: ZeFeng Yin <yinzefeng.yzf@alibaba-inc.com>
Windows.h defines ERROR as a macro which collides with LogLevel::ERROR.
Rename all LogLevel enum values to kDebug/kInfo/kWarn/kError/kFatal style,
and remove the now-unnecessary #undef ERROR workaround in jieba_tokenizer.