egolearner
aad5718112
feat(query): add VectorViewClause zero-copy path and unify validate ( #478 )
...
* feat(query): add VectorViewClause zero-copy path and unify validate
- Add VectorViewClause (string_view-based) as zero-copy counterpart to
VectorClause; variant now holds VectorClause | VectorViewClause | FtsClause
- Add QueryTarget::get_vector_view() unified accessor via std::visit,
returns optional<VectorViewClause> regardless of which variant is held
- Split validate_and_sanitize into QueryTarget::validate (read-only) +
sanitize_sparse_vector (mutate); validate handles both VectorClause and
VectorViewClause via get_vector_view()
- Collection::Query passes original request directly to sqlengine when no
sparse sanitization is needed; only copies when sort is required
- Change build_query_info/BuildSQLInfoFromSearchQuery to take const
SearchQuery& so VectorMatrixNode string_views point to caller's data
2026-06-22 11:32:26 +08:00
feihongxu0824
4679a8f39d
feat: support zvec core-only build ( #481 )
2026-06-09 23:45:42 +08:00
rayx
e720c1fd20
feat: add diskann index ( #369 )
2026-06-04 20:52:43 +08:00
Jalin Wang
f562bdd636
fix(segment): use raw vectors for RaBitQ merge during compaction ( #456 )
...
* fix: rabitQ recall=0
* re-enable rabitQ SegmentCompactReuseTest
* chore: simplify the comment
* chore: simplify the comment
2026-06-03 21:47:12 +08:00
Qinren Zhou
dbea635019
refactor: clarify segment-local row ID handling and tests and fixes bugs ( #432 )
2026-06-03 15:20:45 +08:00
Jalin Wang
95e5ad5105
perf(segment): reuse first vector index file as merge base during compaction ( #440 )
...
Add a fast path that copies the first segment's index file as the merge base and only merges the tail segments into it. Limited to streaming indexes (HNSW, HNSW_RABITQ, FLAT) with matching index_type + quantize_type and no filter; IVF/VAMANA always rebuild (their Merge is dump-then-reopen and would drop the base docs).
Also, fix the incorrect concurrency of the compaction task.
2026-06-03 14:22:07 +08:00
ZeFeng Yin
8ce8e3e228
chore: rm buffer manager ( #437 )
2026-06-02 09:51:41 +08:00
egolearner
02bfb31cf5
feat: add fts support ( #408 )
...
Add BM25-based full-text search with CJK (jieba) tokenization, supporting
query_string and match_string syntax, phrase queries, boolean operators
(AND/OR/NOT/MUST), and hybrid retrieval with existing vector search.
## Core
- BitPacked posting format with block-max WAND pruning
- Tokenizer pipeline: jieba (cut/cut_for_search/hmm/full), whitespace,
lowercase, with extensible pipeline composition
- Query parser: boolean operators, phrase queries, field scoping,
boost, MUST(+) modifier inside OR (ES query_string semantics)
- AST rewriter: dedup repeated terms with linear boost aggregation,
flatten same-type composites, canonicalize OR-with-must_not into AND
wrapper, empty-node propagation, contradiction detection
- FTS reduce/merge integrated into Optimize compaction
- Multi-segment score-descending sort
- Auto-register bundled jieba dict on SDK import
## Performance
- Block-max WAND with cached block_max_info_for (single binary search)
- AVX2/SSE bitpacked encoding with cross-arch scalar fallback
- MultiGet for batch posting retrieval and phrase position verification
- HashSkipList memtable for posting writes
- PinnableSlice zero-copy reads
- Filter pushdown into composite iterators (Disjunction/Conjunction/Phrase)
- Candidate-driven (brute-force) evaluation for selective invert filters
- Precomputed BM25 IDF weights, cached SIMD dispatch pointers
- Shortest-list anchor for phrase position matching
- Single-open per-term posting iterator
## Query
- Tokenize query terms through the same pipeline as indexing
- EmptyNode for zero-token queries (all stop-words / punctuation)
- Backslash unescape after lexing in query parser
- Schema allows collections without vector fields (FTS-only use case)
- Create/Drop Index validates supported index types
- FTS fields disallowed in SQL filter expressions
## Bindings
- C API: fts query params, brute-force ratio config
- Python SDK: FTS search, jieba dict auto-registration
## Internals
- Bypass cppjieba::Jieba to drop KeywordExtractor (~12MB fewer required files)
- Hide tokenizer pipeline from public header (Pimpl-style FtsState)
- ListColumnFamilies to avoid double-open on segment load
- Reorganized fts_column into tokenizer/, posting/, iterator/ subdirs
2026-06-01 15:02:54 +08:00
egolearner
8dcb6cbd7f
refactor: drop VectorQuery, unify single-target query on SearchQuery ( #428 )
2026-05-29 16:36:09 +08:00
ZeFeng Yin
d9b0920ac7
fix: ivf provider sorted by local id ( #422 )
2026-05-26 10:33:44 +08:00
lichen2015
bdf58fc2d2
fix: prevent SIGABRT when adding nullable column to multi-segment collection ( #415 ) ( #416 )
...
When add_column is called on a multi-segment collection with a nullable
field and no expression, segment.cc previously sliced an Arrow ChunkedArray
with an offset that exceeded the array length, triggering SIGABRT in
Arrow's chunked_array.cc:170 assertion.
Fix the slicing logic in segment.cc to materialize null values per segment.
Add comprehensive tests in collection_test.cc and segment_test.cc covering
multi-segment add_column scenarios (nullable/non-nullable, with/without
expression, with/without unflushed data, drop+re-add).
2026-05-24 14:13:57 +08:00
egolearner
9aae7494ad
chore: enable modernize-use-override and fix existing violations ( #419 )
...
Add modernize-use-override to .clang-tidy and apply fixes across
src/ and tests/: replace redundant virtual with override, annotate
missing override on derived methods, and drop redundant virtual on
already-overridden methods.
2026-05-21 19:05:32 +08:00
Jalin Wang
1d4ae0b1b5
refactor: support UTF-8 file paths via std::filesystem ( #359 )
...
Rewrite file/path handling to use std::filesystem and UTF-8-safe helpers.
Switch Windows file open/create paths to wide-char APIs, replace manual
separator concatenation with PathJoin, and enable RocksDB UTF-8 filenames.
Also add UTF-8 path coverage for file IO, version manager recovery, and
collection open/flush/reopen flows.
2026-05-08 17:44:22 +08:00
Qinren Zhou
68a497efdb
fix: sparse vector indices should be ordered ( #382 )
2026-05-07 16:40:49 +08:00
ZeFeng Yin
005680522b
feat: merge vector arrow buffer ( #320 )
2026-04-29 19:24:50 +08:00
kgeg401
1e25294c07
feat(ci): integrate clang-tidy for changed C/C++ files ( #116 )
...
* feat(ci): add clang-tidy checks for changed cpp files
* fix(ci): correct clang-tidy workflow heredoc indentation
* test: stabilize fp16 euclidean matrix comparison
* test: fix fp16 matrix CI and clang-tidy warnings
* ci: scope clang-tidy PR and filter compile-db files
* check nullptr only
* fix nullptr
* fix: format
* chore: ignore c
* fix: tests files
* fix: workflow
* fix: specify clang-tidy version
* fix: filter regex
* fix: more files
* fix: update version
* add thirdparty cache
* add pipeline dependency
* fix clang-tidy trigger
---------
Co-authored-by: kgeg401 <kgeg401@users.noreply.github.com>
Co-authored-by: jiliang.ljl <jiliang.ljl@alibaba-inc.com>
2026-04-20 20:14:21 +08:00
Qinren Zhou
c17bd8876e
fix: validate query_params type ( #351 )
2026-04-20 10:54:06 +08:00
feihongxu0824
6737810190
ci: refact android ci ( #330 )
2026-04-15 15:37:51 +08:00
Jalin Wang
8182cfff93
feat: windows support ( #216 )
...
Preliminary support for Windows with remaining TODOs
2026-03-30 20:47:34 +08:00
Qinren Zhou
ae345ad070
fix: recovery from crash during optimization ( #246 )
2026-03-24 10:53:16 +08:00
egolearner
e5ba11b6fe
feat: add hnsw-rabitq support ( #69 )
...
* feat: add hnsw-rabitq
* undefine transform
* support config rotator type
* support sample_count
* fix ut
* refactor: update interface
* refactor: update interface
* update rabitq index params
* fix interface update
* fix searcher test
* fix streamer test
* streamer support bf and add more ut
* rm env check
* add collection ut
* add schema check
* add rabitq query param binding
* add integration test
* fix local_builder
* cleanup dist calculator
* add RaBitQ-Library submodule
* cleanup rabitq converter/reformer
* add files
* disable build on mac
* disable Feature_Optimize_HNSW_RABITQ
* disable python/tests/test_collection_hnsw_rabitq.py:13
* check avx2/avx512
* fix refine
* fix mac ci
* add dimension check
* fix mac ci
* fix ci
* add missing lib
* fix centroids selection for Cosine/InnerProduct
* rename hnsw-rabitq to hnsw_rabitq
* address comments
* fix compile
* fix rabitqlib name
* fix mac compile
* check avx2/avx512 support
* runtime check again compile-time
* check AUTO_DETECT_ARCH
* rm debug log
* fix
* rabitq use avx2 by default
* add missing file
* fix search_bf/group_by dist
* address comments
* fix typo
* address comments
* rabitq support disable id_map
* address comments
2026-03-19 17:21:42 +08:00
ZeFeng Yin
432de52c5f
[enhance] buffer storage ( #83 )
2026-03-19 15:24:46 +08:00
Qinren Zhou
3c1241f7c6
feat: enlarge indice size limit for sparse vectors ( #229 )
2026-03-16 17:12:24 +08:00
feihongxu0824
fb0b900a80
fix(build): replace CMAKE_SOURCE_DIR with PROJECT_ROOT_DIR for subproject builds ( #195 )
2026-03-04 20:01:52 +08:00
feihongxu0824
24d877da5e
fix: Replace VLAs with std::vector for MSVC compatibility and stack safety ( #190 )
2026-03-03 10:07:21 +08:00
lichen2015
e7ad7cc31e
fix: combined indexer should use key instead of index ( #87 )
...
Co-authored-by: yinzefeng.yzf <yinzefeng.yzf@alibaba-inc.com>
2026-02-12 10:38:05 +08:00
lichen2015
fe90f70a3b
fix: delete_by_filter should use g_doc_id ( #84 )
2026-02-09 17:40:51 +08:00
lichen2015
f1349cc91a
fix: remove unnecessary column_name param from the AddColumn API ( #59 )
2026-02-04 14:44:13 +08:00
feihongxu0824
0ebf5621a6
feat: support core cpp sdk ( #30 )
...
* support core cpp sdk
* fix
* fix
* fix
* fix cr
2026-01-26 20:01:44 +08:00
feihongxu0824
e957442355
feat: refact cpp sdk ( #27 )
...
* refact cpp sdk
---------
Co-authored-by: zhouqinren.zqr <zhouqinren.zqr@alibaba-inc.com>
2026-01-21 20:05:42 +08:00
lichen2015
08d680bd4f
fix: fix and refact segment fetch perf ( #26 )
...
* fix and refact fetch perf
* change git user name
2026-01-19 13:43:30 +08:00
sanyi
6524fabd80
Initial commit
2025-12-30 11:02:17 +08:00