Alibaba lightweight in-process vector database
Go to file
egolearner 443500dc45
feat: add FTS support for Collection::CreateIndex/DropIndex (#445)
* feat: add FTS support for Collection::CreateIndex/DropIndex

Enable dynamic creation and removal of FTS indexes on existing STRING
columns through the standard CreateIndex/DropIndex API, matching the
lifecycle model already used by vector and scalar (invert) indexes.

Key changes:
- New FtsIndexer class (fts_indexer.h/cc) encapsulating per-segment FTS
  RocksDB management: multi-field lifecycle, snapshot, insert, seal
- New BlockType::FTS_INDEX with block_id-based directory naming
  (fts.<block_id>.rocksdb) for crash-safe snapshot-and-swap
- Segment::create_fts_index builds FTS index on a snapshot copy by
  scanning forward store, then outputs new SegmentMeta + FtsIndexer
  for atomic reload (same pattern as create_scalar_index)
- Segment::drop_fts_index snapshots, removes field CFs, outputs updated
  meta (or nullptr when last FTS field is removed)
- Collection layer wires FTS into the existing task dispatch, version
  update, and reload loops alongside vector/invert paths
- CreateIndex/DropIndex reject unsupported index types explicitly
  instead of falling through to the wrong branch

* fixup! feat: add FTS support for Collection::CreateIndex/DropIndex

    fix: address review comments for FTS CreateIndex/DropIndex

    - Reject CreateIndex when column already has a different index type
      (e.g. FTS on an INVERT-indexed column) at both Collection and
      Segment layers
    - Allow same-type different-params CreateIndex to rebuild the index
      (remove old + create new + replay data), aligned with INVERT behavior
    - Return OK when CreateIndex is called with identical params
    - Rename operator[] to get() in both FtsIndexer and InvertedIndexer
    - Add test cases: create→drop→create→drop cycle, and params-change
      rebuild with case-sensitivity verification

* android skip Feature_CreateOrDropFtsIndex

* fixup! feat: add FTS support for Collection::CreateIndex/DropIndex
2026-06-03 17:56:56 +08:00
.github ci(windows): upgrade ilammy/msvc-dev-cmd, remove /MP and add sccache stats (#454) 2026-06-03 14:16:35 +08:00
cmake chore: fix more warnings (#418) 2026-05-21 11:15:41 +08:00
examples refactor: drop VectorQuery, unify single-target query on SearchQuery (#428) 2026-05-29 16:36:09 +08:00
python Refactor CPU feature detection to use an explicit x86 whitelist (#258) 2026-06-02 11:51:25 +08:00
scripts ci: refact android ci (#330) 2026-04-15 15:37:51 +08:00
src feat: add FTS support for Collection::CreateIndex/DropIndex (#445) 2026-06-03 17:56:56 +08:00
tests feat: add FTS support for Collection::CreateIndex/DropIndex (#445) 2026-06-03 17:56:56 +08:00
thirdparty ci(windows): upgrade ilammy/msvc-dev-cmd, remove /MP and add sccache stats (#454) 2026-06-03 14:16:35 +08:00
tools feat: add fts support (#408) 2026-06-01 15:02:54 +08:00
.clang-format Initial commit 2025-12-30 11:02:17 +08:00
.clang-tidy chore: enable modernize-use-override and fix existing violations (#419) 2026-05-21 19:05:32 +08:00
.gitattributes feat: add fts support (#408) 2026-06-01 15:02:54 +08:00
.gitignore feat(ci): integrate clang-tidy for changed C/C++ files (#116) 2026-04-20 20:14:21 +08:00
.gitmodules feat: add fts support (#408) 2026-06-01 15:02:54 +08:00
.pre-commit-config.yaml chore: enable the conventional-pre-commit run sucess and update to latest version (#111) 2026-02-25 18:03:20 +08:00
CMakeLists.txt feat(ci): enable sccache on Windows and propagate compiler launcher to ExternalProject (#439) 2026-06-02 16:23:37 +08:00
CODE_OF_CONDUCT.md Initial commit 2025-12-30 11:02:17 +08:00
CONTRIBUTING.md doc: add v0.3.0 release note (#312) 2026-04-03 15:47:19 +08:00
LICENSE Initial commit 2025-12-30 11:02:17 +08:00
README.md minor: update readme for v0.4.0 (#389) 2026-05-09 11:40:30 +08:00
README_CN.md minor: update readme for v0.4.0 (#389) 2026-05-09 11:40:30 +08:00
pyproject.toml feat(test): enable parallel tests (#384) 2026-05-12 23:14:17 +08:00

README.md

English | 中文

zvec logo

Code Coverage Main License PyPI Release Python Versions npm Release

alibaba%2Fzvec | Trendshift

🚀 Quickstart | 🏠 Home | 📚 Docs | 📊 Benchmarks | 🔎 DeepWiki | 🎮 Discord | 🐦 X (Twitter)

Zvec is an open-source, in-process vector database — lightweight, lightning-fast, and designed to embed directly into applications. Battle-tested within Alibaba Group, it delivers production-grade, low-latency and scalable similarity search with minimal setup.

[!Important] 🚀 v0.4.0 (May 9, 2026)

  • Dart/Flutter SDK: Published the official zvec Flutter package with FFI bindings. Supports Android (arm64-v8a) and iOS (arm64) — no manual native compilation required.
  • iOS Build Support: Added support for building on iOS platforms, expanding cross-platform coverage.
  • Enlarged topK Limit: Relaxed the upper bound on topK to support larger-scale recall scenarios.
  • Bug Fixes: SQ8 quantizer recall drop; Windows path handling; sparse vector index ordering.

👉 Read the Release Notes | View Roadmap 📍

💫 Features

  • Blazing Fast: Searches billions of vectors in milliseconds.
  • Simple, Just Works: Install and start searching in seconds. Pure local, no servers, no config, no fuss.
  • Dense + Sparse Vectors: Work with both dense and sparse embeddings, with native support for multi-vector queries in a single call.
  • Hybrid Search: Combine semantic similarity with structured filters for precise results.
  • Durable Storage: Write-ahead logging (WAL) guarantees persistence — data is never lost, even on process crash or power failure.
  • Concurrent Access: Multiple processes can read the same collection simultaneously; writes are single-process exclusive.
  • Runs Anywhere: As an in-process library, Zvec runs wherever your code runs — notebooks, servers, CLI tools, or even edge devices.

📦 Installation

Python

Requirements: Python 3.10 - 3.14

pip install zvec

Node.js

npm install @zvec/zvec

Supported Platforms

  • Linux (x86_64, ARM64)
  • macOS (ARM64)
  • Windows (x86_64)

🛠️ Building from Source

If you prefer to build Zvec from source, please check the Building from Source guide.

One-Minute Example

import zvec

# Define collection schema
schema = zvec.CollectionSchema(
    name="example",
    vectors=zvec.VectorSchema("embedding", zvec.DataType.VECTOR_FP32, 4),
)

# Create collection
collection = zvec.create_and_open(path="./zvec_example", schema=schema)

# Insert documents
collection.insert([
    zvec.Doc(id="doc_1", vectors={"embedding": [0.1, 0.2, 0.3, 0.4]}),
    zvec.Doc(id="doc_2", vectors={"embedding": [0.2, 0.3, 0.4, 0.1]}),
])

# Search by vector similarity
results = collection.query(
    zvec.VectorQuery("embedding", vector=[0.4, 0.3, 0.3, 0.1]),
    topk=10
)

# Results: list of {'id': str, 'score': float, ...}, sorted by relevance
print(results)

📈 Performance at Scale

Zvec delivers exceptional speed and efficiency, making it ideal for demanding production workloads.

Zvec Performance Benchmarks

For detailed benchmark methodology, configurations, and complete results, please see our Benchmarks documentation.

🤝 Join Our Community

💬 DingTalk 📱 WeChat 🎮 Discord X (Twitter)
DingTalk QR Code WeChat QR Code Discord X (formerly Twitter) Follow
Scan to join Scan to join Click to join Click to follow

❤️ Contributing

We welcome and appreciate contributions from the community! Whether you're fixing a bug, adding a feature, or improving documentation, your help makes Zvec better for everyone.

Check out our Contributing Guide to get started!