Alibaba lightweight in-process vector database
Go to file
feihongxu0824 a9008b0fd2
fix: nullable scalar field filter leaks null documents without inverted index (#410)
When a nullable scalar field has no inverted index, the forward filter path
fails to handle null values from Arrow's filter evaluation:

1. get_forward_bit(): BooleanArray::operator[] returns nullopt for null entries,
   which is_filtered() treats as "no filter" (not filtered), letting null docs
   through. Fix: use value_or(false) to treat null as "not matched".

2. is_matched_by_forward_filter(): reads BooleanScalar.value without checking
   is_valid, which is UB for null scalars. Fix: check is_valid first.

Closes #409

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-05-18 20:25:06 +08:00
.github chore: clang-tidy check fast quits if no cpp files changed (#396) 2026-05-13 10:03:26 +08:00
cmake feat(test): enable parallel tests (#384) 2026-05-12 23:14:17 +08:00
examples minor: fix c example cmake and build c++ dynamic lib (#347) 2026-04-20 22:08:34 +08:00
python enhance: release gil lock in collection bindings (#363) 2026-05-14 10:15:49 +08:00
scripts ci: refact android ci (#330) 2026-04-15 15:37:51 +08:00
src fix: nullable scalar field filter leaks null documents without inverted index (#410) 2026-05-18 20:25:06 +08:00
tests fix: nullable scalar field filter leaks null documents without inverted index (#410) 2026-05-18 20:25:06 +08:00
thirdparty build(thirdparty): replace CRoaring submodule with amalgamation (~292M -> ~1MB) (#381) 2026-05-12 14:55:08 +08:00
tools refactor: remove hamming metric (#365) 2026-05-18 19:31:20 +08:00
.clang-format Initial commit 2025-12-30 11:02:17 +08:00
.clang-tidy feat(ci): integrate clang-tidy for changed C/C++ files (#116) 2026-04-20 20:14:21 +08:00
.gitignore feat(ci): integrate clang-tidy for changed C/C++ files (#116) 2026-04-20 20:14:21 +08:00
.gitmodules build(thirdparty): replace CRoaring submodule with amalgamation (~292M -> ~1MB) (#381) 2026-05-12 14:55:08 +08:00
.pre-commit-config.yaml chore: enable the conventional-pre-commit run sucess and update to latest version (#111) 2026-02-25 18:03:20 +08:00
CMakeLists.txt feat: add iOS build support (#321) 2026-04-09 16:14:26 +08:00
CODE_OF_CONDUCT.md Initial commit 2025-12-30 11:02:17 +08:00
CONTRIBUTING.md doc: add v0.3.0 release note (#312) 2026-04-03 15:47:19 +08:00
LICENSE Initial commit 2025-12-30 11:02:17 +08:00
README.md minor: update readme for v0.4.0 (#389) 2026-05-09 11:40:30 +08:00
README_CN.md minor: update readme for v0.4.0 (#389) 2026-05-09 11:40:30 +08:00
pyproject.toml feat(test): enable parallel tests (#384) 2026-05-12 23:14:17 +08:00

README.md

English | 中文

zvec logo

Code Coverage Main License PyPI Release Python Versions npm Release

alibaba%2Fzvec | Trendshift

🚀 Quickstart | 🏠 Home | 📚 Docs | 📊 Benchmarks | 🔎 DeepWiki | 🎮 Discord | 🐦 X (Twitter)

Zvec is an open-source, in-process vector database — lightweight, lightning-fast, and designed to embed directly into applications. Battle-tested within Alibaba Group, it delivers production-grade, low-latency and scalable similarity search with minimal setup.

[!Important] 🚀 v0.4.0 (May 9, 2026)

  • Dart/Flutter SDK: Published the official zvec Flutter package with FFI bindings. Supports Android (arm64-v8a) and iOS (arm64) — no manual native compilation required.
  • iOS Build Support: Added support for building on iOS platforms, expanding cross-platform coverage.
  • Enlarged topK Limit: Relaxed the upper bound on topK to support larger-scale recall scenarios.
  • Bug Fixes: SQ8 quantizer recall drop; Windows path handling; sparse vector index ordering.

👉 Read the Release Notes | View Roadmap 📍

💫 Features

  • Blazing Fast: Searches billions of vectors in milliseconds.
  • Simple, Just Works: Install and start searching in seconds. Pure local, no servers, no config, no fuss.
  • Dense + Sparse Vectors: Work with both dense and sparse embeddings, with native support for multi-vector queries in a single call.
  • Hybrid Search: Combine semantic similarity with structured filters for precise results.
  • Durable Storage: Write-ahead logging (WAL) guarantees persistence — data is never lost, even on process crash or power failure.
  • Concurrent Access: Multiple processes can read the same collection simultaneously; writes are single-process exclusive.
  • Runs Anywhere: As an in-process library, Zvec runs wherever your code runs — notebooks, servers, CLI tools, or even edge devices.

📦 Installation

Python

Requirements: Python 3.10 - 3.14

pip install zvec

Node.js

npm install @zvec/zvec

Supported Platforms

  • Linux (x86_64, ARM64)
  • macOS (ARM64)
  • Windows (x86_64)

🛠️ Building from Source

If you prefer to build Zvec from source, please check the Building from Source guide.

One-Minute Example

import zvec

# Define collection schema
schema = zvec.CollectionSchema(
    name="example",
    vectors=zvec.VectorSchema("embedding", zvec.DataType.VECTOR_FP32, 4),
)

# Create collection
collection = zvec.create_and_open(path="./zvec_example", schema=schema)

# Insert documents
collection.insert([
    zvec.Doc(id="doc_1", vectors={"embedding": [0.1, 0.2, 0.3, 0.4]}),
    zvec.Doc(id="doc_2", vectors={"embedding": [0.2, 0.3, 0.4, 0.1]}),
])

# Search by vector similarity
results = collection.query(
    zvec.VectorQuery("embedding", vector=[0.4, 0.3, 0.3, 0.1]),
    topk=10
)

# Results: list of {'id': str, 'score': float, ...}, sorted by relevance
print(results)

📈 Performance at Scale

Zvec delivers exceptional speed and efficiency, making it ideal for demanding production workloads.

Zvec Performance Benchmarks

For detailed benchmark methodology, configurations, and complete results, please see our Benchmarks documentation.

🤝 Join Our Community

💬 DingTalk 📱 WeChat 🎮 Discord X (Twitter)
DingTalk QR Code WeChat QR Code Discord X (formerly Twitter) Follow
Scan to join Scan to join Click to join Click to follow

❤️ Contributing

We welcome and appreciate contributions from the community! Whether you're fixing a bug, adding a feature, or improving documentation, your help makes Zvec better for everyone.

Check out our Contributing Guide to get started!