Alibaba lightweight in-process vector database
Go to file
Jalin Wang 78f08bd0cc
feat: auto scalable segment meta section of mmap files (#67)
The underly core module only support a fixed-size segment due to the index file format where the section of segments meta limit the max num of data segments. This commit makes the meta section dynamically allocated and organized in a link-table manner to support much more segments.


* revert local_builder's meta cap setting

* feat: auto scalable segment meta section of mmap files

* Revert "feat: auto scalable segment meta section of mmap files"

This reverts commit 82c11e451544b912b901be33fae79e964a487b1f.

* v2: don't change to absolute offset
2026-02-06 14:58:56 +08:00
.github feat(ci): ci is not triggered when modifying the md file (#55) 2026-02-02 22:18:08 +08:00
cmake feat: refact cpp sdk (#27) 2026-01-21 20:05:42 +08:00
examples/c++ feat: support core cpp sdk (#30) 2026-01-26 20:01:44 +08:00
python fix: remove unnecessary column_name param from the AddColumn API (#59) 2026-02-04 14:44:13 +08:00
scripts Initial commit 2025-12-30 11:02:17 +08:00
src feat: auto scalable segment meta section of mmap files (#67) 2026-02-06 14:58:56 +08:00
tests fix: remove unnecessary column_name param from the AddColumn API (#59) 2026-02-04 14:44:13 +08:00
thirdparty feat: refact cpp sdk (#27) 2026-01-21 20:05:42 +08:00
tools feat: auto scalable segment meta section of mmap files (#67) 2026-02-06 14:58:56 +08:00
.clang-format Initial commit 2025-12-30 11:02:17 +08:00
.gitignore feat: refact cpp sdk (#27) 2026-01-21 20:05:42 +08:00
.gitmodules Initial commit 2025-12-30 11:02:17 +08:00
.pre-commit-config.yaml chore:add git commit msg and branch name in pre-commit and modify org (#1) 2025-12-30 15:39:05 +08:00
CMakeLists.txt feat: refact cpp sdk (#27) 2026-01-21 20:05:42 +08:00
CODE_OF_CONDUCT.md Initial commit 2025-12-30 11:02:17 +08:00
CONTRIBUTING.md doc: update the CMake minimum version requirement (#57) 2026-02-03 17:17:43 +08:00
LICENSE Initial commit 2025-12-30 11:02:17 +08:00
README.md feat(docs): add join us (#54) 2026-02-02 18:49:02 +08:00
pyproject.toml chore(ci): modify setuptools cm for test pypi and support cancel in progress (#21) 2026-01-09 17:48:38 +08:00

README.md

zvec logo

Linux x64 CI macOS ARM64 CI Code Coverage PyPI Release
Python Versions License

🚀 Quickstart | 🏠 Home | 📚 Docs | 📊 Benchmarks | 🎮 Discord | 🐦 X (Twitter)

Zvec is an open-source, in-process vector database — lightweight, lightning-fast, and designed to embed directly into applications. Built on Proxima (Alibaba's battle-tested vector search engine), it delivers production-grade, low-latency, scalable similarity search with minimal setup.

💫 Features

  • Blazing Fast: Searches billions of vectors in milliseconds.
  • Simple, Just Works: Install with pip install zvec and start searching in seconds. No servers, no config, no fuss.
  • Dense + Sparse Vectors: Work with both dense and sparse embeddings, with native support for multi-vector queries in a single call.
  • Hybrid Search: Combine semantic similarity with structured filters for precise results.
  • Runs Anywhere: As an in-process library, Zvec runs wherever your code runs — notebooks, servers, CLI tools, or even edge devices.

📦 Installation

Install Zvec from PyPI with a single command:

pip install zvec

Requirements:

  • Python 3.10 - 3.12
  • Supported platforms:
    • Linux (x86_64)
    • macOS (ARM64)

If you prefer to build Zvec from source, please check the Building from Source guide.

One-Minute Example

import zvec

# Define collection schema
schema = zvec.CollectionSchema(
    name="example",
    vectors=zvec.VectorSchema("embedding", zvec.DataType.VECTOR_FP32, 4),
)

# Create collection
collection = zvec.create_and_open(path="./zvec_example", schema=schema,)

# Insert documents
collection.insert([
    zvec.Doc(id="doc_1", vectors={"embedding": [0.1, 0.2, 0.3, 0.4]}),
    zvec.Doc(id="doc_2", vectors={"embedding": [0.2, 0.3, 0.4, 0.1]}),
])

# Search by vector similarity
results = collection.query(
    zvec.VectorQuery("embedding", vector=[0.4, 0.3, 0.3, 0.1]),
    topk=10
)

# Results: list of {'id': str, 'score': float, ...}, sorted by relevance
print(results)

📈 Performance at Scale

Zvec delivers exceptional speed and efficiency, making it ideal for demanding production workloads.

Zvec Performance Benchmarks

For detailed benchmark methodology, configurations, and complete results, please see our Benchmarks documentation.

🤝 Join Our Community

Stay updated and get support — scan or click:

💬 DingTalk
📱 WeChat
🎮 Discord
Join Server
🐦 X (Twitter)
Follow @zvec_ai

❤️ Contributing

We welcome and appreciate contributions from the community! Whether you're fixing a bug, adding a feature, or improving documentation, your help makes Zvec better for everyone.

Check out our Contributing Guide to get started!