A compact, fast vector search engine written in Go: HNSW and brute-force indexes, int8 quantization, hybrid dense + sparse search, and crash-safe persistence, served over REST and gRPC from a single static binary.
GoVec is inspired by Qdrant, Weaviate and Pinecone, but deliberately smaller. It targets prototyping, local semantic search, and ML apps that need vector search on a single node, without running a distributed platform. It finds the same neighbours as Chroma and Qdrant at the same settings and has the fastest inserts of the three; its queries are slower at 100k vectors, and in memory it uses more RAM (see Performance).
| Repository | Description |
|---|---|
| govec (this repo) | The server: REST + gRPC vector search engine |
| govec-python | Python SDK over REST or gRPC (pip install govec) |
| govec-bench | Benchmark harness comparing GoVec against Chroma and Qdrant |
- Two index types. HNSW for approximate nearest-neighbour search, using the paper's
heuristic neighbour selection and separate
ef_construction/ef_searchknobs. Exact brute-force scan for small collections. - Scalar quantization. Optional int8 storage stores vectors in a quarter of the space; measured at -61% disk and -31% RAM at 100k vectors. Vectors are range-checked before encoding, so quantization can't silently corrupt them.
- Hybrid search. Combine a dense embedding with a sparse keyword vector in one query.
- Metadata filtering. Exact-match filters on arbitrary JSON metadata, with an optional inverted index that pushes selective filters into the graph traversal.
- Durability. A write-ahead log plus atomic snapshots. The server replays the WAL on startup and refuses to start on a log it can't read, rather than silently dropping writes.
- Cosine and Euclidean distance on both index types.
- REST and gRPC with identical semantics, including streaming batch insert over gRPC. A parity test keeps the two from drifting.
- Operable. One binary, YAML config with an environment override for every field, optional bearer-token auth, structured logging with request IDs, and pprof on a separate port.
With Docker:
docker run -p 9697:9697 -v govec-data:/data ghcr.io/pradyothsp/govec:latestOr from source (Go 1.26+):
git clone https://github.com/Pradyothsp/govec.git
cd govec
go run ./cmd/serverThe server listens on http://localhost:9697.
# Insert a vector with metadata
curl -X POST http://localhost:9697/api/v1/vectors \
-H "Content-Type: application/json" \
-d '{"id": "doc-1", "vector": [0.1, 0.9, 0.2], "metadata": {"lang": "en", "year": 2024}}'
# Insert several at once
curl -X POST http://localhost:9697/api/v1/vectors/batch \
-H "Content-Type: application/json" \
-d '{"vectors": [
{"id": "doc-2", "vector": [0.8, 0.1, 0.3], "metadata": {"lang": "de"}},
{"id": "doc-3", "vector": [0.2, 0.8, 0.1], "metadata": {"lang": "en"}}
]}'
# Find the 2 nearest neighbours
curl -X POST http://localhost:9697/api/v1/vectors/search \
-H "Content-Type: application/json" \
-d '{"vector": [0.1, 0.85, 0.15], "k": 2}'{
"success": true,
"data": [
{"id": "doc-1", "score": 0.999, "meta": {"lang": "en", "year": 2024}},
{"id": "doc-3", "score": 0.991, "meta": {"lang": "en"}}
]
}Add a metadata filter (every key must match exactly):
curl -X POST http://localhost:9697/api/v1/vectors/search \
-H "Content-Type: application/json" \
-d '{"vector": [0.1, 0.85, 0.15], "k": 5, "filter": {"lang": "de"}}'For hybrid search, add a sparse_vector ({"indices": [...], "values": [...]}) to both the
insert and the query. Scores are blended as 0.7 × dense + 0.3 × sparse. Hybrid scoring needs
engine.enable_hybrid_search: true, which the Docker image's config sets; without it a sparse
vector is stored but doesn't affect scores.
pip install govec # or: uv add govecSee govec-python for the client API. A minimal
Go client over plain net/http is in examples/client.
Every REST response uses the same envelope:
{"success": bool, "data": ..., "error": "..."}.
| Method | Path | Description |
|---|---|---|
GET |
/health |
Liveness check |
POST |
/api/v1/vectors |
Insert or update (upsert) one vector |
POST |
/api/v1/vectors/batch |
Insert many vectors; reports per-item errors |
GET |
/api/v1/vectors/:id |
Fetch a vector and its metadata |
DELETE |
/api/v1/vectors/:id |
Delete a vector |
POST |
/api/v1/vectors/search |
k-NN search with optional filter and sparse_vector |
GET |
/api/v1/stats |
Vector count |
GET |
/api/v1/info |
Server version and engine configuration (index type, quantization, metric, dimensions) |
POST |
/api/v1/admin/flush |
Write a snapshot now |
POST |
/api/v1/admin/reset |
Delete all vectors |
- Swagger UI: interactive docs at
http://localhost:9697/swagger/index.html. - gRPC: set
grpc.enabled: trueto serve on port 9698. The service definition is inproto/govec/v1. It offers the same operations, plus client-streamingBatchInsert. - Auth: set
server.api_key, and every/api/v1/*request must sendAuthorization: Bearer <key>./healthstays open for probes. - Errors: an invalid vector (empty, or the wrong dimension for the index) returns HTTP 400 /
gRPC
InvalidArgument.
GoVec reads config.yaml from the working directory, or the file named by GOVEC_CONFIG_PATH.
In the Docker image, mount your own file at /app/config.yaml. All data lives in the /data
volume, and gRPC needs -p 9698:9698 once enabled.
Every field also has an environment variable, and the environment wins. The variable name is
GOVEC_ plus the YAML key (storage.wal_path → GOVEC_WAL_PATH). The section is included
only where the key alone would be ambiguous (GOVEC_SERVER_PORT vs GOVEC_GRPC_PORT).
The settings you're most likely to change:
| Setting | Env var | Default | |
|---|---|---|---|
engine.index_type |
GOVEC_INDEX_TYPE |
brute |
hnsw for collections beyond ~10k vectors |
engine.quantization |
GOVEC_QUANTIZATION |
none |
scalar for int8 storage |
engine.distance_metric |
GOVEC_DISTANCE_METRIC |
cosine |
or euclidean |
engine.enable_metadata_index |
GOVEC_ENABLE_METADATA_INDEX |
false |
faster filtered search, more RAM |
storage.data_path |
GOVEC_DATA_PATH |
./govec_data.bin |
snapshot file |
storage.wal_path |
GOVEC_WAL_PATH |
./govec.wal |
write-ahead log |
server.api_key |
GOVEC_API_KEY |
empty (auth off) | bearer token for /api/v1/* |
grpc.enabled |
GOVEC_GRPC_ENABLED |
false |
serve gRPC on grpc.port (9698) |
config.example.yaml documents every option, including the HNSW tuning
parameters (hnsw_m, hnsw_ef_search, hnsw_ef_construction).
Recall tip: both ef defaults are 100, as in Chroma and Qdrant. On large datasets, if recall
matters more than insert speed, raise hnsw_ef_construction (to 200, say: at 100k text
embeddings that looked worth about half a point to a point of recall@10, at slower inserts) or
hnsw_ef_search, which costs query latency instead and can be changed without rebuilding.
Memory tip: set GOMEMLIMIT to 2–3x your expected live data, kept under the container's
memory limit. Under sustained load it cuts resident memory substantially (-46% in benchmarks)
with no latency cost. It isn't set by default because the right value depends on dataset size.
REST (Gin) gRPC
\ /
Engine interface
/ \
VectorIndex[T] HNSWIndex[T] T = float32 | int8
(brute) (ANN)
\ /
WAL + atomic GOB snapshots (+ optional mmap vector store)
- One engine, two transports. REST and gRPC are thin adapters over a single
Engineinterface. Both share input validation and error classification, so they can't disagree. - Generic storage. Both indexes are written once over
T = float32 | int8. Quantization is a config switch, not a separate code path. Under cosine, vectors are unit-normalized before int8 encoding. Under Euclidean, out-of-range vectors are rejected, because normalizing them would change the distance. - HNSW. A fork of coder/hnsw, made generic and extended
with
SELECT-NEIGHBORS-HEURISTIC, a2×Mbase layer, and greedy descent through upper layers on insert. Batch inserts run each round's graph searches in parallel, then apply the mutations serially, so no read ever overlaps a write. - Durability. Every write is appended to the WAL (batch inserts share a single fsync). Snapshots are written to a temp file and atomically renamed. On startup, GoVec loads the snapshot and replays the WAL past it. It refuses to start if either can't be read in full.
Measured with govec-bench on 100k 1536-dimensional
OpenAI text embeddings (DBpedia), one database at a time in Docker with 2 CPUs and 2 GB each.
Every database builds HNSW with M=16, ef_construction=100 and searches with ef 100, cosine;
every number for a database comes from one index, and recall is graded against exact cosine
neighbours.
Medians of three runs (October 7, 2026, GoVec 0.2.1) on a MacBook with Docker Desktop.
| GoVec | GoVec (int8) | GoVec (mmap)† | Chroma | Qdrant | |
|---|---|---|---|---|---|
| Query latency, k=10 (mean) | 6.23 ms | 6.45 ms | 7.93 ms | 5.58 ms | 4.61 ms |
| Query latency over gRPC, k=10 (mean)‡ | 4.10 ms | ||||
| Single insert (mean) | 5.06 ms | 5.16 ms | 5.24 ms | 12.42 ms | 6.04 ms |
| Batch insert, per vector | 1.96 ms | 2.04 ms | 2.12 ms | 2.21 ms | 7.62 ms |
| Recall@10 | 97.7% | 88.4% | 98.1% | 98.1% | 98.6% |
| Recall@50 | 96.2% | 89.1% | 96.4% | 96.4% | 97.8% |
| Restart, empty (median) | 9.0 ms | 10.3 ms | 10.5 ms | 67.2 ms | 12.0 ms |
| Restart with the data loaded (median) | 930 ms | 553 ms | 474 ms | 209 ms | 81 ms |
| RAM | 1,329 MB | 466 MB | 817 MB | 708 MB | 271 MB* |
| Disk (allocated) | 632 MB | 171 MB | 632 MB | 659 MB | 774 MB |
* Qdrant keeps its vectors on disk by default and lets the OS cache only part of them, so its RAM isn't comparable with GoVec's and Chroma's, which hold every vector in memory.
† GoVec with storage.enable_mmap: true (and engine.dimensions set): vectors in memory-mapped
files instead of the Go heap; RAM is what's in use, including those files' cached pages. Its disk
is the same as in memory: each vector is stored once, in the mapped files instead of the snapshot.
‡ The same GoVec index, queried over gRPC instead of REST. Every other figure, Chroma's and Qdrant's included, is over HTTP; Qdrant also has a gRPC API, not measured here. Compare this row with GoVec's REST figure, not with the other databases.
All three find the same neighbours at the same ef, within the point or so that recall moves between index builds. GoVec has the fastest single and batch inserts. Over REST its queries are the slowest at this size, 1.1x Chroma and 1.4x Qdrant: at 10k vectors it was level with Qdrant, and its latency grows faster with the graph, since distance computation doesn't use SIMD yet and neighbour lists are pointers rather than flat ID arrays. Over gRPC the same queries take a third less time. It restarts empty in about 9 ms but takes about 1 s to answer with 100k vectors loaded, because it reads its whole snapshot back first. In memory it holds about 1.9x Chroma's RAM. Memory-mapped (the mmap column), memory in use falls 38%, restarts with data take half as long and recall is the same, but queries are about 27% slower. int8 recall on text embeddings is low for now. SIFT results, the 10k numbers, an ef sweep, the method and how to reproduce are in the govec-bench repo.
GoVec is pre-1.0 (v0.x). The REST and gRPC APIs are stable in practice but may change before 1.0. It is built for prototyping, local semantic search, and embedding in ML applications on a single node.
Limitations to know about:
- Single node, in-memory first. No sharding or replication. The dataset must fit in RAM unless you enable the mmap vector store.
- Filtered HNSW search can return fewer than
kresults when the filter is selective, especially withenable_metadata_index: false. The brute-force index is always exact. - Hybrid search scans all sparse vectors. Sparse scoring is O(n) per query, so it suits small and medium collections.
- Filters are exact-match only. Keys are ANDed together, with no range or
ORoperators.
Planned work and deliberate non-goals are in docs/roadmap.md.
Bug reports and pull requests are welcome. See CONTRIBUTING.md for development setup, the test and lint workflow, and project conventions.
GoVec is licensed under the Apache License 2.0.
The HNSW implementation in internal/hnsw is derived from
coder/hnsw (CC0-1.0).