Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 2 additions & 0 deletions AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -16,6 +16,7 @@ zvec-php/
├── src/ZVecGroupByVectorQuery.php # Group-by vector query builder
├── src/ZVecSchema.php # Schema definition (field types, metrics, vectors)
├── src/ZVecDoc.php # Document handle (getters/setters, serialization)
├── src/ZVecDocIterator.php # Full-scan document iterator (iterDocs)
├── src/ZVecReRanker.php # Base re-ranker class
├── src/ZVecRerankedDoc.php # Reranked document class
├── src/ZVecRrfReRanker.php # RRF re-ranker
Expand Down Expand Up @@ -448,6 +449,7 @@ and runner output.
- `__destruct()` calls `close()` automatically, and never throws — on failure it falls back to dropping the handle so the C++ object is not leaked
- After `destroy()`, any method call causes **segfault** (handle invalidated)
- A **failed** `close()` or `destroy()` must leave the object open and usable. Do not free the handle before the status is known: the C++ object is only safe to drop once upstream says the operation succeeded, or once it says the collection is closed regardless
- While a `ZVecDocIterator` is open, `close()` / `destroy()` / DDL / `optimize()` throw `FAILED_PRECONDITION`. Close iterators first — they auto-close when exhausted

### Memory Leak Regression Tests

Expand Down
11 changes: 11 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -9,6 +9,17 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0

### Added

- **Full-collection scans: `ZVec::iterDocs()` / `ZVecDocIterator`** (#226)
- `iterDocs(?array $outputFields = null, bool $includeVector = true): ZVecDocIterator` returns a forward-only iterator over every document, backed by an isolated snapshot taken at call time. Memory use is constant regardless of collection size. Mirrors the Python SDK's `Collection.iter_docs()` and upstream's `Collection::create_iterator()` (zvec v0.7.0).
- This closes the gap for export, backup and migration: `fetch()` needs primary keys, `queryByFilter()` needs both a filter and a `topk`, and neither can walk a whole collection.
- The iterator closes itself when exhausted, so a completed `foreach` needs no cleanup. `close()` is idempotent for breaking out early, and `isClosed()` reports the state. `rewind()` after the first advance throws, by design — a second `getIterator()` on the same object would otherwise silently yield nothing.
- **While an iterator is open, `close()`, `destroy()`, schema DDL and `optimize()` throw `ZVecException` with code 5 (`FAILED_PRECONDITION`)**; the collection stays open and usable, and writes, `flush()` and queries are unaffected. This depends on the `close()` / `destroy()` behaviour from #222 — a rejected `destroy()` has to leave the handle valid.
- The adapter gives each iterator its own `shared_ptr` to the collection and releases it *after* the iterator. Upstream requires the collection to outlive its iterators, and `~CollectionImpl` blocks until every iterator is closed, so the wrong release order would deadlock a single-threaded PHP process at shutdown.
- `ZVecDocIterator` holds a PHP reference to its `ZVec` object, so the collection cannot be destroyed out from under a live iterator; `tests/test_doc_iterator_lifecycle.phpt` covers dropping the collection variable mid-iteration and the destructor path.
- `outputFields` accepts scalar fields only (`null` = all, `[]` = primary key only); unknown or vector names are rejected upstream. Unset nullable fields come back as `null`, matching `fetch()` since #192 — but only for the fields actually requested, so `outputFields: ['id']` does not make `hasField('weight')` true.
- FFI: `zvec_doc_iterator_t`, `zvec_collection_create_iterator()`, `zvec_doc_iterator_next()`, `zvec_doc_iterator_free()`.
- Tests: `tests/test_doc_iterator.phpt` (empty collection, full scan, vectors on/off, output fields, primary-key-only, nullable normalization, deleted documents skipped, snapshot isolation, auto-close, early close, forward-only, validation, closed collection), `tests/test_doc_iterator_lifecycle.phpt` (the `FAILED_PRECONDITION` guards, writes during iteration, collection lifetime, destructor safety), `tests/test_doc_iterator_shutdown.phpt` (clean process exit with an open iterator — an infinite hang rather than a crash if the release order is wrong), and `tests/test_memory_doc_iterator.phpt` (no handle or C string array leak).

- **Vamana `two_pass_build` and query prefetch** (#220)
- `ZVecIndexParams::forVamana()` takes a trailing `bool $twoPassBuild = false`, which runs a second full-graph Vamana construction pass for better graph quality at the cost of build time. Upstream and Python both default to `false`. Last and optional, so existing positional calls are unaffected.
- `ZVecVectorQuery::setVamanaPrefetch(int $prefetchOffset, int $prefetchLines)` tunes software prefetch during the Vamana graph search, the counterpart of the existing `setHnswPrefetch()`. An offset of `0` disables prefetch; `0` lines means "derive from vector size". Both reject negative values.
Expand Down
33 changes: 33 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -259,6 +259,7 @@ $collection->updateBatch(ZVecDoc ...$docs): array // Returns per-doc status
$collection->delete(string ...$pks): void
$collection->deleteByFilter(string $filter): void
$collection->fetch(string ...$pks): ZVecDoc[] // also fetch(array $pks, ?array $outputFields = null, includeVector: bool = true)
$collection->iterDocs(?array $outputFields = null, bool $includeVector = true): ZVecDocIterator // full scan over a snapshot; foreach ($it as $pk => $doc)
// named `outputFields:` supported; named-arg `pks:` and mixing scalar PKs with an array are rejected with a hint

// Search
Expand Down Expand Up @@ -329,6 +330,36 @@ $schema::METRIC_MIPSL2 = 4 // Modified Inner Product with L2
// that emit E_USER_DEPRECATED warnings.
```

### ZVecDocIterator

Returned by `iterDocs()`. Scans every document with constant memory, over a
snapshot taken when the iterator was created — so documents written afterwards
are not visible to it. This is the full-scan path: no primary keys, no filter,
no `topk`. Useful for export, backup and migration, which `fetch()` (needs PKs)
and `queryByFilter()` (needs a filter and a topk) cannot serve.

```php
foreach ($collection->iterDocs() as $pk => $doc) {
fwrite($fh, json_encode(['pk' => $pk, 'id' => $doc->getInt64('id')]) . "\n");
}
```

- **Forward-only.** `rewind()` after the first advance throws. Call
`iterDocs()` again to scan afresh.
- **Auto-closes on exhaustion.** A completed `foreach` needs no cleanup. Call
`close()` explicitly when breaking out early; it is idempotent, and
`isClosed()` reports the state.
- **Guards the collection.** While one is open, `close()`, `destroy()`,
`addColumn*()` / `alterColumn()` / `dropColumn()` and `optimize()` throw
`ZVecException` with code `5` (`FAILED_PRECONDITION`). Writes, `flush()` and
queries keep working. Close iterators first.
- **Seals a segment.** On a writable collection each call seals the current
writing segment, which may create a small one. A read-only collection is
scanned without writing.
- `outputFields` accepts scalar fields only — `null` means all of them, `[]`
means primary key only. Vector or unknown names are rejected by upstream.
Nullable fields that are unset come back as `null`, matching `fetch()`.

### ZVecDoc

```php
Expand Down Expand Up @@ -868,6 +899,7 @@ zvec-php/
│ ├── ZVecGroupByVectorQuery.php # Group-by vector query builder
│ ├── ZVecSchema.php # Schema definition
│ ├── ZVecDoc.php # Document handle
│ ├── ZVecDocIterator.php # Full-scan document iterator
│ ├── ZVecReRanker.php # Base re-ranker interface
│ ├── ZVecRerankedDoc.php # Reranked document class
│ ├── ZVecRrfReRanker.php # RRF re-ranker
Expand Down Expand Up @@ -933,6 +965,7 @@ See `tasks/done/` for detailed planning documents.
- [x] Group-by vector query builder (`ZVecGroupByVectorQuery`)
- [x] Version API (`getVersion()`, `checkVersion()`)
- [x] DiskANN I/O backend introspection (`getIoBackendType()`, `getIoBackendDescription()`)
- [x] Full-collection scans (`iterDocs()` / `ZVecDocIterator`)
- [x] `allowedBasePath` security restriction in `init()`
- [x] Verbose error details with file/line info
- [x] Collection lifecycle options via `getOptions()`
Expand Down
138 changes: 138 additions & 0 deletions ffi/zvec_ffi.cc
Original file line number Diff line number Diff line change
@@ -1,6 +1,7 @@
#include "zvec_ffi.h"

#include <array>
#include <algorithm>
#include <cstring>
#include <optional>
#include <string>
Expand Down Expand Up @@ -2820,6 +2821,143 @@ zvec_status_t zvec_collection_fetch(zvec_collection_t coll, const char** pks, in
return ok_status();
}

// --- Doc iterator ---

namespace {
struct DocIteratorHolder {
// Declaration order is destruction order reversed, so `iterator` is
// released first -- it hands its slot back to the collection -- and
// `collection` last. Upstream requires the collection to outlive its
// iterators (create_iterator's release_slot captures the raw pointer), and
// ~CollectionImpl blocks until every iterator is closed, which would
// deadlock a single-threaded PHP process. Keeping our own shared_ptr copy
// also means zvec_collection_free() on the PHP side cannot pull the
// collection out from under a live iterator.
std::shared_ptr<Collection> collection;
std::vector<std::string> nullable_fields; // requested nullable fields only
DocIterator::Ptr iterator;
};
} // namespace

zvec_status_t zvec_collection_create_iterator(zvec_collection_t coll,
int has_output_fields,
const char** output_fields,
int output_field_count,
int include_vector,
zvec_doc_iterator_t* out) {
if (!out) {
zvec_status_t st = {1, "null out pointer"};
SET_FFI_ERROR(st);
return st;
}
*out = nullptr;
if (!coll) {
zvec_status_t st = {1, "null handle"};
SET_FFI_ERROR(st);
return st;
}
if (output_field_count < 0 || (output_field_count > 0 && !output_fields)) {
zvec_status_t st = {1, "invalid output fields"};
SET_FFI_ERROR(st);
return st;
}

// Copy the collection out of the registry so the iterator owns a reference.
std::shared_ptr<Collection> owner;
{
std::unique_lock lock(g_collections_mutex);
auto& reg = collections_registry();
auto it = reg.find(static_cast<Collection*>(coll));
if (it == reg.end()) {
zvec_status_t st = {1, "collection handle is not open"};
SET_FFI_ERROR(st);
return st;
}
owner = it->second;
}

IteratorOptions opts;
opts.include_vector_ = include_vector != 0;
if (has_output_fields) {
std::vector<std::string> fields;
fields.reserve(output_field_count);
for (int i = 0; i < output_field_count; i++) {
if (!output_fields[i]) {
zvec_status_t st = {1, "null output field"};
SET_FFI_ERROR(st);
return st;
}
fields.emplace_back(output_fields[i]);
}
// Assigned even when empty: upstream reads an empty vector as "PK only".
opts.output_fields_ = std::move(fields);
}

// Since #192 fetch() reports an absent nullable field as present-and-null.
// Iterate consistently, but only for the fields actually requested -- with
// outputFields: ['id'], hasField('weight') must stay false.
std::vector<std::string> nullable_fields;
auto schema_res = owner->schema();
if (schema_res.has_value()) {
for (const auto& field : schema_res.value().fields()) {
if (field && field->nullable()) {
if (!has_output_fields) {
nullable_fields.push_back(field->name());
} else if (opts.output_fields_ &&
std::find(opts.output_fields_->begin(), opts.output_fields_->end(), field->name()) !=
opts.output_fields_->end()) {
nullable_fields.push_back(field->name());
}
}
}
}

auto res = owner->create_iterator(opts);
if (!res.has_value()) {
return MAKE_STATUS(res.error());
}
auto* h = new DocIteratorHolder{std::move(owner), std::move(nullable_fields), std::move(res).value()};
*out = static_cast<zvec_doc_iterator_t>(h);
return ok_status();
}

zvec_status_t zvec_doc_iterator_next(zvec_doc_iterator_t it, zvec_doc_t* out_doc) {
if (!out_doc) {
zvec_status_t st = {1, "null out pointer"};
SET_FFI_ERROR(st);
return st;
}
*out_doc = nullptr;
if (!it) {
zvec_status_t st = {1, "null handle"};
SET_FFI_ERROR(st);
return st;
}
auto* h = static_cast<DocIteratorHolder*>(it);
auto res = h->iterator->next();
if (!res.has_value()) {
return MAKE_STATUS(res.error());
}
if (!res.value()) {
return ok_status(); // end of iteration
}
// Copy the doc: PHP frees documents with zvec_doc_free, which deletes a
// Doc*, so the shared_ptr cannot be handed out.
auto* doc = new Doc(*res.value());
for (const auto& name : h->nullable_fields) {
if (!doc->has(name)) {
doc->set_null(name);
}
}
*out_doc = static_cast<zvec_doc_t>(doc);
return ok_status();
}

void zvec_doc_iterator_free(zvec_doc_iterator_t it) {
if (!it) return;
delete static_cast<DocIteratorHolder*>(it);
}

// --- Query ---

zvec_status_t zvec_collection_query(zvec_collection_t coll, const char* field_name,
Expand Down
27 changes: 27 additions & 0 deletions ffi/zvec_ffi.h
Original file line number Diff line number Diff line change
Expand Up @@ -384,6 +384,33 @@ zvec_status_t zvec_collection_fetch(zvec_collection_t coll, const char** pks, in
int include_vector,
zvec_query_result_t* result);

// Document iterator (zvec v0.7.0 Collection::create_iterator)
//
// Full scan over an isolated snapshot taken at creation time. The iterator
// holds its own reference to the collection, so zvec_collection_free() may be
// called before zvec_doc_iterator_free(). While an iterator is open,
// close/destroy/DDL/optimize on the collection return code 5
// (FAILED_PRECONDITION) and the collection stays open.
typedef void* zvec_doc_iterator_t;

// has_output_fields = 0: return all scalar fields (output_fields ignored).
// has_output_fields = 1: return only the listed scalar fields; count 0 (and
// output_fields may be NULL) returns no scalar fields, only the PK.
// include_vector: 1 = include vector fields, 0 = skip them.
// On error *out is set to NULL.
zvec_status_t zvec_collection_create_iterator(zvec_collection_t coll,
int has_output_fields,
const char** output_fields,
int output_field_count,
int include_vector,
zvec_doc_iterator_t* out);
// On success *out_doc is a new document owned by the caller (free with
// zvec_doc_free), or NULL at end of iteration. On error *out_doc is NULL.
zvec_status_t zvec_doc_iterator_next(zvec_doc_iterator_t it, zvec_doc_t* out_doc);
// Closes the iterator and releases its snapshot and its collection
// reference. NULL is a no-op.
void zvec_doc_iterator_free(zvec_doc_iterator_t it);

// Query
zvec_status_t zvec_collection_query(zvec_collection_t coll, const char* field_name,
const float* query_vector, uint32_t dim,
Expand Down
5 changes: 5 additions & 0 deletions ffi/zvec_ffi_php.h
Original file line number Diff line number Diff line change
Expand Up @@ -270,6 +270,11 @@ zvec_status_t zvec_collection_fetch(zvec_collection_t coll, const char** pks, in
int include_vector,
zvec_query_result_t* result);

typedef void* zvec_doc_iterator_t;
zvec_status_t zvec_collection_create_iterator(zvec_collection_t coll, int has_output_fields, const char** output_fields, int output_field_count, int include_vector, zvec_doc_iterator_t* out);
zvec_status_t zvec_doc_iterator_next(zvec_doc_iterator_t it, zvec_doc_t* out_doc);
void zvec_doc_iterator_free(zvec_doc_iterator_t it);

zvec_status_t zvec_collection_query(zvec_collection_t coll, const char* field_name,
const float* query_vector, uint32_t dim,
int topk, int include_vector,
Expand Down
62 changes: 62 additions & 0 deletions src/ZVec.php
Original file line number Diff line number Diff line change
Expand Up @@ -18,6 +18,7 @@
require_once __DIR__ . '/ZVecGroupByVectorQuery.php';
require_once __DIR__ . '/ZVecSchema.php';
require_once __DIR__ . '/ZVecDoc.php';
require_once __DIR__ . '/ZVecDocIterator.php';
require_once __DIR__ . '/ZVecReRanker.php';
require_once __DIR__ . '/ZVecRerankedDoc.php';
require_once __DIR__ . '/ZVecRrfReRanker.php';
Expand Down Expand Up @@ -853,6 +854,66 @@ public function fetch(array|string|bool ...$args): array
return self::parseQueryResult($result);
}

/**
* Iterate over every document in the collection.
*
* Unlike fetch() this needs no primary keys, and unlike queryByFilter() no
* filter and no topk — it is the full-scan path, intended for export,
* backup and migration. Memory use is constant regardless of collection
* size.
*
* foreach ($collection->iterDocs() as $pk => $doc) { ... }
*
* The scan runs over an isolated snapshot taken now, so later writes are
* invisible to it. On a writable collection each call seals the current
* writing segment, which may create a small segment.
*
* While the returned iterator is open, close(), destroy(), schema DDL and
* optimize() throw ZVecException (FAILED_PRECONDITION); the iterator closes
* itself when exhausted, so a completed foreach needs no cleanup.
*
* @param string[]|null $outputFields null = all scalar fields, [] = primary
* key only. Vector fields are not allowed
* here and are rejected upstream.
* @param bool $includeVector Whether vector fields are returned
*
* @throws ZVecException On FFI error or when the collection is closed
*/
public function iterDocs(?array $outputFields = null, bool $includeVector = true): ZVecDocIterator
{
$this->checkClosed();
if ($outputFields !== null) {
foreach ($outputFields as $field) {
if (!is_string($field) || $field === '') {
throw new ZVecException('outputFields must contain only non-empty strings');
}
}
}

$ffi = self::ffi();
[$ofArr, $ofCount, $ofCStrings] = self::toCStringArray($ffi, $outputFields ?? []);
$out = $ffi->new('zvec_doc_iterator_t');
try {
self::checkStatus($ffi->zvec_collection_create_iterator(
$this->handle,
$outputFields === null ? 0 : 1,
$ofArr,
$ofCount,
$includeVector ? 1 : 0,
FFI::addr($out),
));
} finally {
self::freeCStringArray($ofCStrings);
// toCStringArray() allocates the char*[N] array unmanaged but frees
// only the strings, so release the array itself here.
if ($ofArr !== null) {
FFI::free($ofArr);
}
}

return new ZVecDocIterator($this, $out);
}

/**
* Index type: HNSW (Hierarchical Navigable Small World).
*
Expand Down Expand Up @@ -2793,6 +2854,7 @@ class_alias(ZVecIndexParams::class, 'ZVecIndexParams');
\class_alias(ZVecQueryInterface::class, 'ZVecQueryInterface');
class_alias(ZVecVectorQuery::class, 'ZVecVectorQuery');
class_alias(ZVecGroupByVectorQuery::class, 'ZVecGroupByVectorQuery');
class_alias(ZVecDocIterator::class, 'ZVecDocIterator');
\class_alias(ZVecReRanker::class, 'ZVecReRanker');
class_alias(ZVecRerankedDoc::class, 'ZVecRerankedDoc');
class_alias(ZVecRrfReRanker::class, 'ZVecRrfReRanker');
Expand Down
Loading
Loading