test(fts): cover the ngram and jieba tokenizers and the stemmer filter (#225) - #233
Merged
Merged
Conversation
zvec picks an I/O backend for DiskANN disk reads on first use: on Linux
it tries io_uring, then libaio, then falls back to synchronous pread();
macOS arm64 always uses pread(). The choice dominates DiskANN
throughput -- pread on Linux usually means libaio simply is not installed
-- and until now there was no way to observe it from PHP.
Three static getters on ZVec, next to getVersion() and following the
same pattern the Python SDK uses for its module-level zvec.io_backend_type():
ZVec::getIoBackendType(): int
ZVec::getIoBackendTypeName(int $type): string
ZVec::getIoBackendDescription(): string
plus IO_BACKEND_PREAD / IO_BACKEND_LIBAIO / IO_BACKEND_IO_URING (0/1/2),
the same values as the upstream C ABI. The description is where upstream
explains how to install an async backend when pread is in use.
These are deliberately static rather than per-collection: the value is
process-wide, it is not part of GlobalConfig, and upstream probes it
lazily, so the getters work without ZVec::init() at all. The test asserts
that.
Adhering to the #215 singleton rule: the adapter includes only the public
zvec/ailego/io/io_backend.h and calls upstream's exported
current_io_backend_type() / current_io_backend_description(), which are
compiled inside libzvec, so the IOBackend singleton is created only there.
The internal io_backend_def.h is not part of the SDK and IOBackend::Instance()
is never referenced from this module. The comment in ffi/zvec_ffi.cc records
why, so it does not get "fixed" later.
$ nm -C ffi/build/libzvec_ffi.so | grep -c 'IOBackend::Instance'
0
$ nm -D --defined-only ffi/build/libzvec_ffi.so | grep io_backend
zvec_get_io_backend_description
zvec_get_io_backend_type
zvec_get_io_backend_type_name
No status to check: these return plain values, matching getVersion(), so
self::checkStatus() is not involved and FFI::free() is not needed -- the
C side owns the memory (a static literal for the name, a thread-local
buffer for the description).
Tests: 192/192, 0 skipped, 0 failed, 2 expected fail. The new
tests/test_io_backend.phpt covers the constants, the name mapping including
the "unknown" fallback, that the value is valid before init(), that the
description mentions the reported backend, the macOS pread rule, and that
the cached value is stable across repeated calls and across init().
tests/test_ffi_load.phpt gained the three symbols (60 -> 63).
Exposes two upstream GlobalConfig::ConfigData fields that had no PHP equivalent. jiebaDictDir is the folder holding jieba.dict.utf8 and hmm_model.utf8 for the jieba FTS tokenizer; ftsBruteForceByKeysRatio (0.0-1.0, default 0.05) is where an FTS query stops walking posting lists and starts scoring candidates one by one, which pays off when the scalar filter is very selective. Read back with getJiebaDictDir() and getFtsBruteForceByKeysRatio(); with no option given the bundled zvec_data/jieba_dict is still found automatically. null rather than 0.0 as the ratio default, because 0.0 is a real value upstream accepts rather than a "use the default" marker like the other ratios. Both values are validated in PHP before any FFI call, which is the only place the check can live. Upstream GlobalConfig::initialize() sets its initialized flag *before* validating, so a value it rejects still leaves the library marked initialized, with defaults, and every later init() returns OK without applying anything -- a silent misconfiguration. NAN also passes the upstream range test, since NAN < 0 and NAN > 1 are both false. A bad jiebaDictDir is rejected for a second, harder reason: cppjieba calls abort() rather than returning a Status, so a wrong folder kills the process with exit 134 later, when a jieba FTS index is created, with no catchable error. The check therefore requires both dictionary files, not just the directory. (Confirmed: the same abort is reachable through per-field extraParams and through ZVEC_JIEBA_DICT_DIR, which is a separate gap.) The getters use global_config_ptr() rather than GlobalConfig::Instance(), per the #215 singleton rule. The ratio message uses var_export() because interpolating NAN emits a PHP warning; the message is the same shape as the other init() validation errors. Tests: 194/194, 0 skipped, 0 failed, 2 expected fail. The two new tests run in separate processes because upstream config is applied only once per process: - test_init_fts_options.phpt: six rejection cases (1.5, -0.1, NAN, empty dir, missing dir, existing dir without the dictionary files) each asserting isInitialized() is still false afterwards, then a successful init against a *copy* of the bundled dictionary plus an end-to-end jieba FTS query that returns j1, proving a custom folder is the one actually used. - test_init_fts_defaults.phpt: the upstream defaults, 0.05 and the bundled dictionary, with no new options passed.
Declaring a full-text index on a STRING field took two steps: create the
collection, then call createIndex(). Upstream puts the index params on the
field itself, so the index can exist from the first insert:
$schema->addString('body', indexParams: ZVecIndexParams::forFts());
$schema->addString('tag', indexParams: ZVecIndexParams::forInvert());
This is what the Python SDK does with FieldSchema(index_param=...), and
what upstream's own hybrid tests do. The new argument is last and
optional, so existing calls are unaffected.
Two validation cases the older signature could not express. Passing both
$withInvertIndex and $indexParams is ambiguous and now throws. And
unlike the other add*() methods, a duplicate field name is reported here
rather than silently dropped: upstream returns AlreadyExists from
add_field(), and zvec_schema_add_field_* discards that Status, so
addString('body') twice currently loses a field without a word. This
function propagates it because it has to build index params anyway.
A mismatched index type is deliberately *not* rejected here. Upstream
reports it at ZVec::create() time with a message naming the field
("scalar field[body] does not support vector index params, but got
index_type ..."), which is clearer than anything the adapter could
produce, so there is one source of truth.
Upstream's FieldSchema constructor clones the index params, so the
ZVecIndexParams object keeps owning its handle and is freed by its own
destructor as usual; nothing new to release on the PHP side.
FFI: zvec_schema_add_field_string_with_index(), declared next to
zvec_collection_create_index() in both headers because the
zvec_index_params_t typedef comes after the schema block.
Tests: 195/195, 0 skipped, 0 failed, 2 expected fail. The new
tests/test_schema_fts_index.phpt queries for a term without any
createIndex() call, checks getIndexType() is 11 before and after a
reopen, covers invert via params (10), both validation errors, and the
pre-existing argument combinations. test_ffi_load.phpt gained the symbol
(67 -> 68).
Note: tests/test_diskann_index.phpt was seen failing once with only
doc2 returned instead of doc1,doc2,doc3, and passed on five subsequent
single runs and four subsequent full-suite runs. Not reproduced and not
caused by this change; a flaky DiskANN query result on a 3-document
collection, worth a separate look if it shows up again.
#225) forFts() forwards the tokenizer name, filter list and extraParams straight to upstream, so every upstream tokenizer and filter already worked. Only the standard tokenizer with the lowercase filter was tested. Three new tests, no C++ and no new PHP method: - test_fts_tokenizer_ngram.phpt: default bigrams, ngram_min/ngram_max, token_chars, and the four limit violations. Queries use AND rather than OR deliberately: with OR, "base" splits into ba/as/se and n2 also matches through "as", which would mask what the tokenizer produced. - test_fts_tokenizer_jieba.phpt: the bundled dictionary, all four cut modes across five terms so the token difference is visible, a mixed Chinese/ASCII document, a user dictionary, and a bad cut_mode. The dictionary is found automatically at init(); no jieba_dict_dir needed. Skips when the bundled dictionary is absent or when ZVEC_JIEBA_DICT_DIR would take precedence. - test_fts_filter_stemmer.phpt: stemming against a lowercase-only baseline, the default english language, the porter language, german with umlauts, and four rejected configurations. Each also reopens the collection and repeats one query, so the configuration is shown to be persisted rather than cached in memory. The docblock fix: forFts() advertised filters ["lowercase", "stemmer_en"] and an extraParams example of "stemmer_lang=en". Neither works. There is no stemmer_en filter, and the language is a JSON key, not a bare string: unknown filter name: stemmer_en failed to parse extra_params JSON: stemmer_lang=en Both bad examples are now asserted as rejected in the stemmer test, so the corrected docblock is backed by a test rather than by assertion alone. A stale duplicated "Create Invert index params" block sitting in front of the real docblock is gone. Rejected cases are logged at ERROR on stderr by upstream, so both tests that expect failures init with LOG_FATAL to keep the expected output clean. Deliberately not tested: a jieba_dict_dir or user_dict_path pointing at a missing file. cppjieba calls abort() inside createIndex(), taking PHP with it at exit 134, so such a case would kill the test runner. Validating those paths is a separate bug; the note is in the README so the trap is documented. Tests: 198/198, 0 skipped, 0 failed, 2 expected fail.
…rs-stemmer # Conflicts: # CHANGELOG.md # README.md # tests/test_ffi_load.phpt
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #225.
Summary
forFts()forwards the tokenizer name, filter list andextraParamsstraight to upstream, so every upstream tokenizer and filter already worked. Onlystandard+lowercasewas tested. This adds coverage, plus a docblock fix.No C++ change and no new PHP method.
The docblock was wrong
It advertised a filter and an
extraParamsexample that do not work:["lowercase", "stemmer_en"]unknown filter name: stemmer_en"stemmer_lang=en"failed to parse extra_params JSON: stemmer_lang=enThere is no
stemmer_enfilter — the language is a JSON key. Both bad examples are now asserted as rejected in the new stemmer test, so the corrected docblock is backed by a test rather than by assertion alone. A stale duplicated "Create Invert index params" block that was sitting in front of the real docblock is gone too.Tests
test_fts_tokenizer_ngram.phpt— default bigrams,ngram_min/ngram_max,token_chars, and the four limit violations.Queries use AND, not OR, deliberately. With OR,
"base"splits intoba/as/seandn2also matches throughas, which would mask what the tokenizer actually produced.test_fts_tokenizer_jieba.phpt— the bundled dictionary, all four cut modes across five terms so the token difference is visible, mixed Chinese/ASCII, a user dictionary, and a badcut_mode:The dictionary is found automatically at
init(); nojieba_dict_dirneeded. Skips when the bundled dictionary is missing or whenZVEC_JIEBA_DICT_DIRwould take precedence.test_fts_filter_stemmer.phpt— stemming measured against a lowercase-only baseline (sofox→d1andfoxes→d2separately, thend1,d2with stemming), the defaultenglish,porter,germanwith umlauts, and four rejected configurations.Each test reopens the collection and repeats one query, showing the tokenizer config is persisted rather than cached in memory.
Two things worth knowing
Rejected cases log to stderr. Upstream logs each rejection at ERROR, which landed in the test output. Both tests that expect failures therefore
initwithLOG_FATAL. Not a library change — a test-hygiene one — but easy to trip over.One case is deliberately not tested. A
jieba_dict_diroruser_dict_pathpointing at a missing file makes cppjieba callabort()insidecreateIndex(), taking PHP down at exit 134:Upstream's
try/catchthere cannot catch it, so such a test would kill the runner. Guarding those paths is a separate bug; I added a note to the README so the trap is documented.Full suite:
test_fts.phptandtest_docs_consistency.phptpass unchanged.test_dbs/is left with only.gitignore.🤖 Generated with Claude Code