Skip to content

test(fts): cover the ngram and jieba tokenizers and the stemmer filter (#225) - #233

Merged
s2x merged 5 commits into
mainfrom
test/225-fts-tokenizers-stemmer
Sep 29, 2026
Merged

s2x merged 5 commits into
mainfrom
test/225-fts-tokenizers-stemmer

Conversation

@s2x

@s2x s2x commented Sep 29, 2026

Copy link
Copy Markdown
Member

Closes #225.

Summary

forFts() forwards the tokenizer name, filter list and extraParams straight to upstream, so every upstream tokenizer and filter already worked. Only standard + lowercase was tested. This adds coverage, plus a docblock fix.

No C++ change and no new PHP method.

The docblock was wrong

It advertised a filter and an extraParams example that do not work:

docblock said upstream says
["lowercase", "stemmer_en"] unknown filter name: stemmer_en
"stemmer_lang=en" failed to parse extra_params JSON: stemmer_lang=en

There is no stemmer_en filter — the language is a JSON key. Both bad examples are now asserted as rejected in the new stemmer test, so the corrected docblock is backed by a test rather than by assertion alone. A stale duplicated "Create Invert index params" block that was sitting in front of the real docblock is gone too.

Tests

test_fts_tokenizer_ngram.phpt — default bigrams, ngram_min/ngram_max, token_chars, and the four limit violations.

Queries use AND, not OR, deliberately. With OR, "base" splits into ba/as/se and n2 also matches through as, which would mask what the tokenizer actually produced.

test_fts_tokenizer_jieba.phpt — the bundled dictionary, all four cut modes across five terms so the token difference is visible, mixed Chinese/ASCII, a user dictionary, and a bad cut_mode:

search 清华: j1        full 清华: j1        mix 清华: (none)     hmm 清华: (none)
search 科学院: j3      full 科学院: j3      mix 科学院: (none)   hmm 科学院: j3

The dictionary is found automatically at init(); no jieba_dict_dir needed. Skips when the bundled dictionary is missing or when ZVEC_JIEBA_DICT_DIR would take precedence.

test_fts_filter_stemmer.phpt — stemming measured against a lowercase-only baseline (so fox → d1 and foxes → d2 separately, then d1,d2 with stemming), the default english, porter, german with umlauts, and four rejected configurations.

Each test reopens the collection and repeats one query, showing the tokenizer config is persisted rather than cached in memory.

Two things worth knowing

Rejected cases log to stderr. Upstream logs each rejection at ERROR, which landed in the test output. Both tests that expect failures therefore init with LOG_FATAL. Not a library change — a test-hygiene one — but easy to trip over.

One case is deliberately not tested. A jieba_dict_dir or user_dict_path pointing at a missing file makes cppjieba call abort() inside createIndex(), taking PHP down at exit 134:

FATAL exp: [ifs.is_open()] false. open /nonexistent/dir/jieba.dict.utf8 failed.

Upstream's try/catch there cannot catch it, so such a test would kill the runner. Guarding those paths is a separate bug; I added a note to the README so the trap is documented.

Full suite:

Number of tests : 198               198
Tests skipped   :   0 (  0.0%)
Tests failed    :   0 (  0.0%)
Expected fail   :   2 (  1.0%)
Tests passed    : 196 ( 99.0%)

test_fts.phpt and test_docs_consistency.phpt pass unchanged. test_dbs/ is left with only .gitignore.

🤖 Generated with Claude Code

s2x added 5 commits September 28, 2026 20:29
zvec picks an I/O backend for DiskANN disk reads on first use: on Linux
it tries io_uring, then libaio, then falls back to synchronous pread();
macOS arm64 always uses pread(). The choice dominates DiskANN
throughput -- pread on Linux usually means libaio simply is not installed
-- and until now there was no way to observe it from PHP.

Three static getters on ZVec, next to getVersion() and following the
same pattern the Python SDK uses for its module-level zvec.io_backend_type():

    ZVec::getIoBackendType(): int
    ZVec::getIoBackendTypeName(int $type): string
    ZVec::getIoBackendDescription(): string

plus IO_BACKEND_PREAD / IO_BACKEND_LIBAIO / IO_BACKEND_IO_URING (0/1/2),
the same values as the upstream C ABI. The description is where upstream
explains how to install an async backend when pread is in use.

These are deliberately static rather than per-collection: the value is
process-wide, it is not part of GlobalConfig, and upstream probes it
lazily, so the getters work without ZVec::init() at all. The test asserts
that.

Adhering to the #215 singleton rule: the adapter includes only the public
zvec/ailego/io/io_backend.h and calls upstream's exported
current_io_backend_type() / current_io_backend_description(), which are
compiled inside libzvec, so the IOBackend singleton is created only there.
The internal io_backend_def.h is not part of the SDK and IOBackend::Instance()
is never referenced from this module. The comment in ffi/zvec_ffi.cc records
why, so it does not get "fixed" later.

    $ nm -C ffi/build/libzvec_ffi.so | grep -c 'IOBackend::Instance'
    0
    $ nm -D --defined-only ffi/build/libzvec_ffi.so | grep io_backend
    zvec_get_io_backend_description
    zvec_get_io_backend_type
    zvec_get_io_backend_type_name

No status to check: these return plain values, matching getVersion(), so
self::checkStatus() is not involved and FFI::free() is not needed -- the
C side owns the memory (a static literal for the name, a thread-local
buffer for the description).

Tests: 192/192, 0 skipped, 0 failed, 2 expected fail. The new
tests/test_io_backend.phpt covers the constants, the name mapping including
the "unknown" fallback, that the value is valid before init(), that the
description mentions the reported backend, the macOS pread rule, and that
the cached value is stable across repeated calls and across init().
tests/test_ffi_load.phpt gained the three symbols (60 -> 63).
Exposes two upstream GlobalConfig::ConfigData fields that had no PHP
equivalent. jiebaDictDir is the folder holding jieba.dict.utf8 and
hmm_model.utf8 for the jieba FTS tokenizer; ftsBruteForceByKeysRatio
(0.0-1.0, default 0.05) is where an FTS query stops walking posting
lists and starts scoring candidates one by one, which pays off when the
scalar filter is very selective. Read back with getJiebaDictDir() and
getFtsBruteForceByKeysRatio(); with no option given the bundled
zvec_data/jieba_dict is still found automatically.

null rather than 0.0 as the ratio default, because 0.0 is a real value
upstream accepts rather than a "use the default" marker like the other
ratios.

Both values are validated in PHP before any FFI call, which is the only
place the check can live. Upstream GlobalConfig::initialize() sets its
initialized flag *before* validating, so a value it rejects still leaves
the library marked initialized, with defaults, and every later init()
returns OK without applying anything -- a silent misconfiguration. NAN
also passes the upstream range test, since NAN < 0 and NAN > 1 are both
false.

A bad jiebaDictDir is rejected for a second, harder reason: cppjieba
calls abort() rather than returning a Status, so a wrong folder kills the
process with exit 134 later, when a jieba FTS index is created, with no
catchable error. The check therefore requires both dictionary files, not
just the directory. (Confirmed: the same abort is reachable through
per-field extraParams and through ZVEC_JIEBA_DICT_DIR, which is a
separate gap.)

The getters use global_config_ptr() rather than GlobalConfig::Instance(),
per the #215 singleton rule.

The ratio message uses var_export() because interpolating NAN emits a PHP
warning; the message is the same shape as the other init() validation
errors.

Tests: 194/194, 0 skipped, 0 failed, 2 expected fail. The two new tests
run in separate processes because upstream config is applied only once
per process:
- test_init_fts_options.phpt: six rejection cases (1.5, -0.1, NAN, empty
  dir, missing dir, existing dir without the dictionary files) each
  asserting isInitialized() is still false afterwards, then a successful
  init against a *copy* of the bundled dictionary plus an end-to-end
  jieba FTS query that returns j1, proving a custom folder is the one
  actually used.
- test_init_fts_defaults.phpt: the upstream defaults, 0.05 and the
  bundled dictionary, with no new options passed.
Declaring a full-text index on a STRING field took two steps: create the
collection, then call createIndex(). Upstream puts the index params on the
field itself, so the index can exist from the first insert:

    $schema->addString('body', indexParams: ZVecIndexParams::forFts());
    $schema->addString('tag',  indexParams: ZVecIndexParams::forInvert());

This is what the Python SDK does with FieldSchema(index_param=...), and
what upstream's own hybrid tests do. The new argument is last and
optional, so existing calls are unaffected.

Two validation cases the older signature could not express. Passing both
$withInvertIndex and $indexParams is ambiguous and now throws. And
unlike the other add*() methods, a duplicate field name is reported here
rather than silently dropped: upstream returns AlreadyExists from
add_field(), and zvec_schema_add_field_* discards that Status, so
addString('body') twice currently loses a field without a word. This
function propagates it because it has to build index params anyway.

A mismatched index type is deliberately *not* rejected here. Upstream
reports it at ZVec::create() time with a message naming the field
("scalar field[body] does not support vector index params, but got
index_type ..."), which is clearer than anything the adapter could
produce, so there is one source of truth.

Upstream's FieldSchema constructor clones the index params, so the
ZVecIndexParams object keeps owning its handle and is freed by its own
destructor as usual; nothing new to release on the PHP side.

FFI: zvec_schema_add_field_string_with_index(), declared next to
zvec_collection_create_index() in both headers because the
zvec_index_params_t typedef comes after the schema block.

Tests: 195/195, 0 skipped, 0 failed, 2 expected fail. The new
tests/test_schema_fts_index.phpt queries for a term without any
createIndex() call, checks getIndexType() is 11 before and after a
reopen, covers invert via params (10), both validation errors, and the
pre-existing argument combinations. test_ffi_load.phpt gained the symbol
(67 -> 68).

Note: tests/test_diskann_index.phpt was seen failing once with only
doc2 returned instead of doc1,doc2,doc3, and passed on five subsequent
single runs and four subsequent full-suite runs. Not reproduced and not
caused by this change; a flaky DiskANN query result on a 3-document
collection, worth a separate look if it shows up again.
#225)

forFts() forwards the tokenizer name, filter list and extraParams straight
to upstream, so every upstream tokenizer and filter already worked. Only
the standard tokenizer with the lowercase filter was tested.

Three new tests, no C++ and no new PHP method:

- test_fts_tokenizer_ngram.phpt: default bigrams, ngram_min/ngram_max,
  token_chars, and the four limit violations. Queries use AND rather than
  OR deliberately: with OR, "base" splits into ba/as/se and n2 also
  matches through "as", which would mask what the tokenizer produced.
- test_fts_tokenizer_jieba.phpt: the bundled dictionary, all four cut
  modes across five terms so the token difference is visible, a mixed
  Chinese/ASCII document, a user dictionary, and a bad cut_mode. The
  dictionary is found automatically at init(); no jieba_dict_dir needed.
  Skips when the bundled dictionary is absent or when
  ZVEC_JIEBA_DICT_DIR would take precedence.
- test_fts_filter_stemmer.phpt: stemming against a lowercase-only
  baseline, the default english language, the porter language, german
  with umlauts, and four rejected configurations.

Each also reopens the collection and repeats one query, so the
configuration is shown to be persisted rather than cached in memory.

The docblock fix: forFts() advertised filters ["lowercase", "stemmer_en"]
and an extraParams example of "stemmer_lang=en". Neither works. There is
no stemmer_en filter, and the language is a JSON key, not a bare string:

    unknown filter name: stemmer_en
    failed to parse extra_params JSON: stemmer_lang=en

Both bad examples are now asserted as rejected in the stemmer test, so the
corrected docblock is backed by a test rather than by assertion alone. A
stale duplicated "Create Invert index params" block sitting in front of the
real docblock is gone.

Rejected cases are logged at ERROR on stderr by upstream, so both tests
that expect failures init with LOG_FATAL to keep the expected output clean.

Deliberately not tested: a jieba_dict_dir or user_dict_path pointing at a
missing file. cppjieba calls abort() inside createIndex(), taking PHP with
it at exit 134, so such a case would kill the test runner. Validating
those paths is a separate bug; the note is in the README so the trap is
documented.

Tests: 198/198, 0 skipped, 0 failed, 2 expected fail.
…rs-stemmer

# Conflicts:
#	CHANGELOG.md
#	README.md
#	tests/test_ffi_load.phpt
@s2x
s2x merged commit 26fd694 into main Sep 29, 2026
15 of 18 checks passed
@s2x
s2x deleted the test/225-fts-tokenizers-stemmer branch September 29, 2026 08:12
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

test(fts): cover ngram and jieba tokenizers and the stemmer filter, and fix the forFts() docblock

1 participant