Skip to content

fix: collate must not drop records of adjacent keys with identical data - #62

Merged
mscasso-scanoss merged 1 commit into
crc64from
fix/collate-dedup-adjacent-keys
Sep 30, 2026
Merged

mscasso-scanoss merged 1 commit into
crc64from
fix/collate-dedup-adjacent-keys

Conversation

@mscasso-scanoss

Copy link
Copy Markdown
Collaborator

Fixes #61

Problem

When a table is collated, ldb_import_list_variable_records() treated a record as a duplicate if its data matched the previous record's data. It did not check whether the previous record belonged to the same key. All keys that share a 4-byte main key are collated together, sorted by key then data. So when the last record of one key and the first record of the next key carried the same data, the second key silently lost its record. The log still reported the full record count.

Fix

  • Compute new_subkey before the duplicate check, and skip a record only when it belongs to the same subkey as the previous record and has the same data.
  • Zero the unused tail of each variable-record slot in the collate buffer. The buffer is reused across keys (and grown with realloc), and the sort compares the whole slot. Leftover bytes from earlier records could keep identical records from sorting next to each other, so some duplicates were not removed. This caused no data loss, only missed deduplication.

Corrections to the issue

  • The check is in ldb_import_list_variable_records, not in ldb_collate_sector.
  • Tables with fixed-size records are not affected. Their nodes do not store subkeys, and the handler gets subkey = NULL, so every record in the buffer belongs to the same 4-byte key. A byte-identical record is therefore a real duplicate, and ldb_eliminate_duplicates is correct.
  • The bug dates from the original collate implementation. It is not a regression.

Tables already collated with the bug may be missing records. Collating them again does not bring the records back; they have to be re-imported.

Tests

  • test_kb_crc64/test_15: the reproduction from the issue with 8-byte keys, plus a real duplicate that must still be removed.
  • test_kb_crc64/test_16: the leftover-bytes case.
  • test_kb/test_11: the reproduction from the issue with 16-byte keys.

Each new test fails without its part of the fix. The full ./run_test.sh suite passes.

🤖 Generated with Claude Code

…ta (#61)

ldb_import_list_variable_records() skipped a record as duplicate by comparing
its data with the previous record only, before checking whether that record
belonged to the same subkey. All subkeys sharing a 4-byte main key are collated
together and sorted by key then data, so when the last record of a key and the
first record of the next key carried the same data, the second key lost its
record silently. Compute new_subkey first and only deduplicate within a subkey.

Also zero the unused tail of each variable-record slot in the collate buffer.
The buffer is reused across keys and grown with realloc, and the sort compares
the whole slot, so stale bytes could keep identical records apart and let
duplicates survive.

Add regression tests for 8-byte (test_kb_crc64) and 16-byte (test_kb) keys,
plus the stale-slot deduplication case.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@mscasso-scanoss mscasso-scanoss self-assigned this Sep 30, 2026
@mscasso-scanoss
mscasso-scanoss merged commit d3ee948 into crc64 Sep 30, 2026
1 check passed
@mscasso-scanoss
mscasso-scanoss deleted the fix/collate-dedup-adjacent-keys branch September 30, 2026 12:19
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant