fix: collate must not drop records of adjacent keys with identical data - #62
Merged
Merged
Conversation
…ta (#61) ldb_import_list_variable_records() skipped a record as duplicate by comparing its data with the previous record only, before checking whether that record belonged to the same subkey. All subkeys sharing a 4-byte main key are collated together and sorted by key then data, so when the last record of a key and the first record of the next key carried the same data, the second key lost its record silently. Compute new_subkey first and only deduplicate within a subkey. Also zero the unused tail of each variable-record slot in the collate buffer. The buffer is reused across keys and grown with realloc, and the sort compares the whole slot, so stale bytes could keep identical records apart and let duplicates survive. Add regression tests for 8-byte (test_kb_crc64) and 16-byte (test_kb) keys, plus the stale-slot deduplication case. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes #61
Problem
When a table is collated,
ldb_import_list_variable_records()treated a record as a duplicate if its data matched the previous record's data. It did not check whether the previous record belonged to the same key. All keys that share a 4-byte main key are collated together, sorted by key then data. So when the last record of one key and the first record of the next key carried the same data, the second key silently lost its record. The log still reported the full record count.Fix
new_subkeybefore the duplicate check, and skip a record only when it belongs to the same subkey as the previous record and has the same data.realloc), and the sort compares the whole slot. Leftover bytes from earlier records could keep identical records from sorting next to each other, so some duplicates were not removed. This caused no data loss, only missed deduplication.Corrections to the issue
ldb_import_list_variable_records, not inldb_collate_sector.subkey = NULL, so every record in the buffer belongs to the same 4-byte key. A byte-identical record is therefore a real duplicate, andldb_eliminate_duplicatesis correct.Tables already collated with the bug may be missing records. Collating them again does not bring the records back; they have to be re-imported.
Tests
test_kb_crc64/test_15: the reproduction from the issue with 8-byte keys, plus a real duplicate that must still be removed.test_kb_crc64/test_16: the leftover-bytes case.test_kb/test_11: the reproduction from the issue with 16-byte keys.Each new test fails without its part of the fix. The full
./run_test.shsuite passes.🤖 Generated with Claude Code