Skip to content

docs: document Unicode behavior for Levenshtein - #70

Open
whitescreendev wants to merge 1 commit into
tdebatty:masterfrom
whitescreendev:docs/levenshtein-unicode
Open

whitescreendev wants to merge 1 commit into
tdebatty:masterfrom
whitescreendev:docs/levenshtein-unicode

Conversation

@whitescreendev

Copy link
Copy Markdown

Summary

Document the Unicode representation used by the Levenshtein implementation.

The implementation uses Java String.length() and charAt(), so the
comparison operates on UTF-16 code units. Supplementary Unicode code points
represented by surrogate pairs can therefore occupy two sequence positions.

The documentation clarifies that this code-unit distance can differ from
Unicode code-point or grapheme-cluster distance.

No implementation behavior is changed.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant