Skip to content
Open
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
21 changes: 21 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -124,6 +124,27 @@ public class MyApp {
}
```

### Unicode representation

This implementation operates on Java `String` values using `String.length()`
and `charAt()`. These APIs expose UTF-16 code units rather than Unicode code
points or user-perceived characters.

As a result, a supplementary Unicode code point represented by a surrogate pair
can occupy two sequence positions in the Levenshtein calculation.

This distinction does not change the Levenshtein recurrence itself; it defines
the sequence elements supplied to the algorithm. Code-unit distance can
therefore differ from Unicode code-point or grapheme-cluster distance for some
text.

Applications that require code-point or grapheme-cluster semantics should
convert or segment the input into the intended sequence units before computing
the distance.

For additional background on runtime string representations and edit-distance
implementations, see [Levenshtein implementations and Unicode sequence units](https://www.levenshtein.net/levenshtein-implementations).

## Normalized Levenshtein
This distance is computed as levenshtein distance divided by the length of the longest string. The resulting value is always in the interval [0.0 1.0] but it is not a metric anymore!

Expand Down