Search code examples
javaunicodecharacterasciiligature

Separating Unicode ligature characters


Throughout the vast number of unicode characters, there are some that actually represent more than one character, like the U+FB00 ligature ff for two 'f' characters. Is there any way easy to convert characters like these into multiple single characters? Preferably something available in the standard Java API, but I can refer to an external library if need be.


Solution

  • U+FB00 is a compatibility character. Normally Unicode doesn't support separate codepoints for ligatures (arguing that it's a layout decision if and when a ligature should be used and should not influence how the data is stored). A few of those still exist to allow round-trip conversion compatibility with older encodings that do represent ligatures as separate entities.

    Luckily, the information which characters the ligature represents is present in the Unicode data file and most capable string handling systems have that data built-in.

    In Java, you'll need to use the Normalizer class and the NFKC form:

    String ff ="\uFB00";
    String normalized = Normalizer.normalize(ff, Form.NFKC);
    System.out.println(ff + " = " + normalized);
    

    This will print

    ff = ff