Search code examples
pythonunicode

Find out the unicode script of a character


Given a unicode character what would be the simplest way to return its script (as "Latin", "Hangul" etc)? unicodedata doesn't seem to provide this kind of feature.


Solution

  • I was hoping someone's done it before, but apparently not, so here's what I've ended up with. The module below (I call it unicodedata2) extends unicodedata and provides script_cat(chr) which returns a tuple (Script name, Category) for a unicode char. Example:

    # coding=utf8
    import unicodedata2
    print unicodedata2.script_cat(u'Ф')  #('Cyrillic', 'L')
    print unicodedata2.script_cat(u'の')  #('Hiragana', 'Lo')
    print unicodedata2.script_cat(u'★')  #('Common', 'So')
    

    The module: https://gist.github.com/2204527