Showing posts with label unicode. Show all posts
Showing posts with label unicode. Show all posts

Saturday, September 7, 2013

Enabling Indic transliteration on web pages

Enable typing in Indian languages in web pages using Pramukh IME. This javascript library is very easy to integrate into your website, and supports 20 Indian languages.

Also, it is open source; so you can add new languages, or change the keyboard layout for existing ones.

Wednesday, June 5, 2013

Installing Python regex on Ubuntu 12.04

To install the alternative regular expression module regex (which has richer Unicode support than re) for Python:
  • Install python-dev from the Software Center.
  • Download the archive, extract, and run sudo python setup.py install.

Thursday, May 2, 2013

Unicode tips for Python

  • To use non-Latin characters in regular expressions, use u'...' instead of r'...', even if you have to escape every backslash; e.g. the regex u'(?u)[०-९]\\s' matches a Devanagari digit followed by whitespace.
  • Remove zero-width joiners/non-joiners from Unicode text to get a normalized representation; otherwise words that are rendered the same in a browser/editor will be stored differently, and will not be equal on comparison; e.g. use the regex u'[\\u200D\\u200C]' and replace all matches with u'' (the empty string).

Friday, August 3, 2012

Displaying Unicode text in the terminal / console

If gnome-terminal has problems displaying Unicode (UTF-8) characters properly, try using konsole (from KDE). It worked for me.

Friday, December 31, 2010

Using Unicode in Latex/Tex

To use Hindi Unicode, do the following -

In the .tex file include the line -
\font\texthi="Lohit Hindi:script=deva,mapping=tex-text" at 11pt.
(Remember to exclude the full stop at the end!)
Enter Hindi unicode text in the document as -
Some English text {\texthi हिन्दी...} more English text...
Compile the document using Xetex or XeLatex.

More information -

"texthi" is the name of the command we just defined to indicate Unicode text. You can give any other name, e.g. \font\abcd="...".

"Lohit Hindi" (remember the space between Lohit and Hindi) is the name of an Open Truetype Font (OTF) installed in the system. You can specify any other font that is installed. In Ubuntu, you can see what fonts are installed as follows -
Goto Main Taskbar -> System -> Preferences -> Appearance -> (Popup opens) -> Click 'Fonts' tab -> Try to change any of the fonts e.g. Application font. You get another popup where the fonts are listed under 'Family'. You can use any of these fonts instead of "Lohit Hindi".

"deva" is a tag that specifies which script (and hence which unicode range) should be used. Hindi uses the devanagari script. For other scripts, find out which script tag to use.

"mapping=tex-text" - I don't know what this. If you do, please tell me.

Thursday, February 4, 2010

Perl tips - Unicode

First, use Encode;

Reading/Writing
  • $string = Encode::decode('UTF-8',$text); (assuming the input file (or STDIN) is encoded in UTF-8).
  • You can now handle $string as you would normal strings (e.g. split(//) will split it at character boundaries)
  • Do $text = Encode::encode('UTF-8', $string); before writing it out to file (assuming you want the output file (or STDOUT) in that encoding).
Regex
  • \p{L} - full glyph (e.g. the letter 'A')
  • \p{M} - partial glyph (e.g. the accent ` on the letter 'A', giving 'À')
  • \p{N} - digit
  • \p{P} - punctuation
  • \p{kannada} - any Kannada character
  • \P{} - invert the condition
E.g. to match a line if it contains no numerals and no punctuation do
$line ~= m/\p{N}|\p{P}/