IndicPhone

IndicPhone is a universal algorithm for phonetically hashing Indic-language words written in all Brahmic scripts in the Unicode table, similar to the Metaphone algorithm for English. For a given word, it generates three Romanized phonetic keys (hashes), each representing an increasing degree of affinity to various phonetic properties of the original word. It was originally written in 2013 for enabling fuzzy search in the Olam Malayalam dictionary.

The algorithm generates phonetic hashes for Indic words by accounting for linguistic features such as compounding and gemination. It exploits the consistent phonetic ordering of glyphs (derived from the ISCII-based Unicode layout) in Brahmic scripts in their respective Unicode blocks, to transform all Indic glyphs in a single pass using a universal mapping of phonetic offsets to alphanumeric Roman characters.

Demo

Examples

Input key0 key1 key2
Kamala
कमलKMLKMLKML
কমলKMLKMLKML
ಕಮಲKMLKMLKML
കമലKMLKMLKML
Svatantra
स्वतन्त्रSV0N0RSV0N0RSV0N0R
স্বতন্ত্রSB0N0RSB0N0RSB0N0R
ಸ್ವತನ್ತ್ರSV0N0RSV0N0RSV0N0R
സ്വതന്ത്രSV0N0RSV0N0RSV0N0R
Namaste
नमस्तेNMS0NMS0NMS06
নমস্তেNMS0NMS0NMS06
ನಮಸ್ತೇNMS0NMS0NMS06
നമസ്തേNMS0NMS0NMS06
Nīlakkuyil
नीलक्कुयिल्NLKYLNLKYLN4LK25Y4L
নীলক্কুযিল্NLKYLNLKYLN4LK25Y4L
ನೀಲಕ್ಕುಯಿಲ್NLKYLNLKYLN4LK25Y4L
നീലക്കുയിൽNLKYLNLKYLN4LK25Y4L
Gaṅgā
गंगK3KK3KK3K
গংগK3KK3KK3K
ಗಂಗK3KK3KK3K
ഗംഗK3KK3KK3K
Himālaya
हिमालयHMLYHMLYH4MLY
হিমালয়HMLYHMLYH4MLY
ಹಿಮಾಲಯHMLYHMLYH4MLY
ഹിമാലയHMLYHMLYH4MLY

The algorithm

IndicPhone converts a word in any Indic/Brahmic script into three phonetic hashes in a single left-to-right pass. The crux of the algorithm is a universal phonetic mapping table that is shared across every script, which is used to map sounds to specific Roman characters.

1. Universal translation table

Each Brahmic script has its own distinctive 128-point Unicode block, where all blocks follow the same ISCII-derived phonetic ordering. A given offset (the position of a glyph within a block) represents the same sound in every script. For example, offset 0x15 is क in Devanagari, ক in Bengali, ಕ in Kannada, and ക in Malayalam, all the same K sound.

offsetDevanagariBengaliKannadaMalayalamCode
Vowels
0x5अঅಅഅA
0x6आআಆആA
0x7इইಇഇI
0x8ईঈಈഈI
0x9उউಉഉU
0xaऊঊಊഊU
0xbऋঋಋഋR
0xeऎ—ಎഎE
0xfएএಏഏE
0x10ऐঐಐഐAI
0x12ऒ—ಒഒO
0x13ओওಓഓO
0x14औঔಔഔO
0x60ॠৠೠൠR
Consonants
0x15कকಕകK
0x16खখಖഖK
0x17गগಗഗK
0x18घঘಘഘK
0x19ङঙಙങNG
0x1aचচಚചC
0x1bछছಛഛC
0x1cजজಜജJ
0x1dझঝಝഝJ
0x1eञঞಞഞNJ
0x1fटটಟടT
0x20ठঠಠഠT
0x21डডಡഡT
0x22ढঢಢഢT
0x23णণಣണN1
0x24तতತത0
0x25थথಥഥ0
0x26दদದദ0
0x27धধಧധ0
0x28नনನനN
0x2aपপಪപP
0x2bफফಫഫF
0x2cबবಬബB
0x2dभভಭഭB
0x2eमমಮമM
0x2fयযಯയY
0x30रরರരR
0x31ऱ—ಱറR1
0x32लলಲലL
0x33ळ—ಳളL1
0x34ऴ——ഴZ
0x35व—ವവV
0x36शশಶശS1
0x37षষಷഷS1
0x38सসಸസS
0x39हহಹഹH

Blank cell = no glyph at the particular Unicode offset.

2. Vowels and modifiers

As Brahmic scripts are abugidas, a consonant carries the inherent a sound. This is assumed to be the default. Other phonetic properties are denoted by appending a specific modifier digit to the consonant's code (see §3). The broader keys are derived by dropping specific digit classes (see §3.1).

InputRule→ key2
mlക
hiक
bnক
knಕ
inherent 'a'K
mlകി
hiकि
bnকি
knಕಿ
vowel sign (mātrā)K4
mlകു
hiकु
bnকু
knಕು
vowel signK5
mlക്ക
hiक्क
bnক্ক
knಕ್ಕ
gemination (doubled glyphs)K2
mlകം
hiकं
bnকং
knಕಂ
anusvāraK3

3. Digit codes

Phonetic details beyond the mapped letter in the table is recorded by appending a modifier digit to that letter's Roman code. The finest key, key2, retains all properties; the coarser key1 and key0 are formed by stripping specific digit classes from them.

DigitDescriptionExampleExists in
1"Hard" sound (retroflex vs. sibilant distinction)ण → N1,   ष → S1key1, key2
2Gemination (doubled consonant)ക്ക → K2key2
3Anusvāraകം → K3key0, key1, key2
4–9 Mātrā (vowel signs)
4 = i/ī
5 = u/ū
6 = e/ē
7 = ai
8 = o/ō
9 = au
കി → K4key2

0 is not in any digit classes. 0 is not a modifier digit, but is a base sound coded after the Greek θ's "th" for the dental plosives ('dantya', eg: त थ द ध). While the anusvāra 3 is a modifier, nasalisation inherently changes a word's identity too much, so it is not discarded like other digit classes. So, both these digits are retained in every key, while every other digit class ([1], 02, 4–9) are dropped in classes to form the coarser keys.

3.1. Digit class transformation

KeyTypeResult
key2Full encodingRetain all digits
key1key2 sans [2] and [4-9]Drops gemination and vowel signs. Keeps hard sounds [1] and anusvāra [3]
key0key1 sans [1]Drops hard sounds and keeps only anusvāra [3]

3.1.2. Example

The algorithm parses a word character by character, left to right, in one pass. It holds at most one consonant at a time and tracks three states (idle, holding a consonant, and inside a conjunct after a virama), deciding at each character whether to emit the held consonant, merge a group, or attach a modifier. Below is an example of the state transformation with the Malayalam word നീലക്കുയിൽ (nīlakkuyil, a cuckoo).

CharacterTypeStateCodekey2
നconsonantHold consonant N— 
◌ീvowel sign (ī)Emit held N and then the vowelN4N4
ലconsonantHold consonant L—N4
കconsonantEmit held L and hold KLN4L
◌്viramaBegin conjunct while K still held—N4L
കconsonantSame consonant after virama, so it's a geminateK2N4LK2
◌ുvowel sign (u)Attach the vowel5N4LK25
യconsonantHold consonant Y—N4LK25
◌ിvowel sign (i)Emit held Y and then the vowelY4N4LK25Y4
ൽchillu (ḷ)Emit the bare consonant directlyLN4LK25Y4L

The final key2 is then reduced to two coarser keys by stripping away classes of digits:

KeyHashDigits stripped
key2N4LK25Y4L—
key1NLKYLVowel/gemination digits [2], [4-9]
key0NLKYLHard-sound marker [1]

Source code

See main.js for the Javascript implementation of the algorithm used in the demo of this page.