World Lexer Features

Like MULTI_LEXER, the WORLD_LEXER lexer enables you to index documents that contain different languages. It automatically detects the languages of a document and, therefore, does not require you to create a language column in the base table.

WORLD_LEXER processes all database character sets and supports the Unicode 5.0 standard. For WORLD_LEXER to be effective with documents that use multiple languages, AL32UTF-8 or UTF8 Oracle character set encoding must be specified. This includes supplementary, or “surrogate-pair,” characters. Table D-2 and Table D-3 show the languages supported by WORLD_LEXER. This list may change as the Unicode standard changes, and in any case should not be considered exhaustive. (Languages are grouped by Unicode writing system, not by natural language groupings.)

Table 2 Languages Supported by the World Lexer (Space-separated)

Language Group Languages Include
Arabic Arabic, Farsi, Kurdish, Pashto, Sindhi, Urdu
Armenian Armenian
Bengali Assamese, Bengali
Bopomofo Hakka Chinese, Minnan Chinese
Cyrillic Over 50 languages, including Belorussian, Bulgarian, Macedonian, Moldavian, Russian, Serbian, Serbo-Croatian, Ukrainian
Devenagari Bhojpuri, Bihari, Hindi, Kashmiri, Marathi, Nepali, Pali, Sanskrit
Ethiopic Amharic, Ge’ez, Tigrinya, Tigre
Georgian Georgian
Greek Greek
Gujarati Gujarati, Kacchi
Gurmukhi Punjabi
Hebrew Hebrew, Ladino, Yiddish
Kaganga Redjang
Kannada Kanarese, Kannada
Korean Korean, Hanja Hangul
Latin Afrikaans, Albanian, Basque, Breton, Catalan, Croatian, Czech, Danish, Dutch, English, Esperanto, Estonian, Faeroese, Fijian, Finnish, Flemish, French, Frisian, German, Hawaiian, Hungarian, Icelandic, Indonesian, Irish, Italian, Lappish, Classic Latin, Latvian, Lithuanian, Malay, Maltese, Pinyin Mandarin, Maori, Norwegian, Polish, Portuguese, Provencal, Romanian, Rumanian, Samoan, Scottish Gaelic, Slovak, Slovene, Slovenian, Sorbian, Spanish, Swahili, Swedish, Tagalog, Turkish, Vietnamese, Welsh
Malayalam Malayalam
Mongolian Mongolian
Oriya Oriya
Sinhalese, Sinhala Pali, Sinhalese
Syriac Aramaic, Syriac
Tamil Tamil
Telugu Telugu
Thaana Dhiveli, Divehi, Maldivian

Table 3 Languages Supported by the World Lexer (Non-space-separated)

Language Group Languages Include
Chinese Cantonese, Mandarin, Pinyin phonograms
Japanese Japanese (Hiragana, Kanji, Katakana)
Khmer Cambodian, Khmer
Lao Lao
Myanmar Burmese
Thai Thai
Tibetan Dzongkha, Tibetan

Table D-4 shows languages not supported by the World Lexer.

Table 4 Languages Not Supported by the World Lexer

Language Group Languages Include
Buhid Buhid
Canadian Syllabics Blackfoot, Carrier, Cree, Dakhelh, Inuit, Inuktitut, Naskapi, Nunavik, Nunavut, Ojibwe, Sayisi, Slavey
Cherokee Cherokee
Cypriot Cypriot
Limbu Limbu
Ogham Ogham
Runic Runic
Tai Le (Tai Lu, Lue, Dai Le) Tai Le
Ugaritic Ugaritic
Yi Yi
Yi Jang Hexagram Yi Jang