CoLI-Machine Learning Approaches for Code-mixed Language Identification at the Word Level in Kannada-English Texts

Shashirekha Hosahalli Lakshmaiah; Fazlourrahman Balouchzahi; Mudoor Devadas Anusha; Grigori Sidorov

doi:10.12700/APH.19.10.2022.10.8

CoLI-Machine Learning Approaches for Code-mixed Language Identification at the Word Level in Kannada-English Texts

Shashirekha Hosahalli Lakshmaiah, Fazlourrahman Balouchzahi, Mudoor Devadas Anusha, Grigori Sidorov

Centro de Investigación en Computación (CIC)

Producción científica: Contribución a una revista › Artículo › revisión exhaustiva

4 Citas (Scopus)

Resumen

The task of automatically identifying a language used in a given text is called Language Identification (LI). India is a multilingual country and many Indians especially youths are comfortable with Hindi and English, in addition to their local languages. Hence, they often use more than one language to post their comments on social media. Texts containing more than one language are called “code-mixed texts” and are a good source of input for LI. Languages in these texts may be mixed at sentence level, word level or even at sub-word level. LI at word level is a sequence labeling problem where each and every word in a sentence is tagged with one of the languages in the predefined set of languages. For many NLP applications, using code-mixed texts, the first but very crucial preprocessing step will be identifying the languages in a given text. In order to address word level LI in code-mixed Kannada-English (Kn-En) texts, this work presents i) the construction of code-mixed Kn-En dataset called CoLI-Kenglish dataset, ii) code-mixed Kn-En embedding and iii) learning models using Machine Learning (ML), Deep Learning (DL) and Transfer Learning (TL) approaches. Code-mixed Kn-En texts are extracted from Kannada YouTube video comments to construct CoLI-Kenglish dataset and code-mixed Kn-En embedding. The words in CoLI-Kenglish dataset are grouped into six major categories, namely, “Kannada”, “English”, “Mixed-language”, “Name”, “Location” and “Other”. Code-mixed embeddings are used as features by the learning models and are created for each word, by merging the word vectors with sub-words vectors of all the sub-words in each word and character vectors of all the characters in each word. The learning models, namely, CoLI-vectors and CoLI-ngrams based on ML, CoLI-BiLSTM based on DL and CoLI-ULMFiT based on TL approaches are built and evaluated using CoLI-Kenglish dataset. The performances of the learning models illustrated, the superiority of CoLI-ngrams model, compared to other models with a macro average F1-score of 0.64. However, the results of all the learning models were quite competitive with each other.

Idioma original	Inglés
Páginas (desde-hasta)	123-141
Número de páginas	19
Publicación	Acta Polytechnica Hungarica
Volumen	19
N.º	10
DOI	https://doi.org/10.12700/APH.19.10.2022.10.8
Estado	Publicada - 2022

Acceder al documento

10.12700/APH.19.10.2022.10.8

Otros archivos y enlaces

Enlace a la publicación en Scopus

Citar esto

@article{9ecad1f177b44f2d8d63cf4a34c0931e,

title = "CoLI-Machine Learning Approaches for Code-mixed Language Identification at the Word Level in Kannada-English Texts",

abstract = "The task of automatically identifying a language used in a given text is called Language Identification (LI). India is a multilingual country and many Indians especially youths are comfortable with Hindi and English, in addition to their local languages. Hence, they often use more than one language to post their comments on social media. Texts containing more than one language are called “code-mixed texts” and are a good source of input for LI. Languages in these texts may be mixed at sentence level, word level or even at sub-word level. LI at word level is a sequence labeling problem where each and every word in a sentence is tagged with one of the languages in the predefined set of languages. For many NLP applications, using code-mixed texts, the first but very crucial preprocessing step will be identifying the languages in a given text. In order to address word level LI in code-mixed Kannada-English (Kn-En) texts, this work presents i) the construction of code-mixed Kn-En dataset called CoLI-Kenglish dataset, ii) code-mixed Kn-En embedding and iii) learning models using Machine Learning (ML), Deep Learning (DL) and Transfer Learning (TL) approaches. Code-mixed Kn-En texts are extracted from Kannada YouTube video comments to construct CoLI-Kenglish dataset and code-mixed Kn-En embedding. The words in CoLI-Kenglish dataset are grouped into six major categories, namely, “Kannada”, “English”, “Mixed-language”, “Name”, “Location” and “Other”. Code-mixed embeddings are used as features by the learning models and are created for each word, by merging the word vectors with sub-words vectors of all the sub-words in each word and character vectors of all the characters in each word. The learning models, namely, CoLI-vectors and CoLI-ngrams based on ML, CoLI-BiLSTM based on DL and CoLI-ULMFiT based on TL approaches are built and evaluated using CoLI-Kenglish dataset. The performances of the learning models illustrated, the superiority of CoLI-ngrams model, compared to other models with a macro average F1-score of 0.64. However, the results of all the learning models were quite competitive with each other.",

keywords = "Code-mixed texts, Deep Learning, Language Identification, Machine Learning, Transfer Learning",

author = "Lakshmaiah, {Shashirekha Hosahalli} and Fazlourrahman Balouchzahi and Anusha, {Mudoor Devadas} and Grigori Sidorov",

year = "2022",

doi = "10.12700/APH.19.10.2022.10.8",

language = "Ingl{\'e}s",

volume = "19",

pages = "123--141",

journal = "Acta Polytechnica Hungarica",

issn = "1785-8860",

number = "10",

}

CoLI-Machine Learning Approaches for Code-mixed Language Identification at the Word Level in Kannada-English Texts. / Lakshmaiah, Shashirekha Hosahalli; Balouchzahi, Fazlourrahman; Anusha, Mudoor Devadas et al.
En: Acta Polytechnica Hungarica, Vol. 19, N.º 10, 2022, p. 123-141.

Producción científica: Contribución a una revista › Artículo › revisión exhaustiva

TY - JOUR

T1 - CoLI-Machine Learning Approaches for Code-mixed Language Identification at the Word Level in Kannada-English Texts

AU - Lakshmaiah, Shashirekha Hosahalli

AU - Balouchzahi, Fazlourrahman

AU - Anusha, Mudoor Devadas

AU - Sidorov, Grigori

PY - 2022

Y1 - 2022

N2 - The task of automatically identifying a language used in a given text is called Language Identification (LI). India is a multilingual country and many Indians especially youths are comfortable with Hindi and English, in addition to their local languages. Hence, they often use more than one language to post their comments on social media. Texts containing more than one language are called “code-mixed texts” and are a good source of input for LI. Languages in these texts may be mixed at sentence level, word level or even at sub-word level. LI at word level is a sequence labeling problem where each and every word in a sentence is tagged with one of the languages in the predefined set of languages. For many NLP applications, using code-mixed texts, the first but very crucial preprocessing step will be identifying the languages in a given text. In order to address word level LI in code-mixed Kannada-English (Kn-En) texts, this work presents i) the construction of code-mixed Kn-En dataset called CoLI-Kenglish dataset, ii) code-mixed Kn-En embedding and iii) learning models using Machine Learning (ML), Deep Learning (DL) and Transfer Learning (TL) approaches. Code-mixed Kn-En texts are extracted from Kannada YouTube video comments to construct CoLI-Kenglish dataset and code-mixed Kn-En embedding. The words in CoLI-Kenglish dataset are grouped into six major categories, namely, “Kannada”, “English”, “Mixed-language”, “Name”, “Location” and “Other”. Code-mixed embeddings are used as features by the learning models and are created for each word, by merging the word vectors with sub-words vectors of all the sub-words in each word and character vectors of all the characters in each word. The learning models, namely, CoLI-vectors and CoLI-ngrams based on ML, CoLI-BiLSTM based on DL and CoLI-ULMFiT based on TL approaches are built and evaluated using CoLI-Kenglish dataset. The performances of the learning models illustrated, the superiority of CoLI-ngrams model, compared to other models with a macro average F1-score of 0.64. However, the results of all the learning models were quite competitive with each other.

AB - The task of automatically identifying a language used in a given text is called Language Identification (LI). India is a multilingual country and many Indians especially youths are comfortable with Hindi and English, in addition to their local languages. Hence, they often use more than one language to post their comments on social media. Texts containing more than one language are called “code-mixed texts” and are a good source of input for LI. Languages in these texts may be mixed at sentence level, word level or even at sub-word level. LI at word level is a sequence labeling problem where each and every word in a sentence is tagged with one of the languages in the predefined set of languages. For many NLP applications, using code-mixed texts, the first but very crucial preprocessing step will be identifying the languages in a given text. In order to address word level LI in code-mixed Kannada-English (Kn-En) texts, this work presents i) the construction of code-mixed Kn-En dataset called CoLI-Kenglish dataset, ii) code-mixed Kn-En embedding and iii) learning models using Machine Learning (ML), Deep Learning (DL) and Transfer Learning (TL) approaches. Code-mixed Kn-En texts are extracted from Kannada YouTube video comments to construct CoLI-Kenglish dataset and code-mixed Kn-En embedding. The words in CoLI-Kenglish dataset are grouped into six major categories, namely, “Kannada”, “English”, “Mixed-language”, “Name”, “Location” and “Other”. Code-mixed embeddings are used as features by the learning models and are created for each word, by merging the word vectors with sub-words vectors of all the sub-words in each word and character vectors of all the characters in each word. The learning models, namely, CoLI-vectors and CoLI-ngrams based on ML, CoLI-BiLSTM based on DL and CoLI-ULMFiT based on TL approaches are built and evaluated using CoLI-Kenglish dataset. The performances of the learning models illustrated, the superiority of CoLI-ngrams model, compared to other models with a macro average F1-score of 0.64. However, the results of all the learning models were quite competitive with each other.

KW - Code-mixed texts

KW - Deep Learning

KW - Language Identification

KW - Machine Learning

KW - Transfer Learning

UR - http://www.scopus.com/inward/record.url?scp=85159043247&partnerID=8YFLogxK

U2 - 10.12700/APH.19.10.2022.10.8

DO - 10.12700/APH.19.10.2022.10.8

M3 - Artículo

AN - SCOPUS:85159043247

SN - 1785-8860

VL - 19

SP - 123

EP - 141

JO - Acta Polytechnica Hungarica

JF - Acta Polytechnica Hungarica

IS - 10

ER -

CoLI-Machine Learning Approaches for Code-mixed Language Identification at the Word Level in Kannada-English Texts

Resumen

Acceder al documento

Otros archivos y enlaces

Huella

Citar esto