Hindi-English code-mixing, the alternation between the two languages within a single utterance, accounts for a substantial share of user-generated content in India. Because such text is written in the Latin script without a standardised transliteration convention, one underlying word surfaces in many spellings and many tokens fall outside English lexica. This paper addresses emotion detection in this setting as a four-way classification problem over anger, fear, happiness and sadness, a finer-grained task than the polarity classification targeted by most prior work. We make three contributions. First, we construct a corpus of 1,589 code-mixed sentences drawn from Twitter and from video-platform comment sections, annotated by two bilingual speakers with an inter-annotator agreement of kappa = 0.94; the video-comment portion covers a longer, less hashtag-dominated register than the Twitter-only corpora used previously. Second, we propose a normalisation procedure that clusters transliteration variants by combining distributional similarity over skip-gram vectors with a hard constraint on consonant identity, since variation is carried almost entirely by vowels. Third, we compare five baselines under a common five-fold cross-validation protocol: Naive Bayes over character and over word n-grams, a word-level LSTM, a sub-word LSTM with convolved character embeddings, and an SVM over frequency-based word vectors. The sub-word LSTM attains the highest accuracy, 76.6%, against a majority-class floor of 30.8%, and the controlled comparison against an otherwise identical word-level LSTM isolates sub-word representation as the source of that gain. On macro-averaged F1, however, it ties with word n-gram Naive Bayes at 0.77, a margin that five-fold cross-validation on 1,589 instances cannot resolve. We report a per-class error analysis and discuss the limitations imposed by corpus scale.
Detection of Emotions in Hindi-English Code Mixed Text Data
Hindi-English code-mixing, the alternation between the two languages within a single utterance, accounts for a substantial share of user-generated content in India. Because such text is written in the Latin script without a standardised transliteration convention, one underlying…
- Year
- 2021
- Hosting
- Full text hostedCC-BY-4.0
Cite
Notes
Only stored in your browser.
Attribution
- Abstract & full text
- arxiv.org/abs/2105.09226CC-BY-4.0
- TL;DR
- Semantic Scholar