NLP for Nepali
I TRIED!!
we started with a relatively simple question.
can we build sentiment analysis for nepali?
it was our major project. we wanted to work with natural language processing, and nepali felt like the obvious problem to work on. there are millions of people using the language, but the amount of language technology available around it is nowhere close to what exists for english.
so we started looking at what actually exists.
and very quickly, the problem stopped being sentiment analysis.
the thing about low-resource languages
when you build nlp for english, you inherit an entire ecosystem.
datasets, pretrained models, benchmarks, tokenizers, libraries, research, tooling.
you can spend most of your time thinking about the model.
with nepali, you start noticing everything that is missing.
less annotated data.
smaller language resources.
different writing styles.
romanized nepali.
spelling variations.
mixed languages.
emojis.
informal text.
and suddenly the question isn't just:
"which model should we use?"
it becomes:
"what exactly are we training this model to understand?"
our paper describes nepali as a low-resource language where limitations still exist around annotation scale, vocabulary coverage and practical integration.
that sounds like a research problem.
when you actually try building something, it becomes an engineering problem.
then came the data
our sentiment pipeline had 7,989 core rows, another 1,885 supplementary rows and 1,199 held-out samples.
it sounds like a lot until you start thinking about what those numbers actually represent.
a dataset is not just a number of rows.
it is the language contained inside those rows.
how people express themselves.
how people spell.
how people mix languages.
how people express emotion.
how many different ways someone can say essentially the same thing.
and whether the dataset actually represents the people who will eventually use the system.
this is where low-resource nlp becomes difficult.
you can have enough data to train a model.
but not necessarily enough data to understand the language.
sentiment is not just positive or negative
we decided to work with five sentiment classes.
negative.
semi-negative.
neutral.
semi-positive.
positive.
and this made the problem more interesting.
because the difficult part wasn't identifying obviously positive or obviously negative text.
the difficult part was the boundary.
our model performed particularly well on positive sentiment, with an f1 score of 0.9206.
but neutral was much harder, with an f1 score of 0.5899.
semi-negative was also difficult, with an f1 score of 0.6563.
the confusion wasn't completely random.
the model struggled mostly with neighbouring sentiment categories.
and that tells you something important about language.
people don't communicate in five clean categories.
machines do.
humans don't.
the model wasn't the interesting part
we used word-level tf-idf.
character-level tf-idf.
affective signals.
and multiclass logistic regression.
nothing about that sounds particularly futuristic.
and that was actually one of the things i found interesting.
there is a tendency in ai to assume that every problem needs a bigger model.
more parameters.
more compute.
more layers.
more complicated architectures.
but when you are working with a low-resource language, the equation changes.
you have limited data.
you have limited compute.
you need something that can actually be deployed.
and sometimes a relatively simple model with the right features is more useful than a much larger model that you cannot properly control, interpret or deploy.
then we realized sentiment wasn't enough
real nepali text doesn't arrive as a perfectly formatted sentence waiting to be classified.
people make mistakes.
they misspell words.
they write in romanized nepali.
they mix scripts.
they use emojis.
they type the same word in several different ways.
so we started thinking about the system as more than sentiment analysis.
what if the same pipeline could also correct spelling?
what if it could predict the next word?
that became the architecture.
one api.
one input.
different tasks routed through the same system.
sentiment analysis.
spelling correction.
next-word prediction.
building around the language
for spelling correction, we used edit distance to generate candidates and then reranked them using contextual and frequency information.
for next-word prediction, we used trigram and bigram statistics with lexical priors and recency-aware ranking.
again, these aren't necessarily the most fashionable approaches.
but they fit the problem.
the spelling system achieved 0.9880 detection accuracy and a 0.9760 exact correction rate.
the next-word prediction system achieved 0.7350 top-1 precision, 0.9010 hit@3 and 0.9413 hit@5.
the number i cared about most
our sentiment model achieved 0.7139 accuracy and 0.7180 macro f1.
those numbers were good enough to show that the approach worked.
but the number itself wasn't the biggest takeaway.
the bigger takeaway was understanding what the number actually means.
it doesn't mean:
"the machine understands nepali sentiment 71% of the time."
it means that under a particular dataset, task definition and evaluation setup, the model achieved that performance.
that's a much more honest way of looking at machine learning.
because the moment you move outside the dataset, everything can change.
the real problem is generalization
our own paper acknowledges this.
the spelling evaluation used artificial typo data.
sentiment ambiguity remained between neighbouring classes.
and the system still needed more testing across completely new domains.
this is probably one of the hardest parts of building language technology for nepali.
you can build something that works.
you can measure it.
you can put the number in a paper.
and you can still have no idea how well it will behave when thousands of people start using it differently from your dataset.
that's not a nepali problem.
that's machine learning.
but low-resource languages make the problem much more obvious.
so what did i actually learn?
i started this project thinking i was going to learn sentiment analysis.
i ended up learning about data.
about language.
about evaluation.
about feature engineering.
about deployment.
and about the difference between building a model and building a language system.
the biggest lesson was probably this:
the model is only one part of the problem.
the real challenge is creating enough language resources around it for the model to have something meaningful to learn from.
and this is where nepali nlp gets interesting
we don't necessarily need to wait until nepali has the same amount of data as english.
we can build with what we have.
hybrid approaches.
better datasets.
better benchmarks.
better normalization.
real-world user data.
open-source resources.
systems designed specifically around the way nepali is actually written.
our paper proposed future work around collecting real typo data, improving normalization for different romanized spellings, and allowing models to continuously learn and detect drift over time.
there is still a lot left to build.
and maybe that's the interesting part.
because when you work on a language that already has an enormous research ecosystem, you are often improving something that already exists.
when you work on a low-resource language, sometimes you are still building the ecosystem itself.
the datasets.
the benchmarks.
the tools.
the systems.
the knowledge.
and eventually, the infrastructure that makes it possible for the next person to build something that we couldn't.
that's probably what this major project gave me.
not a perfect sentiment analysis model.
but a much better understanding of what it actually means to build ai for a language that the ai ecosystem hasn't fully caught up with yet.