Building the Romanian NLP API that should already exist
LexicRo — open-core Romanian language intelligence infrastructure. Looking for early feedback, collaborators, and anchor users.
/analyze returns lemma, part of speech and full morphology for every token in context - 98.14% POS accuracy, 95.50% lemma accuracy (UD Romanian test split, gold tokenisation). Conjugation is live too. Free tier: 1,000 requests/day, no credit card.If you've ever tried to do anything programmatic with Romanian text — parse a sentence, get the correct inflected form of a noun, check whether a verb conjugation is right — you've probably hit the same wall.
There's no clean API for it. Verbecc handles conjugation, but there's nothing for morphology or lemmatisation. This is not good enough for production.
pip install. LexicRo is that as a hosted, versioned HTTP contract for Romanian: callable without shipping a model, stamping the version it answered with, and telling you per token whether the answer came from a lexicon or a prediction.
Romanian NLP tooling lags well behind English. The academic resources are there — DEXonline (313k+ lemmas), RoLEX (330k morphosyntactic entries), the Universal Dependencies Romanian Treebank — they're just not packaged in a way developers can actually use.
That's what LexicRo is for.
What we're building
A hosted REST API — with an open-source core — covering the endpoints Romanian developers actually need:
POST /analyze
→ lemma, POS, case, gender, number, person, tense per token, plus where each answer came from.
# Verb conjugation
GET /conjugate/{verb}
→ seven moods, including perfect simplu and viitor I. For a verb the conjugator does not recognise, returns a predicted paradigm and marks it as predicted.
The core is open source (MIT). It is built on openly licensed Romanian language resources — see attribution — and licensing terms for the model weights are still being worked out. The hosted API is available now either way, with a generous free tier (1,000 req/day, no credit card).
Built on what already exists
We're not starting from scratch. The data and models are there — they just need engineering:
Data: the MULTEXT-East Romanian word-form lexicon (428k entries, CC BY-SA 4.0), and the UD Romanian RRT treebank (9.5k annotated sentences, CC BY-SA 4.0, developed at RACAI). Full citations on the attribution page.
Models: bert-base-romanian-cased-v1 fine-tuned for morphological tagging, paired with lexicon lookup — the model resolves what the dictionary cannot, the dictionary supplies what the model would only guess at. verbecc for conjugation, extended.
Infrastructure: FastAPI, Docker, full OpenAPI spec.
/analyze handles contextual disambiguation, not just
dictionary lookup — it knows that sare is the noun sare ("salt") in
Pune sare în mâncare and the verb sări ("to jump") in
Pisica sare pe masă. Measured on the UD Romanian RRT test split with gold tokenisation:
98.14% part-of-speech accuracy, 98.43% morphological features
(F1), 95.50% lemma accuracy. Full numbers and caveats in the
documentation.
What we're looking for right now
Three ways in.
Everything described here is live. Read the docs, or get in touch if you want to talk about your use case.