estonian-mcp

Estonian NLP tools so AI agents stop hallucinating Estonian.

Community: Submitted by a user or imported; check the owner before granting accessOnlineNo sign-inGlobalFreeRead-only

What it can do

  • Tokenize: Split Estonian text into sentences and words. Returns a dict with `sentences` (list of strings) and `words` (list of strings). Input is capped at 100,000 characters.
  • Analyze Morphology: Run full morphological analysis on Estonian text. For each word returns lemma(s), part-of-speech, grammatical form, root, ending, clitic, compound parts, ambiguity info, and a usag
  • Paradigm: Generate the full inflection paradigm for an Estonian word. For nominals (nouns, adjectives, pronouns, cardinals, ordinals, comparatives, superlatives): produces all 14 cases × 2 numbers = u

What data it sees

Do you need an account

No: the server works without sign-in

Estonian NLP tools so AI agents stop hallucinating Estonian. 26 read-only tools wrapping EstNLTK, EKI Reeglid, and Riigi Teataja: spell-check, morphological analysis, lemmatization, full inflection paradigms, WordNet synonyms, fastText related words, named-entity recognition, orthography/grammar checks (capitalization, compounds, commas, numbers, abbreviation hyphenation, object case, register, style, redundancy, calque risk), editorial checks that catch kantseliit in reports and academic prose and flag a document that names one thing three ways, plus legal-Estonian tools (legalese simplification that protects terms of art, defined-term + cross-reference tracking, and canonical legal-usage collocations from public-domain legislation). Paradigms cover ordinals, comparatives and superlatives, include the short illative (majja beside majasse), and lead every slot with the form built on that paradigm's own stem, so raamatutele comes before the literary raamatuile. A word with two inflection types (kott inflects as either koti or kota, two different words sharing a nominative) returns one consistent table per type instead of a merged one. Offline and PII-free, so confidential text never leaves the machine. Works in Claude, ChatGPT and Codex, and is listed in Anthropic's official Connectors Directory. Scores 99.1% on TalTech's inflection_et and 100% once the 13 gold rows that contradict EKI are adjudicated, plus 99.3% first-form over all fourteen cases on 11,011 rows of real Estonian text. Use it as ground truth so the model stops inventing Estonian spelling and case forms and stops guessing whether a word fits its register or its domain.

Server tool list (26)

Raw names from tools/list. Only developers need these.

tokenizeSplit Estonian text into sentences and words. Returns a dict with `sentences` (list of strings) and `words` (list of strings). Input is capped at 100,000 characters.
analyze_morphologyRun full morphological analysis on Estonian text. For each word returns lemma(s), part-of-speech, grammatical form, root, ending, clitic, compound parts, ambiguity info, and a usage note flagging archaic / foreign / abbreviation / interjection / proper-noun cases. By default returns the first (most likely) analysis per word; set `all_analyses=True` to return every ambiguous analysis. Each word's response includes: - lemma, partofspeech, form, root, ending, clitic, root_tokens - analyses_count: how many alternative analyses Vabamorf produced for this surface form (>1 means the word is morphologically ambiguous) - is_ambiguous: shorthand for analyses_count > 1 - usage_note: machine code (None if neutral) — "archaic" / "foreign" / "abbreviation" / "interjection" / "proper-noun" - usage_note_estonian: human-readable Estonian rendering of the same flag (quote this verbatim in Estonian replies; do NOT translate the English usage_note yourself) - indeclinable: True for words that stay in base form when used attributively (lexical indeclinables like `täis`, -tud/-nud past participles like `tuntud`, and the -mata form like `täitmata`) — i.e. they do NOT take the noun's case ending in agreement. Use this before inflecting a noun phrase so you don't wrongly decline an invariant adjective. Input is capped at 100,000 characters.
paradigmGenerate the full inflection paradigm for an Estonian word. For nominals (nouns, adjectives, pronouns, cardinals, ordinals, comparatives, superlatives): produces all 14 cases × 2 numbers = up to 28 forms, plus the SHORT ILLATIVE (`adt`, ainsuse lühike sisseütlev) for the words that have one: `majja` beside `majasse`, `kätte` beside `käesse`. Both are correct illatives and the short one is often the commoner, so neither should be "corrected" into the other. For most words that have one, though, the short illative is spelled exactly like the singular partitive (`vend`: `venda` is both), so a surface appearing in this table does not confirm the case is right where it was used. For verbs: produces infinitives, the indicative present and past, the conditional, the imperative, the participles, and the umbisikuline tegumood's own finite forms (`kasutatakse`, `kasutati`, `kasutataks`), about 39 in all. Other parts of speech (adverbs, conjunctions, particles) don't inflect, so `forms` is empty. Each form entry has the Vabamorf form code (e.g. `sg p`, `ksin`), its Estonian label (e.g. `ainsuse osastav`, `tingiv 1.p ainsus`), and the surface form Vabamorf generated. Use `form_estonian` verbatim in Estonian replies — don't translate the English `form` code. AMBIGUOUS LEMMAS. Some lemmas belong to more than one inflection type (`kott` inflects as either `koti` or `kota`, two different words that share a nominative). `forms` is then one internally consistent paradigm, `paradigm_key` names it by its singular genitive, `paradigm_count` says how many exist, and the rest are in `other_paradigms`. Read `ambiguity_estonian` before quoting a form. **Pass an inflected form (`koti`) rather than the bare lemma when you know which word you mean**: that selects the paradigm exactly. Phase-1 scope: covers the most commonly-needed forms per word class, not every theoretical form Vabamorf can produce. Single-word input, capped at 200 characters.
lemmatizeReturn lemma (dictionary form) for each word in the text. Concise output: `[{"word": ..., "lemma": ...}, ...]`. Input is capped at 100,000 characters.
pos_tagReturn part-of-speech tag for each word. POS tag set: S=noun, V=verb, A=adj, P=pron, D=adv, K=adp, J=conj, N=numeral, I=interj, Y=abbrev, X=foreign, Z=punct, etc. Input is capped at 100,000 characters.
spell_checkCheck Estonian spelling for each word and optionally return suggestions. Returns one entry per word with `text`, `spelling` (bool), and `suggestions` (list of correction candidates) when `suggestions=True`. Input is capped at 100,000 characters. CAVEAT: Vabamorf accepts ANY morphologically well-formed word, including compounds you just invented (e.g. `toortõlkeoht`) — it splits them into valid roots and reports `spelling: true`. So passing spell_check does NOT mean a word is real, attested Estonian. For a coined or unusual compound, confirm it with `check_compound_familiarity` before trusting it.
syllabifySplit a single Estonian word into syllables with quantity and accent. Each syllable entry: `{"syllable": str, "quantity": int, "accent": int}`. Input is capped at 200 characters and must contain no whitespace.
named_entitiesExtract named entities (PER/LOC/ORG) using EstNLTK's CRF model. Returns `[{"text": ..., "type": ..., "start": ..., "end": ...}, ...]`. Input is capped at 100,000 characters.
find_related_wordsFind Estonian words semantically similar to the input via fastText. Returns the top-n nearest neighbours by cosine similarity over a pre-trained Estonian fastText model (Common Crawl + Wikipedia, 2018). Useful for breaking repetition, finding alternative phrasings, or expanding vocabulary when WordNet's exact-meaning synonyms aren't enough. Distinct from `synonyms`: that one returns WordNet synsets — words with the same meaning. This one returns words that *pattern* with the input in real Estonian text, which can include near-synonyms, related concepts, and (sometimes) antonyms. Known quirks of the embedding model: - **Inflections crowd the top results** for some words. fastText sees `kasutama` and `kasutada` as related because the surface forms share subword n-grams; you may want to lemmatize matches yourself to dedupe. - **Antonyms can appear** because antonyms occur in similar contexts (`tark` may surface `loll`). Treat the list as "semantically nearby" rather than "synonymous." - **Polysemy is not disambiguated.** `lahe` (which means both "bay" and the colloquial "cool") will return whichever sense dominates the training data. Single-word input only, capped at 200 characters.
synonymsLook up Estonian synonyms via WordNet. Returns synsets (groups of synonymous lemmas) for the input word, each with its definition and example usages. Useful when you want Claude to pick a different word with the same meaning, e.g. swap an over-used verb in marketing copy. Word-sense ambiguity is preserved: a polysemous word returns multiple synsets, one per meaning. Input capped at 200 characters. WORD-FIT CHECK: when the question is "is this the right word here?" rather than "give me an alternative", READ EACH `definition` and test it against the user's actual context — do not just harvest `lemmas`. Estonian glosses routinely carry a domain constraint that decides the answer: `korpus` returns the sense "kirjaliku või suulise teksti elektrooniline kogu", so calling a set of IMAGES a `korpus` is wrong however natural it sounds in ML jargon; `andmestik` carries no such constraint. A gloss naming a medium, field, or material is a constraint on where the word may be used. Note also that a word can be well-formed, correctly spelled and still the wrong register — for that, check_officialese and classify_register, not this tool.
check_compoundsHeuristic Estonian compound-word check (liitsõnaõigekiri). Scans for common AI-generated splits of words that should be written as a single compound — `kooli maja` (wrong) → `koolimaja` (right), `nädala vahetus` (wrong) → `nädalavahetus` (right), etc. Uses a curated bigram lexicon (~30 entries covering the highest-frequency AI mistakes); not a full liitsõnaõigekiri solver. Phase-1 limitations: only catches the bigrams in the lexicon. Estonian compounding is highly productive and most valid compounds aren't enumerated here. Treat hits as high-confidence; absence of hits does not prove the compound writing is correct everywhere. Input capped at 100,000 characters.
check_punctuationHeuristic Estonian punctuation check — comma-before-clause rule. Flags missing commas before subordinating conjunctions where Estonian rules require one: et (that/in order to), kuna (because), sest (because), kuigi (although), kuid (but), vaid (rather), nagu (like), mistõttu (because of which), millepärast, kuhu. Phase-1 limitations: only the comma-before-clause-conjunction rule is covered. `kui`, `mis`, `kes` are deliberately excluded because their function is contextual (kui = than/as in comparisons doesn't need a comma). Listing commas, apposition commas, dash and colon rules — all out of scope for phase 1. Input capped at 100,000 characters.
check_hyphenationReturn safe line-break positions for an Estonian word (poolitamine). Different from `syllabify` (which is phonological): this returns character offsets where a typesetter can legally break the word across lines. Applies the no-orphan-edge rule (don't leave fewer than 2 characters before or after the break point). Phase-1 limitation: pure syllable-boundary based. Compound-boundary preference (Estonian poolitamine prefers `kooli-maja` over `koo-limaja`) is not yet applied. Input must be a single word with no whitespace, capped at 200 characters.
check_numbersHeuristic Estonian number-writing check. Flags two clear-cut cases per EKI Reeglid: - Decimal separator: Estonian uses a comma (3,14), not a period (3.14). - Thousands separator: Estonian uses a space (1 000 000), not a comma (1,000,000). Phase-1 limitations: spell-out-vs-digits guidance (the one-to-ten-spelled-out convention) is intentionally not implemented — it requires distinguishing measurements, dates, years, and lists from running prose, and naive flagging produces too many false positives. Input capped at 100,000 characters.
check_capitalizationHeuristic Estonian capitalization checker (Algustäheortograafia). Scans Estonian text for the most common AI-generated capitalization errors per EKI's Reeglid: - Weekday names capitalized mid-sentence (Esmaspäeval → esmaspäeval) - Month names capitalized mid-sentence (Jaanuaris → jaanuaris) - Nationality names capitalized mid-sentence (Eestlane → eestlane) - Country/language adjectives capitalized before a culture or language noun (Eesti keel → eesti keel; Eesti köök → eesti köök). The bare capitalized form on its own (Eesti, Eestis) is left alone because it's a valid country proper-noun usage. Sentence-initial capitalization is always allowed. All-caps acronyms are ignored. Returns each issue with rule code, an Estonian rule label (`rule_estonian` — quote this verbatim in Estonian replies, don't translate the English `rule`), a user-facing explanation, and a suggested correction. Input capped at 100,000 characters. PHASE-1 LIMITATION: this is a lexicon-based heuristic, not a full EÕS implementation. Compound-word capitalization, punctuation rules, and hyphenation are NOT covered by this tool (separate check_compounds / check_punctuation / check_hyphenation tools may follow).
check_compound_familiarityfastText-based diagnostic for compound-noun familiarity in Estonian. For each compound noun (root_tokens length >= 2), returns its top fastText neighbours, a `top_score` similarity, a `neighbour_quality` breakdown, and `is_suspect: true` + human-readable `reasons` when the compound is out-of-vocab AND its top similarity is below 0.60 OR its neighbours are mostly scrape-artifact tokens. This catches both `toortõlkeoht` (OOV, top 0.571 — over the old 0.55 gate but a coinage) and `mõtteliin` (literal English "train of thought"; real Estonian is `mõttekäik`). Output is diagnostic, not authoritative. Even with the 100K-vocab medium model, some legitimate but rare compounds (e.g. `tervisekindlustus`) can still be OOV; the rule favours recall, so a flagged real compound just earns a second look. Judge by the included neighbours: semantically coherent neighbours (related real words) mean the compound is fine; neighbours that recycle the input's morphemes or are junk tokens mean a likely coinage. Input capped at 100,000 characters.
check_abbreviation_hyphenationHeuristic check for the EKI Reeglid rule that case endings on Latin-letter / all-caps abbreviations are separated by a hyphen. Catches the common AI mistake of writing `MCPst`, `APIga`, `OÜle` instead of `MCP-st`, `API-ga`, `OÜ-le`. Uses Vabamorf's POS+form analysis to identify tokens recognised as abbreviations carrying a case ending; only flags those that aren't already hyphenated. Phase-1 scope: matches what Vabamorf tags as `Y` (abbreviation). Custom acronyms Vabamorf doesn't know (your brand acronym, niche industry shorthand) won't be flagged because Vabamorf doesn't see them as abbreviations. Input capped at 100,000 characters.
check_object_caseHeuristic Estonian object-case-government check. Catches the single biggest class of confidently-wrong Estonian that AI agents produce: direct objects in the wrong case after negation or after partitive-governing verbs. Two rules in phase 1: - **Negation → partitive**: any sentence containing 'ei', 'pole', 'ära', 'ärge', 'ärgu', 'ärgem', or 'mitte' must have direct objects in partitive. Flags nominative / genitive nouns. - **Partitive-only verbs**: the verbs `armastama`, `vihkama`, `vajama`, `soovima`, `ootama`, `austama`, `kartma`, `puudutama`, `tundma` always take partitive direct objects. Flags any noun in nominative/genitive in the same sentence. Phase-1 limitation: no syntactic parser, so we can't perfectly distinguish subject from object. Subjects in negation/partitive-verb sentences may be flagged as false positives. Treat hits as "worth a second look", not authoritative. Proper nouns are skipped. Input capped at 100,000 characters.
check_redundancyHeuristic Estonian pleonasm / semantic-doubling check. Flags phrasing that is grammatically valid but reads redundant to a native speaker — the class of error AI agents produce when they stack synonyms. Phase-1 rules, all high-precision: - **Doubled 'also' particles**: `samuti ka`, `ka samuti`, `ühtlasi ka` — both words mean "also/too", so together they're a tautology. (This is the exact `samuti ka suvesärgid` case.) - **Double superlative**: `kõige` before an already-absolute adjective (`optimaalne`, `ideaalne`, `maksimaalne`, `täiuslik`, `ainus`, …) — like English "most optimal". Lemma-matched, so all inflected forms count. - **Fixed pleonasm phrases**: a small curated set (`ajaline periood`, `väike nüanss`, `üldine konsensus`, …). Conservative by design — it catches the obvious, high-confidence cases, not every redundancy. Absence of flags is not proof the prose is tight. Input capped at 100,000 characters.
check_legaleseAid for simplifying Estonian legal text WITHOUT losing legal precision. Returns two things: - `issues`: archaic 'kantseliit' filler with plain equivalents (`käesolev` → `see`, `juhul kui` → `kui`), plus over-long / heavily subordinated sentences worth splitting. - `terms_of_art`: specialised legal terms detected in the text that MUST be kept verbatim when rewriting — swapping `hagi` or `vastutus` for a general synonym changes the legal meaning. Use this as a do-not-touch list while you simplify. Heuristic, precision-first, backed by curated starter lexicons (not exhaustive). Input capped at 100,000 characters.
check_defined_termsStructural map of a long Estonian legal document. Extracts every term defined with `(edaspidi «X»)`, counts how often each is actually used, lists `§` / `lõige` / `punkt` / `artikkel` cross-references, and flags defined-but-unused or doubly-defined terms — the consistency errors that creep into long contracts and statutes. Regex-based and PII-free (nothing is stored). Input cap is raised to 500,000 characters so a whole contract fits in one call; the echoed `text` is truncated to a 2,000-character preview.
common_legal_usageCanonical legal collocations for a term, from an offline corpus index. Answers "what's the standard legal phrasing" — returns how often the term occurs in Estonian legal text and the words most frequently seen directly before/after it (`hagi` → `esitama` before it = 'esitama hagi'; `kohustus` → `täitmine` after it = 'kohustuse täitmine'). Use it so the AI picks real, idiomatic legalese instead of inventing collocations. A frequency signal, not prescriptive; coverage is bounded by the corpus. The bundled index is a proof-of-concept sample; the full-corpus artifact is loaded via ESTNLTK_MCP_LEGAL_INDEX. Input is a single word.
check_styleHeuristic Estonian style metrics for newsletter / ad / email copy. Returns four metrics that flag common writing issues, each with an Estonian-language summary line for quoting verbatim: - repetition: lemma-aware (so 'kasutab' and 'kasutamine' both count under 'kasutama'). Threshold scales with text length so short replies don't fire on natural repeats. - passive_voice: ratio of Estonian -takse/-ti/-tud/-tav forms over total verbs. Newsletter copy usually wants <15%. - sentence_length: mean, stddev, min, max in content words. Low stddev = monotonous rhythm. - hedging: density of hedging words (võib-olla, vist, pigem, ehk, ilmselt, …). >5% reads wishy-washy. Phase-1 limitation: heuristic only. No detection of cliché phrases, weasel-words beyond the curated 15 lemmas, or genre-specific style drift. Input capped at 100,000 characters.
check_officialeseFlag Estonian kantseliit in reports, academic and business prose. check_legalese is the legal-text sibling; use THIS one for anything that is not a statute or contract, where check_legalese finds nothing because its lexicon and length gate are tuned for legislation. Returns `issues` (each with an Estonian `rule_estonian` label, an explanation and a concrete suggestion) plus `metrics`: - nominalisation: -mine verbal nouns per 100 words, each paired with the verb to use instead (`hindamine` → `hindama`) - noun_verb_ratio: nimisõnastiil density; over ~2.0 reads heavy - impersonal_voice: umbisikuline tegumood, correctly counted (negated impersonals included, `ei`/`ära` and attributive -tud participles excluded) - clause-stacking: 3+ subordinate-clause openers in one sentence, the 'mille käigus … ning …' pile-up that a word count alone misses - long-sentence: 25+ content words, calibrated for Estonian - officialese filler: `omama` → `olema`, `kujutab endast` → `on`, `viidi läbi` → `tehti`, `X-i poolt tehtud` → `X-i tehtud` Heuristic and precision-first — no flags does not prove the text is plain. Input capped at 100,000 characters.
check_term_consistencyFlag a document that calls the same thing several different names. The classic long-document defect, and the one a model editing paragraph-by-paragraph reliably misses: a dataset that is `andmestik` on page 1, `teadusandmestik` on page 2 and `korpus` on page 3. Two precision-first rules: - `shared-compound-head`: a bare noun and a compound built on it both occur (`andmestik` + `pildiandmestik`), or 3+ lemmas share one head. - `shared-wordnet-synset`: two lemmas sit in one Estonian WordNet synset, i.e. WordNet calls them synonyms. Each group lists its variants with occurrence counts and the dominant one, so you can standardise on the most-used term. The tool does not decide which variant is right — some groups are genuinely distinct concepts, so read them before rewriting. CHECK `degraded` BEFORE TRUSTING AN EMPTY RESULT. When Estonian WordNet is not installed, the `shared-wordnet-synset` rule cannot run; the tool then returns `degraded: true`, says so in `summary_estonian`, and marks the rule false in `rules_run`. "No groups found" from a degraded run means "the compound-head rule found nothing", NOT "the terminology is consistent". Known gap: synonyms sharing neither a head nor a synset (korpus / andmestik) are not caught. Input capped at 100,000 characters.
classify_registerHeuristic register classifier for Estonian (formal vs colloquial). Returns a tier label (English in `tier`, correct Estonian in `tier_estonian` — quote that field verbatim when composing an Estonian-language reply rather than translating `tier` yourself, to avoid mistranslations like "formalne" instead of the correct "formaalne"), a normalised score in [-1, 1] (positive = formal, negative = colloquial), and the matched formal/colloquial markers found in the text. Useful for sanity-checking that marketing copy hasn't drifted into officialese, or that a contract draft hasn't slipped into chat tone. The lexicon covers legal-administrative AND academic/report vocabulary. `structure` adds two syntactic signals — umbisikuline tegumood ratio and noun/verb density — bounded at +0.4, applied only from 25 words up and only when the lexicon is not net-colloquial. LIMITATION: still a heuristic, not a trained model. Address forms and finer syntax go uncaught, and most newsletter prose scores 'neutral'. Use the result as a directional hint, not a verdict; for a full kantseliit breakdown with per-issue suggestions, call check_officialese. Input capped at 100,000 characters.