What Is Named Entity Recognition (NER)? A Complete Guide

Analyzing text data on a screen, illustrating named entity recognition

Named entity recognition (NER) is a natural language processing task that automatically finds and classifies real-world entities — people, organizations, locations, dates, monetary amounts and more — inside unstructured text. It is the step that turns a wall of words into structured, queryable facts: given the sentence "Apple opened a new lab in Zurich in March," a NER system tells you that Apple is an organization, Zurich is a location and March is a date. That deceptively simple ability powers search engines, media monitoring platforms, document anonymization and knowledge graphs. This guide covers what NER is, the main entity types, how the technology evolved from hand-written rules to transformers and LLMs, how entity linking resolves ambiguity, where NER is used in production, which tools to reach for, and the challenges that still make it a hard problem.

Table
  1. What exactly does named entity recognition do?
  2. The main types of named entities
  3. How NER works: from rules to LLMs
    1. 1. Rule-based and gazetteer systems
    2. 2. Classical machine learning
    3. 3. Deep learning and transformers
    4. 4. Large language models
  4. Entity linking and disambiguation: from mentions to meaning
  5. What NER is used for in the real world
    1. Media monitoring and brand intelligence
    2. Search and recommendation
    3. Anonymization and PII redaction
    4. Knowledge graphs and analytics
    5. Document processing pipelines
  6. Tools and libraries for NER
  7. Why NER is still hard
  8. Frequently asked questions about named entity recognition
    1. What is the difference between NER and entity linking?
    2. Do I still need to train a NER model, or can I just use an LLM?
    3. How accurate is named entity recognition?
    4. Which entity types can NER detect?

What exactly does named entity recognition do?

NER performs two jobs at once. First, detection: it locates the exact span of text that mentions an entity ("Barack Obama", "European Central Bank", "$4.2 billion"). Second, classification: it assigns each span a category such as PERSON, ORG or MONEY. The output is a set of labeled spans that downstream software can index, count, link or redact.

The task was formalized in the 1990s at the MUC-6 evaluation conference, where "named entity" originally covered names of people, organizations and locations plus numeric expressions like dates and amounts. Three decades later the core idea is unchanged, but the category inventories, the methods and the accuracy have all moved dramatically.

It helps to see NER as one layer in a larger language-understanding stack. Below it sit tokenization and part-of-speech tagging; above it sit relation extraction ("who acquired whom?"), event detection and full natural language understanding. NER is often the first semantic layer — the point where the machine stops seeing strings and starts seeing things.

Readers of this site may remember that we explored this question long before it was fashionable: our earlier essay on how machines identify who is who looked at the philosophical and practical puzzle of teaching software to recognize identity in text. This guide is the modern, systematic companion to that piece.

The main types of named entities

There is no single universal tag set — each corpus and product defines its own — but most systems build on a common core:

  • PERSON — names of real or fictional people ("Marie Curie", "Sherlock Holmes").
  • ORG — companies, institutions, agencies, sports teams ("UNESCO", "Bayern Munich").
  • LOC / GPE — physical locations and geopolitical entities ("Alps", "Portugal", "Brooklyn").
  • DATE / TIME — absolute and relative temporal expressions ("July 2026", "last Tuesday").
  • MONEY / PERCENT / QUANTITY — numeric expressions with units.
  • PRODUCT / WORK_OF_ART / EVENT / LAW / LANGUAGE — extended types found in richer schemes such as OntoNotes, which defines 18 categories.

Specialized domains extend this much further. Biomedical NER tags genes, proteins, diseases and drug names; legal NER tags statutes, courts and case citations; financial NER tags tickers, instruments and regulatory filings. A common design decision in real projects is whether to use a coarse scheme (4–7 types, easier to annotate consistently) or a fine-grained one (dozens or hundreds of types, more useful downstream but harder to train). Fine-grained typing — deciding that "Apple" is not just an ORG but a technology company — remains an active research area.

How NER works: from rules to LLMs

The history of NER is a compressed history of NLP itself. Four generations of methods coexist in production today, and each still has a legitimate niche.

1. Rule-based and gazetteer systems

The earliest systems combined gazetteers (long lists of known names — countries, cities, company registries) with hand-written patterns: capitalized word sequences following "Mr.", tokens matching a date regex, strings ending in "Inc." or "GmbH". Rules are transparent, fast, cheap and easy to fix — and they still shine for closed, well-formatted categories such as dates, IBANs, email addresses or product codes. Their weakness is recall on anything unseen: a gazetteer cannot recognize a startup founded yesterday, and patterns collapse when text is lowercase, noisy or in a language with no capitalization cues.

2. Classical machine learning

From the early 2000s NER was reframed as sequence labeling: each token gets a tag from a scheme like BIO (B-PER = beginning of a person, I-PER = inside, O = outside). Models such as Hidden Markov Models and especially Conditional Random Fields (CRFs) learned from annotated corpora, using engineered features — capitalization, suffixes, neighboring words, gazetteer membership. These systems generalized far better than rules but depended heavily on feature engineering and thousands of hand-labeled sentences per domain.

3. Deep learning and transformers

Neural models removed the feature engineering. First came BiLSTM-CRF architectures over word embeddings; then, from 2018, transformer encoders like BERT and RoBERTa reset the state of the art. Pretrained on billions of words, a transformer produces contextual representations, so the same string "Washington" gets different vectors in "Washington signed the bill" and "flew to Washington" — resolving ambiguity that stumped every earlier method. Fine-tuned transformers routinely exceed 92–94% F1 on the classic CoNLL-2003 benchmark and remain the default choice for high-volume production NER: accurate, fast on a GPU (or even CPU with distilled models), and cheap per document.

4. Large language models

LLMs such as GPT-4-class and open-weight instruction models changed the economics of getting started. With zero-shot or few-shot prompting ("Extract all people, organizations and locations from this text as JSON") they perform respectable NER with no training data at all, handle brand-new entity types on request, and can explain their decisions. The trade-offs are real: higher latency and cost per document, output-format drift, occasional hallucinated spans, and — on well-defined standard categories — accuracy that still typically trails a fine-tuned transformer. In practice LLMs excel at low-volume, high-variety extraction and at bootstrapping: generating silver-standard training data used to fine-tune a small, cheap specialist model.

ApproachHow it worksStrengthsWeaknessesBest for
Rules + gazetteersPatterns, regexes, name listsTransparent, fast, no training data, easy to patchPoor recall on unseen names; brittle on noisy textDates, IDs, closed vocabularies, compliance patterns
Classical ML (CRF)Sequence labeling with engineered featuresGeneralizes beyond lists; lightweight; well understoodNeeds labeled data and feature engineeringConstrained domains, low-resource hardware
Transformers (BERT-style)Fine-tuned contextual encodersState-of-the-art accuracy; cheap at scale; handles ambiguityNeeds annotated data per domain; retraining to add typesHigh-volume production pipelines
LLMsZero/few-shot prompting or fine-tuningNo training data needed; flexible types; fast prototypingCost, latency, format drift, occasional hallucinationPrototyping, rare types, data generation, long-tail domains

Entity linking and disambiguation: from mentions to meaning

Recognizing that "Paris" is a location is only half the battle. Which Paris — the French capital, Paris, Texas, or Paris Hilton (not a location at all)? This is the job of named entity disambiguation and entity linking (NEL): mapping each recognized mention to a unique entry in a knowledge base such as Wikidata, Wikipedia or an internal company registry.

A typical linking pipeline has three stages:

  • Candidate generation — retrieve plausible knowledge-base entries for the mention ("Paris" → ~40 candidates), using alias tables and fuzzy matching.
  • Disambiguation — score candidates against the surrounding context. If the sentence mentions the Seine and croissants, the French capital wins; if it mentions Texas highways, the other one does. Modern systems embed mention context and candidate descriptions in the same vector space and pick the nearest match.
  • NIL detection — recognize when the entity is not in the knowledge base at all (a new startup, a private individual) and either create a new entry or flag it.

Closely related is coreference resolution: understanding that "the company", "it" and "the Cupertino giant" all refer to the same Apple mentioned two paragraphs earlier. Together, NER + linking + coreference turn a document into a small graph of uniquely identified entities and relations — the raw material of every knowledge graph.

What NER is used for in the real world

Media monitoring and brand intelligence

The flagship application. Monitoring platforms ingest millions of news articles, broadcasts and social posts per day, and NER is what lets them answer "show me every mention of my brand, my executives and my competitors — and who else appears alongside them." Entity-level tracking, rather than plain keyword matching, is what separates modern tools from a saved search: it distinguishes Orange the telecom from orange the fruit. We cover this ecosystem in depth in our guide to what media monitoring is and how it works.

Search and recommendation

Search engines run NER on both queries and documents. Recognizing that the query "jaguar speed" probably concerns the animal while "jaguar F-Pace price" concerns the carmaker lets the engine route each to the right results, build entity panels, and answer questions directly. E-commerce and enterprise search use the same trick to map free-text queries onto product and document catalogs.

Anonymization and PII redaction

Privacy regulations such as GDPR and HIPAA require removing personal data — names, addresses, phone numbers, ID numbers — before documents are shared, stored or used to train models. NER-based redaction automates this at a scale no human team can match. It is a high-stakes use case: a single missed name is a compliance breach, so these pipelines favor recall and usually keep a human review step.

Knowledge graphs and analytics

Every large knowledge graph — from Google's to a pharma company's internal drug-interaction graph — is fed by extraction pipelines in which NER identifies the nodes and relation extraction supplies the edges. Financial firms mine filings and news for company–event pairs; intelligence and OSINT analysts map people–organization networks; scientific curation teams extract genes and diseases from millions of papers.

Document processing pipelines

A huge share of business text starts life as scans and PDFs. The standard pattern is a two-stage pipeline: optical character recognition converts the image to text, then NER pulls out the parties, dates, amounts and clauses that matter — turning invoices, contracts and forms into database rows. OCR quality directly caps NER quality, which is why choosing the right engine matters; see our roundup of the best OCR software for the front half of that pipeline.

Tools and libraries for NER

The ecosystem is mature; the practical question is usually build-vs-call.

  • spaCy — the industrial-strength Python default. Ships pretrained pipelines for 25+ languages, transformer-backed accuracy modes, and a clean API for training custom entity types. Excellent speed/accuracy balance for production.
  • Hugging Face Transformers — thousands of ready token-classification models (BERT, RoBERTa, XLM-R, DeBERTa fine-tunes) plus the tooling to fine-tune your own. The route to state-of-the-art accuracy and to multilingual coverage via XLM-R-style models.
  • Stanza — Stanford's neural NLP library, notable for consistent, linguistically careful pipelines across 60+ languages and for biomedical/clinical NER models.
  • Flair, GLiNER and friends — Flair popularized contextual string embeddings; GLiNER-style models perform zero-shot NER for arbitrary type names with a compact model, a very practical middle ground between fixed-schema transformers and full LLMs.
  • Cloud APIs — Google Cloud Natural Language, Amazon Comprehend and Azure AI Language offer NER (and PII detection) as a pay-per-call service. Zero infrastructure and solid quality, at the cost of per-document pricing, data-residency questions and limited custom types (though all three now support custom-entity training).
  • Annotation tooling — if you train your own models, labeling tools such as Prodigy, Label Studio or doccano are where projects actually succeed or fail: consistent guidelines and a few thousand well-annotated sentences beat any architecture change.

A sensible default in 2026: prototype with an LLM or GLiNER to validate the schema, then fine-tune a spaCy or Hugging Face transformer on a few thousand examples for the production workload, and keep a rule layer for the closed categories.

Why NER is still hard

Benchmark scores above 90% can suggest a solved problem. Practitioners know better — the gap between CoNLL-2003 and your documents is where projects stall.

  • Ambiguity is everywhere. "Jordan" is a country, a river, a surname and a brand. Context usually disambiguates, but short texts (tweets, queries, table cells) offer little context to work with.
  • Domain shift. A model trained on news collapses on clinical notes, legal contracts or game chat. Vocabulary, entity types and writing conventions all differ; every serious deployment budgets for domain adaptation.
  • Language inequality. English enjoys lavish training data; most of the world's languages do not. Capitalization cues vanish in German (all nouns capitalized) and in scripts with no case at all; word segmentation complicates Chinese and Japanese; morphology explodes surface forms in Turkish or Finnish. Multilingual transformers narrow the gap but do not close it.
  • Emerging entities. New companies, products, people and events appear daily, and by definition no training corpus contains them. Systems that lean on memorized names fail here; robust ones rely on contextual patterns and refreshable knowledge bases, and linking pipelines need NIL detection for entities not yet in any registry.
  • Nested and discontinuous mentions. In "University of California, Berkeley", a location sits inside an organization. Standard BIO tagging cannot represent nesting, so specialized architectures are needed where it matters (biomedicine is full of it).
  • Noisy input. OCR errors, ASR transcripts, typos and creative social-media spelling ("gooogle", "@elonmusk") degrade accuracy sharply. Garbage in, mislabeled entities out.
  • Evaluation subtlety. Exact-span F1 punishes a system that finds "Barack Obama" when the gold label is "President Barack Obama". Deciding what counts as correct is a policy choice, not a formula.

None of these is fatal; all of them are why NER remains an engineering discipline rather than a checkbox.

Frequently asked questions about named entity recognition

What is the difference between NER and entity linking?

NER finds and classifies mentions in text ("Paris" → LOCATION). Entity linking goes one step further and maps each mention to a unique record in a knowledge base ("Paris" → Wikidata Q90, the French capital). NER answers "what kind of thing is this?"; linking answers "which specific thing is this?". Production systems that need to aggregate information across documents almost always need both.

Do I still need to train a NER model, or can I just use an LLM?

For prototypes, low volumes or unusual entity types, prompting an LLM is often enough and requires zero training data. For high-volume pipelines, a fine-tuned transformer is usually more accurate on standard categories and one to two orders of magnitude cheaper per document. A common pattern is to use the LLM to generate training data, then distill that knowledge into a small specialist model for production.

How accurate is named entity recognition?

On clean English news text, fine-tuned transformers reach roughly 92–94% F1 on standard benchmarks like CoNLL-2003 — near human agreement. On real-world material — other languages, specialized domains, noisy OCR or social text — accuracy can drop to 70–85% without adaptation. The honest answer is: measure on your data before trusting any published number.

Which entity types can NER detect?

Classic systems detect people, organizations, locations, dates, times, money and percentages. Richer schemes like OntoNotes add products, events, works of art, laws and languages, and domain-specific models detect genes, diseases, drugs, statutes, financial instruments and more. With zero-shot models and LLMs you can now define almost any type on the fly — "extract all programming languages and cloud services" — without training a new model.

Named entity recognition sits at the point where language becomes data. If this guide answered the "what" and "how", explore the rest of our Language & Document AI section for the tools and pipelines that put it to work.

Recommended:

Go up

This web uses cookies More info