spaCy Tutorial: Industrial-Strength NLP in Python

# spaCy: NLP Kelas Industri di Python spaCy adalah pustaka open-source untuk pemrosesan bahasa alami (NLP) yang dirancang untuk penggunaan produksi. Jika large language model unggul dalam menghasilka...

By Ruby Abdullah · · tutorial
spaCyNLPNamed Entity RecognitionText ProcessingPythonMachine Learning

spaCy: Industrial-Strength NLP in Python

spaCy is an open-source library for natural language processing (NLP) built for production use. Where large language models excel at open-ended generation, spaCy focuses on fast, deterministic, and structured text analysis: tokenization, part-of-speech tagging, dependency parsing, and named entity recognition. This tutorial walks through spaCy from installation to training a custom model, using a coherent business-text example throughout.

What spaCy Is and Where It Fits

spaCy is a Python library that turns raw text into structured linguistic data. It ships with pre-trained statistical models for many languages and exposes a clean object model (Doc, Token, Span) for working with the results. The library is written with performance in mind, so it can process large volumes of text quickly on a CPU.

spaCy vs. Large Language Models

LLMs and spaCy solve different problems, and in practice they are often combined.

  • Determinism: spaCy produces the same output for the same input every time. This matters for pipelines that feed downstream systems or that must be auditable.
  • Speed and cost: spaCy runs on commodity CPUs and processes thousands of documents per second for many tasks. There is no per-token API cost.
  • Structure: spaCy returns typed, span-level annotations (entities, tokens, dependencies) rather than free text you have to parse again.
  • Privacy: text never has to leave your infrastructure.

A common pattern is to use spaCy for high-volume extraction and routing, and reserve an LLM for the smaller subset of cases that genuinely need open-ended reasoning. spaCy also integrates LLM prompts as pipeline components through the spacy-llm package when you want both in one workflow.

Installation and Downloading Models

Install spaCy with pip. Using a virtual environment is recommended so that model versions stay pinned to your project.

python -m venv .venv

source .venv/bin/activate # on Windows: .venv\Scripts\activate

pip install spacy

Models are distributed separately from the library. Download a small English model to get started:

python -m spacy download encorewebsm

The English models follow a naming convention: en (language), core (general-purpose pipeline), web (trained on web text), and a size suffix:

  • encorewebsm — small, fast, no word vectors.
  • encorewebmd — medium, includes word vectors.
  • encoreweblg — large, more vectors.
  • encorewebtrf — transformer-based, highest accuracy, needs more compute.

You can confirm what is installed and validate compatibility:

python -m spacy info

python -m spacy validate

The nlp Pipeline and Core Objects

Loading a model gives you an nlp object. Calling it on a string runs the full pipeline and returns a Doc.

import spacy

nlp = spacy.load("encorewebsm")

text = "Acme Corp acquired Globex Ltd for $4.5 billion in March 2023."

doc = nlp(text)

print(type(doc)) #

print(len(doc)) # number of tokens

print([token.text for token in doc])

A Doc is a sequence of Token objects. A contiguous slice of tokens is a Span. These three objects are the foundation for everything else.

# Token: a single word, punctuation mark, or symbol

first = doc[0]

print(first.text, first.idx) # "Acme" 0

Span: a slice of the Doc

span = doc[0:2]

print(span.text) # "Acme Corp"

Sentences are Spans too

for sent in doc.sents:

print(sent.text)

Inspecting the Pipeline

The nlp object holds an ordered list of components. You can inspect and modify it.

print(nlp.pipenames)

['tok2vec', 'tagger', 'parser', 'attributeruler', 'lemmatizer', 'ner']

Each component adds annotations to the Doc as it passes through.

Tokenization, POS Tagging, and Lemmatization

Tokenization splits text into tokens using language-specific rules. spaCy handles contractions, punctuation, and special cases without losing the original character offsets.

doc = nlp("Acme didn't expand into Europe until 2024.")

for token in doc:

print(f"{token.text:>10} | {token.lemma:>8} | {token.pos:>6} | {token.tag}")

Useful token attributes:

  • token.text — the original text.
  • token.lemma — the base form (for example, "acquired" becomes "acquire").
  • token.pos — coarse-grained part of speech (NOUN, VERB, PROPN).
  • token.tag — fine-grained tag (VBD, NNP).
  • token.isstop — whether it is a stop word.
  • token.isalpha, token.likenum, token.ispunct — useful boolean flags.

contentwords = [t.lemma for t in doc if not t.isstop and not t.ispunct]

print(contentwords)

Dependency Parsing

The dependency parser assigns syntactic relations between tokens, forming a tree. This is how you discover which noun is the subject of a verb, or which token a modifier attaches to.

doc = nlp("Acme Corp acquired Globex Ltd for $4.5 billion.")

for token in doc:

print(f"{token.text:>10} --{token.dep}--> {token.head.text}")

Each token has a head (its syntactic parent) and a dep label describing the relation (nsubj, dobj, prep). You can also navigate children:

verb = [t for t in doc if t.lemma == "acquire"][0]

subjects = [child for child in verb.children if child.dep == "nsubj"]

objects = [child for child in verb.children if child.dep == "dobj"]

print("subject:", subjects)

print("object:", objects)

This gives you a simple foundation for relationship extraction, which we return to later.

Named Entity Recognition (NER)

NER identifies real-world objects such as organizations, people, money, and dates. Recognized entities appear in doc.ents as Span objects with a label.

doc = nlp("Acme Corp acquired Globex Ltd for $4.5 billion in March 2023.")

for ent in doc.ents:

print(f"{ent.text:>15} | {ent.label:>10} | {ent.startchar}-{ent.endchar}")

Typical output labels include ORG, MONEY, and DATE. To understand a label, ask spaCy:

print(spacy.explain("ORG"))     # "Companies, agencies, institutions, etc."

Visualizing with displaCy

spaCy ships with a built-in visualizer. In a script it serves a small web page; in a notebook it renders inline.

from spacy import displacy

doc = nlp("Acme Corp acquired Globex Ltd for $4.5 billion in March 2023.")

Entity highlighting

displacy.serve(doc, style="ent")

Dependency tree

displacy.serve(doc, style="dep")

In a Jupyter notebook, use displacy.render(doc, style="ent") instead of serve. You can also export the rendered SVG to a file for reports.

Rule-Based Matching: Matcher and PhraseMatcher

Statistical models are powerful, but some patterns are better expressed as rules: product codes, units, or fixed phrases. spaCy provides two matchers for this.

The Matcher

The token Matcher matches sequences described by token attributes. Each pattern is a list of dictionaries, one per token.

from spacy.matcher import Matcher

matcher = Matcher(nlp.vocab)

Match " acquired "

pattern = [

{"POS": "PROPN", "OP": "+"},

{"LEMMA": "acquire"},

{"POS": "PROPN", "OP": "+"},

]

matcher.add("ACQUISITION", [pattern])

doc = nlp("Acme Corp acquired Globex Ltd last year.")

for matchid, start, end in matcher(doc):

span = doc[start:end]

print(nlp.vocab.strings[matchid], "->", span.text)

The OP key controls quantifiers: ! (negate), ? (optional), + (one or more), * (zero or more).

The PhraseMatcher

When you already have a list of exact terms, such as a product catalog or a watchlist of company names, the PhraseMatcher is faster and simpler.

from spacy.matcher import PhraseMatcher

phrasematcher = PhraseMatcher(nlp.vocab, attr="LOWER")

terms = ["Acme Corp", "Globex Ltd", "Initech"]

patterns = [nlp.makedoc(t) for t in terms]

phrasematcher.add("COMPANIES", patterns)

doc = nlp("Both acme corp and Initech reported strong quarters.")

for matchid, start, end in phrasematcher(doc):

print(doc[start:end].text)

Custom Entities with the EntityRuler

The EntityRuler lets you add entities through rules and merge them with the statistical NER. This is the right tool when you have domain terms the model does not know.

ruler = nlp.addpipe("entityruler", before="ner")

patterns = [

{"label": "PRODUCT", "pattern": "Widget Pro"},

{"label": "PRODUCT", "pattern": [{"LOWER": "widget"}, {"LOWER": "lite"}]},

{"label": "ORG", "pattern": "Initech"},

]

ruler.addpatterns(patterns)

doc = nlp("Initech launched Widget Pro and widget lite this quarter.")

for ent in doc.ents:

print(ent.text, ent.label)

Placing the ruler before="ner" lets your rules take priority; placing it after lets the statistical model win on conflicts. Choose based on how much you trust each source.

Word Vectors and Similarity

The medium and large models include word vectors, which let you measure semantic similarity. The small model does not include them, so install encorewebmd for this section.

python -m spacy download encorewebmd

nlpmd = spacy.load("encorewebmd")

doc = nlpmd("revenue profit lettuce")

tokens = [t for t in doc]

for t in tokens:

print(t.text, t.hasvector, round(t.vectornorm, 2))

print(tokens[0].similarity(tokens[1])) # revenue vs profit (higher)

print(tokens[0].similarity(tokens[2])) # revenue vs lettuce (lower)

Doc and Span objects also expose .similarity, computed from the average of their token vectors. These vectors are static (context-free), so for context-sensitive similarity you would use a transformer model or a dedicated embedding model.

Efficient Processing with nlp.pipe

Calling nlp(text) in a loop is fine for a handful of documents, but for a corpus it is slow. nlp.pipe processes texts as a stream, batches them, and can use multiple processes.

texts = [

"Acme Corp acquired Globex Ltd for $4.5 billion.",

"Initech hired 200 engineers in Berlin.",

"Umbrella Inc reported a loss in Q3 2024.",

]

for doc in nlp.pipe(texts, batchsize=50):

orgs = [ent.text for ent in doc.ents if ent.label == "ORG"]

print(orgs)

Two important options:

  • batchsize — how many texts to buffer per batch. Tune it to your document size.
  • nprocess — number of worker processes. Use -1 for all cores, but measure first; the overhead can outweigh the benefit for short texts.

If you only need part of the pipeline, disable components you do not use to save time:

with nlp.selectpipes(enable=["tok2vec", "ner"]):

for doc in nlp.pipe(texts):

print(doc.ents)

You can also attach your own metadata to each text and get it back alongside the Doc:

data = [("doc-1", "Acme Corp grew."), ("doc-2", "Initech shrank.")]

for doc, ctx in nlp.pipe(data, astuples=True):

print(ctx, "->", [t.text for t in doc])

Custom Pipeline Components

You can insert your own logic into the pipeline with the @Language.component decorator. A component receives a Doc, modifies it, and returns it. The example below extracts simple acquirer/target relationships using the dependency parse and stores them as a custom attribute.

from spacy.language import Language

from spacy.tokens import Doc

Register a custom extension attribute on Doc

if not Doc.hasextension("acquisitions"):

Doc.setextension("acquisitions", default=[])

@Language.component("acquisitionextractor")

def acquisitionextractor(doc):

results = []

for token in doc:

if token.lemma == "acquire" and token.pos == "VERB":

subj = [c for c in token.children if c.dep == "nsubj"]

obj = [c for c in token.children if c.dep == "dobj"]

if subj and obj:

results.append({

"acquirer": subj[0].text,

"target": obj[0].text,

})

doc..acquisitions = results

return doc

nlp.addpipe("acquisitionextractor", last=True)

doc = nlp("Acme acquired Globex. Initech acquired Hooli.")

print(doc..acquisitions)

Custom extension attributes are accessed through the . namespace, which keeps them separate from spaCy's built-in attributes. Components can be ordered with first, last, before, or after.

Training a Custom NER Model

When the built-in entities are not enough, you can train your own. Modern spaCy uses a configuration-driven workflow centered on a config.cfg file and the spacy train command, rather than hand-written training loops.

Step 1: Prepare Training Data as a DocBin

Training data is stored in the binary .spacy format using DocBin. Each example is a Doc with gold-standard entity spans.

import spacy

from spacy.tokens import DocBin

nlp = spacy.blank("en")

TRAINDATA = [

("Acme Corp launched Widget Pro.", {"entities": [(0, 9, "ORG"), (19, 29, "PRODUCT")]}),

("Initech ships Widget Lite today.", {"entities": [(0, 7, "ORG"), (14, 25, "PRODUCT")]}),

]

def makedocbin(data):

db = DocBin()

for text, ann in data:

doc = nlp.makedoc(text)

ents = []

for start, end, label in ann["entities"]:

span = doc.charspan(start, end, label=label, alignmentmode="contract")

if span is not None:

ents.append(span)

doc.ents = ents

db.add(doc)

return db

makedocbin(TRAINDATA).todisk("./train.spacy")

makedocbin(TRAINDATA).todisk("./dev.spacy") # use a real held-out set in practice

Always check that charspan returns a span; misaligned offsets silently produce None and would otherwise drop your annotation.

Step 2: Create a Config File

spaCy can generate a sensible starting config for you, then fill in defaults.

python -m spacy init config config.cfg --lang en --pipeline ner

python -m spacy init fill-config config.cfg config.cfg

The generated config.cfg declares the pipeline, the model architecture, and training hyperparameters. The key sections to know are [paths] (where your data lives), [nlp] (pipeline definition), and [training] (optimizer, batching, evaluation).

Step 3: Train and Evaluate

Point the trainer at your data and an output directory.

python -m spacy train config.cfg \

--output ./output \

--paths.train ./train.spacy \

--paths.dev ./dev.spacy

spaCy writes the best and last models to ./output. You can evaluate on a held-out set and then load the trained pipeline like any other model.

python -m spacy evaluate ./output/model-best ./dev.spacy

trained = spacy.load("./output/model-best")

doc = trained("Globex Inc released Widget Pro.")

print([(e.text, e.label) for e in doc.ents])

For real projects, use more examples (hundreds to thousands per label), keep a genuine held-out dev set, and consider spacy debug data to catch annotation problems before training.

Transformer-Based Models

For the highest accuracy, spaCy offers transformer pipelines such as encorewebtrf, which use a model like RoBERTa under the hood through the spacy-transformers package.

pip install spacy-transformers

python -m spacy download encorewebtrf

nlptrf = spacy.load("encorewebtrf")

doc = nlptrf("Acme Corp acquired Globex Ltd for $4.5 billion.")

print([(e.text, e.label) for e in doc.ents])

Transformer pipelines give context-sensitive representations and usually better accuracy on hard cases, at the cost of speed and memory. They benefit greatly from a GPU. A practical approach is to prototype with the small model, measure accuracy on your data, and move up to md/lg/trf only if the numbers justify it.

Best Practices

  • Load the model once. spacy.load is expensive; load at startup and reuse the nlp object.
  • Use nlp.pipe for batches. It is the single biggest speedup for corpus processing.
  • Disable unused components. If you only need NER, do not pay for the parser.
  • Pin model versions. Models are tied to spaCy versions; record them in your requirements so results stay reproducible.
  • Combine rules and statistics. Use EntityRuler and matchers for known, high-precision patterns and the statistical model for the long tail.
  • Validate offsets when building training data. Always confirm char_span is not None.
  • Measure before scaling up. Start with sm, and only adopt transformers when accuracy gains are real and justify the cost.
  • Keep a held-out evaluation set. Without it you cannot tell whether changes actually help.

Conclusion and Key Takeaways

spaCy is a practical choice when you need fast, structured, and repeatable text processing in production. It complements large language models rather than competing with them: spaCy handles high-volume extraction and structure, while an LLM can take the smaller set of cases that need open reasoning.

Key takeaways:

  • The nlp pipeline turns text into a Doc of Token and Span objects carrying tokenization, POS, lemma, dependency, and entity annotations.
  • Rule-based tools (Matcher, PhraseMatcher, EntityRuler) add high-precision, domain-specific behavior alongside the statistical models.
  • Word vectors in the md/lg models enable similarity comparisons; transformers in trf push accuracy further.
  • nlp.pipe and component selection are the main levers for processing large corpora efficiently.
  • The config-driven training workflow (DocBin data, config.cfg, spacy train) lets you build custom NER models in a reproducible way.

With these building blocks you can assemble robust NLP pipelines that are fast to run, easy to audit, and straightforward to extend.

Related Articles

Complete Comet ML Tutorial: MLOps Platform for Experiment Tracking and Model Management

Tutorial Lengkap Comet ML: Platform MLOps untuk Experiment Tracking dan Model Management Dalam dunia machine learning mo...

MLX Tutorial: Apple's Machine Learning Framework for Apple Silicon

Tutorial MLX: Framework Machine Learning Apple untuk Apple Silicon MLX adalah framework machine learning open-source dar...

SHAP Tutorial: Explainable AI and Model Interpretability

SHAP - Panduan Praktis Explainable AI dan Interpretabilitas Model Model machine learning makin sering dipakai untuk meng...

PyOD Tutorial: Anomaly and Outlier Detection in Python

Deteksi Anomali di Python dengan PyOD: Panduan Praktis Sebagian besar dataset di dunia nyata mengandung sebagian kecil d...