BERTopic Tutorial: Modern Topic Modeling with Embeddings

# BERTopic: Pemodelan Topik Modern dengan Embedding BERTopic adalah library pemodelan topik yang menggabungkan embedding transformer, reduksi dimensi, clustering berbasis kepadatan, dan representasi...

By Ruby Abdullah · · tutorial
BERTopicTopic ModelingNLPClusteringEmbeddingsPython

BERTopic: Modern Topic Modeling with Embeddings

BERTopic is a topic modeling library that combines transformer embeddings, dimensionality reduction, density-based clustering, and a class-based TF-IDF representation to discover themes in text collections. Instead of treating documents as bags of words like classical methods, it works on the semantic level, which usually produces more coherent and interpretable topics. This tutorial walks through the full BERTopic pipeline using a practical example: clustering a corpus of customer support tickets.

What Topic Modeling Is and Why BERTopic Helps

Topic modeling is the task of automatically grouping a collection of documents into themes ("topics") and describing each theme with a set of representative words. It is unsupervised: you do not provide labels, the model finds structure on its own.

The traditional approach is Latent Dirichlet Allocation (LDA). LDA treats each document as a mixture of topics and each topic as a distribution over words. It works on raw word counts, so it ignores word order and meaning. In practice this leads to a few recurring problems:

  • It struggles with short texts (tweets, tickets, reviews) where word co-occurrence is sparse.
  • It requires you to fix the number of topics up front, which is hard to guess.
  • The topics it produces often mix unrelated words because it has no notion of semantic similarity.

BERTopic takes a different route. It represents each document as a dense embedding from a language model, so documents that mean similar things sit close together in vector space even when they share no exact words. It then clusters those embeddings and, only at the end, derives keywords per cluster. The key insight is that clustering and keyword extraction are separated, which makes the whole process modular and easy to customize.

The default pipeline has five stages:

  • Embed documents with a sentence-transformer model.
  • Reduce the embedding dimensionality with UMAP.
  • Cluster the reduced embeddings with HDBSCAN.
  • Tokenize the documents per cluster with a vectorizer (CountVectorizer).
  • Weight the tokens with class-based TF-IDF (c-TF-IDF) to find topic keywords.
  • Each stage is a swappable component, which is the main reason BERTopic is flexible.

    Installation

    pip install bertopic
    

    BERTopic pulls in sentence-transformers, UMAP, HDBSCAN, and scikit-learn automatically. For optional features install the extras you need:

    # Saving models in the safetensors format
    

    pip install bertopic[safetensors]

    Visualization support (plotly is included, but datamapplot is extra)

    pip install bertopic[visualization]

    LLM-based topic labels via the OpenAI backend

    pip install bertopic[openai]

    If you have a GPU, install a CUDA build of PyTorch first so the embedding step runs faster.

    A First Model

    Let's build a model on a small set of support tickets. In a real project you would load thousands of documents; the API is identical.

    from bertopic import BERTopic
    
    

    docs = [

    "My invoice shows a charge I did not authorize this month",

    "I was billed twice for the same subscription, please refund",

    "The app crashes every time I open the reports page",

    "Reports tab freezes and then the application closes",

    "How do I reset my password? The reset email never arrives",

    "I cannot log in, the password reset link is broken",

    "Can you explain the difference between the Pro and Team plans?",

    "What features are included in the Team plan pricing?",

    # ... thousands more in practice

    ]

    topicmodel = BERTopic()

    topics, probs = topicmodel.fittransform(docs)

    fittransform returns two arrays. topics holds the topic id assigned to each document, and probs holds the probability that each document belongs to its assigned topic.

    Inspecting Topics

    gettopicinfo() returns a DataFrame summarizing every topic: its id, how many documents it contains, a short auto-generated name, and the top keywords.
    info = topicmodel.gettopicinfo()
    

    print(info[["Topic", "Count", "Name"]])

    To see the keywords and their c-TF-IDF weights for one topic, use gettopic:

    print(topicmodel.gettopic(0))
    

    [('refund', 0.041), ('billed', 0.038), ('invoice', 0.033), ...]

    To find which topic a phrase is closest to, use findtopics:

    similartopics, similarity = topicmodel.findtopics("billing problem", topn=3)
    

    print(similartopics, similarity)

    Topic -1: The Outliers

    BERTopic uses HDBSCAN by default, and HDBSCAN does not force every document into a cluster. Documents that do not fit any dense region are assigned to topic -1, the outlier topic. This is a feature, not a bug: it keeps noisy or one-off documents from polluting real topics. You can reduce or reassign these outliers later (covered below), but you should never interpret topic -1 as a meaningful theme.

    The Modular Pipeline

    Every stage of the default pipeline can be replaced. You pass the components into the BERTopic constructor. Below we build each one explicitly and assemble them.

    1. Embedding Model

    By default BERTopic uses all-MiniLM-L6-v2, a fast English model. You can pass any sentence-transformers model by name, or a model object.

    from sentencetransformers import SentenceTransformer
    
    

    embeddingmodel = SentenceTransformer("all-MiniLM-L6-v2")

    topicmodel = BERTopic(embeddingmodel=embeddingmodel)

    For non-English or mixed-language corpora, use a multilingual model:

    embeddingmodel = SentenceTransformer("paraphrase-multilingual-MiniLM-L12-v2")
    

    topicmodel = BERTopic(embeddingmodel=embeddingmodel)

    There is also a shortcut for multilingual support without choosing a model yourself:

    topicmodel = BERTopic(language="multilingual")
    

    2. Dimensionality Reduction with UMAP

    Embeddings are high-dimensional (384 or 768 dimensions). Clustering works better in a lower-dimensional space, so BERTopic reduces them with UMAP. The most important parameters are nneighbors (local vs. global structure) and ncomponents (target dimensions).

    from umap import UMAP
    
    

    umapmodel = UMAP(

    nneighbors=15,

    ncomponents=5,

    mindist=0.0,

    metric="cosine",

    randomstate=42, # set this for reproducible results

    )

    Setting randomstate is essential if you want the same topics across runs. UMAP is stochastic, so without it the clustering will shift slightly each time.

    3. Clustering with HDBSCAN

    HDBSCAN finds clusters of varying density and labels low-density points as outliers. minclustersize controls the smallest allowed topic; larger values produce fewer, broader topics.

    from hdbscan import HDBSCAN
    
    

    hdbscanmodel = HDBSCAN(

    minclustersize=15,

    metric="euclidean",

    clusterselectionmethod="eom",

    predictiondata=True, # needed for transform() and outlier reduction

    )

    Set predictiondata=True so the model can later assign new documents and reduce outliers.

    4. Vectorizer (CountVectorizer)

    After clustering, BERTopic tokenizes the documents to build topic keywords. A CountVectorizer gives you control over stop words, n-grams, and minimum frequency. Removing stop words here, after clustering, is safe because it does not affect how documents were grouped, only how topics are labeled.

    from sklearn.featureextraction.text import CountVectorizer
    
    

    vectorizermodel = CountVectorizer(

    stopwords="english",

    ngramrange=(1, 2), # include bigrams like "password reset"

    mindf=2,

    )

    5. c-TF-IDF Representation

    c-TF-IDF (class-based TF-IDF) treats all documents in a cluster as a single document and computes TF-IDF across clusters. Words frequent in one cluster but rare across others get high scores, which makes good keywords. You can tune it with ClassTfidfTransformer.

    from bertopic.vectorizers import ClassTfidfTransformer
    
    

    ctfidfmodel = ClassTfidfTransformer(

    reducefrequentwords=True, # down-weight very common words

    )

    Assembling the Pipeline

    topicmodel = BERTopic(
    

    embeddingmodel=embeddingmodel,

    umapmodel=umapmodel,

    hdbscanmodel=hdbscanmodel,

    vectorizermodel=vectorizermodel,

    ctfidfmodel=ctfidfmodel,

    verbose=True,

    )

    topics, probs = topicmodel.fittransform(docs)

    Representation Models: Better Topic Labels

    The c-TF-IDF keywords are a solid baseline, but they can be redundant or generic. Representation models refine the keyword list after the fact without re-clustering. You pass them via representationmodel.

    KeyBERTInspired

    This re-ranks candidate keywords by their semantic similarity to the topic, producing labels that read more naturally.

    from bertopic.representation import KeyBERTInspired
    
    

    representationmodel = KeyBERTInspired()

    topicmodel = BERTopic(representationmodel=representationmodel)

    MaximalMarginalRelevance

    MMR reduces redundancy by penalizing keywords that are too similar to ones already chosen, giving a more diverse set. The diversity parameter ranges from 0 (relevance only) to 1 (maximum diversity).

    from bertopic.representation import MaximalMarginalRelevance
    
    

    representationmodel = MaximalMarginalRelevance(diversity=0.3)

    Combining Representations

    You can apply several representations and compare them by passing a dictionary. BERTopic computes each and stores them under the given keys.

    from bertopic.representation import KeyBERTInspired, MaximalMarginalRelevance
    
    

    representationmodel = {

    "KeyBERT": KeyBERTInspired(),

    "MMR": MaximalMarginalRelevance(diversity=0.3),

    }

    topicmodel = BERTopic(representationmodel=representationmodel)

    LLM-Based Labels (Optional)

    For human-readable topic titles you can ask a large language model to summarize each cluster. BERTopic sends a few representative documents and keywords to the model and uses the response as the label.

    from bertopic.representation import OpenAI
    

    import openai

    client = openai.OpenAI(apikey="YOURKEY")

    prompt = "Topic keywords: [KEYWORDS]\nDocuments: [DOCUMENTS]\nGive a short topic label."

    representationmodel = OpenAI(client, model="gpt-4o-mini", prompt=prompt, chat=True)

    topicmodel = BERTopic(representationmodel=representationmodel)

    This is optional and makes network calls, so keep it out of offline or cost-sensitive pipelines.

    Controlling the Number of Topics

    BERTopic does not require you to fix the topic count, but you often want fewer topics than HDBSCAN produces by default.

    mintopicsize is the simplest lever. It maps to HDBSCAN's minclustersize and sets how many documents a topic needs at minimum. Larger values mean fewer, broader topics.
    topicmodel = BERTopic(mintopicsize=20)
    

    nrtopics lets you ask for a target count or "auto" to let BERTopic merge similar topics automatically.
    topicmodel = BERTopic(nrtopics=10)        # merge down to ~10 topics
    

    topicmodel = BERTopic(nrtopics="auto") # merge based on similarity

    You can also reduce topics after fitting, which avoids re-embedding:

    topicmodel.reducetopics(docs, nrtopics=8)
    

    Reducing Outliers

    If topic -1 holds too many documents, you can reassign them to their nearest real topic. reduceoutliers supports several strategies; c-tf-idf and embeddings are common choices.

    newtopics = topicmodel.reduceoutliers(docs, topics, strategy="c-tf-idf")
    
    

    Persist the new assignments so labels reflect them

    topicmodel.updatetopics(docs, topics=newtopics)

    Be deliberate here: forcing outliers into topics can blur otherwise clean clusters. Reduce them only when downstream consumers cannot handle an "unassigned" bucket.

    Visualizations

    BERTopic ships interactive Plotly visualizations. Each returns a figure you can show or save to HTML.

    # Inter-topic distance map (topics as bubbles in 2D)
    

    topicmodel.visualizetopics()

    Top keywords per topic as bar charts

    topicmodel.visualizebarchart(topntopics=8)

    Hierarchical relationships between topics

    topicmodel.visualizehierarchy()

    Documents projected into 2D, colored by topic

    topicmodel.visualizedocuments(docs)

    For visualizedocuments it is faster to pass precomputed embeddings (see below) and reduced 2D coordinates so the plot does not recompute them.

    fig = topicmodel.visualizebarchart(topntopics=8)
    

    fig.writehtml("topicsbarchart.html")

    Advanced Topics

    Topics Over Time (Dynamic Topic Modeling)

    If your documents have timestamps, you can track how topics rise and fall. You first fit a normal model, then compute representations per time bin.

    topicsovertime = topicmodel.topicsovertime(
    

    docs,

    timestamps, # list of dates aligned with docs

    nrbins=20,

    )

    topicmodel.visualizetopicsovertime(topicsovertime, topntopics=10)

    This reuses the same global topics and only recomputes c-TF-IDF within each time slice, so the topic definitions stay consistent across the timeline.

    Topics per Class

    When documents carry a category (product line, region, channel), you can see how each topic is expressed within each class.

    topicsperclass = topicmodel.topicsperclass(docs, classes=ticketchannels)
    

    topicmodel.visualizetopicsperclass(topicsperclass, topntopics=10)

    Hierarchical Topic Modeling

    To understand how fine-grained topics merge into broader themes, build a hierarchy.

    hierarchicaltopics = topicmodel.hierarchicaltopics(docs)
    

    topicmodel.visualizehierarchy(hierarchicaltopics=hierarchicaltopics)

    Guided and Seeded Topic Modeling

    If you already know some themes should exist, pass seed words to nudge the model toward them.

    seedtopiclist = [
    

    ["billing", "invoice", "refund", "charge"],

    ["login", "password", "reset", "access"],

    ]

    topicmodel = BERTopic(seedtopiclist=seedtopiclist)

    Zero-Shot Topic Modeling

    Zero-shot lets you predefine topic labels and assign documents to them, while still discovering new topics for documents that match none.

    zeroshottopiclist = ["billing issues", "app crashes", "account access"]
    

    topicmodel = BERTopic(

    zeroshottopiclist=zeroshottopiclist,

    zeroshotminsimilarity=0.85,

    )

    topics, probs = topicmodel.fittransform(docs)

    Predicting on New Documents

    Once fitted, transform assigns topics to unseen documents without retraining. This requires predictiondata=True on the HDBSCAN model.

    newdocs = [
    

    "I keep getting charged after canceling my plan",

    "The dashboard goes blank when I export to PDF",

    ]

    newtopics, newprobs = topicmodel.transform(newdocs)

    print(newtopics)

    Using Precomputed Embeddings for Speed

    Embedding is the slowest stage. If you experiment with UMAP or HDBSCAN settings repeatedly, compute the embeddings once and reuse them.

    from sentencetransformers import SentenceTransformer
    
    

    embeddingmodel = SentenceTransformer("all-MiniLM-L6-v2")

    embeddings = embeddingmodel.encode(docs, showprogressbar=True)

    Pass the array directly; BERTopic skips the embedding step

    topicmodel = BERTopic(embeddingmodel=embeddingmodel)

    topics, probs = topicmodel.fittransform(docs, embeddings=embeddings)

    You can then tweak clustering parameters and refit on the same embeddings array in seconds rather than minutes.

    Saving and Loading Models

    BERTopic supports three serialization options. The recommended approach is safetensors, which stores the topic weights compactly and lets you reload the embedding model by name rather than pickling it.

    # Safetensors (recommended): small, safe, portable
    

    topicmodel.save(

    "supportmodel",

    serialization="safetensors",

    savectfidf=True,

    saveembeddingmodel="all-MiniLM-L6-v2",

    )

    PyTorch format

    topicmodel.save("supportmodel", serialization="pytorch")

    Pickle (stores everything, but ties you to library versions)

    topicmodel.save("supportmodel.pkl", serialization="pickle")

    Loading mirrors the save call:

    from bertopic import BERTopic
    
    

    loaded = BERTopic.load("supportmodel")

    Prefer safetensors or pytorch for production. Pickle is convenient but brittle across version upgrades and unsafe to load from untrusted sources.

    Best Practices

    • Preprocess lightly. Do not lowercase or strip stop words before embedding; transformer models use that context. Apply stop-word removal in the CountVectorizer step instead, where it only affects labels.
    • Tune mintopicsize first. It is the single most influential setting. Start around 1-2% of your corpus size and adjust based on whether topics feel too granular or too broad.
    • Set randomstate in UMAP. Without it, results change every run, which makes debugging and stakeholder communication difficult.
    • Treat topic -1 honestly. Inspect outliers before reducing them; a large outlier group often signals that mintopicsize is too high or the corpus is genuinely noisy.
    • Cache embeddings during experimentation. Reusing a precomputed array turns parameter sweeps from minutes into seconds.
    • Validate topics with humans. Coherence scores help, but reading the top documents per topic is the most reliable check that the topics are useful.

    Conclusion

    BERTopic reframes topic modeling as a pipeline of embedding, reduction, clustering, and keyword extraction, with each stage independently swappable. That modularity is its real strength: you can start with the defaults, then improve embeddings for your language, tune clustering granularity, and refine labels with representation models or an LLM, all without rewriting your code.

    Key Takeaways

    • BERTopic uses semantic embeddings instead of word counts, so it handles short and noisy text better than LDA.
    • The five-stage pipeline (embed, reduce, cluster, vectorize, c-TF-IDF) is fully customizable.
    • Topic -1 is the outlier bucket, not a real topic; reduce it deliberately.
    • Control topic granularity with mintopicsize, nrtopics, and reducetopics.
    • Representation models like KeyBERTInspired and MMR sharpen topic labels.
    • Advanced modes cover topics over time, per class, hierarchy, guided, and zero-shot modeling.
    • For reproducibility set UMAP's random_state; for speed reuse precomputed embeddings; for portability save with safetensors.

    Related Articles

    Sentence Transformers Tutorial: Embeddings, Similarity, and Rerankers

    Sentence Transformers: Embedding, Kemiripan Semantik, dan Reranker Sentence Transformers (sering disebut SBERT) adalah p...

    Complete txtai Tutorial: All-in-One Embeddings Database for Semantic Search and LLM Workflows

    Tutorial Lengkap txtai: Database Embeddings All-in-One untuk Semantic Search dan LLM Workflows txtai adalah framework Py...

    spaCy Tutorial: Industrial-Strength NLP in Python

    spaCy: NLP Kelas Industri di Python spaCy adalah pustaka open-source untuk pemrosesan bahasa alami (NLP) yang dirancang ...

    FAISS Tutorial: Efficient Vector Similarity Search at Scale

    FAISS: Pencarian Kemiripan Vektor yang Efisien dalam Skala Besar FAISS (Facebook AI Similarity Search) adalah library C+...