Sentence Transformers: Embeddings, Semantic Similarity, and Rerankers
Sentence Transformers (often called SBERT) is a Python library for turning text into dense vector embeddings that capture meaning rather than surface words. In this tutorial we build a small retrieval system over a toy corpus, measure semantic similarity, add a cross-encoder reranker, and then fine-tune our own embedding model with the modern training API. The goal is a practical, end-to-end view of how the library fits into a real search or RAG pipeline.
What Sentence Transformers Is
The sentence-transformers library wraps transformer models (BERT, RoBERTa, MPNet, and many others) and adds pooling so that an entire sentence or paragraph maps to a single fixed-length vector. Two texts with similar meaning produce vectors that are close together, which lets you compare them with cosine similarity instead of keyword matching.
The library is maintained alongside the Hugging Face ecosystem, so models load from the Hub, datasets use the datasets format, and trained models push back to the Hub with one call. It is the standard tool for building embedding-based search, clustering, deduplication, and the retrieval stage of RAG systems.
Bi-Encoders vs Cross-Encoders
There are two model families in the library, and choosing correctly is the single most important design decision.
A bi-encoder encodes each text independently into a vector. You embed your whole corpus once, store the vectors, and at query time you embed only the query and compare it against the stored vectors. This is fast and scales to millions of documents because comparison is just a dot product. The trade-off is accuracy: the model never sees the query and document together, so it can miss subtle interactions.
A cross-encoder takes a pair of texts at once (query and candidate) and outputs a single relevance score. Because the model attends across both texts jointly, it is far more accurate. The cost is that you cannot precompute anything: every query-document pair must be run through the model. Scoring a query against a million documents is not feasible.
The standard pattern combines both. The bi-encoder retrieves a few dozen candidates quickly, then the cross-encoder reranks just those candidates for precision. This is the retrieve-then-rerank pattern we build later.
Query --> [Bi-encoder] --> top 50 candidates --> [Cross-encoder] --> top 5 reranked
(fast, approximate) (slow, precise)
Installation
Install the library with pip. It pulls in PyTorch, transformers, and datasets as dependencies.
pip install -U sentence-transformers
For training and evaluation you may also want a few extras. Installing accelerate enables faster and multi-GPU training, and datasets is required for the training API (it usually comes in already).
pip install -U accelerate datasets
Verify the install and check which device is available.
import torch
from sentencetransformers import SentenceTransformer
print("sentence-transformers ready")
print("CUDA available:", torch.cuda.isavailable())
Loading a Model and Encoding Text
The core class is SentenceTransformer. Pass it a model name from the Hub and it downloads and caches the weights. A good general-purpose starting model is all-MiniLM-L6-v2: small, fast, 384-dimensional, and strong on English semantic similarity.
from sentencetransformers import SentenceTransformer
model = SentenceTransformer("all-MiniLM-L6-v2")
sentences = [
"How do I reset my account password?",
"Steps to recover a forgotten login",
"The weather in Jakarta is hot today.",
]
embeddings = model.encode(sentences)
print(embeddings.shape) # (3, 384)
encode() returns a NumPy array by default. Useful arguments:
batchsizecontrols how many texts are processed at once. Larger batches are faster on a GPU but use more memory.converttotensor=Truereturns a PyTorch tensor, which is convenient when you keep everything on the GPU for similarity math.normalizeembeddings=Truescales each vector to unit length, so a dot product equals cosine similarity. Do this when your downstream index uses dot-product scoring.showprogressbar=Trueis helpful for large corpora.
embeddings = model.encode(
sentences,
batchsize=32,
converttotensor=True,
normalizeembeddings=True,
showprogressbar=True,
)
Multilingual Models
all-MiniLM-L6-v2 is English-centric. For Indonesian, mixed-language, or cross-lingual search, load a multilingual model. paraphrase-multilingual-MiniLM-L12-v2 covers 50+ languages and maps translations of the same sentence to nearby vectors.
multi = SentenceTransformer("paraphrase-multilingual-MiniLM-L12-v2")
pairs = [
"How do I reset my password?",
"Bagaimana cara mengatur ulang kata sandi saya?",
]
emb = multi.encode(pairs, converttotensor=True, normalizeembeddings=True)
print(multi.similarity(emb[0], emb[1])) # high score across languages
Prompts and promptname
Some newer models (for example the e5 and bge families, and Nomic models) were trained with instruction prefixes. They behave best when queries and documents get different prefixes such as "query: " and "passage: ". The library supports this through the prompts configuration and the promptname argument so you do not hardcode the strings everywhere.
model = SentenceTransformer(
"intfloat/multilingual-e5-small",
prompts={
"query": "query: ",
"passage": "passage: ",
},
)
queryemb = model.encode("reset password", promptname="query")
docemb = model.encode("Steps to recover a forgotten login", promptname="passage")
Always check a model's card to learn whether it expects prompts. Using the wrong prompt convention quietly degrades quality.
Computing Similarity
Once you have embeddings, model.similarity computes pairwise scores between two sets of embeddings. By default it uses cosine similarity, matching how most models were trained.
emb = model.encode(sentences, converttotensor=True)
Pairwise matrix: every row vs every column
scores = model.similarity(emb, emb)
print(scores)
The result is a matrix where entry [i][j] is the similarity between sentence i and sentence j. You can also change the metric on the model with model.similarityfnname = "dot" if your model was trained for dot-product scoring.
For one-off comparisons without a model object, the util module exposes helpers.
from sentencetransformers import util
cos = util.cos
sim(emb[0], emb[1])
print(float(cos))
util.cossim is handy when you have raw embeddings from anywhere and just need a cosine score.
Building a Semantic Search System
Let us build the retrieval core over a small corpus. The bi-encoder embeds the corpus once; queries are embedded on demand and compared with util.semanticsearch, which handles top-k selection efficiently.
from sentencetransformers import SentenceTransformer, util
model = SentenceTransformer("all-MiniLM-L6-v2")
corpus = [
"Reset your password from the account settings page.",
"Enable two-factor authentication for extra security.",
"Refunds are processed within 5 business days.",
"Update your billing address before the next invoice.",
"Contact support if your login is locked after failed attempts.",
"Export your data as CSV from the dashboard.",
]
corpus
emb = model.encode(corpus, converttotensor=True, normalizeembeddings=True)
query = "I forgot my login credentials"
query
emb = model.encode(query, converttotensor=True, normalizeembeddings=True)
hits = util.semantic
search(queryemb, corpusemb, topk=3)
for hit in hits[0]:
print(f"{hit['score']:.3f} {corpus[hit['corpus
id']]}")
util.semanticsearch accepts a batch of query embeddings and returns, for each query, a list of hits sorted by score. Each hit has a corpusid and a score. For production scale you would push corpusemb into a vector database, but the scoring logic is the same.
Paraphrase Mining and Clustering
Beyond search, embeddings power two common batch tasks.
Paraphrase mining finds the most similar pairs across a list without computing the full N-by-N matrix in memory.util.paraphrasemining is optimized for this.
from sentencetransformers import util
candidates = [
"How do I reset my password?",
"What are the steps to recover my login?",
"Where can I export my data?",
"How can I download my data as a file?",
"When will my refund arrive?",
]
pairs = util.paraphrase
mining(model, candidates, topk=2)
for score, i, j in pairs[:5]:
print(f"{score:.3f} | {candidates[i]} <> {candidates[j]}")
Clustering groups semantically related texts. You can feed normalized embeddings into scikit-learn's KMeans, or use util.communitydetection for fast, threshold-based clustering without choosing the number of clusters in advance.
emb = model.encode(candidates, converttotensor=True, normalizeembeddings=True)
clusters = util.community
detection(emb, threshold=0.6, mincommunitysize=2)
for cid, cluster in enumerate(clusters):
print(f"Cluster {cid}:", [candidates[idx] for idx in cluster])
Reranking with a Cross-Encoder
The bi-encoder is fast but can rank a loosely related document above the best one. A cross-encoder fixes the ordering of the top candidates. The CrossEncoder class loads a model trained to score (query, document) pairs; ms-marco-MiniLM-L-6-v2 is a solid default for English passage ranking.
from sentencetransformers import CrossEncoder
reranker = CrossEncoder("cross-encoder/ms-marco-MiniLM-L-6-v2")
query = "I forgot my login credentials"
Take the candidates the bi-encoder retrieved
candidate
texts = [corpus[hit["corpusid"]] for hit in hits[0]]
pairs = [(query, text) for text in candidate
texts]
rerankscores = reranker.predict(pairs)
reranked = sorted(zip(rerankscores, candidatetexts), reverse=True)
for score, text in reranked:
print(f"{score:.3f} {text}")
The Retrieve-Then-Rerank Pipeline
Putting it together: retrieve a wide net cheaply, then rerank precisely. A convenience method reranker.rank does the pairing and sorting for you.
def search(query, topkretrieve=20, topkfinal=3):
q
emb = model.encode(query, converttotensor=True, normalizeembeddings=True)
hits = util.semantic
search(qemb, corpusemb, topk=topkretrieve)[0]
documents = [corpus[h["corpus
id"]] for h in hits]
ranked = reranker.rank(query, documents, topk=topkfinal)
return [(r["score"], documents[r["corpusid"]]) for r in ranked]
for score, text in search("how to get money back"):
print(f"{score:.3f} {text}")
Retrieve more candidates than you ultimately show (20 to 100 is typical) so the reranker has enough material to find the best answer.
Fine-Tuning a Custom Embedding Model
Pretrained models are a strong baseline, but a model tuned on your own domain usually retrieves noticeably better. The modern API mirrors the Hugging Face Trainer: you provide a dataset, a loss, training arguments, and an optional evaluator.
Choosing a Loss and Dataset Format
The loss dictates the shape of your data.
MultipleNegativesRankingLossis the workhorse for retrieval. It trains on positive pairs(anchor, positive)and treats all other examples in the batch as negatives. Your dataset needs two (or three, with a hard negative) text columns and no explicit labels. Larger batches give more in-batch negatives and usually help.CosineSimilarityLosstrains on pairs with a float label between 0 and 1, useful when you have graded similarity scores rather than clean positive pairs.
Datasets are Hugging Face Dataset objects; column order matters and label columns must be named label or score.
from datasets import Dataset
For MultipleNegativesRankingLoss: (anchor, positive) pairs
traindataset = Dataset.fromdict({
"anchor": [
"How do I reset my password?",
"Where can I export my data?",
"When will my refund arrive?",
],
"positive": [
"Reset your password from the account settings page.",
"Export your data as CSV from the dashboard.",
"Refunds are processed within 5 business days.",
],
})
Setting Up the Trainer
SentenceTransformerTrainer ties everything together. Training arguments come from SentenceTransformerTrainingArguments, which extends the familiar TrainingArguments.
from sentencetransformers import (
SentenceTransformer,
SentenceTransformerTrainer,
SentenceTransformerTrainingArguments,
)
from sentence
transformers.losses import MultipleNegativesRankingLoss
model = SentenceTransformer("all-MiniLM-L6-v2")
loss = MultipleNegativesRankingLoss(model)
args = SentenceTransformerTrainingArguments(
outputdir="models/support-embeddings",
numtrainepochs=3,
perdevicetrainbatchsize=32,
learningrate=2e-5,
warmupratio=0.1,
fp16=True, # set bf16=True on newer GPUs instead
evalstrategy="steps",
evalsteps=100,
savestrategy="steps",
savesteps=100,
loggingsteps=20,
)
trainer = SentenceTransformerTrainer(
model=model,
args=args,
traindataset=traindataset,
loss=loss,
)
trainer.train()
Evaluating Retrieval Quality
To know whether fine-tuning helped, measure retrieval metrics on held-out data. InformationRetrievalEvaluator reports metrics such as Recall@k, MRR, and NDCG. It needs three mappings: queries, the corpus, and which corpus ids are relevant for each query.
from sentencetransformers.evaluation import InformationRetrievalEvaluator
queries = {"q1": "how to reset password", "q2": "download my data"}
corpus
map = {str(i): text for i, text in enumerate(corpus)}
relevant = {"q1": {"0"}, "q2": {"5"}} # corpus ids that answer each query
irevaluator = InformationRetrievalEvaluator(
queries=queries,
corpus=corpusmap,
relevantdocs=relevant,
name="support-eval",
)
Run standalone, or pass evaluator=irevaluator to the Trainer
results = irevaluator(model)
print(results)
Pass the evaluator to the trainer with evaluator=irevaluator to track these metrics during training and pick the best checkpoint.
Training a Cross-Encoder
You can fine-tune a reranker the same way with CrossEncoderTrainer. Cross-encoder training data is (query, document, label) where the label is a relevance score, often binary (1 for relevant, 0 for not).
from datasets import Dataset
from sentencetransformers.crossencoder import (
CrossEncoder,
CrossEncoderTrainer,
CrossEncoderTrainingArguments,
)
from sentencetransformers.crossencoder.losses import BinaryCrossEntropyLoss
cedata = Dataset.fromdict({
"query": ["reset password", "reset password", "export data"],
"document": [
"Reset your password from the account settings page.",
"Refunds are processed within 5 business days.",
"Export your data as CSV from the dashboard.",
],
"label": [1.0, 0.0, 1.0],
})
cemodel = CrossEncoder("microsoft/MiniLM-L12-H384-uncased", numlabels=1)
celoss = BinaryCrossEntropyLoss(cemodel)
ceargs = CrossEncoderTrainingArguments(
outputdir="models/support-reranker",
numtrainepochs=2,
perdevicetrainbatchsize=16,
)
cetrainer = CrossEncoderTrainer(
model=cemodel, args=ceargs, traindataset=cedata, loss=celoss,
)
cetrainer.train()
In practice, mining hard negatives (wrong documents that look plausible) makes a much stronger reranker than random negatives.
Saving, Loading, and Sharing
Save a trained model to disk and load it back like any other model.
model.savepretrained("models/support-embeddings")
reloaded = SentenceTransformer("models/support-embeddings")
To share on the Hugging Face Hub, log in first (huggingface-cli login) and push. This uploads the weights, configuration, and an auto-generated model card.
model.pushtohub("your-username/support-embeddings")
Speeding Up Inference: Quantization and ONNX
For production serving, two backends reduce latency without retraining. The library can load models with the ONNX or OpenVINO runtime via the backend argument, and supports quantized (int8) weights for further speedups on CPU.
# ONNX backend for faster CPU inference
model = SentenceTransformer(
"all-MiniLM-L6-v2",
backend="onnx",
modelkwargs={"filename": "onnx/modelqint8avx512.onnx"},
)
Quantized int8 models are smaller and faster on CPU with a small accuracy cost; measure on your own data before committing. On GPU, using fp16/bf16 and larger batch sizes is usually the simplest win.
Best Practices
- Match the model to your language and domain. Use multilingual models for Indonesian or mixed-language corpora, and check the model card for expected prompts.
- Normalize embeddings when your index scores by dot product, and keep the same normalization at query and index time.
- Always retrieve more candidates than you display, then rerank with a cross-encoder for the final ordering.
- Build an evaluation set early. Without
InformationRetrievalEvaluatormetrics you cannot tell whether fine-tuning helped or hurt. - Prefer
MultipleNegativesRankingLosswith as large a batch as memory allows; in-batch negatives are what make it effective. - Cache corpus embeddings. Re-encoding a large corpus on every restart is wasteful; store the vectors.
- For reranker training, invest in hard negatives rather than collecting more random pairs.
Conclusion and Key Takeaways
Sentence Transformers gives you a complete toolkit: fast bi-encoders for retrieval, accurate cross-encoders for reranking, and a modern training API to adapt both to your domain.
- Bi-encoders embed texts independently and scale to large corpora; cross-encoders score pairs jointly for higher precision but cannot precompute.
- The retrieve-then-rerank pattern combines the strengths of both and is the default architecture for serious search and RAG.
encode(),model.similarity, and theutilhelpers (semanticsearch,cossim,paraphrasemining,communitydetection) cover most everyday tasks.- Fine-tuning with
SentenceTransformerTrainer, a loss likeMultipleNegativesRankingLoss, and anInformationRetrievalEvaluatortypically yields the largest quality gains on domain data. - For production, consider ONNX or quantized backends to cut latency, and always cache your corpus embeddings.
Start from a strong pretrained model, measure on your own evaluation set, and fine-tune only where the numbers justify it.