Loci Similes
LociSimiles is a Python package for finding intertextual links in Latin literature using pre-trained language models.
Basic Usage
# Load example query and source documents
query_doc = Document("../data/hieronymus_samples.csv")
source_doc = Document("../data/vergil_samples.csv")
# Load the pipeline with pre-trained models
pipeline = ClassificationPipelineWithCandidategeneration(
classification_name="...",
embedding_model_name="...",
device="cpu",
)
# Run the pipeline with the query and source documents
results = pipeline.run(
query=query_doc, # Query document
source=source_doc, # Source document
top_k=3 # Number of top similar candidates to classify
)
pretty_print(results)
# Save results to CSV or JSON
pipeline.to_csv("results.csv")
pipeline.to_json("results.json")
Multiclass Classifier Inference
Binary classifiers remain supported. If you load a trained multiclass
sequence-classification model, LociSimiles preserves the usual
judgment_score as the total probability of an intertextual link while also
returning the predicted class label and full class probabilities.
pipeline = ClassificationPipelineWithCandidateGeneration(
classification_name="path-or-hf-id-for-trained-multiclass-model",
embedding_model_name="julian-schelb/multilingual-e5-large-emb-lat-intertext-v1",
label_names=["no_match", "cit", "cf"],
positive_labels=["cit", "cf"],
device="cpu",
)
results = pipeline.run(query=query_doc, source=source_doc, top_k=20)
first = results["query-id"][0]
print(first.judgment_score) # P(cit) + P(cf)
print(first.predicted_label) # e.g. "cit" or "cf"
print(first.class_probabilities) # {"no_match": ..., "cit": ..., "cf": ...}
Command-Line Interface
LociSimiles provides a command-line tool for running the pipeline directly from the terminal:
Basic Usage
locisimiles query.csv source.csv -o results.csv
Two-Stage Pipeline Example
locisimiles query.csv source.csv -o results.csv \
--pipeline two-stage \
--classification-model julian-schelb/xlm-roberta-large-class-lat-intertext-v1 \
--embedding-model julian-schelb/multilingual-e5-large-emb-lat-intertext-v1 \
--top-k 20 \
--threshold 0.85 \
--device cuda \
--verbose
Word2Vec Retrieval Example
locisimiles query.csv source.csv -o results.csv \
--pipeline word2vec-retrieval \
--word2vec-model-path ./models/latin_w2v_bamman_lemma300_100_1.model \
--word2vec-interval 2 \
--word2vec-order-free \
--top-k 20 \
--threshold 0.85
Latin BERT Retrieval Example (Gong-Style)
locisimiles query.csv source.csv -o results.csv \
--pipeline latin-bert-retrieval \
--latin-bert-model ashleygong03/bamman-burns-latin-bert \
--top-k 20 \
--threshold 0.85
BM25 Retrieval Example
BM25 is the benchmark's best single retriever, and requires no trained model:
locisimiles query.csv source.csv -o results.csv \
--pipeline bm25-retrieval \
--bm25-k1 1.5 \
--bm25-b 0.75 \
--top-k 20 \
--threshold 0.85
BM25 + Lexical Classifier Example (Best Non-Neural)
Combines BM25 retrieval with a trained LogReg/GBDT classifier
(LexicalClassifierTrainer) — no neural model required end to end:
locisimiles query.csv source.csv -o results.csv \
--pipeline bm25-lexical-two-stage \
--lexical-classifier-path ./models/lexical_classifier.joblib \
--top-k 20 \
--threshold 0.85
TF-IDF retrieval (--pipeline tfidf-retrieval) and BM25 + cross-encoder
classification (--pipeline bm25-two-stage, the benchmark's "best combined"
configuration) are also available; see the CLI docs
for the full set of pipelines and options.
If --word2vec-model-path is not provided, the CLI expects a local model at:
models/latin_w2v_bamman_lemma300_100_1.model
Word2Vec mode requires pre-lemmatized input in the same CSV format (seg_id, text).
Options
-
Input/Output:
query: Path to query document CSV file (columns:seg_id,text)source: Path to source document CSV file (columns:seg_id,text)-o, --output: Path to output CSV file for results (required)
-
Models:
--classification-model: HuggingFace model for classification (default: xlm-roberta-large-class-lat-intertext-v1)--embedding-model: HuggingFace model for embeddings (default: multilingual-e5-large-emb-lat-intertext-v1)--word2vec-model-path: Local path to a gensim.modelfile (Word2Vec pipeline)--lexical-classifier-path: Local path to a.joblibartifact fromLexicalClassifierTrainer(required forbm25-lexical-two-stage)
-
Pipeline Parameters:
--pipeline: Selecttwo-stage,word2vec-retrieval,latin-bert-retrieval,latin-bert-two-stage,tfidf-retrieval,bm25-retrieval,bm25-two-stage, orbm25-lexical-two-stage(default:two-stage)-k, --top-k: Number of top candidates to retrieve per query segment (default: 10)-t, --threshold: Decision threshold for output filtering (default: 0.85)--word2vec-interval: Max token gap for Word2Vec bigrams (default: 0)--word2vec-order-free: Enable order-insensitive Word2Vec bigrams--lexical-disable-lemmatize: Disable CLTK lemmatization for TF-IDF/BM25/lexical-classifier pipelines--tfidf-ngram-max: Maximum lemma n-gram size for TF-IDF (default: 1)--bm25-k1/--bm25-b: BM25 term-frequency saturation / length-normalization parameters (defaults: 1.5 / 0.75)
-
Device:
--device: Chooseauto,cuda,mps, orcpu(default: auto-detect)
-
Other:
-v, --verbose: Enable detailed progress output-h, --help: Show help message
Output Format
The CLI saves results to a CSV file with the following columns:
query_id: Query segment identifierquery_text: Query text contentsource_id: Source segment identifiersource_text: Source text contentsimilarity: Cosine similarity score (0-1)probability: Link confidence (0-1); for multiclass classifiers this is the summed probability of positive classesabove_threshold: "Yes" if probability ≥ threshold, otherwise "No"
When a multiclass classifier returns class metadata, the CLI also writes
predicted_class_id, predicted_label, and class_probabilities.
Training
LociSimiles also ships trainers for every trainable approach in the
benchmark: LexicalClassifierTrainer, Word2VecTrainer,
ClassificationTrainer, and EmbeddingTrainer. The three pair/label
trainers share one input type, TrainingData, which bundles a query/source
Document pair with a GroundTruth of labeled pairs and offers all four of
the paper's negative-sampling methods as chainable methods:
from locisimiles.document import Document
from locisimiles.ground_truth import GroundTruth
from locisimiles.training.data import TrainingData
from locisimiles.training.classification import ClassificationTrainer, ClassificationTrainerConfig
query_doc = Document("query.csv")
source_doc = Document("source.csv")
positives = GroundTruth("known_positives.csv")
data = TrainingData(query_doc, source_doc, positives).sample_random_negatives(n_per_query=5)
config = ClassificationTrainerConfig(
output_dir="models/classifier",
label_names={0: "no_match", 1: "cit", 2: "cf"},
)
trainer = ClassificationTrainer(config)
trainer.fit(data=data)
model_path = trainer.save()
See the Training module docs for the full API, including threshold tuning/application for the classifier and all four negative-sampling methods.
Optional Gradio GUI
Install the optional GUI extra to experiment with a minimal Gradio front end:
pip install locisimiles[gui]
Launch the interface from the command line:
locisimiles-gui
In the GUI, choose Word2Vec Retrieval (Burns-Style) in Pipeline Configuration to enable Word2Vec controls:
- Word2Vec Model Path: local gensim
.modelfile - Bigram Interval: token gap for bigram generation
- Order-Free Bigrams: optional order-insensitive matching
If the model path is invalid or missing, processing fails with a clear error message.
Metadata
Release files for locisimiles 2.1.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| locisimiles-2.1.1.tar.gz | 101.8 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| locisimiles-2.1.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 236.7 kB
Release files / locisimiles-2.1.1.tar.gz
| Download URL | locisimiles-2.1.1.tar.gz |
|---|---|
| Size | 101.8 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
421671a6885e7938023d509379d68f15a30b0249798d3e8113f4d4361e3f5824
|
|
BLAKE2b-256 checksum How to use checksums |
d3576b083cdffc6bbc5c40ec385f8ab32405b4468edb40021534632c688c319b
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Aug 24, 2026.
Transparency logRelease files / locisimiles-2.1.1-py3-none-any.whl
| Download URL | locisimiles-2.1.1-py3-none-any.whl |
|---|---|
| Size | 134.9 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
08613ac38ca28a9162e3293ed4192b9697f0172f527391a5420b4e3a9b18632b
|
|
BLAKE2b-256 checksum How to use checksums |
c25efa5509664f6b9cdca6a7103c62778d9913389428c317f72ca915f4acd014
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Aug 24, 2026.
Transparency log