Skip to main content

Loci Similes

LociSimiles is a Python package for finding intertextual links in Latin literature using pre-trained language models.

Basic Usage

# Load example query and source documents
query_doc = Document("../data/hieronymus_samples.csv")
source_doc = Document("../data/vergil_samples.csv")

# Load the pipeline with pre-trained models
pipeline = ClassificationPipelineWithCandidategeneration(
    classification_name="...",
    embedding_model_name="...",
    device="cpu",
)

# Run the pipeline with the query and source documents
results = pipeline.run(
    query=query_doc,    # Query document
    source=source_doc,  # Source document
    top_k=3             # Number of top similar candidates to classify
)

pretty_print(results)

# Save results to CSV or JSON
pipeline.to_csv("results.csv")
pipeline.to_json("results.json")

Multiclass Classifier Inference

Binary classifiers remain supported. If you load a trained multiclass sequence-classification model, LociSimiles preserves the usual judgment_score as the total probability of an intertextual link while also returning the predicted class label and full class probabilities.

pipeline = ClassificationPipelineWithCandidateGeneration(
  classification_name="path-or-hf-id-for-trained-multiclass-model",
  embedding_model_name="julian-schelb/multilingual-e5-large-emb-lat-intertext-v1",
  label_names=["no_match", "cit", "cf"],
  positive_labels=["cit", "cf"],
  device="cpu",
)

results = pipeline.run(query=query_doc, source=source_doc, top_k=20)
first = results["query-id"][0]
print(first.judgment_score)       # P(cit) + P(cf)
print(first.predicted_label)      # e.g. "cit" or "cf"
print(first.class_probabilities)  # {"no_match": ..., "cit": ..., "cf": ...}

Command-Line Interface

LociSimiles provides a command-line tool for running the pipeline directly from the terminal:

Basic Usage

locisimiles query.csv source.csv -o results.csv

Two-Stage Pipeline Example

locisimiles query.csv source.csv -o results.csv \
  --pipeline two-stage \
  --classification-model julian-schelb/xlm-roberta-large-class-lat-intertext-v1 \
  --embedding-model julian-schelb/multilingual-e5-large-emb-lat-intertext-v1 \
  --top-k 20 \
  --threshold 0.85 \
  --device cuda \
  --verbose

Word2Vec Retrieval Example

locisimiles query.csv source.csv -o results.csv \
  --pipeline word2vec-retrieval \
  --word2vec-model-path ./models/latin_w2v_bamman_lemma300_100_1.model \
  --word2vec-interval 2 \
  --word2vec-order-free \
  --top-k 20 \
  --threshold 0.85

Latin BERT Retrieval Example (Gong-Style)

locisimiles query.csv source.csv -o results.csv \
  --pipeline latin-bert-retrieval \
  --latin-bert-model ashleygong03/bamman-burns-latin-bert \
  --top-k 20 \
  --threshold 0.85

BM25 Retrieval Example

BM25 is the benchmark's best single retriever, and requires no trained model:

locisimiles query.csv source.csv -o results.csv \
  --pipeline bm25-retrieval \
  --bm25-k1 1.5 \
  --bm25-b 0.75 \
  --top-k 20 \
  --threshold 0.85

BM25 + Lexical Classifier Example (Best Non-Neural)

Combines BM25 retrieval with a trained LogReg/GBDT classifier (LexicalClassifierTrainer) — no neural model required end to end:

locisimiles query.csv source.csv -o results.csv \
  --pipeline bm25-lexical-two-stage \
  --lexical-classifier-path ./models/lexical_classifier.joblib \
  --top-k 20 \
  --threshold 0.85

TF-IDF retrieval (--pipeline tfidf-retrieval) and BM25 + cross-encoder classification (--pipeline bm25-two-stage, the benchmark's "best combined" configuration) are also available; see the CLI docs for the full set of pipelines and options.

If --word2vec-model-path is not provided, the CLI expects a local model at:

models/latin_w2v_bamman_lemma300_100_1.model

Word2Vec mode requires pre-lemmatized input in the same CSV format (seg_id, text).

Options

  • Input/Output:

    • query: Path to query document CSV file (columns: seg_id, text)
    • source: Path to source document CSV file (columns: seg_id, text)
    • -o, --output: Path to output CSV file for results (required)
  • Models:

    • --classification-model: HuggingFace model for classification (default: xlm-roberta-large-class-lat-intertext-v1)
    • --embedding-model: HuggingFace model for embeddings (default: multilingual-e5-large-emb-lat-intertext-v1)
    • --word2vec-model-path: Local path to a gensim .model file (Word2Vec pipeline)
    • --lexical-classifier-path: Local path to a .joblib artifact from LexicalClassifierTrainer (required for bm25-lexical-two-stage)
  • Pipeline Parameters:

    • --pipeline: Select two-stage, word2vec-retrieval, latin-bert-retrieval, latin-bert-two-stage, tfidf-retrieval, bm25-retrieval, bm25-two-stage, or bm25-lexical-two-stage (default: two-stage)
    • -k, --top-k: Number of top candidates to retrieve per query segment (default: 10)
    • -t, --threshold: Decision threshold for output filtering (default: 0.85)
    • --word2vec-interval: Max token gap for Word2Vec bigrams (default: 0)
    • --word2vec-order-free: Enable order-insensitive Word2Vec bigrams
    • --lexical-disable-lemmatize: Disable CLTK lemmatization for TF-IDF/BM25/lexical-classifier pipelines
    • --tfidf-ngram-max: Maximum lemma n-gram size for TF-IDF (default: 1)
    • --bm25-k1 / --bm25-b: BM25 term-frequency saturation / length-normalization parameters (defaults: 1.5 / 0.75)
  • Device:

    • --device: Choose auto, cuda, mps, or cpu (default: auto-detect)
  • Other:

    • -v, --verbose: Enable detailed progress output
    • -h, --help: Show help message

Output Format

The CLI saves results to a CSV file with the following columns:

  • query_id: Query segment identifier
  • query_text: Query text content
  • source_id: Source segment identifier
  • source_text: Source text content
  • similarity: Cosine similarity score (0-1)
  • probability: Link confidence (0-1); for multiclass classifiers this is the summed probability of positive classes
  • above_threshold: "Yes" if probability ≥ threshold, otherwise "No"

When a multiclass classifier returns class metadata, the CLI also writes predicted_class_id, predicted_label, and class_probabilities.

Training

LociSimiles also ships trainers for every trainable approach in the benchmark: LexicalClassifierTrainer, Word2VecTrainer, ClassificationTrainer, and EmbeddingTrainer. The three pair/label trainers share one input type, TrainingData, which bundles a query/source Document pair with a GroundTruth of labeled pairs and offers all four of the paper's negative-sampling methods as chainable methods:

from locisimiles.document import Document
from locisimiles.ground_truth import GroundTruth
from locisimiles.training.data import TrainingData
from locisimiles.training.classification import ClassificationTrainer, ClassificationTrainerConfig

query_doc = Document("query.csv")
source_doc = Document("source.csv")
positives = GroundTruth("known_positives.csv")

data = TrainingData(query_doc, source_doc, positives).sample_random_negatives(n_per_query=5)

config = ClassificationTrainerConfig(
    output_dir="models/classifier",
    label_names={0: "no_match", 1: "cit", 2: "cf"},
)
trainer = ClassificationTrainer(config)
trainer.fit(data=data)
model_path = trainer.save()

See the Training module docs for the full API, including threshold tuning/application for the classifier and all four negative-sampling methods.

Optional Gradio GUI

Install the optional GUI extra to experiment with a minimal Gradio front end:

pip install locisimiles[gui]

Launch the interface from the command line:

locisimiles-gui

In the GUI, choose Word2Vec Retrieval (Burns-Style) in Pipeline Configuration to enable Word2Vec controls:

  • Word2Vec Model Path: local gensim .model file
  • Bigram Interval: token gap for bigram generation
  • Order-Free Bigrams: optional order-insensitive matching

If the model path is invalid or missing, processing fails with a clear error message.

Metadata

Release files for locisimiles 2.1.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for locisimiles 2.1.1
File Size Uploaded
locisimiles-2.1.1.tar.gz 101.8 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for locisimiles 2.1.1
File Interpreter ABI Platform
locisimiles-2.1.1-py3-none-any.whl Python 3 none any Details

Total release size: 236.7 kB

Release files / locisimiles-2.1.1.tar.gz

Download URL locisimiles-2.1.1.tar.gz
Size 101.8 kB
Tags Source
SHA-256 checksum
How to use checksums
421671a6885e7938023d509379d68f15a30b0249798d3e8113f4d4361e3f5824
BLAKE2b-256 checksum
How to use checksums
d3576b083cdffc6bbc5c40ec385f8ab32405b4468edb40021534632c688c319b
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 24, 2026.

Transparency log

Release files / locisimiles-2.1.1-py3-none-any.whl

Download URL locisimiles-2.1.1-py3-none-any.whl
Size 134.9 kB
Tags Python 3
SHA-256 checksum
How to use checksums
08613ac38ca28a9162e3293ed4192b9697f0172f527391a5420b4e3a9b18632b
BLAKE2b-256 checksum
How to use checksums
c25efa5509664f6b9cdca6a7103c62778d9913389428c317f72ca915f4acd014
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 24, 2026.

Transparency log

Release history Release notifications | RSS feed

2.1.3

2 release files

2.1.2

2 release files

This release

2.1.1 This release

2 release files

2.1.0

2 release files

2.0.0

2 release files

1.8.0

2 release files

1.7.0

2 release files

1.6.0

2 release files

1.5.0

2 release files

1.4.0

2 release files

1.3.0

2 release files

1.2.0

2 release files

1.1.0

2 release files

1.0.1

2 release files

0.3.6

2 release files

0.3.5

2 release files

0.3.4

2 release files

0.3.3

2 release files

0.3.2

2 release files

0.3.1

2 release files

0.3.0

2 release files

0.2.9

2 release files

0.2.8

2 release files

0.2.7

2 release files

0.2.6

2 release files

0.2.5

2 release files

0.2.4

2 release files

0.2.3

2 release files

0.2.2

2 release files

0.2.1

2 release files

0.2.0

2 release files

0.1.4

2 release files

0.1.3

2 release files

0.1.2

2 release files

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page