Skip to main content

Loci Similes

LociSimiles is a Python package for finding intertextual links in Latin literature using pre-trained language models.

Basic Usage

# Load example query and source documents
query_doc = Document("../data/hieronymus_samples.csv")
source_doc = Document("../data/vergil_samples.csv")

# Load the pipeline with pre-trained models
pipeline = ClassificationPipelineWithCandidategeneration(
    classification_name="...",
    embedding_model_name="...",
    device="cpu",
)

# Run the pipeline with the query and source documents
results = pipeline.run(
    query=query_doc,    # Query document
    source=source_doc,  # Source document
    top_k=3             # Number of top similar candidates to classify
)

pretty_print(results)

# Save results to CSV or JSON
pipeline.to_csv("results.csv")
pipeline.to_json("results.json")

Multiclass Classifier Inference

Binary classifiers remain supported. If you load a trained multiclass sequence-classification model, LociSimiles preserves the usual judgment_score as the total probability of an intertextual link while also returning the predicted class label and full class probabilities.

pipeline = ClassificationPipelineWithCandidateGeneration(
  classification_name="path-or-hf-id-for-trained-multiclass-model",
  embedding_model_name="julian-schelb/multilingual-e5-large-emb-lat-intertext-v1",
  label_names=["no_match", "cit", "cf"],
  positive_labels=["cit", "cf"],
  device="cpu",
)

results = pipeline.run(query=query_doc, source=source_doc, top_k=20)
first = results["query-id"][0]
print(first.judgment_score)       # P(cit) + P(cf)
print(first.predicted_label)      # e.g. "cit" or "cf"
print(first.class_probabilities)  # {"no_match": ..., "cit": ..., "cf": ...}

Command-Line Interface

LociSimiles provides a command-line tool for running the pipeline directly from the terminal:

Basic Usage

locisimiles query.csv source.csv -o results.csv

Two-Stage Pipeline Example

locisimiles query.csv source.csv -o results.csv \
  --pipeline two-stage \
  --classification-model julian-schelb/xlm-roberta-large-class-lat-intertext-v1 \
  --embedding-model julian-schelb/multilingual-e5-large-emb-lat-intertext-v1 \
  --top-k 20 \
  --threshold 0.85 \
  --device cuda \
  --verbose

Word2Vec Retrieval Example

locisimiles query.csv source.csv -o results.csv \
  --pipeline word2vec-retrieval \
  --word2vec-model-path ./models/latin_w2v_bamman_lemma300_100_1.model \
  --word2vec-interval 2 \
  --word2vec-order-free \
  --top-k 20 \
  --threshold 0.85

Latin BERT Retrieval Example (Gong-Style)

locisimiles query.csv source.csv -o results.csv \
  --pipeline latin-bert-retrieval \
  --latin-bert-model ashleygong03/bamman-burns-latin-bert \
  --top-k 20 \
  --threshold 0.85

BM25 Retrieval Example

BM25 is the benchmark's best single retriever, and requires no trained model:

locisimiles query.csv source.csv -o results.csv \
  --pipeline bm25-retrieval \
  --bm25-k1 1.5 \
  --bm25-b 0.75 \
  --top-k 20 \
  --threshold 0.85

BM25 + Lexical Classifier Example (Best Non-Neural)

Combines BM25 retrieval with a trained LogReg/GBDT classifier (LexicalClassifierTrainer) — no neural model required end to end:

locisimiles query.csv source.csv -o results.csv \
  --pipeline bm25-lexical-two-stage \
  --lexical-classifier-path ./models/lexical_classifier.joblib \
  --top-k 20 \
  --threshold 0.85

TF-IDF retrieval (--pipeline tfidf-retrieval) and BM25 + cross-encoder classification (--pipeline bm25-two-stage, the benchmark's "best combined" configuration) are also available; see the CLI docs for the full set of pipelines and options.

If --word2vec-model-path is not provided, the CLI expects a local model at:

models/latin_w2v_bamman_lemma300_100_1.model

Word2Vec mode requires pre-lemmatized input in the same CSV format (seg_id, text).

Options

  • Input/Output:

    • query: Path to query document CSV file (columns: seg_id, text)
    • source: Path to source document CSV file (columns: seg_id, text)
    • -o, --output: Path to output CSV file for results (required)
  • Models:

    • --classification-model: HuggingFace model for classification (default: xlm-roberta-large-class-lat-intertext-v1)
    • --embedding-model: HuggingFace model for embeddings (default: multilingual-e5-large-emb-lat-intertext-v1)
    • --word2vec-model-path: Local path to a gensim .model file (Word2Vec pipeline)
    • --lexical-classifier-path: Local path to a .joblib artifact from LexicalClassifierTrainer (required for bm25-lexical-two-stage)
  • Pipeline Parameters:

    • --pipeline: Select two-stage, word2vec-retrieval, latin-bert-retrieval, latin-bert-two-stage, tfidf-retrieval, bm25-retrieval, bm25-two-stage, or bm25-lexical-two-stage (default: two-stage)
    • -k, --top-k: Number of top candidates to retrieve per query segment (default: 10)
    • -t, --threshold: Decision threshold for output filtering (default: 0.85)
    • --word2vec-interval: Max token gap for Word2Vec bigrams (default: 0)
    • --word2vec-order-free: Enable order-insensitive Word2Vec bigrams
    • --lexical-disable-lemmatize: Disable CLTK lemmatization for TF-IDF/BM25/lexical-classifier pipelines
    • --tfidf-ngram-max: Maximum lemma n-gram size for TF-IDF (default: 1)
    • --bm25-k1 / --bm25-b: BM25 term-frequency saturation / length-normalization parameters (defaults: 1.5 / 0.75)
  • Device:

    • --device: Choose auto, cuda, mps, or cpu (default: auto-detect)
  • Other:

    • -v, --verbose: Enable detailed progress output
    • -h, --help: Show help message

Output Format

The CLI saves results to a CSV file with the following columns:

  • query_id: Query segment identifier
  • query_text: Query text content
  • source_id: Source segment identifier
  • source_text: Source text content
  • similarity: Cosine similarity score (0-1)
  • probability: Link confidence (0-1); for multiclass classifiers this is the summed probability of positive classes
  • above_threshold: "Yes" if probability ≥ threshold, otherwise "No"

When a multiclass classifier returns class metadata, the CLI also writes predicted_class_id, predicted_label, and class_probabilities.

Optional Gradio GUI

Install the optional GUI extra to experiment with a minimal Gradio front end:

pip install locisimiles[gui]

Launch the interface from the command line:

locisimiles-gui

In the GUI, choose Word2Vec Retrieval (Burns-Style) in Pipeline Configuration to enable Word2Vec controls:

  • Word2Vec Model Path: local gensim .model file
  • Bigram Interval: token gap for bigram generation
  • Order-Free Bigrams: optional order-insensitive matching

If the model path is invalid or missing, processing fails with a clear error message.

Metadata

Release files for locisimiles 1.8.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for locisimiles 1.8.0
File Size Uploaded
locisimiles-1.8.0.tar.gz 82.2 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for locisimiles 1.8.0
File Interpreter ABI Platform
locisimiles-1.8.0-py3-none-any.whl Python 3 none any Details

Total release size: 192.3 kB

Release files / locisimiles-1.8.0.tar.gz

Download URL locisimiles-1.8.0.tar.gz
Size 82.2 kB
Tags Source
SHA-256 checksum
How to use checksums
070ce01719ab39d24cacf33607b2ba5d31c48be53a052c5bb8e50557482386d3
BLAKE2b-256 checksum
How to use checksums
9832a7dfcdf73edf27c6d25a45a69913f3c2dfc4a91af7f4c233e2d854bfc48e
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 23, 2026.

Transparency log

Release files / locisimiles-1.8.0-py3-none-any.whl

Download URL locisimiles-1.8.0-py3-none-any.whl
Size 110.0 kB
Tags Python 3
SHA-256 checksum
How to use checksums
6045b31d6211d0fd2705e5e73e094d2cb65c8a1f6fd9de7c6ca947987b7ebd06
BLAKE2b-256 checksum
How to use checksums
f8dd75327057ca1067a79ebede6d41844b2c15d367e615bea6d4f512a976c8cb
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 23, 2026.

Transparency log

Release history Release notifications | RSS feed

2.1.3

2 release files

2.1.2

2 release files

2.1.1

2 release files

2.1.0

2 release files

2.0.0

2 release files

This release

1.8.0 This release

2 release files

1.7.0

2 release files

1.6.0

2 release files

1.5.0

2 release files

1.4.0

2 release files

1.3.0

2 release files

1.2.0

2 release files

1.1.0

2 release files

1.0.1

2 release files

0.3.6

2 release files

0.3.5

2 release files

0.3.4

2 release files

0.3.3

2 release files

0.3.2

2 release files

0.3.1

2 release files

0.3.0

2 release files

0.2.9

2 release files

0.2.8

2 release files

0.2.7

2 release files

0.2.6

2 release files

0.2.5

2 release files

0.2.4

2 release files

0.2.3

2 release files

0.2.2

2 release files

0.2.1

2 release files

0.2.0

2 release files

0.1.4

2 release files

0.1.3

2 release files

0.1.2

2 release files

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page