Skip to main content

llm-forensic-timeline

Tests Documentation Status PyPI

A work in progress repository for DFRWS APAC 2025 paper: A standardized methodology and dataset for evaluating LLM-based digital forensic timeline analysis.

The package covers the two halves of the methodology:

  • Run — upload a log2timeline/Plaso CSV timeline to an LLM, replay the prompts of a task, and save every answer the model generates.
  • Evaluate — score those answers against a manually validated ground truth with BLEU, ROUGE, and BERTScore, and write one CSV row per file.

Documentation: https://llm-forensic-timeline.readthedocs.io

Requirements

  1. Create a virtual environment using venv (Python 3.10 or newer; BERTScore additionally needs 3.12+):

    python3.13 -m venv .venv

  2. Activate the environment:

    source .venv/bin/activate (Windows: .venv\Scripts\activate)

  3. Install the package:

    pip install -e ".[all]"

    Or, without cloning: pip install llm-forensic-timeline[all]

  4. Open dataset repository on Zenodo

  5. Download scenario-1.zip and unzip it

  6. Copy all files from scenario-1/timeline/events directory to dataset/event-summarization directory

The extras split the optional halves: bertscore adds modern-bert-score, llm adds openai and python-dotenv, and all adds both. requirements.txt still holds the full environment used for the paper.

How to run

  1. Add .env file in src directory. Add your OpenAI API keys in the .env file, something like: OPENAI_API_KEY=sk-proj-xxx

  2. Run the task

    llm-forensic-timeline run --task event-summarization

    The original script still works: python src/run.py

How to evaluate

llm-forensic-timeline evaluate compares the LLM answers against the ground truth and writes the scores to a CSV file:

llm-forensic-timeline evaluate \
    --groundtruth-dir event-reconstruction/ground-truth-multiple \
    --llm-dir event-reconstruction/chatgpt-with-knowledge-multiple \
    --output event-reconstruction/chatgpt-with-knowledge-multiple-results.csv

The original script still works: edit the task and prediction variables at the top of src/evaluation-list.py, then run it from the directory that contains the task directory.

Three metrics are reported per file:

  1. BLEU and ROUGE (ROUGE-1, ROUGE-2, ROUGE-L, ROUGE-Lsum) through the Hugging Face evaluate library. Both are n-gram overlap metrics, so they only measure lexical similarity.

  2. BERTScore through modern-bert-score, which is semantic-aware and therefore credits an answer that is worded differently but carries the same meaning. This addresses the limitation of BLEU and ROUGE discussed in Sec. 4.4.2 of the paper. The columns bertscore_precision, bertscore_recall, and bertscore_f1 hold the mean over all entries of a file.

Notes on BERTScore:

  • The model is LazerLambda/ModernBERT-large-ModBERTScore-19, the strongest checkpoint shipped with the package. It is the largest one, and its 8192-token context keeps long JSON entries of a forensic timeline intact, while the RoBERTa checkpoints truncate at 512 tokens. The weights (about 1.6 GB) are downloaded from Hugging Face on the first run and cached afterwards.
  • Scores are rescaled with the baseline of the model, as recommended by the BERTScore authors. Without rescaling, scores of unrelated texts sit near 0.8; with rescaling, they sit near 0 and can be slightly negative, which makes the results easier to read.
  • The script selects CUDA or Apple Silicon (MPS) automatically and falls back to the CPU. Evaluating a large timeline on the CPU is slow.
  • Pass --no-bertscore to score with BLEU and ROUGE only, without loading a transformer model.

Development

Run the test suite:

pip install -e ".[test]"
pytest

Build the documentation locally:

pip install -r docs/requirements.txt
python -m sphinx -W -b html docs/source docs/_build/html

The Run Tests workflow runs pytest on Python 3.10 through 3.13 and builds the documentation on every push and pull request to main. Pushing a v*.*.* tag builds the distributions and publishes them to PyPI through OIDC Trusted Publishing, without any stored credentials.

Metadata

Release files for llm-forensic-timeline 0.0.2

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for llm-forensic-timeline 0.0.2
File Size Uploaded
llm_forensic_timeline-0.0.2.tar.gz 22.0 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for llm-forensic-timeline 0.0.2
File Interpreter ABI Platform
llm_forensic_timeline-0.0.2-py3-none-any.whl Python 3 none any Details

Total release size: 38.6 kB

Release files / llm_forensic_timeline-0.0.2.tar.gz

Download URL llm_forensic_timeline-0.0.2.tar.gz
Size 22.0 kB
Tags Source
SHA-256 checksum
How to use checksums
073cd83893ace08a1e563003b002e9e5bfc2563e6c1da325aa3d06b88fb07e12
BLAKE2b-256 checksum
How to use checksums
b8773ebb596e9fd9f84e3df54454dec0ea7cf093a40669dcc35f7a3d8473bf26
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 7, 2026.

Transparency log

Release files / llm_forensic_timeline-0.0.2-py3-none-any.whl

Download URL llm_forensic_timeline-0.0.2-py3-none-any.whl
Size 16.5 kB
Tags Python 3
SHA-256 checksum
How to use checksums
6649ab82b5aba592ea960922dc6b42891c340f0635eab72bcc08a345b6de1e14
BLAKE2b-256 checksum
How to use checksums
a4bff6db3248f549102aaaa97080d918fd9d776714b25281d86109ff363de341
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 7, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.0.2 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page