llm-forensic-timeline
A work in progress repository for DFRWS APAC 2025 paper: A standardized methodology and dataset for evaluating LLM-based digital forensic timeline analysis.
The package covers the two halves of the methodology:
- Run — upload a log2timeline/Plaso CSV timeline to an LLM, replay the prompts of a task, and save every answer the model generates.
- Evaluate — score those answers against a manually validated ground truth with BLEU, ROUGE, and BERTScore, and write one CSV row per file.
Documentation: https://llm-forensic-timeline.readthedocs.io
Requirements
-
Create a virtual environment using
venv(Python 3.10 or newer; BERTScore additionally needs 3.12+):python3.13 -m venv .venv -
Activate the environment:
source .venv/bin/activate(Windows:.venv\Scripts\activate) -
Install the package:
pip install -e ".[all]"Or, without cloning:
pip install llm-forensic-timeline[all] -
Open dataset repository on Zenodo
-
Download
scenario-1.zipand unzip it -
Copy all files from
scenario-1/timeline/eventsdirectory todataset/event-summarizationdirectory
The extras split the optional halves: bertscore adds modern-bert-score, llm adds openai and python-dotenv, and all adds both. requirements.txt still holds the full environment used for the paper.
How to run
-
Add
.envfile insrcdirectory. Add your OpenAI API keys in the.envfile, something like:OPENAI_API_KEY=sk-proj-xxx -
Run the task
llm-forensic-timeline run --task event-summarizationThe original script still works:
python src/run.py
How to evaluate
llm-forensic-timeline evaluate compares the LLM answers against the ground truth and writes the scores to a CSV file:
llm-forensic-timeline evaluate \
--groundtruth-dir event-reconstruction/ground-truth-multiple \
--llm-dir event-reconstruction/chatgpt-with-knowledge-multiple \
--output event-reconstruction/chatgpt-with-knowledge-multiple-results.csv
The original script still works: edit the task and prediction variables at the top of src/evaluation-list.py, then run it from the directory that contains the task directory.
Three metrics are reported per file:
-
BLEU and ROUGE (ROUGE-1, ROUGE-2, ROUGE-L, ROUGE-Lsum) through the Hugging Face
evaluatelibrary. Both are n-gram overlap metrics, so they only measure lexical similarity. -
BERTScore through
modern-bert-score, which is semantic-aware and therefore credits an answer that is worded differently but carries the same meaning. This addresses the limitation of BLEU and ROUGE discussed in Sec. 4.4.2 of the paper. The columnsbertscore_precision,bertscore_recall, andbertscore_f1hold the mean over all entries of a file.
Notes on BERTScore:
- The model is
LazerLambda/ModernBERT-large-ModBERTScore-19, the strongest checkpoint shipped with the package. It is the largest one, and its 8192-token context keeps long JSON entries of a forensic timeline intact, while the RoBERTa checkpoints truncate at 512 tokens. The weights (about 1.6 GB) are downloaded from Hugging Face on the first run and cached afterwards. - Scores are rescaled with the baseline of the model, as recommended by the BERTScore authors. Without rescaling, scores of unrelated texts sit near 0.8; with rescaling, they sit near 0 and can be slightly negative, which makes the results easier to read.
- The script selects CUDA or Apple Silicon (MPS) automatically and falls back to the CPU. Evaluating a large timeline on the CPU is slow.
- Pass
--no-bertscoreto score with BLEU and ROUGE only, without loading a transformer model.
Development
Run the test suite:
pip install -e ".[test]"
pytest
Build the documentation locally:
pip install -r docs/requirements.txt
python -m sphinx -W -b html docs/source docs/_build/html
The Run Tests workflow runs pytest on Python 3.10 through 3.13 and builds the documentation on every push and pull request to main. Pushing a v*.*.* tag builds the distributions and publishes them to PyPI through OIDC Trusted Publishing, without any stored credentials.
Metadata
Release files for llm-forensic-timeline 0.0.2
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| llm_forensic_timeline-0.0.2.tar.gz | 22.0 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| llm_forensic_timeline-0.0.2-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 38.6 kB
Release files / llm_forensic_timeline-0.0.2.tar.gz
| Download URL | llm_forensic_timeline-0.0.2.tar.gz |
|---|---|
| Size | 22.0 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
073cd83893ace08a1e563003b002e9e5bfc2563e6c1da325aa3d06b88fb07e12
|
|
BLAKE2b-256 checksum How to use checksums |
b8773ebb596e9fd9f84e3df54454dec0ea7cf093a40669dcc35f7a3d8473bf26
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 7, 2026.
Transparency logRelease files / llm_forensic_timeline-0.0.2-py3-none-any.whl
| Download URL | llm_forensic_timeline-0.0.2-py3-none-any.whl |
|---|---|
| Size | 16.5 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
6649ab82b5aba592ea960922dc6b42891c340f0635eab72bcc08a345b6de1e14
|
|
BLAKE2b-256 checksum How to use checksums |
a4bff6db3248f549102aaaa97080d918fd9d776714b25281d86109ff363de341
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 7, 2026.
Transparency log