Skip to main content

prompt_miner

Data-mining tools for the Cirilla project — gathering and shaping training data for a budget LLM.

[!NOTE] This package is a part of the Cirilla project

Fandom Scraper

A lightweight Fandom wiki scraper built on the MediaWiki API.

No GPU required. No heavy ML dependencies. Works with any language supported by the target wiki.

How it works

  1. The scraper enumerates every page on the wiki via the MediaWiki list=allpages API, across all content namespaces (auto-detected — e.g. articles and Maps). This includes subpages (e.g. Geralt of Rivia/Netflix series).
  2. Each page is fetched through api.php and split into three artifacts:
    • the article body → plain .txt (tables are kept as readable rows, image captions and trailing sections like Trivia are preserved),
    • the infobox (race, profession, family, abilities, …) → structured .json,
    • the "Quick Answers" Q&A widget, when present → instruction .json.
  3. The async crawl runs across many concurrent fetchers and stops on its own once every discovered page has been processed.

Discovery does not rely on an NER model, so it works regardless of language and requires no GPU. Beyond full-wiki enumeration you can also seed the crawl from specific categories, search queries, or JSON seed lists, and optionally follow in-article links (BFS crawl).

User-agent strings are rotated on every request (pool of 7 real browser UAs) with matching Accept-Language headers per language, so the scraper blends in with normal browser traffic.

Installation

# from the repo root
uv sync          # or: pip install -e .

Usage

Full-wiki dump (default)

from prompt_miner.fandom_scraping import scrape_fandom

scrape_fandom(
    out_path="./pages",          # article body .txt files
    instruct_path="./instruct",  # Quick-Answers Q&A .json files (when found)
    infobox_path="./infobox",    # infobox .json files (when present)
    n_workers=50,                # concurrent async fetchers (raise for speed)
    wiki="Witcher",              # fandom wiki subdomain
    lang="en",                   # language code
    namespaces="content",        # auto-detect content namespaces (e.g. 0 + Map)
    use_allpages=True,           # enumerate the whole wiki via list=allpages
    follow_links=False,          # redundant when use_allpages is on
)

This discovers and scrapes every article on the wiki and terminates when the frontier is drained.

Targeted scraping

Disable use_allpages and seed from any combination of sources instead:

# Specific categories (reusable across wikis) — names with or without "Category:"
scrape_fandom(out_path="./pages", instruct_path="./instruct", infobox_path="./infobox",
              wiki="Witcher", lang="en",
              use_allpages=False, categories=["Subpages", "Characters"])

# Search-bar queries, then follow links from the matches
scrape_fandom(out_path="./pages", instruct_path="./instruct", infobox_path="./infobox",
              wiki="Witcher", lang="en",
              use_allpages=False, search_queries=["witcher", "monster"], follow_links=True)

# JSON seed-title files (a folder of JSON arrays of page names)
scrape_fandom(out_path="./pages", instruct_path="./instruct", infobox_path="./infobox",
              wiki="Witcher", lang="en",
              use_allpages=False, in_path="./seeds", follow_links=True)

A seed file is just a JSON array of starting page names:

[
    "Geralt of Rivia", "Triss Merigold", "Vesemir", "Lambert",
    "Eskel", "Shani", "Zoltan Chivay", "Dandelion (Jaskier)"
]

Parameters

Parameter Default Description
out_path / instruct_path / infobox_path Output folders for body text, Q&A JSON, infobox JSON.
in_path None Optional folder of JSON seed-title lists.
n_workers 10 Concurrent async fetchers.
wiki / lang "Witcher" / "en" Wiki subdomain and language code.
namespaces "content" "content" (auto-detect), an int, or a list of namespace ids.
use_allpages True Enumerate the entire wiki via list=allpages.
categories None List of category names to seed from.
search_queries None List of search-bar queries to seed from.
follow_links False Also crawl in-article links (BFS).
include_redirects False Include redirect pages (duplicate content).

Language support

The wiki URL is built from wiki + lang: https://<wiki>.fandom.com for English and https://<wiki>.fandom.com/<lang> otherwise. For the Polish Witcher wiki (witcher.fandom.com/pl):

scrape_fandom(
    out_path="./pages_pl",
    instruct_path="./instruct_pl",
    infobox_path="./infobox_pl",
    wiki="Witcher",
    lang="pl",
)

Languages with native Accept-Language headers: en, pl, de, fr, es, ru, pt, it, zh, ja, ko, cs, hu, tr, uk. Any other IETF language code is passed through with a generic fallback header.

Converting instructions to JSONL

If the scraper finds any Q&A pairs, convert them to the Cirilla .jsonl format:

from prompt_miner.fandom_scraping import instructions_into_conv

instructions_into_conv('./instruct', './fandom_instruct.jsonl')

Output format:

{"subject": "Shani", "text": [{"role": "user", "content": "What role did Shani play in the Battle of Brenna?"}, {"role": "assistant", "content": "Shani played a vital role..."}], "data type": "conv", "source": "fandom"}

Effectiveness

Full-wiki enumeration covers the entire site rather than only the pages reachable from a seed list. For the English Witcher wiki this is 12,307 articles (including 427 subpages) across the content namespaces.

Training Data Pipeline

After scraping, prompt_miner turns the raw articles into scored, structured instruct (Q&A) training data through a chain of LLM-prompted stages. This part needs a GPU and vLLM (already installed by uv sync) — every stage below loads a real quantized model (speakleash/Bielik-*, CYFRAGOVPL/Llama-PLLuM-*).

Stages

scraped .txt (+ infobox)
  → quality_check                    score 1-10, keep only pages above a threshold
  → extract_topics                   break the article into topics + supporting info
  → generate_subjects                group topics into broader subjects
  → create_syllabus                  design an ordered lesson outline per subject
  → extract_concepts                 break each syllabus section into atomic concept + key_info pairs
  → create_instruct_from_concept     generate a multi-hop Q&A conversation per concept
  → judge_instruct                   verify each conversation is grounded in its concept's key_info

Each stage is a script in scripts/, run in order (every script reads the previous stage's .jsonl and writes its own):

python scripts/assess_quality.py                 # -> prompt_miner/data/pl/async_fandom_scores.jsonl
python scripts/extract_topics.py                 # -> prompt_miner/data/pl/async_fandom_topics.jsonl
python scripts/generate_subjects.py              # -> prompt_miner/data/pl/async_fandom_subjects.jsonl
python scripts/create_syllabus.py                # -> prompt_miner/data/pl/async_fandom_syllabus.jsonl
python scripts/extract_concepts.py               # -> prompt_miner/data/pl/async_fandom_concepts.jsonl
python scripts/generate_instruct_from_concept.py # -> prompt_miner/data/pl/async_fandom_concept_instructs.jsonl
python scripts/judge_instruct.py                 # -> prompt_miner/data/pl/async_fandom_judged.jsonl

Edit the paths/models at the top of each script to point at a different scrape or language.

There's also an older, simpler path, scripts/generate_instruct.py, which generates instruct conversations directly from the raw article + infobox, skipping the topic/subject/syllabus/concept breakdown — useful for a quick pass, but without the per-concept grounding the judge stage checks for.

Prompts

All prompts live in prompt_miner/prompts/ as Jinja2 templates and are registered in prompt_manager.py. Every stage has a system prompt (task instructions + examples) and a user prompt (just the data), in both English and Polish (_pl suffix). Prompts are looked up by class name via AVAILABLE_PROMPTS:

from prompt_miner.prompts.prompt_manager import AVAILABLE_PROMPTS
AVAILABLE_PROMPTS["ExtractTopicsSystemPL"].render()

Language support

Like the scraper, the pipeline is duplicated per language — the scripts above point at prompt_miner/data/pl/... and use the *PL prompt classes; swap both for English (prompt_miner/data/en/..., drop the PL suffix).

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

prompt_miner-0.2.0.tar.gz (19.4 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

prompt_miner-0.2.0-py3-none-any.whl (21.6 kB view details)

Uploaded Python 3

File details

Details for the file prompt_miner-0.2.0.tar.gz.

File metadata

  • Download URL: prompt_miner-0.2.0.tar.gz
  • Upload date:
  • Size: 19.4 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: uv/0.9.26 {"installer":{"name":"uv","version":"0.9.26","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

File hashes

Hashes for prompt_miner-0.2.0.tar.gz
Algorithm Hash digest
SHA256 6d2f446df92227d6ad13edad46226f2595dea602bc676865ed093971cbbcc364
MD5 936496ca939e412223ff38168ab01176
BLAKE2b-256 d6c922f9427c032f5513b9bb1fbe52533d15e5824d3f3ee62f7096e519d820eb

See more details on using hashes here.

File details

Details for the file prompt_miner-0.2.0-py3-none-any.whl.

File metadata

  • Download URL: prompt_miner-0.2.0-py3-none-any.whl
  • Upload date:
  • Size: 21.6 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: uv/0.9.26 {"installer":{"name":"uv","version":"0.9.26","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

File hashes

Hashes for prompt_miner-0.2.0-py3-none-any.whl
Algorithm Hash digest
SHA256 b50b99cbaf7ba8cb523f5a67b6f2f672dcd958d1584d40801b751225afebc1fb
MD5 5d84b75b474a2fd6784cb899bb1c0c1b
BLAKE2b-256 0cde11154f8706737f95573c9707940a68ffb7e80f22ea231035c0e808c622bd

See more details on using hashes here.

Release history Release notifications | RSS feed

This release

0.2.0 This release

2 files

0.1.0

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page