Transaction pattern utilities and dataset for statement generators
Project description
Solyanka
Toolkit + dataset for transaction-pattern driven synthetic statements and downstream LLM fine-tuning. This package ships the curated YAML files, their schema, and a tiny loader that apps or notebooks can use without worrying about file layout.
Install
pip install solyanka # consumers
pip install -e ".[dev]" # local hacking (tests + linters)
Runtime use
from solyanka import PatternsService
svc = PatternsService() # auto-discovers packaged data
general = svc.load_general_patterns()
eea = svc.load_eea_patterns()
thailand = svc.load_country_patterns("Thailand")
# Recommended helper: general + (EEA) + country
bundle = svc.get_country_patterns("Germany")
# Fine-grained slices (e.g., validation scripts)
custom = svc.get_patterns(country="Germany", include="general,eea")
# API-ready dictionaries
payload = svc.get_pattern_dicts(country="Spain")
Override the dataset path (e.g., while editing YAML) via PatternsService(base_dir=Path("./transaction_patterns"))
or the TRANSACTION_PATTERNS_DIR environment variable.
Layout
solyanka/transaction_patterns/data/*.yml— curated pattern files (general.yml,eea.yml,<country>.yml).solyanka/transaction_patterns/data/schema.json— JSON Schema enforced by tests/CI.solyanka/transaction_patterns/service.py— public loader API (keep backward compatible).tests/— schema regression + loader behaviour.
Field spec
Required fields
| Field | Meaning |
|---|---|
title |
Merchant label. Plain string or template object. |
prettyTitle |
Short merchant label for UI use. Strip cities/countries/punctuation manually (no scripts); follow the documented brand rules (Uber/Amazon/Airbnb/Youtube/Godaddy/Myprotein/Zenni/Bolt/Iherb/Lotus, etc.). |
currency |
Uppercase ISO 4217 (EUR, USD, GBP, ...). |
amountRange |
{min, max} floats describing the observed local-currency range. Mutually exclusive with amounts. |
amounts |
List of specific float amounts (e.g. [9.99, 19.99]). Use this instead of amountRange for fixed price points. |
amountFormat |
Rounding strategy: n>0 decimals, 0 whole units, n<0 powers of ten (e.g., -2 rounds to 100s). |
types |
Non-empty list of lowercase tags (shopping, restaurant, transportation, ...). |
prettyTitle |
Short merchant label for UI use. Strip cities/countries/punctuation manually (no scripts); follow the documented brand rules (Uber/Amazon/Airbnb/Youtube/Godaddy/Myprotein/Zenni/Bolt/Iherb/Lotus, etc.). |
Optional fields
| Field | Why / how |
|---|---|
weight |
Relative selection probability. 100 = baseline, 120–150 very common, 50 niche. |
refundProbability |
Chance (0–1) that the generator emits a CARD_REFUND for this pattern. |
refundDelayMinHours/Max |
Boundaries for automatic refund timing (defaults: 72 / 288 hours). |
numberOfOccurrences |
Global cap per statement (useful for rare, one-off merchants). |
subscriptionFrequencyDays |
Frequency for recurring charges (e.g., 30 for monthly subscriptions). |
country |
Required when using region. Valid values defined in schema.json. |
region |
Geographic region within a country. Valid values defined in schema.json. |
Regions
The region field allows you to specify the geographic region within a country where a transaction pattern is localized. When using region, you must also specify country — the JSON schema validates that the region matches the country.
Valid countries and their regions are defined in schema.json (the allOf section with if/then rules). The schema is the single source of truth for allowed values.
When to use regions:
- Use
regiononly when the transaction clearly belongs to a specific geographic area (e.g., a local restaurant, a regional shop). - Do not set
regionfor online services, nationwide chains, or when the location is ambiguous. - When specifying
region, you must also setcountryto the matching country name.
Example with region:
- title: "Patong Beach Hotel PHUKET"
prettyTitle: "Patong Beach Hotel"
currency: "THB"
amountRange: {min: 2000, max: 8000}
amountFormat: 0
types: ["hotel", "accommodation"]
country: "Thailand"
region: "Phuket"
Template titles
title:
type: template
template: "Revolut**{num}* DUBLIN"
prettyTitle: "Revolut"
params:
num:
generator: random_digits
length: 4
zero_pad: true
globalConstant: true
transform:
case: upper
Generators & parameters
| Generator | Required params | Optional params | Notes |
|---|---|---|---|
random_digits |
length |
zero_pad (default true) |
digits only; zero_pad keeps leading zeroes |
random_alnum |
length |
charset |
mix of letters/digits; charset restricts symbols. The default is abcdefghijklmnopqrstuvwxyz0123456789. |
choice |
options |
weights (same length) |
uniform when weights omitted |
Extras:
globalConstant: true— reuse the same generated value across the statement (great for IDs).transform.case:upper,lower, ortitle.
Examples
Simple grocery merchant
- title: "Tesco Express"
prettyTitle: "Tesco Express"
currency: "GBP"
amountRange: {min: 5.0, max: 50.0}
amountFormat: 2
types: ["groceries", "shopping"]
weight: 120
Subscription service
- title: "Netflix.com"
prettyTitle: "Netflix"
currency: "EUR"
amountRange: {min: 13.49, max: 13.49}
amountFormat: 2
subscriptionFrequencyDays: 30
numberOfOccurrences: 10
types: ["entertainment", "subscription"]
weight: 300
Template with refund metadata
- title:
type: template
prettyTitle: "Airbnb"
template: "Airbnb * {code} 662-105-6167"
params:
code:
generator: random_alnum
length: 12
charset: "abcdefghijklmnopqrstuvwxyz0123456789"
transform:
case: lower
globalConstant: true
currency: "USD"
amountRange: {min: 70.0, max: 900.0}
amountFormat: 2
refundProbability: 0.4
types: ["housing"]
weight: 700
Exact amounts (Fixed price points)
- title: "Spotify Premium"
prettyTitle: "Spotify"
currency: "EUR"
amounts: [4.99, 9.99, 14.99]
amountFormat: 2
types: ["subscription", "entertainment"]
weight: 200
Pattern authoring workflow
- Pick the right file (
general.yml,eea.yml, or<country>.yml). - Study existing entries (Thailand’s file is a good reference for tone + “uglified” merchant names).
- Choose realistic
amountRange,amountFormat, tags, and weights. - Use templates when merchants expose reference numbers.
- Annotate generated blocks with comments (e.g.,
# Generated transaction pattern - online food). - Refresh
prettyTitleafter title tweaks by applying the manual derivation rules (strip city/country noise, drop IDs, canonicalize big brands). - Run
pytestto validate againstschema.jsonbefore committing/publishing.
Tests & release
pytest # validates YAML + loader invariants
python -m build # optional local artifact check
- CI:
.github/workflows/ci.ymlruns pytest on push/PR. - Release automation: merge PRs into
mainwithmajor release,minor release, orpatch releaselabels to control how.github/workflows/release-tagger.ymlbumps the version after CI finishes green. No label defaults to a build bump (v1.2.3→v1.2.3.1, etc.). The workflow updatespyproject.tomland tags the commit asv<version>. - PyPI publish: semantic tags (
vMAJOR.MINOR.PATCHwith optional.<build_or_label>) trigger.github/workflows/release.yml. The release tagger simply creates the tag, so publishing is entirely driven by tag pushes (manual or automated). - Need to generate new country patterns for a task? See
AGENTS.mdfor the full enrichment workflow.
Pattern preview workflow
Pull requests that touch solyanka/transaction_patterns/data/** automatically run
.github/workflows/pattern-preview.yml. The workflow uses python -m solyanka.pattern_preview
to diff the branch against the PR base, synthesize up to three example transactions from the
touched patterns, and posts a Markdown table comment (including short/pretty titles) back onto the PR so reviewers can eyeball
the new merchants. Preview-only fixtures live under tests/pattern_preview/ and are injected
via the workflow using the --extra-patterns flag so they stay separate from the shipped data.
If the rendered tables grow beyond GitHub’s comment limit, the workflow automatically splits
the output into sequential comments while keeping each pattern block intact.
Run the same command locally to preview the output before pushing changes:
python -m solyanka.pattern_preview \
--base-ref origin/main \
--head-ref HEAD \
--samples-per-pattern 3 \
--extra-patterns tests/pattern_preview
Conventions: keep YAML human-readable (sorted keys, helpful comments), avoid UUID-looking titles, and update schema/tests whenever the structure changes.
Purpose recap
Solyanka is the single source of truth for transaction-pattern assets used by the bank-statement generator and any LLM training pipelines. Treat it like a dataset project: tight validation, small focused API surface, deterministic releases.
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file solyanka-0.3.3.2.tar.gz.
File metadata
- Download URL: solyanka-0.3.3.2.tar.gz
- Upload date:
- Size: 66.3 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.1.0 CPython/3.13.7
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
5132bc979b34f216770f8b96156fb9c69f06ecdd9e52bba18d3bc4c4d97c04a2
|
|
| MD5 |
b7e7756d9c2a7559e59b9bafb88979ad
|
|
| BLAKE2b-256 |
cab742c2b2e1fc04131a4c8e411a8776d98b41b1300a289df95a1c12ca644ced
|
File details
Details for the file solyanka-0.3.3.2-py3-none-any.whl.
File metadata
- Download URL: solyanka-0.3.3.2-py3-none-any.whl
- Upload date:
- Size: 78.4 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.1.0 CPython/3.13.7
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
5e2b4c0ac10922d324167eddd5926bd202b6a4cbedd6cfebf829f5b75eacf3e0
|
|
| MD5 |
126dcdb6c594f0c64ed260ab14c8086c
|
|
| BLAKE2b-256 |
557e9fab9fedf2696b4b0eeed0266eb8e7ec0cb08dd2666dde688029d29bc1f7
|