Skip to main content

secure-log2test

CI CodeQL codecov PyPI Downloads Python License: MIT Last commit OpenSSF Best Practices Tessl

Turn a Kibana or Splunk API log export into an executable pytest suite. Auth headers and secret-looking body fields redacted before they reach the output.

secure-log2test demo

Status: v1.2.0 on PyPI. Stable per semver. Active roadmap, see open issues.

📖 Read the design write-up on Dev.to: privacy constraint, three-layer redaction, the v1.0.0 to v1.0.1 user-feedback story.

Why

You have Kibana logs from staging or production. Each entry is a real request: method, URL, status, duration, headers, body. That's a regression suite waiting to happen. Most teams either ignore it, screenshot interesting failures into Jira, or hand-write pytest cases from log entries one at a time.

I needed a faster path. secure-log2test reads a Kibana JSON export and writes a pytest module you can run and commit. Auth values get replaced with ***REDACTED*** before they ever touch the output, so a generated suite is safe to push to a public repo.

The tool exists because at work I kept doing the same five steps by hand for every production incident: open Kibana, scroll, copy the failing request, paste into a new test, repeat. Five minutes per request times ten requests means an hour gone before any actual debugging starts.

Quickstart

pip install secure-log2test

secure-log2test data/sample_kibana_export.json --output tests_generated.py
pytest tests_generated.py -v

A sample export ships with the repo (data/sample_kibana_export.json), so you can see real output without setting up a Kibana instance first. Grab it from the GitHub repo if you installed from PyPI.

For local development:

git clone https://github.com/golikovichev/secure-log2test
cd secure-log2test
python -m venv .venv && source .venv/bin/activate  # or .venv\Scripts\activate on Windows
pip install -e ".[dev]"
pytest tests/ -v

How it works

Two stages, kept separate.

Parse (core/parser.py). Reads the Kibana JSON and validates each entry through Pydantic v2. Two layers of redaction run before any further processing:

  • A static list of well-known headers (authorization, proxy-authorization, proxy-authenticate, cookie, set-cookie, x-api-key, x-auth-token, x-csrf-token, x-access-token, refresh-token, id-token, x-amz-security-token, authentication, dpop, x-hub-signature, x-hub-signature-256). The last three carry credential material (a DPoP proof JWT, webhook HMAC signatures) whose names the regex below would otherwise miss.
  • A regex pattern (auth|token|secret|key|session|cookie|credential|bearer|password|passwd|pwd) that catches custom header names and body field names project teams invent.

The same logic walks request bodies recursively, so {"password": "..."}, {"client_secret": "..."}, OAuth {"refresh_token": "..."} all get scrubbed at parse time. It also runs over URL query strings, so ?access_token=... or ?api_key=... are redacted while the path and non-sensitive parameters stay intact. Name matching is case-insensitive. Values get replaced with ***REDACTED***. The original input dict is not mutated.

Generate (core/generator.py). Takes the cleaned entries and renders a Jinja2 template (templates/test_module.py.j2) into a pytest module. Each log entry becomes one test_* function. The slug filter turns /api/v1/users/42 into a stable function name.

Options

Flag Required Description
--input Path to a Kibana JSON, Splunk CSV/JSON, or Grafana Loki CSV/JSON export (positional)
--source Input log source: auto (default, detect), kibana, splunk, or loki
--output Output path for generated file
--format Output format: pytest (default), json, or csv
--base-url Base URL prefix for generated requests (pytest only)
--templates Custom templates directory (pytest only)
--redact-marker Replacement string for redacted secrets (default ***REDACTED***)
--assert-config JSON config that adds response-body assertions per endpoint (pytest only)

Pick a marker your downstream pipeline expects:

secure-log2test data/sample_kibana_export.json --redact-marker "[SCRUBBED]"

Response body assertions

By default a generated test checks the status code. Log exports rarely carry the response body, so richer checks are user-declared in a JSON config passed with --assert-config. Each rule targets an endpoint by method and url (matched exactly against the request as it appears in the generated test, url compared after redaction) and adds field-level equality and/or a JSON Schema match:

{
  "rules": [
    {
      "method": "GET",
      "url": "/api/v1/users",
      "expect_fields": {"page": 1, "total": 42},
      "schema": "schemas/users.json"
    }
  ]
}
secure-log2test data/sample_kibana_export.json --assert-config assertions.json
  • expect_fields compares top-level JSON keys for equality.
  • schema points to a JSON Schema file relative to the config; it is read at generation time and inlined into the test, so the emitted suite needs no schema file alongside it at run time.
  • A schema rule makes the generated test import jsonschema; install it (pip install jsonschema) only when you use one.
  • Every check runs behind a content-type guard, so a non-JSON response skips the body assertion instead of failing on response.json().

The split lets you reuse the parser for other formats. If you want to generate Locust scripts, k6 scenarios, or an OpenAPI spec from the same logs, the parser stays. Only the template changes.

Custom redaction rules

The built-in blacklist covers the common credential names. When your team uses its own (an internal tenant token, a bespoke secret field, a national ID field), drop a secure-log2test.toml in the directory you run the tool from:

[redaction]
extra_header_names = ["x-tenant-ref", "x-internal-token"]
extra_field_patterns = ["ssn", "account_number"]
extra_field_paths = ["data.user.note", "items.ref"]
  • extra_header_names are exact header names, matched case-insensitively.
  • extra_field_patterns are regexes matched as a substring against header, body-field, and URL-parameter names, the same surface the built-in matcher covers.
  • extra_field_paths are a.b.c body dict-key paths, for a field whose key is innocuous but whose value is sensitive (a free-text note, an internal id) so the name-based rules miss it. A path matches in full only, a * segment matches any one key (data.*.token), and lists are transparent so items.ref reaches every element. Keys match case-insensitively, like the rest of the redactor; a key that literally contains a dot cannot be targeted.
  • The built-in defaults always stay on; config only adds. A pattern or path that is malformed stops the run with a clear error rather than silently passing secrets through.

No config file means the built-in behaviour is unchanged. On Python 3.10 the file is parsed with tomli; 3.11+ uses the standard-library tomllib.

Sample output

Given this Kibana log entry:

{
  "method": "POST",
  "url": "/api/v1/users",
  "status": 201,
  "headers": {"Authorization": "Bearer abc.xyz", "Content-Type": "application/json"},
  "body": {"name": "Test", "email": "test@example.com"}
}

The generator emits something like:

def test_post_api_v1_users():
    response = requests.post(
        f"{BASE_URL}/api/v1/users",
        headers={"Authorization": "***REDACTED***", "Content-Type": "application/json"},
        json={"name": "Test", "email": "test@example.com"},
    )
    assert response.status_code == 201, (
        f"Expected 201, got {response.status_code}: {response.text[:200]}"
    )

The Authorization value never leaves the parser intact. You set the real token in your environment at run time.

Limitations

What v1.0.1 does not handle yet. Calling them out so the tool stays trustworthy.

  • Input shapes other than Kibana (Elasticsearch hits), Splunk (CSV / JSON), and Grafana Loki Explore (CSV / JSON) search exports.
  • Single-file input. Multi-file batch mode is on the roadmap.
  • Output format: pytest, JSON, or CSV.
  • Nested and repeated JSON fields in body assertions. expect_fields compares top-level keys only (see Response body assertions); deeper paths are not matched yet.
  • JSON-path body-field redaction. Custom header names and field-name patterns now load from secure-log2test.toml (see Custom redaction rules); redacting a specific nested JSON path is still open on #2.
  • OAuth replay. Only static Authorization headers, redacted to a placeholder.
  • Multipart bodies and file uploads.
  • Streaming responses or chunked transfer.

If something on this list blocks you, open an issue.

Roadmap

Version Tracks Adds
Unreleased #1 Response body assertions plus optional schema match (landed, see Response body assertions).
Unreleased #2 Custom redaction rules via config file. Config-driven header names and field-name patterns landed (see Custom redaction rules); JSON-path body redaction still open.

Open the issue tracker for the live picture; two good first issue slots are currently open if you want to jump in.

Tests

pytest tests/ -v

183 tests, covering:

  • Parser unit tests for valid input, malformed input, header redaction, body redaction walker, empty bodies.
  • Edge cases for 5xx responses, missing fields, custom auth header patterns, OAuth refresh tokens in request bodies.
  • CI smoke test that runs the CLI end-to-end on the sample export and parses the generated Python with ast.parse.

CI runs on Python 3.10, 3.11, 3.12, and 3.13 via GitHub Actions.

Security note

The redaction layer catches the well-known auth headers plus anything whose name contains auth, token, secret, key, session, cookie, credential, bearer, password, passwd, or pwd. This works for both header names and JSON body field names. If your team uses something the pattern misses (a truly opaque internal name), add it in a secure-log2test.toml (see Custom redaction rules), or to SENSITIVE_HEADERS in core/parser.py, before generating output. PRs welcome.

Never commit a generated suite that includes real production tokens. The redaction layer is a safety net, not a substitute for review.

Related projects and patterns

Once secure-log2test has produced your replay suite, the next layer of work is usually fixture organisation, shared auth, and the parametrize patterns that scale across hundreds of generated cases. The tessl-labs/pytest-api-testing skill on the Tessl Registry collects those follow-on conventions: httpx AsyncClient setup, conftest.py fixture shape, database isolation, parametrize for edge cases, and auth-flow handling. Useful reference when the generated suite starts growing its own test infrastructure.

Sister projects in the same workspace:

Contributing

Issue templates and PR guidance live in CONTRIBUTING.md. Bug reports with a redacted sample log are the most useful kind.

Licence

MIT. See LICENSE.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

secure_log2test-1.3.0.tar.gz (48.5 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

secure_log2test-1.3.0-py3-none-any.whl (28.1 kB view details)

Uploaded Python 3

File details

Details for the file secure_log2test-1.3.0.tar.gz.

File metadata

  • Download URL: secure_log2test-1.3.0.tar.gz
  • Upload date:
  • Size: 48.5 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.13

File hashes

Hashes for secure_log2test-1.3.0.tar.gz
Algorithm Hash digest
SHA256 46c7fb0da18867497d619d9be0253a97044fe7749741b8b58c6d80b412d9d1e0
MD5 2196514cb42e11fe8dbf43642138f49d
BLAKE2b-256 f5934238f41078dfd3b2f744429228be527fe48492d09e532e19c044503a307d

See more details on using hashes here.

Provenance

The following attestation bundles were made for secure_log2test-1.3.0.tar.gz:

Publisher: publish.yml on golikovichev/secure-log2test

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file secure_log2test-1.3.0-py3-none-any.whl.

File metadata

  • Download URL: secure_log2test-1.3.0-py3-none-any.whl
  • Upload date:
  • Size: 28.1 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.13

File hashes

Hashes for secure_log2test-1.3.0-py3-none-any.whl
Algorithm Hash digest
SHA256 0db12c47bac76ca6e8919024df245a79184d49ca9fb8bafeb3d0e44ab0a72105
MD5 6f5108ec2d822aa6ec1b815b5e1aa940
BLAKE2b-256 9d60ea422728547945bda0d83fddec51ad2476f643e1f6cb74d016594be5815c

See more details on using hashes here.

Provenance

The following attestation bundles were made for secure_log2test-1.3.0-py3-none-any.whl:

Publisher: publish.yml on golikovichev/secure-log2test

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page