This release is a pre-release and may not be stable for production use.
mockpyspark
A drop-in mock for Apache PySpark 4.x for unit and integration testing. Runs entirely in Python — no JVM, no Hadoop, no cluster — using PyArrow as the compute engine and DuckDB for SQL.
pip install mock-pyspark
mockpyspark (the PyPI distribution is mock-pyspark; you still import mockpyspark) installs alongside real pyspark; there is no namespace collision. Tests opt in via a one-liner so that from pyspark.sql import ... in your production code transparently resolves to the mock.
Quick start
# conftest.py
import mockpyspark
mockpyspark.activate()
# tests/test_my_pipeline.py
from pyspark.sql import SparkSession, functions as F
def test_normalize(spark=None):
spark = SparkSession.builder.getOrCreate()
df = spark.createDataFrame([
{"id": 1, "name": " alice ", "score": 10},
{"id": 2, "name": "BOB", "score": 20},
])
out = (
df.withColumn("name", F.initcap(F.trim(F.col("name"))))
.filter(F.col("score") > 15)
.select("id", "name", "score")
.collect()
)
assert out == [(2, "Bob", 20)]
No changes to your Spark logic. The alias just makes pyspark.* imports resolve to mockpyspark.* for the test process.
Version targeting
mockpyspark supports both PySpark 4.0.x and 4.1.x. APIs that only exist in 4.1 (e.g. current_time, TimeType) raise AttributeError when running in 4.0 mode.
It picks a target version automatically, in order:
PYSPARK_VERSIONenv var- The
pysparkpin inpyproject.toml,requirements*.txt,poetry.lock,uv.lock,Pipfile.lock, orsetup.cfg(walking up from the working directory) SparkSession.builder.config("spark.mock.targetVersion", "4.0").getOrCreate()— per-session override- If nothing is found, defaults to the latest supported version and emits a
UserWarning. SetMOCKPYSPARK_STRICT_VERSION=1to raise instead.
So if your project pins pyspark==4.0.1 in pyproject.toml, the mock will refuse 4.1-only APIs without any configuration on your part.
Usage modes
| Mode | What to do |
|---|---|
| Mock for tests (default) | mockpyspark.activate() in conftest. |
| Pytest plugin | pytest_plugins = ["mockpyspark.pytest_plugin"] in conftest. |
| Real PySpark for a run | Pass --no-mock-pyspark to pytest, or don't call activate(). |
The pytest plugin also accepts --no-mock-pyspark to bypass activation for a single invocation, useful for running the same suite against real PySpark for integration coverage.
See docs/usage.md for the full pattern, including switching between mock and real PySpark in the same suite, and known limitations.
Detecting the mock at runtime
import pyspark
if getattr(pyspark, "__mock__", False):
print("running against mockpyspark")
Contributing
Run unit tests (no Java, no Docker):
git clone https://github.com/PeterDowdy/mock-pyspark.git && cd mock-pyspark
pip install pyarrow duckdb pandas numpy pytest
PYTHONPATH=. pytest -q -m unit
Run integration tests against real PySpark (requires Java 17):
# Via Docker (easiest)
./start.sh 4.1 --integration-test
# Or locally with a PySpark venv
python -m venv /opt/spark-venv && /opt/spark-venv/bin/pip install pyspark==4.1.0
PYTHONPATH=. PYSPARK_VERSION=4.1 PYSPARK_PYTHON=/opt/spark-venv/bin/python pytest -q -m integration
See docs/contributing.md for the full setup guide.
API coverage
Supports PySpark 4.0.x and 4.1.x — 439+ functions, full DataFrame/Column/Window/Catalog API, CSV/JSON/Parquet IO.
- STATUS-4.1.md — PySpark 4.1 coverage
- STATUS-4.0.md — PySpark 4.0 coverage
Release files for mock-pyspark 0.0.1.dev20260816
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| mock_pyspark-0.0.1.dev20260816.tar.gz | 260.2 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| mock_pyspark-0.0.1.dev20260816-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 550.2 kB
Release files / mock_pyspark-0.0.1.dev20260816.tar.gz
| Download URL | mock_pyspark-0.0.1.dev20260816.tar.gz |
|---|---|
| Size | 260.2 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
042f3a080957f31b284642316dddb933ef485a9ea21c319ff6bb031a17563275
|
|
BLAKE2b-256 checksum How to use checksums |
aac3fcc2ee601e1d031ea0ddc03235190e013212a249dbf0b3b99b98d9044065
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Aug 16, 2026.
Transparency logRelease files / mock_pyspark-0.0.1.dev20260816-py3-none-any.whl
| Download URL | mock_pyspark-0.0.1.dev20260816-py3-none-any.whl |
|---|---|
| Size | 290.0 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
65a692058b4b1db0aa785053841e237681083a9121e8e1e20b6f03ec0055c0ee
|
|
BLAKE2b-256 checksum How to use checksums |
aa7424d55b23e8db056d879672e99f6da58e4758361bfded867f2ecadac987a9
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Aug 16, 2026.
Transparency log