praw-cli
A composable Reddit crawler for the command line. Fetch posts and comments from any source, shape the data through a lazy filter pipeline, and emit it in any format — all without writing a single line of Python.
# stream the top 500 posts from r/MachineLearning as JSONL
praw-cli posts r/MachineLearning --sort top --time year --limit 500 --format jsonl
# search across r/programming, keep only high-signal posts, project to four fields
praw-cli search "New web framework" --sub r/programming \
--filter "score>=100" --filter "num_comments>=10" \
--fields id,title,score,url --format csv --output results.csv
# full comment tree of a submission, depth-first, minimum score 5
praw-cli comments https://reddit.com/r/Python/comments/xyz/ \
--depth 10 --min-score 5 --format jsonl
# fetch a user's recent posts as a terminal table
praw-cli user spez --mode posts --limit 20 --format table
# re-process a saved dataset offline — no API, no credentials needed
praw-cli input posts.jsonl --filter "score>=500" --format csv --output filtered.csv
Key Features & How They Solve Your Problems
Smart Memory & Reliability
- Constant Memory Footprint (Lazy Pipelines): Typical scrapers buffer huge arrays in memory, causing out-of-memory crashes on large crawls. praw-cli processes data lazily item-by-item (
Iterator[Record]). Streaming 100,000 posts uses the same constant memory as streaming 10. - Resumable Extractions (Checkpointing): Long-running scrapes often fail halfway due to network drops or API limits. praw-cli checkpoints progress, letting you
--resumeinterrupted sessions without refetching from scratch. - Offline Re-processing (
input): Apply new filters, change output formats, or extract field subsets from any previously saved.jsonl,.json, or.csvfile — without touching the API or needing credentials. Pipe from stdin too.
Zero-Code Data Preparation
- Data Science-Focused DSL: Stop writing custom Python scripts just to filter text. Chain
--filterconditions directly in the CLI:- NLP Cleaning: Keep long-form content using
selftext len>= 500. - Temporal Sorting: Restrict dates natively using ISO-8601 strings, like
created_utc >= 2024-01-01. - Targeted Mining: Target specific topics with keyword groupings (
has,has_all) or regular expressions (title ~= \bbot\b).
- NLP Cleaning: Keep long-form content using
- Field Projections: Keep output datasets lightweight and clean by extracting only the schema columns you need (e.g.,
--fields id,title,score,author).
Instant Pandas & R Integrations
- Diverse Output Formats: Stream outputs directly to
jsonl(preferred for streaming/Pandas), standardjson,csv(for R/Excel), or rich console tables and Markdown reports. - Multi-Sink Pipeline: Write the raw data to a
.jsonldatabase while simultaneously writing a preview to a.csvsummary in a single pass.
Built-in Scientific Rigor
- Crawl Manifests: Every execution automatically generates a
.manifest.jsondetailing exact parameters, versioning, records filtered, and a config fingerprint, preventing configuration drift in research environments. - Deterministic Sampling: Extract reproducible subsets of huge subreddits using systematic or randomized sampling (e.g., Bernoulli trial at
rate = 0.1with a fixed seed).
Installation
Requires Python 3.12 or later and a Reddit API application (free, read-only access is sufficient for most use cases).
Via PyPI (Recommended)
Install the package directly using pip or pipx:
pip install praw-cli
From Source (Using uv)
If you want to run it locally or contribute to development:
git clone https://github.com/othonhugo/praw-cli
cd praw-cli
uv sync
source .venv/bin/activate
Documentation
Full documentation is available in the docs/ directory.
- Getting Started: Installation and Authentication
- Usage Guide:
- Development:
Metadata
Release files for praw-cli 0.7.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| praw_cli-0.7.0.tar.gz | 55.4 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| praw_cli-0.7.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 79.0 kB
Release files / praw_cli-0.7.0.tar.gz
| Download URL | praw_cli-0.7.0.tar.gz |
|---|---|
| Size | 55.4 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
98b30685dac110abb64bafbefc7f6e923bf01b2c4a8d86b25547c262f0c358e3
|
|
BLAKE2b-256 checksum How to use checksums |
0b15fa6d80c2c7daab0071ee14b704356966ec2fc6dcd5c18d30b4e36b662ffe
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/6.1.0 CPython/3.13.12
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Jul 15, 2026.
Transparency logRelease files / praw_cli-0.7.0-py3-none-any.whl
| Download URL | praw_cli-0.7.0-py3-none-any.whl |
|---|---|
| Size | 23.5 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
9ce9220dc967b6fec312a28a45239c9029ea0d11d508c78732d073d7fc99c397
|
|
BLAKE2b-256 checksum How to use checksums |
1f8d949737a98997b9831ecb959aee08de9a08ece57958d2a032ecd53b600e70
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/6.1.0 CPython/3.13.12
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Jul 15, 2026.
Transparency log