Skip to main content

praw-cli

Python 3.11+ Code style: ruff

A composable Reddit crawler for the command line. Fetch posts and comments from any source, shape the data through a lazy filter pipeline, and emit it in any format — all without writing a single line of Python.

# stream the top 500 posts from r/MachineLearning as JSONL
praw-cli posts r/MachineLearning --sort top --time year --limit 500 --format jsonl

# search across r/programming, keep only high-signal posts, project to four fields
praw-cli search "New web framework" --sub r/programming \
  --filter "score>=100" --filter "num_comments>=10" \
  --fields id,title,score,url --format csv --output results.csv

# full comment tree of a submission, depth-first, minimum score 5
praw-cli comments https://reddit.com/r/Python/comments/xyz/ \
  --depth 10 --min-score 5 --format jsonl

# fetch a user's recent posts as a terminal table
praw-cli user spez --mode posts --limit 20 --format table

# re-process a saved dataset offline — no API, no credentials needed
praw-cli input posts.jsonl --filter "score>=500" --format csv --output filtered.csv

Key Features & How They Solve Your Problems

Smart Memory & Reliability

  • Constant Memory Footprint (Lazy Pipelines): Typical scrapers buffer huge arrays in memory, causing out-of-memory crashes on large crawls. praw-cli processes data lazily item-by-item (Iterator[Record]). Streaming 100,000 posts uses the same constant memory as streaming 10.
  • Resumable Extractions (Checkpointing): Long-running scrapes often fail halfway due to network drops or API limits. praw-cli checkpoints progress, letting you --resume interrupted sessions without refetching from scratch.
  • Offline Re-processing (input): Apply new filters, change output formats, or extract field subsets from any previously saved .jsonl, .json, or .csv file — without touching the API or needing credentials. Pipe from stdin too.

Zero-Code Data Preparation

  • Data Science-Focused DSL: Stop writing custom Python scripts just to filter text. Chain --filter conditions directly in the CLI:
    • NLP Cleaning: Keep long-form content using selftext len>= 500.
    • Temporal Sorting: Restrict dates natively using ISO-8601 strings, like created_utc >= 2024-01-01.
    • Targeted Mining: Target specific topics with keyword groupings (has, has_all) or regular expressions (title ~= \bbot\b).
  • Field Projections: Keep output datasets lightweight and clean by extracting only the schema columns you need (e.g., --fields id,title,score,author).

Instant Pandas & R Integrations

  • Diverse Output Formats: Stream outputs directly to jsonl (preferred for streaming/Pandas), standard json, csv (for R/Excel), or rich console tables and Markdown reports.
  • Multi-Sink Pipeline: Write the raw data to a .jsonl database while simultaneously writing a preview to a .csv summary in a single pass.

Built-in Scientific Rigor

  • Crawl Manifests: Every execution automatically generates a .manifest.json detailing exact parameters, versioning, records filtered, and a config fingerprint, preventing configuration drift in research environments.
  • Deterministic Sampling: Extract reproducible subsets of huge subreddits using systematic or randomized sampling (e.g., Bernoulli trial at rate = 0.1 with a fixed seed).

Installation

Requires Python 3.12 or later and a Reddit API application (free, read-only access is sufficient for most use cases).

Install the package directly using pip or pipx:

pip install praw-cli

From Source (Using uv)

If you want to run it locally or contribute to development:

git clone https://github.com/othonhugo/praw-cli
cd praw-cli
uv sync
source .venv/bin/activate

Documentation

Full documentation is available in the docs/ directory.

Metadata

Release files for praw-cli 0.7.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for praw-cli 0.7.0
File Size Uploaded
praw_cli-0.7.0.tar.gz 55.4 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for praw-cli 0.7.0
File Interpreter ABI Platform
praw_cli-0.7.0-py3-none-any.whl Python 3 none any Details

Total release size: 79.0 kB

Release files / praw_cli-0.7.0.tar.gz

Download URL praw_cli-0.7.0.tar.gz
Size 55.4 kB
Tags Source
SHA-256 checksum
How to use checksums
98b30685dac110abb64bafbefc7f6e923bf01b2c4a8d86b25547c262f0c358e3
BLAKE2b-256 checksum
How to use checksums
0b15fa6d80c2c7daab0071ee14b704356966ec2fc6dcd5c18d30b4e36b662ffe
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.12

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Jul 15, 2026.

Transparency log

Release files / praw_cli-0.7.0-py3-none-any.whl

Download URL praw_cli-0.7.0-py3-none-any.whl
Size 23.5 kB
Tags Python 3
SHA-256 checksum
How to use checksums
9ce9220dc967b6fec312a28a45239c9029ea0d11d508c78732d073d7fc99c397
BLAKE2b-256 checksum
How to use checksums
1f8d949737a98997b9831ecb959aee08de9a08ece57958d2a032ecd53b600e70
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.12

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Jul 15, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.7.0 This release

2 release files

0.6.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page