Skip to main content

scraper-for-x

Read-only scraping of logged-in X/Twitter data via a harvest-then-replay hybrid: a stealth browser (or a cookie import) logs you in once and harvests the session, then every read afterward is a plain httpx GraphQL request — no browser in the loop, no x-client-transaction-id replay.

Read DISCLAIMER.md before using this. Using this tool violates X's Terms of Service, publishing it exposes its maintainer, and scraping other people's tweets can make you a data controller over their personal data under GDPR. Use a dedicated/throwaway account, not your primary one.

Installation

Base install — cookie-import login only, no browser dependency:

pip install scraper-for-x

The [browser] extra — adds a stealth browser for scrape-x login:

pip install "scraper-for-x[browser]"

If you only ever import cookies from a session you already have (e.g. exported from your own logged-in browser), the base install is all you need.

Quick Start

# 1. One-time interactive login — opens a real browser window, you log in by hand.
scrape-x login

# 2. Fetch a profile's tweets.
scrape-x fetch nasa --limit 50

CLI overview

Command Purpose
scrape-x login One-time login: headed stealth browser by default, or --cookies FILE to import an existing session
scrape-x status Check whether the persisted session is logged in, expired, or rate-limited
scrape-x setup Provision the login browser into an isolated cache (requires [browser])
scrape-x doctor Authenticated round-trip + query-id freshness check (--refresh re-anchors query-ids from x.com's main.js, browser-free)
scrape-x fetch <identifier> A profile's tweets/media (--limit, --since, --until, --by screen_name|id) — --replies not yet implemented, see below
scrape-x search <query> Not yet implemented, see below
scrape-x tweet <identifier> A single tweet plus its reply/conversation thread (--replies)

fetch/search/tweet all share --format json|ndjson, --output PATH, --profile NAME, --profile-dir PATH, --wait-on-limit, --max-wait, --raw (+ --no-redact), and -v/--verbose. See the CLI Reference for every flag and exit code.

Known limitation: search and fetch --replies (v0.1.0)

Live-verified 2026-07-05: X's SearchTimeline and UserTweetsAndReplies GraphQL operations require a fresh, single-use x-client-transaction-id header on every request — unlike plain profile fetches and single-tweet lookups, which this package proved work over plain httpx replay with no such header. A captured transaction-id cannot be harvested once and reused like a session cookie or query-id; reproducing X's generator for it is exactly the fragility this package's harvest-then-replay architecture was built to avoid (see twikit#408).

Both scrape-x search and scrape-x fetch --replies fail fast with a clear FeatureNotImplementedError (exit code 1) rather than a confusing network error. A browser-observe fallback for these two operations specifically (the same approach scrape-fb uses) is on the roadmap — see wiki/FAQ-and-Troubleshooting.md.

Python API

from scraper_for_x import XScraper

XScraper(profile="default").login()  # one-time, opens a headed browser

with XScraper(profile="default") as x:
    tweets = x.fetch_user_tweets("nasa", limit=50)
    for tweet in x.iter_user_tweets("nasa", limit=50):
        ...  # must be consumed inside the `with` block

    x.fetch_tweet("https://x.com/nasa/status/1234567890")
    # x.search(...) raises FeatureNotImplementedError -- see "Known limitation" above.

Documentation

This README covers the essentials. For everything else, see the wiki:

License

MIT — see LICENSE. The license covers the code; it does not cover what you do with the data you collect (see DISCLAIMER.md).

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

scraper_for_x-0.2.0.tar.gz (118.2 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

scraper_for_x-0.2.0-py3-none-any.whl (54.5 kB view details)

Uploaded Python 3

File details

Details for the file scraper_for_x-0.2.0.tar.gz.

File metadata

  • Download URL: scraper_for_x-0.2.0.tar.gz
  • Upload date:
  • Size: 118.2 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.13

File hashes

Hashes for scraper_for_x-0.2.0.tar.gz
Algorithm Hash digest
SHA256 5627bdfd26dda2148a618c2b1e894bed6f9532fb74487c047911102378e62021
MD5 7ffc8b20115e54303c93d61e618b9b15
BLAKE2b-256 6aadd07a1d5ce7dcb6ec167d0806d4678e6c4a211e97d81e8434270cb42b72e5

See more details on using hashes here.

Provenance

The following attestation bundles were made for scraper_for_x-0.2.0.tar.gz:

Publisher: publish.yml on tjdwls101010/Scraper-for-X

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file scraper_for_x-0.2.0-py3-none-any.whl.

File metadata

  • Download URL: scraper_for_x-0.2.0-py3-none-any.whl
  • Upload date:
  • Size: 54.5 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.13

File hashes

Hashes for scraper_for_x-0.2.0-py3-none-any.whl
Algorithm Hash digest
SHA256 f4248259d1aea8123baa8638d42cdb349685db854fccd41e8eb9795a696cc284
MD5 c9af3dfa69e86c56b042ba241049ed4a
BLAKE2b-256 2c27938b6a5e240567254be3f5896ba218166c0c1a08c8a4f52a616e5834a87b

See more details on using hashes here.

Provenance

The following attestation bundles were made for scraper_for_x-0.2.0-py3-none-any.whl:

Publisher: publish.yml on tjdwls101010/Scraper-for-X

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

0.3.2

2 files

0.3.1

2 files

0.3.0

2 files

This release

0.2.0 This release

2 files

0.1.0

2 files

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page