Skip to main content

scraper-for-x

scraper-for-x

Read-only scraping of logged-in X/Twitter data via a harvest-then-replay hybrid: a stealth browser (or a cookie import) logs you in once and harvests the session, then every read afterward is a plain httpx GraphQL request — no browser in the loop.

Reads your home feed, any profile's tweets and replies, a tweet's thread, search, and the social graph (following / followers / retweeters). Each command is a single-target primitive; chaining them into multi-hop exploration is a caller's job, not the CLI's.

Read DISCLAIMER.md before using this. Using this tool violates X's Terms of Service, publishing it exposes its maintainer, and scraping other people's tweets can make you a data controller over their personal data under GDPR. Use a dedicated/throwaway account, not your primary one.

Installation

Base install — cookie-import login only, no browser dependency:

pip install scraper-for-x

The [browser] extra — adds a stealth browser for scrape-x login:

pip install "scraper-for-x[browser]"

If you only ever import cookies from a session you already have (e.g. exported from your own logged-in browser), the base install is all you need.

Quick Start

# 1. One-time interactive login — opens a real browser window, you log in by hand.
scrape-x login

# 2. Fetch a profile's tweets.
scrape-x fetch nasa --limit 50

CLI overview

Command Purpose
scrape-x login One-time login: headed stealth browser by default, or --cookies FILE to import an existing session
scrape-x status Check whether the persisted session is logged in, expired, or rate-limited
scrape-x setup Provision the login browser into an isolated cache (requires [browser])
scrape-x doctor Authenticated round-trip + query-id freshness check (--refresh re-anchors query-ids from x.com's main.js, browser-free)
scrape-x catalog Machine-readable JSON description of every command, argument and exit code (offline)
scrape-x schema The output object schema; --json emits JSON Schema (offline)
scrape-x feed Your home feed — takes no target, the feed belongs to the session
scrape-x fetch <identifier> A profile's tweets/media (--limit, --since, --until, --by screen_name|id, --replies†)
scrape-x search <query> Tweets matching a query or advanced operators (--product latest|top)†
scrape-x tweet <identifier> A single tweet plus its reply/conversation thread (--replies)
scrape-x following <identifier> Accounts a user follows — emits User objects
scrape-x followers <identifier> Accounts following a user — emits User objects†
scrape-x retweeters <tweet> Accounts that retweeted a tweet — emits User objects

† Needs a generated x-client-transaction-id — see below.

All read commands share --format json|ndjson, --output PATH, --profile NAME, --profile-dir PATH, --wait-on-limit, --max-wait, --raw (+ --no-redact), and -v/--verbose. See the CLI Reference for every flag and exit code — or just run scrape-x catalog for the same thing as JSON.

The transaction-id wall, and how this package gets past it

Three operations — SearchTimeline (search), UserTweetsAndReplies (fetch --replies) and Followers — reject any request without a fresh x-client-transaction-id header. The header is single-use: replaying even a real captured one fails, because X's own client already spent it. So it cannot be harvested once and reused the way session cookies and query-ids are.

Since v0.3.0 this package generates that header per request, in pure Python, from ingredients x.com serves on its own home page (algorithm ported from the MIT-licensed XClientTransaction). Verified working 2026-07-20.

Be honest with yourself about this: it is reverse-engineered, and X can invalidate it with any client deploy. It is the one part of this package where "it worked yesterday" is not evidence it works today. When it breaks, those three commands exit 4 with a clear message; everything else keeps working. search and fetch --replies additionally fall back to driving the stealth browser and reading the response X's own client receives (requires the [browser] extra) — that returns only the first page, and says so.

What X no longer exposes

likers is not implemented and will not be: X has removed the likers list entirely. /status/<id>/likes redirects to the tweet, and the operation name appears in none of the JavaScript chunks x.com serves today (checked 2026-07-20, all 685 of them). For quoters, use scrape-x search "quoted_tweet_id:<id>" — that is exactly what X's own /quotes tab does.

Python API

from scraper_for_x import XScraper

XScraper(profile="default").login()  # one-time, opens a headed browser

with XScraper(profile="default") as x:
    tweets = x.fetch_user_tweets("nasa", limit=50)
    for tweet in x.iter_user_tweets("nasa", limit=50):
        ...  # must be consumed inside the `with` block

    x.fetch_tweet("https://x.com/nasa/status/1234567890")
    x.fetch_home(limit=20)                  # your home feed
    x.search("artemis", product="Latest")   # needs a generated transaction id
    x.fetch_user_tweets("nasa", replies=True, limit=50)

    # Social graph -- these return `User`, not `Tweet`.
    x.fetch_following("nasa", limit=100)
    x.fetch_followers("nasa", limit=100)    # needs a generated transaction id
    x.fetch_retweeters("1234567890", limit=100)

Documentation

This README covers the essentials. For everything else, see the wiki:

License

MIT — see LICENSE. The license covers the code; it does not cover what you do with the data you collect (see DISCLAIMER.md).

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

scraper_for_x-0.3.0.tar.gz (158.3 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

scraper_for_x-0.3.0-py3-none-any.whl (72.1 kB view details)

Uploaded Python 3

File details

Details for the file scraper_for_x-0.3.0.tar.gz.

File metadata

  • Download URL: scraper_for_x-0.3.0.tar.gz
  • Upload date:
  • Size: 158.3 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.13

File hashes

Hashes for scraper_for_x-0.3.0.tar.gz
Algorithm Hash digest
SHA256 d2c5a3c626f9da4fe1b2ace267fbb9b3d00e0c0a9a4a873b28e80a26b8543466
MD5 b926613371584bd7527773048e3b2782
BLAKE2b-256 32614cc0447b8f8b496d5951655d6866b9ca84e48fc372d94a8c7dd414d2b004

See more details on using hashes here.

Provenance

The following attestation bundles were made for scraper_for_x-0.3.0.tar.gz:

Publisher: publish.yml on tjdwls101010/Scraper-for-X

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file scraper_for_x-0.3.0-py3-none-any.whl.

File metadata

  • Download URL: scraper_for_x-0.3.0-py3-none-any.whl
  • Upload date:
  • Size: 72.1 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.13

File hashes

Hashes for scraper_for_x-0.3.0-py3-none-any.whl
Algorithm Hash digest
SHA256 046d4e0caf7761832c97f4d9a03f6b7347863e0e5515aeb45adbfbca46a7b115
MD5 4695e1bf98952b00b68b787ca610e0c9
BLAKE2b-256 6f2ad6f0ec6cffddbadc59f5cb0a4641a627ff0b3756614a9280a3b6b9805f35

See more details on using hashes here.

Provenance

The following attestation bundles were made for scraper_for_x-0.3.0-py3-none-any.whl:

Publisher: publish.yml on tjdwls101010/Scraper-for-X

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

0.3.2

2 files

0.3.1

2 files

This release

0.3.0 This release

2 files

0.2.0

2 files

0.1.0

2 files

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page