Skip to main content

scraper-for-x

scraper-for-x

Read-only scraping of logged-in X/Twitter data via a harvest-then-replay hybrid: a stealth browser (or a cookie import) logs you in once and harvests the session, then every read afterward is a plain httpx GraphQL request — no browser in the loop.

Reads your home feed, any profile's tweets and replies, a tweet's thread, search, and the social graph (following / followers / retweeters). Each command is a single-target primitive; chaining them into multi-hop exploration is a caller's job, not the CLI's.

Read DISCLAIMER.md before using this. Using this tool violates X's Terms of Service, publishing it exposes its maintainer, and scraping other people's tweets can make you a data controller over their personal data under GDPR. Use a dedicated/throwaway account, not your primary one.

Installation

Base install — cookie-import login only, no browser dependency:

pip install scraper-for-x

The [browser] extra — adds a stealth browser for scrape-x login:

pip install "scraper-for-x[browser]"

If you only ever import cookies from a session you already have (e.g. exported from your own logged-in browser), the base install is all you need.

Quick Start

# 1. One-time interactive login — opens a real browser window, you log in by hand.
scrape-x login

# 2. Fetch a profile's tweets.
scrape-x fetch nasa --limit 50

CLI overview

Command Purpose
scrape-x login One-time login: headed stealth browser by default, or --cookies FILE to import an existing session
scrape-x status Check whether the persisted session is logged in, expired, or rate-limited
scrape-x setup Provision the login browser into an isolated cache (requires [browser])
scrape-x doctor Authenticated round-trip + query-id freshness check (--refresh re-anchors query-ids from x.com's main.js, browser-free)
scrape-x catalog Machine-readable JSON description of every command, argument and exit code (offline)
scrape-x schema The output object schema; --json emits JSON Schema (offline)
scrape-x feed Your home feed — takes no target, the feed belongs to the session
scrape-x fetch <identifier> A profile's tweets/media (--limit, --since, --until, --by screen_name|id, --replies†)
scrape-x search <query> Tweets matching a query or advanced operators (--product latest|top)†
scrape-x tweet <identifier> A single tweet plus its reply/conversation thread (--replies)
scrape-x following <identifier> Accounts a user follows — emits User objects
scrape-x followers <identifier> Accounts following a user — emits User objects†
scrape-x retweeters <tweet> Accounts that retweeted a tweet — emits User objects

† Needs a generated x-client-transaction-id — see below.

All read commands share --format json|ndjson, --output PATH, --profile NAME, --profile-dir PATH, --wait-on-limit, --max-wait, --raw (+ --no-redact), and -v/--verbose. See the CLI Reference for every flag and exit code — or just run scrape-x catalog for the same thing as JSON.

The transaction-id wall, and how this package gets past it

Three operations — SearchTimeline (search), UserTweetsAndReplies (fetch --replies) and Followers — reject any request without a fresh x-client-transaction-id header. The header is single-use: replaying even a real captured one fails, because X's own client already spent it. So it cannot be harvested once and reused the way session cookies and query-ids are.

Since v0.3.0 this package generates that header per request, in pure Python, from ingredients x.com serves on its own home page (algorithm ported from the MIT-licensed XClientTransaction). Verified working 2026-07-20.

Be honest with yourself about this: it is reverse-engineered, and X can invalidate it with any client deploy. It is the one part of this package where "it worked yesterday" is not evidence it works today. When it breaks, those three commands exit 4 with a clear message; everything else keeps working. search and fetch --replies additionally fall back to driving the stealth browser and reading the response X's own client receives (requires the [browser] extra) — that returns only the first page, and says so.

What X no longer exposes

likers is not implemented and will not be: X has removed the likers list entirely. /status/<id>/likes redirects to the tweet, and the operation name appears in none of the JavaScript chunks x.com serves today (checked 2026-07-20, all 685 of them). For quoters, use scrape-x search "quoted_tweet_id:<id>" — that is exactly what X's own /quotes tab does.

Python API

from scraper_for_x import XScraper

XScraper(profile="default").login()  # one-time, opens a headed browser

with XScraper(profile="default") as x:
    tweets = x.fetch_user_tweets("nasa", limit=50)
    for tweet in x.iter_user_tweets("nasa", limit=50):
        ...  # must be consumed inside the `with` block

    x.fetch_tweet("https://x.com/nasa/status/1234567890")
    x.fetch_home(limit=20)                  # your home feed
    x.search("artemis", product="Latest")   # needs a generated transaction id
    x.fetch_user_tweets("nasa", replies=True, limit=50)

    # Social graph -- these return `User`, not `Tweet`.
    x.fetch_following("nasa", limit=100)
    x.fetch_followers("nasa", limit=100)    # needs a generated transaction id
    x.fetch_retweeters("1234567890", limit=100)

Documentation

This README covers the essentials. For everything else, see the wiki:

License

MIT — see LICENSE. The license covers the code; it does not cover what you do with the data you collect (see DISCLAIMER.md).

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

scraper_for_x-0.3.1.tar.gz (167.4 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

scraper_for_x-0.3.1-py3-none-any.whl (72.3 kB view details)

Uploaded Python 3

File details

Details for the file scraper_for_x-0.3.1.tar.gz.

File metadata

  • Download URL: scraper_for_x-0.3.1.tar.gz
  • Upload date:
  • Size: 167.4 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.13

File hashes

Hashes for scraper_for_x-0.3.1.tar.gz
Algorithm Hash digest
SHA256 5847afa93dd0269f6e4b55beb064bd02360fbea964ef5b619becea2bbb1522ee
MD5 6c90919944f987648f3cb9e1c164563d
BLAKE2b-256 545ae8d4ce43105697ca6ccae812001f127ca65d5130dfae5d7a633592f8bf47

See more details on using hashes here.

Provenance

The following attestation bundles were made for scraper_for_x-0.3.1.tar.gz:

Publisher: publish.yml on tjdwls101010/Scraper-for-X

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file scraper_for_x-0.3.1-py3-none-any.whl.

File metadata

  • Download URL: scraper_for_x-0.3.1-py3-none-any.whl
  • Upload date:
  • Size: 72.3 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.13

File hashes

Hashes for scraper_for_x-0.3.1-py3-none-any.whl
Algorithm Hash digest
SHA256 2488f2548a737ef0898a08515a5dc65bc8df24fc164312a3557f57adfa3fbc39
MD5 f2840559f4122662dd5571ca7bca8c79
BLAKE2b-256 10f694f68bf986cbaba3aa150b4b98477b62f0c54cab9767a0e167d3a7e09467

See more details on using hashes here.

Provenance

The following attestation bundles were made for scraper_for_x-0.3.1-py3-none-any.whl:

Publisher: publish.yml on tjdwls101010/Scraper-for-X

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

0.3.2

2 files

This release

0.3.1 This release

2 files

0.3.0

2 files

0.2.0

2 files

0.1.0

2 files

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page