pyXcom
使用前请自行打开 Chrome 并登录 X。pyXcom 不会打开或控制浏览器,也不会替你登录。
Collect public X user profiles, main posts, authored replies and keyword matches using your existing Chrome or Edge session. Outputs are three linked CSV tables: users.csv, posts.csv and comments.csv.
The interface follows PykTok's approachable get/save convention, with explicit parameters and resumable datasets. pyXcom is an independent implementation; it does not depend on PykTok or twikit.
Install
Python 3.10 or newer:
python -m pip install pyXcom
From a cloned repository, use python -m pip install ..
Or in a development environment:
python3 -m venv .venv
.venv/bin/python -m pip install -e .
The distribution name is pyXcom; the import and command are pyxcom.
Quick start
- Open Chrome yourself and sign in to X.
- Find your profile at
chrome://version→ Profile Path. Use its folder name, such asDefaultorProfile 3. - Run:
from pyxcom import XClient
with XClient(profile="Default") as client:
user = client.get_user("thsottiaux")
print(user.name, user.followers_count)
result = client.save_user_activity(
"thsottiaux",
output_dir="output/tibo_year",
since="2025-09-29",
until="2026-09-30",
max_pages=2,
)
print(result.to_dict())
Repeat the same call and output directory to resume. max_pages=2 fetches at most two pages per timeline per call. Remove this option to continue paging until the source ends or the date boundary is reached. A saved partial result reports its stop reason.
Optional constructor settings include browser="edge", cookie_db=..., proxy=..., delay=1.0, and timeout=30. Cookies are read into memory and sent only to X. No passwords or exported cookie files are needed.
Python functions
Methods belong to XClient. Every method lists its accepted parameters explicitly.
| Task | Get into memory | Stream records | Save a dataset |
|---|---|---|---|
| User profile | get_user(handle) |
— | — |
| Post by ID or URL | get_post(post_id_or_url) |
— | — |
| Multiple post IDs | get_posts(post_ids) |
— | — |
| Authored main posts | get_user_posts(handle, ...) |
iter_user_posts(handle, ...) |
save_user_posts(handle, output_dir, ...) |
| Authored replies | get_user_replies(handle, ...) |
iter_user_replies(handle, ...) |
save_user_replies(handle, output_dir, ...) |
| Replies below a main post | get_post_comments(post_id_or_url, ...) |
iter_post_comments(post_id_or_url, ...) |
save_post_comments(post_id_or_url, output_dir, ...) |
| Keyword matches | get_search_posts(keyword, ...) |
iter_search_posts(keyword, ...) |
save_search_posts(keyword, output_dir, ...) |
| Main posts + replies | — | — | save_user_activity(handle, output_dir, ...) |
| Multiple accounts | — | — | save_users_activity(handles, output_dir, ...) |
get_*returns aProfile,Post, orlist[Post]; it does not write a dataset.iter_*returns an iterator. Requests happen as it is consumed.save_*writes the standard dataset and returnsCollectionResultwithoutput_dir,post_count,pages_fetched,complete, andreason.post_countcounts all collected observations, including replies; use the output manifest for separate table counts.- Parameters:
handlefor one account,handlesfor several,keywordfor search,output_dirfor saved datasets, andsince/untilfor dates. Options are keyword-only. - Dates are UTC:
sinceinclusive,untilexclusive.max_pagesandlimitmust be positive integers orNone. Saved timeline/mirrorlimitis a page-boundary stopping threshold, so a final page can exceed it; iterator limits are exact.
See API reference for signatures, return types and compatibility names.
Search one account
with XClient(profile="Default") as client:
matches = client.get_search_posts(
"reset", handle="thsottiaux",
since="2025-09-29", until="2026-09-30", limit=20,
)
result = client.save_search_posts(
"reset", output_dir="output/tibo_reset",
handle="thsottiaux", since="2025-09-29", until="2026-09-30",
)
Direct account search scans main posts and authored replies, then matches a literal word/phrase without case sensitivity. Matching posts do not represent deduplicated real-world events. Cross-user search requires an explicit mirror_base; it sends search terms to that mirror and retrieves discovered post details from X. No mirror is enabled by default.
Collect several accounts
with XClient(profile="Default") as client:
result = client.save_users_activity(
["OpenAI", "AnthropicAI", "thsottiaux", "sama", "alexalbert__", "bcherny"],
output_dir="output/ai_year",
since="2025-09-29", until="2026-09-30",
pages_per_round=5, rounds=1,
)
Repeat to resume. rounds=0, wait_on_rate_limit=True keeps running and waits for rate windows when needed. progress=callback receives status dictionaries.
Collect a main post and nested comments
with XClient(profile="Default") as client:
result = client.save_post_comments(
"1973931546550894681",
output_dir="output/tibo_conversation",
max_depth=2,
max_comments=100,
max_pages=20,
)
print(result.complete, result.reason)
pyxcom post-comments 1973931546550894681 --profile Default \
--max-depth 2 --max-comments 100 --max-pages 20 \
--output-dir output/tibo_conversation
The root belongs in posts.csv, replies from all visible authors in comments.csv, and author profiles in users.csv. The root must be a main post. Level 1 replies directly to it; level 2 replies to those comments. Quoted content, unrelated recommendations and comments beyond the requested depth are excluded. Profile details absent from the response remain unavailable.
max_comments is an exact total comment cap, excluding the root. max_pages caps conversation requests per call. Repeat the same root/depth/output to resume saved branches and cursors; raise the comment cap to continue a capped dataset. Pending page records are retained, so stopping midway through a page does not skip them. Changing depth requires a separate dataset. Python accepts None for unlimited comment/page budgets; depth must be positive.
Stop reasons include comment_limit, page_limit, rate_limited, partial_conversation, and visible_source_end. complete=true means the visible queue was exhausted within the requested depth, not that every comment on X was recovered. Missing ancestors and stalled pagination are reported as partial. iter_post_comments / get_post_comments return reply records without writing a dataset; use save_post_comments when you need coverage status and resumability.
Output
output/tibo_year/
├── users.csv
├── posts.csv
├── comments.csv
├── manifest.json
└── .pyxcom/ # Checkpoints, observations, source streams and reports
| Table | Unit | Main keys |
|---|---|---|
users.csv |
One user | user_id, username, display_name |
posts.csv |
One original or quote main post | post_id, author_id, post_type, quoted_post_id |
comments.csv |
One reply at any observed depth | comment_id, author_id, root_post_id, parent_post_id, depth |
Replies to oneself remain comments. A quote within a reply retains its quote link in comments.csv. Direct replies are depth 1; replies to those are depth 2. Missing ancestry leaves depth empty with a depth_status; parent_in_dataset and root_in_dataset indicate observed relationships. Do not interpret missing parents as first-level replies.
user_replies means replies authored by the selected account. Use *_post_comments to collect replies below a particular main post from all visible authors, including nested replies. Neither endpoint establishes full conversation coverage.
CSV conventions: UTF-8 with BOM, snake_case columns, string IDs (import as text in Excel), JSON arrays for list fields, empty values for unavailable scalars, UTC timestamp fields. Users without a saved profile have profile_available=false. Engagement counts are snapshots; unknown values are not replaced with zero.
If source records include pure reposts, an additional interactions.csv preserves them. The current authored-account collector does not establish complete repost activity. manifest.json records schema/layout versions, table counts, collection scope, relationship coverage and hashes. .pyxcom/ must be retained for resuming; the three CSV files can be shared independently for analysis.
Command line
CLI task names correspond to the Python methods:
pyxcom user thsottiaux --profile Default
pyxcom user-posts thsottiaux --profile Default --output-dir output/tibo_posts
pyxcom user-replies thsottiaux --profile Default --output-dir output/tibo_replies
pyxcom user-activity thsottiaux --profile Default \
--since 2025-09-29 --until 2026-09-30 --output-dir output/tibo_year
pyxcom search-posts reset --handle thsottiaux --profile Default \
--since 2025-09-29 --until 2026-09-30 --output-dir output/tibo_reset
pyxcom users-activity --handles OpenAI AnthropicAI --profile Default \
--since 2025-09-29 --until 2026-09-30 --output-dir output/ai_year
Add --proxy URL if your network requires one. Use pyxcom COMMAND --help for task options. Individual user, post, and raw-post commands accept --output FILE for a JSON file.
Export, migration and verification
These operations work offline, without browser credentials:
from pyxcom import export_tables, validate_collection, finalize_collection
export_tables("output/ai_year")
print(validate_collection("output/ai_year"))
# Rebuild an interrupted dataset from its saved observations:
finalize_collection("output/ai_year")
pyxcom export --output-dir output/ai_year
pyxcom validate --output-dir output/ai_year
pyxcom finalize --output-dir output/ai_year
pyxcom schema --output-dir output/ai_year
Opening/exporting a legacy dataset migrates its recognized internal files into .pyxcom/, with originals backed up under .pyxcom/legacy/. It publishes the three tables at the root; the previous mixed posts.csv is retained internally. Unrelated user files are left in place. Conflicting migration destinations fail rather than overwrite. schema --apply also backfills classification fields in older records; field reports live under .pyxcom/.
Compatibility and limitations
Old Python names get_search, iter_search, save_search (with user=), and save_accounts remain supported. Prefer *_search_posts (with handle=) and save_users_activity for new code. timeline="replies" still works, but the explicit *_user_replies methods are clearer. Timeline outputs exclude same-author context of the other role. Resuming an older timeline applies the same rule and preserves its original observations in an internal compressed backup. Combined activity keeps both roles.
Old CLI names profile, posts, activity, search, batch, --user and dataset --output remain aliases. Existing code reading root posts.csv as a mixed table must adapt: it now contains main posts only; replies are in comments.csv.
X endpoints may change or impose limits. pyXcom discovers current GraphQL query IDs from X's web bundle, while response parsers still need maintenance. A completed run indicates the requested date boundary or visible source end was reached, not proof of all historical content. Deleted, protected or unavailable posts cannot be recovered. Output migration alone does not fetch comments; call save_post_comments to acquire them from X.
License
pyXcom is released under the MIT License. Copyright © 2026 Haocheng Wang.
Citation
Author: Haocheng Wang, Communication University of China.
If you use pyXcom in a paper, thesis, dataset, or other research output, please cite the software and report the version used. GitHub's Cite this repository menu reads the machine-readable CITATION.cff.
Suggested reference
Wang, H. (2026). pyXcom: Structured and auditable X data collection for communication research (Version 0.6.0) [Computer software]. https://github.com/haochengw372-hash/pyXcom
BibTeX
@software{wang2026pyxcom,
author = {Wang, Haocheng},
title = {{pyXcom}: Structured and Auditable X Data Collection for Communication Research},
year = {2026},
version = {0.6.0},
url = {https://github.com/haochengw372-hash/pyXcom}
}
For reproducible reporting, also describe the collection dates, account or keyword scope, comment-depth limits, package version, and coverage/stop reasons recorded in the output manifest. No DOI or published-paper citation is currently assigned; this reference cites the software itself. Citation is appreciated and does not add a condition to the MIT license.
Metadata
Release files for pyXcom 0.6.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| pyxcom-0.6.0.tar.gz | 41.5 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| pyxcom-0.6.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 87.6 kB
Release files / pyxcom-0.6.0.tar.gz
| Download URL | pyxcom-0.6.0.tar.gz |
|---|---|
| Size | 41.5 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
045b3ab3a2e62e0a41390fc870dcda3b05ba17c5d49a092d0e61f4e1544d4daf
|
|
BLAKE2b-256 checksum How to use checksums |
da90f3dadbcc730125c3c764e32c909a139bae1acea585d695305b5e9e77a250
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 30, 2026.
Transparency logRelease files / pyxcom-0.6.0-py3-none-any.whl
| Download URL | pyxcom-0.6.0-py3-none-any.whl |
|---|---|
| Size | 46.0 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
f281e2a61e7b160ee6030e0dc271745e11e6de1e1de971976724a890fb038517
|
|
BLAKE2b-256 checksum How to use checksums |
762a0b3aa5489b906811f49cd9713428d9e9099ae9f067eee8240a5b00a84d65
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 30, 2026.
Transparency log