Skip to main content

Tokopaedi Async 🚀

High-performance, serverless-safe async Python scraper for Tokopedia.

PyPI PyPI Downloads GitHub Repo stars GitHub forks License: MIT Python Version Tests

tokopaedi-async is an asynchronous fork of Tokopaedi. It talks to Tokopedia's internal GraphQL API through curl_cffi, which reproduces a real iOS TLS/JA3 fingerprint, and pulls many products and reviews at once instead of one at a time.

It is built to run inside a function runtime as well as a long-lived process: importing the package has no side effects, each request owns its HTTP session, and every call has a timeout.

Terminal run of example.py

import asyncio
from tokopaedi_async import browse, SearchFilters

async def main():
    filters = SearchFilters(pmin=15_000_000, pmax=25_000_000, rt=4.5)
    results = await browse(
        "Asus Zenbook S14 32GB",
        max_result=10,
        filters=filters,
        reviews_per_product=20,
        concurrency=8,          # bounded fan-out
    )
    for product in list(results)[:3]:
        print(product.product_id, product.product_name, product.price_text)

asyncio.run(main())

Jupyter usage


Table of contents


Why this fork

The original library is synchronous: one product per call, one request at a time. Enriching 100 products meant 100 sequential round-trips.

This fork makes the network layer concurrent and puts a seam around it:

Original tokopaedi-async
I/O model requests, blocking curl_cffi.AsyncSession, non-blocking
Bulk enrichment sequential loop ("one at a time") asyncio fan-out with a semaphore
Pagination recursion per page one bounded loop
Failure mode raised to the caller logged, returned as data, batch survives
Testing needs the network offline via injected session

The fingerprint handling (Fingerprint-Data, JA3/TLS) is unchanged — it is still what keeps the requests looking like the real iOS app.


Installation

With uv (recommended — fastest, lockfile-aware):

uv add tokopaedi-async

With pip:

pip install tokopaedi-async

With Poetry:

poetry add tokopaedi-async

Requires Python 3.9+. The only runtime dependency is curl-cffi.

To run it once without adding it to a project:

uvx --from tokopaedi-async python -c "import tokopaedi_async; print(tokopaedi_async.__version__)"

Quick start

One call: search and enrich everything

browse() is the convenience path — it searches, then enriches each result in parallel.

import asyncio
from tokopaedi_async import browse, SearchFilters

async def main():
    results = await browse(
        "Asus Zenbook S14 32GB",
        max_result=10,
        filters=SearchFilters(pmin=15_000_000, rt=4.5, shop_tier=2),
        reviews_per_product=20,
        concurrency=8,
    )
    print(f"{len(results)} products")
    print(results.json()[0])          # plain dicts, ready for json.dump / pandas

asyncio.run(main())

Step by step

When you want to inspect the intermediate state — for example to filter cheaply on search results before paying for detail fetches — call the pieces yourself.

import asyncio
from tokopaedi_async import search, SearchFilters

async def main():
    # 1. Search. Cheap: one request per page, ~20 products per page.
    results = await search("Asus Zenbook S14 32GB", max_result=40,
                           filters=SearchFilters(rt=4.5))

    # 2. Narrow before enriching, so you don't fetch details you'll discard.
    keep = [p for p in results if p.rating and p.rating >= 4.7]
    print(f"{len(keep)} of {len(results)} kept")

    # 3. Enrich in parallel. Both calls return a list of failures, not raising.
    detail_failures = await results.enrich_details(concurrency=8)
    review_failures = await results.enrich_reviews(max_result=10, concurrency=8)
    print(f"{len(detail_failures)} detail failures, {len(review_failures)} review failures")

    # 4. Export.
    import json
    with open("result.json", "w", encoding="utf-8") as f:
        json.dump(results.json(), f, indent=2, ensure_ascii=False)

asyncio.run(main())

A single product

import asyncio
from tokopaedi_async import get_product, get_reviews

async def main():
    product = await get_product(product_id=123456789)
    # or: await get_product(url="https://www.tokopedia.com/shop/product-name")

    reviews = await get_reviews(product_id=product.product_id, max_result=50)
    print(product.product_name, product.price_text, len(reviews), "reviews")

asyncio.run(main())

Core concepts

Everything is async. await every network call: search, get_product, get_reviews, browse, and the enrich_* methods. There is no synchronous fallback.

Paginated calls return the full page set, not the first page. search() and get_reviews() keep requesting pages until max_result is reached or the API runs out of results. The pagination loop is bounded (max_pages), so a cursor that never advances cannot hang your process.

Enrichment is bounded and reports failures as data. enrich_details() and enrich_reviews() fan out with an asyncio.Semaphore(concurrency). Each returns a list of (item, exception) pairs for the items that failed — an empty list means everything succeeded. Nothing is printed and nothing is swallowed, so one throttled request cannot discard a batch of 100.

Extraction is separate from transport. Each scraper splits into a pure parser (parse_search, parse_product, parse_reviews) and a network function. The parser is what tests use.

Nothing happens at import time. No logging configuration, no session creation, no environment reads. This is what makes the package safe on a serverless host.


API reference

Full parameter and data-model reference: docs/API.md.

Network calls

Function Returns Notes
await search(keyword, max_result=100, filters=None, debug=False, session=None) SearchResults Paginated, deduped by product id.
await get_product(product_id=None, url=None, debug=False, session=None) ProductData | None One PDP fetch. Returns None if the URL cannot be parsed or the request fails.
await get_reviews(product_id=None, url=None, max_result=10, debug=False, session=None) List[ProductReview] Paginated. Resolves a URL to an id exactly once.
await browse(keyword, max_result=20, filters=None, reviews_per_product=0, concurrency=10, session=None) SearchResults search + enrichment in one call.

session is optional. Pass one to reuse connections in a long-lived process; omit it and each request creates and closes its own session — the correct default on serverless.

Enrichment

failures = await results.enrich_details(concurrency=8)          # -> list[(item, exc)]
failures = await results.enrich_reviews(max_result=10, concurrency=8)
failures = await product.enrich_details()                        # single product
failures = await product.enrich_reviews(max_result=10)

Both are idempotent: a second call is a no-op once the data is attached. has_detail and has_reviews record which enrichment has run.

Pure parsers (no network)

Function Input Output
parse_search(payload) decoded searchProductV5 response List[ProductData]
parse_product(payload) decoded pdpGetLayout response ProductData | None
parse_reviews(payload) decoded review-list response List[ProductReview]

These are the seams to test against. Capture a real response once, store it as a fixture, and assert on the dataclasses.

Data models

ProductData, ProductReview, TokopaediShop, ProductMedia, ProductVariant, ProductOption — all @dataclass, all with .json() returning a plain dict.

SearchResults is the list-like container returned by search() and browse():

len(results)          # number of products
results[0]            # first product
for p in results: ... # iterate
results.json()        # list[dict]
results_a + results_b # concatenate

Search filters

Every SearchFilters field is optional; only the fields you set are sent to the API.

Field Type Meaning
pmin / pmax int Price range in IDR.
condition int 1 = new, 2 = used.
shop_tier int 2 = Mall / Official Store, 3 = Power Shop / Power Merchant.
rt float Minimum average rating, 0.0–5.0.
is_discount bool Discounted products only.
bebas_ongkir_extra bool "Bebas Ongkir Extra" products only.
is_fulfillment bool Fulfilled by Tokopedia only.
is_plus bool Tokopedia Plus products only.
cod bool Cash-on-delivery eligible only.
latest_product int Product age in days: 7, 30, or 90.

Unrecognised shop tiers resolve to None rather than being reported as "Normal", so a missing badge does not become a wrong answer.


Running on serverless (Vercel, Lambda, Cloud Run)

A full working example lives in examples/serverless_app.py. It is a FastAPI app that runs unchanged locally under uvicorn and on Vercel's Python runtime.

uv add 'tokopaedi-async[serverless]'   # or: pip install 'tokopaedi-async[serverless]'
uvicorn examples.serverless_app:app --reload
curl "http://127.0.0.1:8000/search?q=zenbook&limit=5&reviews=10"

Four things make this work, and each is a deliberate property of the library:

1. No import-time side effects. A cold start is just an import. The package does not call logging.basicConfig(), open sockets, or read configuration at import. Verify it yourself:

import logging
before = (logging.getLogger().level, len(logging.getLogger().handlers))
import tokopaedi_async
after = (logging.getLogger().level, len(logging.getLogger().handlers))
assert before == after       # import changed nothing

2. Per-request sessions, closed deterministically. Do not cache a module-level AsyncSession on serverless — the worker can be frozen or recycled between invocations, and a session that outlives its event loop raises on reuse. The default (omit session=) creates and closes one per request, on both the success and failure paths.

# Just omit session= on serverless.
results = await browse(q, max_result=20, concurrency=8)

3. Every request has a timeout. Serverless platforms kill the invocation, not the socket; a hung read would otherwise consume the whole time budget. DEFAULT_TIMEOUT is 20 seconds. Lower it when your platform's limit is tighter.

4. Bound your work per invocation. Clamp max_result and concurrency from user input so one caller cannot request 10 000 products and blow the function's time limit:

MAX_PRODUCTS = 40
MAX_REVIEWS = 20
CONCURRENCY = 8

limit = min(user_limit, MAX_PRODUCTS)
reviews = min(user_reviews, MAX_REVIEWS)
results = await browse(q, max_result=limit, reviews_per_product=reviews, concurrency=CONCURRENCY)

examples/serverless_app.py applies all four and returns plain JSON from results.json().

Vercel deployment — point vercel.json at the ASGI app:

{
  "builds": [{ "src": "examples/serverless_app.py", "use": "@vercel/python" }],
  "routes": [{ "src": "/(.*)", "dest": "examples/serverless_app.py" }]
}

Running on a platform with a short wall-clock limit (Vercel Hobby: 10s) means keeping max_result small or splitting the scrape across requests. The library's job is to make each request bounded and leak-free; the budget is yours to set.


Testing a scraper without the network

The point of the transport seam is that tests never hit Tokopedia. Write one fake session and drive the real pagination, dedupe, and merge logic through it.

import asyncio
import tokopaedi_async as t

class FakeResponse:
    def __init__(self, payload): self._payload = payload
    def json(self): return self._payload
    def raise_for_status(self): pass

class FakeSession:
    def __init__(self, responses): self.responses = responses; self.calls = 0
    async def post(self, url, headers=None, json=None, timeout=None):
        self.calls += 1
        return FakeResponse(self.responses[min(self.calls - 1, len(self.responses) - 1)])
    async def close(self): pass

def test_search_dedupes_across_pages():
    page = {"data": {"searchProductV5": {"header": {"additionalParams": None},
            "data": {"products": [{"id": 1, "name": "A", "url": "u", "mediaURL": {},
                                   "shop": {}, "badge": {}, "price": {}, "stock": {}}]}}}}
    results = asyncio.run(t.search("x", session=FakeSession([page])))
    assert [p.product_id for p in results] == [1]

For pure parser tests, no fake is needed at all — pass a captured payload straight to parse_product / parse_search / parse_reviews.

The repository ships 35 such tests, all offline:

pytest              # 35 passed
pytest --cov=tokopaedi_async

Architecture

Seven small modules, one seam between "build a payload" and "send it".

src/tokopaedi_async/
├── transport.py       ← the seam: one POST, headers, fingerprint, session lifecycle, timeout
├── pagination.py      ← one bounded cursor loop, used by search and get_reviews
├── concurrency.py     ← one bounded fan-out, used by both enrich_* paths
├── search.py          ← search() + parse_search()
├── get_product.py     ← get_product() + parse_product()
├── get_reviews.py     ← get_reviews() + parse_reviews()
├── tokopaedi_types.py ← dataclasses, SearchResults, shop tier mapping
├── get_fingerprint.py ← randomized iOS device fingerprint
└── custom_logging.py  ← opt-in log levels; configures nothing on import

Each scraper is split the same way: a pure parser (fixture JSON → dataclasses) and a network function (build payload → transport.post_graphql → parser). The pure half is unit-testable; the network half is tested by injecting a fake session at the same seam.

If Tokopedia's response shape changes, the change is contained in one parser. If the request shape changes, it is contained in one build_headers/payload pair.


Development

git clone https://github.com/pandamoon21/tokopaedi-async.git
cd tokopaedi-async
uv sync                    # creates .venv from uv.lock, incl. dev deps
uv run pytest
uv run python example.py   # hits the real API

Or with Poetry:

poetry install
poetry run pytest
poetry run python example.py     # hits the real API

The project uses PEP 621 metadata in pyproject.toml, so uv sync, pip install ., and poetry build all read the same source of truth. uv.lock is committed for reproducible dev environments.

Conventions:

  • Async everywhere. All network functions are coroutines; never block the loop.
  • Failures are values. Network functions log and return None/[]/failure lists; they do not raise for a single bad request. Argument errors (ValueError) still raise.
  • Parsers stay pure. No I/O, no globals inside parse_*.
  • No import-time side effects. Adding one breaks serverless safety — the test suite asserts the root logger is untouched after import.

Disclaimer

This is an unofficial library, not affiliated with, endorsed by, or supported by Tokopedia. It is intended for educational and research purposes. Use it responsibly and respect Tokopedia's Terms of Service.

Credits & license

Metadata

Release files for tokopaedi-async 0.2.2

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for tokopaedi-async 0.2.2
File Size Uploaded
tokopaedi_async-0.2.2.tar.gz 31.8 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for tokopaedi-async 0.2.2
File Interpreter ABI Platform
tokopaedi_async-0.2.2-py3-none-any.whl Python 3 none any Details

Total release size: 62.9 kB

Release files / tokopaedi_async-0.2.2.tar.gz

Download URL tokopaedi_async-0.2.2.tar.gz
Size 31.8 kB
Tags Source
SHA-256 checksum
How to use checksums
eb1a8e9d4a76c4d52b494b5163c6bab943a16e789121728c15aaf3f2c9ef6dbc
BLAKE2b-256 checksum
How to use checksums
11af9a280e4b959c07a66b80afa6407ba46bc960159afc0cfd1f32ba909f8cfc
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.12

Release files / tokopaedi_async-0.2.2-py3-none-any.whl

Download URL tokopaedi_async-0.2.2-py3-none-any.whl
Size 31.1 kB
Tags Python 3
SHA-256 checksum
How to use checksums
38f538abe95d9c1c11016c56339c057d3569464be64b85ee951419ed7ec4ef70
BLAKE2b-256 checksum
How to use checksums
b61cd2ed8fce09c76ec74e951add47f77a46cdf12b5ad9924b87a264b973fc2d
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.12

Release history Release notifications | RSS feed

This release

0.2.2 This release

2 release files

0.2.1

2 release files

0.1.2

2 release files

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page