Skip to main content

pddiktiaw

PyPI PyPI Downloads GitHub Repo stars GitHub forks License: MIT Python Version

Unofficial Python wrapper for the PDDIKTI public API — Indonesia's national higher-education database (Pangkalan Data Pendidikan Tinggi): students, lecturers, institutions and study programmes.

FastAPI server (Swagger UI), async library, and a one-shot CLI — in one file. No API key, no account, no configuration.

⚠️ Unofficial. Not affiliated with Kementerian Pendidikan Tinggi, Sains dan Teknologi. Public government data; use it for verification and research, not for building profiles of individuals.

pip install -r requirements.txt
uvicorn pddiktiaw:app --port 8000     # → http://localhost:8000/docs

✨ Why pddiktiaw?

Web scrapers pddiktiaw
Data source HTML scraping of the PDDIKTI console The console's own JSON API
Auth n/a None exists — no key, no account
Response HTML you then parse Clean JSON records
Speed Server-rendered page per request ~1 ms cached, ~0.6 s cold
Rate discipline Up to you Throttled to the upstream's own 60/min

The PDDIKTI web console is a React SPA talking to a JSON API on the same host. pddiktiaw calls that API directly and hands you the records.


📊 Real-world benchmark (measured, not estimated)

Live hits from one machine, median of 3 runs (2026-09):

Test Latency Payload
GET /api/v1/pt?q=univ (cached) ~1 ms 21 institutions · 3 KB JSON
GET /api/v1/pt?q=univ (cold — hits upstream) ~0.6 s "
GET /api/v1/search?q=universitas&kind=all&pages=2 ~2.5 s up to 800 records

scripts/transport_bench.py measures the four candidate HTTP stacks against the live host. Your numbers will vary with network; the order of magnitude is the point.


🤖 AI-agent instructions

CLAUDE.md is a purpose-built instruction file for AI coding agents (and new developers): how to install, verify, extend and deploy pddiktiaw without tripping on the throttle, the sampling behaviour, or the route quirks.

  • File: CLAUDE.md
  • Covers: commands, the three-layer architecture in pddiktiaw.py, the measured upstream facts that look like bugs but aren't, the cache-key rule, and where to add a new endpoint/transport.

Point any agent (or your future self) at CLAUDE.md first.


🚀 Quickstart

Option 1 — pip:

pip install pddiktiaw
pddiktiaw                            # serves http://0.0.0.0:8000 (Swagger at /docs)
pddiktiaw self                       # offline self-check
pddiktiaw pt "universitas"           # one-shot JSON commands (see below)

Option 2 — from source:

git clone https://github.com/pandamoon21/PDDIKTI-API-Wrapper && cd PDDIKTI-API-Wrapper
python3 -m venv .venv && . .venv/bin/activate
pip install -r requirements.txt
uvicorn pddiktiaw:app --port 8000

Option 3 — CLI / module (no install):

python3 pddiktiaw.py serve          # same server (honours $PORT)
python3 pddiktiaw.py self           # offline self-check, no network

Verify:

curl http://localhost:8000/api/v1/health
# {"status":"ok","version":"1.1.0","upstream":"ok","transport":"httpx"}

curl "http://localhost:8000/api/v1/pt?q=universitas&limit=5"

Interactive docs (Swagger UI): http://localhost:8000/docs


🖥 CLI

pddiktiaw is also a one-shot CLI — no server needed.

python3 pddiktiaw.py --help

python3 pddiktiaw.py mhs "budi"              # students (default command)
python3 pddiktiaw.py mahasiswa "budi"        # alias
python3 pddiktiaw.py dosen "teknik informatika"
python3 pddiktiaw.py pt "universitas"        # institutions
python3 pddiktiaw.py prodi "informatika"     # study programmes
python3 pddiktiaw.py all "universitas"       # every kind at once

python3 pddiktiaw.py search "universitas" --kind pt --pages 3 --limit 50
python3 pddiktiaw.py detail <id>             # full mahasiswa record
python3 pddiktiaw.py stats                   # cache hit/miss counters
python3 pddiktiaw.py status                  # upstream auth + caveats
python3 pddiktiaw.py dashboard               # version, uptime, cache, routes
python3 pddiktiaw.py self                    # offline self-check
python3 pddiktiaw.py serve                   # start the HTTP server
$ python3 pddiktiaw.py pt univ --limit 3
3 record(s) for 'univ' (kind=pt)
  UNIVERSITAS SURYA                            041056           UNIVERSITAS SURYA
  UNIVERSITAS ROYAL                            011069           UNIVERSITAS ROYAL
  UNIVERSITAS SENIOR MEDAN                     011073           UNIVERSITAS SENIOR MEDAN

Errors print a clean error: ... message to stderr with exit code 1 — no tracebacks. Add --json to any search command for raw output.


🐍 Use as a Python package

All methods are async and share the same cache, throttle and transport as the HTTP API:

import asyncio
from pddiktiaw import PDDIKTI

async def main():
    api = PDDIKTI()

    hits = await api.pt("universitas", limit=5)        # institutions
    rec  = await api.detail(hits["records"][0]["id"])  # full mahasiswa record

    # Addressable lookups — no detail route exists upstream for these, so they
    # resolve a query to one record instead.
    pt    = await api.pt_by_code("041056")             # by PDDIKTI code
    dosen = await api.dosen_by_nidn("9905536334")      # by NIDN
    pr    = await api.prodi_by_name("INFORMATIKA", pt="UNIVERSITAS SURYA")

    print(api.stats())                                 # cache counters
    await api.get("/api/pencarian/pt/univ")            # raw passthrough
    await api.close()

asyncio.run(main())

Every PDDIKTI method: search(), mahasiswa(), dosen(), pt(), prodi(), search_all_kinds(), detail(), pt_by_code(), prodi_by_name(), dosen_by_nidn(), get(), post(), stats(), close().

Prefer await api.get(path, ttl=…) / post(path, body, ttl=…) for ad-hoc calls to routes that aren't wrapped yet — they go through the same throttle + cache.


🎯 Search semantics — sampling, not pagination

This is the single most important thing to know, and it shapes the whole API.

The upstream ignores pagination. page, limit, offset and size are accepted and all return the same thing: up to 100 records per call. There is no real paging.

Those 100 records are a random sample, re-drawn on every request. Repeating an identical URL returns a different 100 ids. A single call is a sample, not a page one.

So pddiktiaw gives you pages instead:

await api.search("universitas", kind="pt", pages=1)   # one sample, ≤100
await api.search("universitas", kind="pt", pages=5)   # union of 5 samples

pages=N repeats the call N times and unions the samples by id. Measured: kind=all&pages=2 yielded 800 records (4 categories × 100 × 2, with overlap collapsed); kind=pt&pages=3 yielded 300. That is the only way to see past ~100.

Two details make it work, and both are load-bearing:

  • A multi-page search bypasses the response cache while it expands. Serving a cached sample would hand back the same 100 ids and the union would never grow.
  • Only the all/mhs/pt/prodi envelope kinds share one upstream route; each page widens all four categories at once, so a page counts as "exhausted" only when the union grew by nothing. dosen has its own flat route and stops as soon as it returns fewer than 100 rows.
Want Do
A quick look pages=1
Broad coverage of a common term pages=5–10
Exhaustive pages=20 (the API cap) and accept ~1 s per page

Politeness is built in and non-negotiable: 1 request/second, matching the upstream's own ratelimit-limit: 60 per 60 s. pages=20 therefore costs ~20 s.

Kinds and routes

kind Upstream route Envelope key
mhs (alias mahasiswa) enc/all mahasiswa
dosen enc/dosen (flat list)
pt enc/all pt
prodi enc/all prodi
all enc/all all four merged, each row tagged _kind

Why mhs/pt/prodi all use enc/all: measured, the enc path's middle segment only matters for dosen — enc/all, enc/mhs, enc/pt, enc/prodi and even enc/bogus all return the identical 4-category envelope. route_for() exploits that; see scripts/pencarian_probe.py to re-verify it live.


📡 Endpoints

All responses are the raw upstream JSON (/api/v1/detail/{id} unwraps data). Rate limits are per-client, enforced server-side.

Method Route Upstream Cache TTL
GET / — (local) —
GET /api/v1/health — (live probe) local
GET /api/v1/status — (local) —
GET /api/v1/dashboard — (local) —
GET /api/v1/cache/stats — (local) —
GET /api/v1/search?q=&kind=&pages=&limit= enc/<kind> 5m
GET /api/v1/mahasiswa?q= enc/all → mahasiswa 5m
GET /api/v1/dosen?q= enc/dosen 5m
GET /api/v1/pt?q= enc/all → pt 5m
GET /api/v1/prodi?q= enc/all → prodi 5m
GET /api/v1/detail/{id} POST /api/detail/mhs 15m

Full schema at /docs. kind is one of all · mhs · dosen · pt · prodi.

Example responses (captured live)

GET /api/v1/pt?q=univ&limit=2 → 200 · Cache-Control: public, max-age=300

{
  "query": "univ",
  "kind": "pt",
  "count": 2,
  "pages": 1,
  "records": [
    { "id": "Abtmnf-sQB5hELv…==", "kode": "041056",
      "nama_singkat": "UNIV SURYA", "nama": "UNIVERSITAS SURYA" }
  ]
}

GET /api/v1/dosen?q=teknik informatika&limit=1 → 200 (flat lecturer list)

{
  "query": "teknik informatika", "kind": "dosen", "count": 1, "pages": 1,
  "records": [
    { "id": "e2Nwi5EYQJg2oSYz…==", "nama": "SUPRAPTO", "nidn": "9905536334",
      "nuptk": "", "nama_pt": "AKADEMI TEKNIK PIRI",
      "sinkatan_pt": "", "nama_prodi": "TEKNIK INFORMATIKA" }
  ]
}

GET /api/v1/search?q=universitas&kind=all&limit=1 → 200 (each row tagged)

{
  "query": "universitas", "kind": "all", "count": 1, "pages": 1,
  "records": [
    { "id": "L-eBRygtawo3FDUw…==", "nama": "MUHAMMAD UNI", "nim": "0515010042",
      "nama_pt": "UNIVERSITAS SERAMBI MEKKAH", "nama_prodi": "MANAJEMEN",
      "_kind": "mahasiswa" }
  ]
}

GET /api/v1/detail/{id} → 200 · Cache-Control: public, max-age=900

{
  "id": "…", "nim": "…", "nama": "…", "jenjang": "S1", "prodi": "…",
  "nama_pt": "…", "kode_pt": "…", "id_pt": "…", "id_sms": "…",
  "jenis_daftar": "…", "jenis_kelamin": "…", "kode_prodi": "…"
}

Errors

Case Status Body
Upstream 4xx same {"error": true, "code": <status>, "detail": …}
Cloudflare challenge 403 detail names curl_cffi as the fix
Upstream timeout 504 {"error": true, "code": 504, "detail": "upstream timeout"}
Upstream unreachable 502 {"error": true, "code": 502, "detail": "upstream error (…)": …}
Non-mahasiswa id → detail 500 explains that only the mhs route exists
Bad path / type 404 / 422 FastAPI default

⚡ Performance

Measured against one live host (scripts/scenarios.py):

Scenario Result
Cold request (hits upstream) ~0.6–3.4 s
Warm request (cached) ~0.01 ms (280 000× faster than cold)
10 concurrent identical requests 1 upstream call (single-flight)
20 concurrent distinct requests throttled to the 60/min budget, never above it
  • Async + connection pooling — one shared client (httpx / curl_cffi / wreq) reuses TCP/TLS connections.
  • Single-flight de-duplication — concurrent callers asking for the same thing share one upstream call. Without it, 10 parallel identical searches were 10 round trips and 10 slots of the 60/min budget for one answer (measured).
  • Bounded TTL cache — PDDIKTIAW_CACHE_MAX_ENTRIES (default 2048) caps memory. Expired entries are reclaimed on write, so a long-running process that sees many distinct q values cannot grow without bound (measured: 20 000 distinct terms → ~1 800 entries, not 20 000).
  • Cache-Control on every response → CDN/proxy caching works too.
  • Outbound throttle — 1 req/s, 2 concurrent, matching the upstream's own limit. This is what keeps the wrapper safe to run against a government API.

Note: the in-memory cache is per-process. Change detection (changed=1) is only meaningful on a persistent SQL backend.


💾 Response cache (DB)

By default responses are cached in memory (TTL). For persistence across restarts or multiple instances, switch to a SQL backend:

Env Value Notes
PDDIKTIAW_CACHE_BACKEND memory (default) · sqlite · mysql · postgres memory needs no setup
PDDIKTIAW_CACHE_DB_URL sqlite:///pddiktiaw_cache.db · mysql://user:pass@host/db · postgresql://user:pass@host/db required when backend ≠ memory
# SQLite (zero setup — stdlib driver)
export PDDIKTIAW_CACHE_BACKEND=sqlite
export PDDIKTIAW_CACHE_DB_URL=sqlite:///pddiktiaw_cache.db
uvicorn pddiktiaw:app --port 8000

# MySQL
pip install pymysql
export PDDIKTIAW_CACHE_BACKEND=mysql PDDIKTIAW_CACHE_DB_URL=mysql://user:pass@localhost:3306/pddiktiaw

# PostgreSQL
pip install "psycopg[binary]"
export PDDIKTIAW_CACHE_BACKEND=postgres PDDIKTIAW_CACHE_DB_URL=postgresql://user:pass@localhost:5432/pddiktiaw

The table pddiktiaw_cache stores each response as JSON + a sha256 hash.

🔄 How response changes are detected

  1. A response is stored with its sha256 hash.
  2. When its TTL expires, the upstream is re-fetched.
  3. The new hash is compared to the stored one:
    • Different → the row is updated and flagged changed=1, updated_at set.
    • Same → the row is refreshed (TTL extended), changed=0.

Inspect live:

curl http://localhost:8000/api/v1/cache/stats
# {"backend": "TTLCache", "hits": 2, "misses": 2, "hit_rate": 0.5, "entries": 2}

Change detection is only meaningful on a persistent backend. An in-memory cache resets on restart, so there is nothing to compare against.


⚙️ Configuration

Everything is optional — see .env.example.

Variable Required? What it does Default
PDDIKTIAW_TRANSPORT No httpx · curl_cffi · wreq httpx
PDDIKTIAW_WREQ_EMULATION No wreq emulation target Chrome136
PDDIKTIAW_CACHE_BACKEND No memory · sqlite · mysql · postgres memory
PDDIKTIAW_CACHE_DB_URL Only if backend ≠ memory DSN, e.g. sqlite:///pddiktiaw_cache.db —
PDDIKTIAW_CACHE_MAX_ENTRIES No Memory-backend entry cap (evicts nearest-to-expiry past it) 2048
PDDIKTIAW_API_KEY No Does nothing — kept only as a marker (see below) —

🔑 There is no API key

Both upstream routes are unauthenticated, and the one header that looks like access control isn't:

Route Measured behaviour
GET /api/pencarian/<kind>/<term> open — no headers beyond User-Agent
POST /api/detail/mhs {"id": …} 403 without Origin, 200 with it
x-recaptcha-token header not enforced — missing, forged and empty all return the same bytes

So the only gate is the Origin header, which the wrapper always sets. An Origin check only stops browser-mediated cross-site reads; it is not access control. This also corrects a widely repeated reading: the web detail route is not "reCAPTCHA-verified server-side".

default_headers() therefore sends no token header at all — sending a stale one could only hurt.

🔀 Transport

The default is httpx: plain TLS, no extra dependency.

The upstream serves no Cloudflare challenge. Measured (python3 scripts/transport_bench.py, 4 requests each): plain httpx, curl_cffi, primp and wreq all returned 4/4 HTTP 200 with cf-cache-status: DYNAMIC and no cf-mitigated.

Transport Result Median Note
httpx (default) 4/4 ok 172 ms no extra dependency
curl_cffi 4/4 ok 175 ms libcurl impersonation
primp 4/4 ok 272 ms sync-only → thread hop; not wired in
wreq 4/4 ok 239 ms BoringSSL, full JA3/JA4/HTTP2

So impersonation buys nothing here today. The others stay selectable for the day that changes:

PDDIKTIAW_TRANSPORT=curl_cffi uvicorn pddiktiaw:app --port 8000
PDDIKTIAW_TRANSPORT=wreq      uvicorn pddiktiaw:app --port 8000

wreq is the strongest option (BoringSSL + full fingerprint control); its response API is not httpx-shaped and is normalised inside _request().


📚 API docs & playground

URL What it is
/ Service info (name, version, docs pointer)
/docs Swagger UI — interactive playground. "Try it out" on any endpoint.
/redoc ReDoc — readable reference docs
/openapi.json Machine-readable OpenAPI spec (for codegen / clients)

🚢 Deploy

pddiktiaw works with zero configuration — no required env var.

Option A — Fly.io

fly launch --no-deploy      # reads fly.toml
fly deploy

fly.toml is preconfigured: port 8000, region sin, autoscaling, HTTPS forced, health check on /api/v1/health.

Option B — Vercel

npx vercel

api/index.py + vercel.json are included. Caveats: cold starts (~2–5 s) and a per-instance, ephemeral cache — fine for demos, not for latency-critical use.

Option C — Docker

docker build -t pddiktiaw .
docker run -p 8000:8000 pddiktiaw

Option D — Docker Compose

docker compose up -d

Prebuilt images are published to GHCR on every v* tag (ghcr.io/pandamoon21/pddiktiaw, linux/amd64 + linux/arm64) by .github/workflows/docker.yml.


🧪 Tests

Offline — no network needed:

python3 -m pytest tests/ -q      # 52 passed, 1 skipped
python3 pddiktiaw.py self        # fast offline sanity check

Covers: the Origin/Referer header contract, the no-token-header rule, route mapping, envelope flattening (including the all-null no-match shape), cache TTL and expiry, SQL change detection, error mapping, transport normalisation, the route registry, the async client surface, and the CLI.

Every assertion here is a fact measured against the live upstream, pinned so a future upstream change fails loudly instead of silently returning wrong data.

Live probes (hit the real host, respect the 1 req/s throttle):

python3 scripts/doc_examples.py        # re-run every documented example
python3 scripts/pencarian_probe.py     # re-verify the enc-route behaviour
python3 scripts/transport_bench.py     # re-verify all four HTTP transports

doc_examples.py exists because tests/ is offline and so cannot catch a README example that has drifted — and one had: a pages=N snippet that silently returned 100 records.


📦 Project layout

pddiktiaw/
├── pddiktiaw.py             # the whole API: config, client, cache, routes, CLI
├── pyproject.toml           # packaging (console script: pddiktiaw)
├── CLAUDE.md                # AI-agent / dev instructions
├── CHANGELOG.md             # version history
├── api/index.py             # Vercel serverless entrypoint
├── vercel.json              # Vercel config
├── requirements.txt         # pinned for deploys (pyproject has >= bounds)
├── Dockerfile               # python:3.12-slim, non-root
├── docker-compose.yml       # one-command deploy, with healthcheck
├── fly.toml                 # Fly.io config
├── pytest.ini               # pythonpath for tests
├── .env.example             # every env var documented
├── scripts/
│   ├── doc_examples.py      # live: re-run every example the README prints
│   ├── pencarian_probe.py   # live: re-verify the enc-route quirk
│   ├── transport_bench.py   # live: compare the four HTTP transports
│   └── scenarios.py         # real-life scenario audit (offline / --live)
├── tests/
│   └── test_pddiktiaw.py    # 52 offline checks
└── README.md

⚠️ Known limits

  • ~100 records per call, sampled at random. See Search semantics. There is no page 2.
  • No detail route except mahasiswa. POST /api/detail/mhs is the only one; a dosen/pt/prodi id 500s there. Use pt_by_code() / dosen_by_nidn() / prodi_by_name() to resolve those to a single record.
  • No aggregate statistics endpoint. The "Statistik Pendidikan Tinggi" publications on the console are PDFs on external hosts, not API resources.
  • Rate limit is hard. 60/min. The wrapper throttles to 1 req/s; don't remove it, and don't point heavy crawlers at this.
  • Version drift. If the console's API changes, endpoints may start returning 403/404. route_for(), default_headers() and _upstream_error() are the constants to update — all in one file.
  • Unofficial and unaffiliated. For research and personal use.

🛠 Development

python3 -m pytest tests/ -q                  # offline
python3 pddiktiaw.py self                    # offline
python3 -m ruff check pddiktiaw.py tests/ scripts/
uvicorn pddiktiaw:app --reload --port 8000   # dev server

Conventions (full detail in CLAUDE.md):

  • Single-file rule — new endpoints go in pddiktiaw.py. Avoid new dependencies; fastapi + httpx + stdlib cover almost everything, and the optional transports / DB drivers stay optional.
  • Never remove the throttle or the cache. 1 req/s is the upstream's own ceiling.
  • Keep tests offline. tests/ must never touch the network.
  • Comments carry the measurement. # measured: … / # ponytail: … — a claim without a measurement is how a wrapper like this goes wrong.

📜 License & disclaimer

MIT. Unofficial and not affiliated with Kementerian Pendidikan Tinggi, Sains dan Teknologi. This is a public government database containing personal academic records:

  • Search endpoints are aggregate; detail() resolves one person's full record (name, NIM, institution, programme, enrolment date, status). Use it for verification or research, not to build profiles of individuals.
  • The wrapper throttles to 1 req/s to stay inside the upstream's own 60/min limit. Do not remove it.
  • No personal data is cached by default beyond the in-process TTL response cache.

Metadata

Release files for pddiktiaw 1.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for pddiktiaw 1.1.0
File Size Uploaded
pddiktiaw-1.1.0.tar.gz 36.2 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for pddiktiaw 1.1.0
File Interpreter ABI Platform
pddiktiaw-1.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 64.0 kB

Release files / pddiktiaw-1.1.0.tar.gz

Download URL pddiktiaw-1.1.0.tar.gz
Size 36.2 kB
Tags Source
SHA-256 checksum
How to use checksums
98d6f0a216121001c0d61bb74ede4a98dcdfef465d2249120fceefde118f64bf
BLAKE2b-256 checksum
How to use checksums
55be5cd6cdd5417eda01b46c17c021bbcc7f071d512192b980820fe89bccae52
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.12.10

Release files / pddiktiaw-1.1.0-py3-none-any.whl

Download URL pddiktiaw-1.1.0-py3-none-any.whl
Size 27.8 kB
Tags Python 3
SHA-256 checksum
How to use checksums
e3bcf5241c4cf3a3f0a4360c88810572f700eda0784ed7b4e473be112a49dc53
BLAKE2b-256 checksum
How to use checksums
070f396b71497fa364f901ead550b41b7416ea7ed4b6785e2f9a027dda80ea9c
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.12.10

Release history Release notifications | RSS feed

This release

1.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page