pddiktiaw
Unofficial Python wrapper for the PDDIKTI public API — Indonesia's national higher-education database (Pangkalan Data Pendidikan Tinggi): students, lecturers, institutions and study programmes.
FastAPI server (Swagger UI), async library, and a one-shot CLI — in one file. No API key, no account, no configuration.
⚠️ Unofficial. Not affiliated with Kementerian Pendidikan Tinggi, Sains dan Teknologi. Public government data; use it for verification and research, not for building profiles of individuals.
pip install -r requirements.txt
uvicorn pddiktiaw:app --port 8000 # → http://localhost:8000/docs
✨ Why pddiktiaw?
| Web scrapers | pddiktiaw | |
|---|---|---|
| Data source | HTML scraping of the PDDIKTI console | The console's own JSON API |
| Auth | n/a | None exists — no key, no account |
| Response | HTML you then parse | Clean JSON records |
| Speed | Server-rendered page per request | ~1 ms cached, ~0.6 s cold |
| Rate discipline | Up to you | Throttled to the upstream's own 60/min |
The PDDIKTI web console is a React SPA talking to a JSON API on the same host.
pddiktiaw calls that API directly and hands you the records.
📊 Real-world benchmark (measured, not estimated)
Live hits from one machine, median of 3 runs (2026-09):
| Test | Latency | Payload |
|---|---|---|
GET /api/v1/pt?q=univ (cached) |
~1 ms | 21 institutions · 3 KB JSON |
GET /api/v1/pt?q=univ (cold — hits upstream) |
~0.6 s | " |
GET /api/v1/search?q=universitas&kind=all&pages=2 |
~2.5 s | up to 800 records |
scripts/transport_bench.pymeasures the four candidate HTTP stacks against the live host. Your numbers will vary with network; the order of magnitude is the point.
🤖 AI-agent instructions
CLAUDE.md is a purpose-built instruction file for AI coding
agents (and new developers): how to install, verify, extend and deploy
pddiktiaw without tripping on the throttle, the sampling behaviour, or the
route quirks.
- File:
CLAUDE.md - Covers: commands, the three-layer architecture in
pddiktiaw.py, the measured upstream facts that look like bugs but aren't, the cache-key rule, and where to add a new endpoint/transport.
Point any agent (or your future self) at CLAUDE.md first.
🚀 Quickstart
Option 1 — pip:
pip install pddiktiaw
pddiktiaw # serves http://0.0.0.0:8000 (Swagger at /docs)
pddiktiaw self # offline self-check
pddiktiaw pt "universitas" # one-shot JSON commands (see below)
Option 2 — from source:
git clone https://github.com/pandamoon21/PDDIKTI-API-Wrapper && cd PDDIKTI-API-Wrapper
python3 -m venv .venv && . .venv/bin/activate
pip install -r requirements.txt
uvicorn pddiktiaw:app --port 8000
Option 3 — CLI / module (no install):
python3 pddiktiaw.py serve # same server (honours $PORT)
python3 pddiktiaw.py self # offline self-check, no network
Verify:
curl http://localhost:8000/api/v1/health
# {"status":"ok","version":"1.1.0","upstream":"ok","transport":"httpx"}
curl "http://localhost:8000/api/v1/pt?q=universitas&limit=5"
Interactive docs (Swagger UI): http://localhost:8000/docs
🖥 CLI
pddiktiaw is also a one-shot CLI — no server needed.
python3 pddiktiaw.py --help
python3 pddiktiaw.py mhs "budi" # students (default command)
python3 pddiktiaw.py mahasiswa "budi" # alias
python3 pddiktiaw.py dosen "teknik informatika"
python3 pddiktiaw.py pt "universitas" # institutions
python3 pddiktiaw.py prodi "informatika" # study programmes
python3 pddiktiaw.py all "universitas" # every kind at once
python3 pddiktiaw.py search "universitas" --kind pt --pages 3 --limit 50
python3 pddiktiaw.py detail <id> # full mahasiswa record
python3 pddiktiaw.py stats # cache hit/miss counters
python3 pddiktiaw.py status # upstream auth + caveats
python3 pddiktiaw.py dashboard # version, uptime, cache, routes
python3 pddiktiaw.py self # offline self-check
python3 pddiktiaw.py serve # start the HTTP server
$ python3 pddiktiaw.py pt univ --limit 3
3 record(s) for 'univ' (kind=pt)
UNIVERSITAS SURYA 041056 UNIVERSITAS SURYA
UNIVERSITAS ROYAL 011069 UNIVERSITAS ROYAL
UNIVERSITAS SENIOR MEDAN 011073 UNIVERSITAS SENIOR MEDAN
Errors print a clean error: ... message to stderr with exit code 1 — no
tracebacks. Add --json to any search command for raw output.
🐍 Use as a Python package
All methods are async and share the same cache, throttle and transport as the HTTP API:
import asyncio
from pddiktiaw import PDDIKTI
async def main():
api = PDDIKTI()
hits = await api.pt("universitas", limit=5) # institutions
rec = await api.detail(hits["records"][0]["id"]) # full mahasiswa record
# Addressable lookups — no detail route exists upstream for these, so they
# resolve a query to one record instead.
pt = await api.pt_by_code("041056") # by PDDIKTI code
dosen = await api.dosen_by_nidn("9905536334") # by NIDN
pr = await api.prodi_by_name("INFORMATIKA", pt="UNIVERSITAS SURYA")
print(api.stats()) # cache counters
await api.get("/api/pencarian/pt/univ") # raw passthrough
await api.close()
asyncio.run(main())
Every PDDIKTI method: search(), mahasiswa(), dosen(), pt(), prodi(),
search_all_kinds(), detail(), pt_by_code(), prodi_by_name(),
dosen_by_nidn(), get(), post(), stats(), close().
Prefer await api.get(path, ttl=…) / post(path, body, ttl=…) for ad-hoc calls
to routes that aren't wrapped yet — they go through the same throttle + cache.
🎯 Search semantics — sampling, not pagination
This is the single most important thing to know, and it shapes the whole API.
The upstream ignores pagination. page, limit, offset and size are
accepted and all return the same thing: up to 100 records per call. There is no
real paging.
Those 100 records are a random sample, re-drawn on every request. Repeating an identical URL returns a different 100 ids. A single call is a sample, not a page one.
So pddiktiaw gives you pages instead:
await api.search("universitas", kind="pt", pages=1) # one sample, ≤100
await api.search("universitas", kind="pt", pages=5) # union of 5 samples
pages=N repeats the call N times and unions the samples by id. Measured:
kind=all&pages=2 yielded 800 records (4 categories × 100 × 2, with overlap
collapsed); kind=pt&pages=3 yielded 300. That is the only way to see past
~100.
Two details make it work, and both are load-bearing:
- A multi-page search bypasses the response cache while it expands. Serving a cached sample would hand back the same 100 ids and the union would never grow.
- Only the
all/mhs/pt/prodienvelope kinds share one upstream route; each page widens all four categories at once, so a page counts as "exhausted" only when the union grew by nothing.dosenhas its own flat route and stops as soon as it returns fewer than 100 rows.
| Want | Do |
|---|---|
| A quick look | pages=1 |
| Broad coverage of a common term | pages=5–10 |
| Exhaustive | pages=20 (the API cap) and accept ~1 s per page |
Politeness is built in and non-negotiable: 1 request/second, matching the
upstream's own ratelimit-limit: 60 per 60 s. pages=20 therefore costs ~20 s.
Kinds and routes
kind |
Upstream route | Envelope key |
|---|---|---|
mhs (alias mahasiswa) |
enc/all |
mahasiswa |
dosen |
enc/dosen |
(flat list) |
pt |
enc/all |
pt |
prodi |
enc/all |
prodi |
all |
enc/all |
all four merged, each row tagged _kind |
Why mhs/pt/prodi all use enc/all: measured, the enc path's middle
segment only matters for dosen — enc/all, enc/mhs, enc/pt, enc/prodi
and even enc/bogus all return the identical 4-category envelope. route_for()
exploits that; see scripts/pencarian_probe.py to re-verify it live.
📡 Endpoints
All responses are the raw upstream JSON (/api/v1/detail/{id} unwraps
data). Rate limits are per-client, enforced server-side.
| Method | Route | Upstream | Cache TTL |
|---|---|---|---|
| GET | / |
— (local) | — |
| GET | /api/v1/health |
— (live probe) | local |
| GET | /api/v1/status |
— (local) | — |
| GET | /api/v1/dashboard |
— (local) | — |
| GET | /api/v1/cache/stats |
— (local) | — |
| GET | /api/v1/search?q=&kind=&pages=&limit= |
enc/<kind> |
5m |
| GET | /api/v1/mahasiswa?q= |
enc/all → mahasiswa |
5m |
| GET | /api/v1/dosen?q= |
enc/dosen |
5m |
| GET | /api/v1/pt?q= |
enc/all → pt |
5m |
| GET | /api/v1/prodi?q= |
enc/all → prodi |
5m |
| GET | /api/v1/detail/{id} |
POST /api/detail/mhs |
15m |
Full schema at /docs. kind is one of all · mhs · dosen · pt · prodi.
Example responses (captured live)
GET /api/v1/pt?q=univ&limit=2 → 200 · Cache-Control: public, max-age=300
{
"query": "univ",
"kind": "pt",
"count": 2,
"pages": 1,
"records": [
{ "id": "Abtmnf-sQB5hELv…==", "kode": "041056",
"nama_singkat": "UNIV SURYA", "nama": "UNIVERSITAS SURYA" }
]
}
GET /api/v1/dosen?q=teknik informatika&limit=1 → 200 (flat lecturer list)
{
"query": "teknik informatika", "kind": "dosen", "count": 1, "pages": 1,
"records": [
{ "id": "e2Nwi5EYQJg2oSYz…==", "nama": "SUPRAPTO", "nidn": "9905536334",
"nuptk": "", "nama_pt": "AKADEMI TEKNIK PIRI",
"sinkatan_pt": "", "nama_prodi": "TEKNIK INFORMATIKA" }
]
}
GET /api/v1/search?q=universitas&kind=all&limit=1 → 200 (each row tagged)
{
"query": "universitas", "kind": "all", "count": 1, "pages": 1,
"records": [
{ "id": "L-eBRygtawo3FDUw…==", "nama": "MUHAMMAD UNI", "nim": "0515010042",
"nama_pt": "UNIVERSITAS SERAMBI MEKKAH", "nama_prodi": "MANAJEMEN",
"_kind": "mahasiswa" }
]
}
GET /api/v1/detail/{id} → 200 · Cache-Control: public, max-age=900
{
"id": "…", "nim": "…", "nama": "…", "jenjang": "S1", "prodi": "…",
"nama_pt": "…", "kode_pt": "…", "id_pt": "…", "id_sms": "…",
"jenis_daftar": "…", "jenis_kelamin": "…", "kode_prodi": "…"
}
Errors
| Case | Status | Body |
|---|---|---|
| Upstream 4xx | same | {"error": true, "code": <status>, "detail": …} |
| Cloudflare challenge | 403 |
detail names curl_cffi as the fix |
| Upstream timeout | 504 |
{"error": true, "code": 504, "detail": "upstream timeout"} |
| Upstream unreachable | 502 |
{"error": true, "code": 502, "detail": "upstream error (…)": …} |
| Non-mahasiswa id → detail | 500 |
explains that only the mhs route exists |
| Bad path / type | 404 / 422 |
FastAPI default |
⚡ Performance
Measured against one live host (scripts/scenarios.py):
| Scenario | Result |
|---|---|
| Cold request (hits upstream) | ~0.6–3.4 s |
| Warm request (cached) | ~0.01 ms (280 000× faster than cold) |
| 10 concurrent identical requests | 1 upstream call (single-flight) |
| 20 concurrent distinct requests | throttled to the 60/min budget, never above it |
- Async + connection pooling — one shared client (httpx / curl_cffi / wreq) reuses TCP/TLS connections.
- Single-flight de-duplication — concurrent callers asking for the same thing share one upstream call. Without it, 10 parallel identical searches were 10 round trips and 10 slots of the 60/min budget for one answer (measured).
- Bounded TTL cache —
PDDIKTIAW_CACHE_MAX_ENTRIES(default 2048) caps memory. Expired entries are reclaimed on write, so a long-running process that sees many distinctqvalues cannot grow without bound (measured: 20 000 distinct terms → ~1 800 entries, not 20 000). Cache-Controlon every response → CDN/proxy caching works too.- Outbound throttle — 1 req/s, 2 concurrent, matching the upstream's own limit. This is what keeps the wrapper safe to run against a government API.
Note: the in-memory cache is per-process. Change detection (
changed=1) is only meaningful on a persistent SQL backend.
💾 Response cache (DB)
By default responses are cached in memory (TTL). For persistence across restarts or multiple instances, switch to a SQL backend:
| Env | Value | Notes |
|---|---|---|
PDDIKTIAW_CACHE_BACKEND |
memory (default) · sqlite · mysql · postgres |
memory needs no setup |
PDDIKTIAW_CACHE_DB_URL |
sqlite:///pddiktiaw_cache.db · mysql://user:pass@host/db · postgresql://user:pass@host/db |
required when backend ≠ memory |
# SQLite (zero setup — stdlib driver)
export PDDIKTIAW_CACHE_BACKEND=sqlite
export PDDIKTIAW_CACHE_DB_URL=sqlite:///pddiktiaw_cache.db
uvicorn pddiktiaw:app --port 8000
# MySQL
pip install pymysql
export PDDIKTIAW_CACHE_BACKEND=mysql PDDIKTIAW_CACHE_DB_URL=mysql://user:pass@localhost:3306/pddiktiaw
# PostgreSQL
pip install "psycopg[binary]"
export PDDIKTIAW_CACHE_BACKEND=postgres PDDIKTIAW_CACHE_DB_URL=postgresql://user:pass@localhost:5432/pddiktiaw
The table pddiktiaw_cache stores each response as JSON + a sha256 hash.
🔄 How response changes are detected
- A response is stored with its sha256 hash.
- When its TTL expires, the upstream is re-fetched.
- The new hash is compared to the stored one:
- Different → the row is updated and flagged
changed=1,updated_atset. - Same → the row is refreshed (TTL extended),
changed=0.
- Different → the row is updated and flagged
Inspect live:
curl http://localhost:8000/api/v1/cache/stats
# {"backend": "TTLCache", "hits": 2, "misses": 2, "hit_rate": 0.5, "entries": 2}
Change detection is only meaningful on a persistent backend. An in-memory cache resets on restart, so there is nothing to compare against.
⚙️ Configuration
Everything is optional — see .env.example.
| Variable | Required? | What it does | Default |
|---|---|---|---|
PDDIKTIAW_TRANSPORT |
No | httpx · curl_cffi · wreq |
httpx |
PDDIKTIAW_WREQ_EMULATION |
No | wreq emulation target | Chrome136 |
PDDIKTIAW_CACHE_BACKEND |
No | memory · sqlite · mysql · postgres |
memory |
PDDIKTIAW_CACHE_DB_URL |
Only if backend ≠ memory | DSN, e.g. sqlite:///pddiktiaw_cache.db |
— |
PDDIKTIAW_CACHE_MAX_ENTRIES |
No | Memory-backend entry cap (evicts nearest-to-expiry past it) | 2048 |
PDDIKTIAW_API_KEY |
No | Does nothing — kept only as a marker (see below) | — |
🔑 There is no API key
Both upstream routes are unauthenticated, and the one header that looks like access control isn't:
| Route | Measured behaviour |
|---|---|
GET /api/pencarian/<kind>/<term> |
open — no headers beyond User-Agent |
POST /api/detail/mhs {"id": …} |
403 without Origin, 200 with it |
x-recaptcha-token header |
not enforced — missing, forged and empty all return the same bytes |
So the only gate is the Origin header, which the wrapper always sets. An
Origin check only stops browser-mediated cross-site reads; it is not access
control. This also corrects a widely repeated reading: the web detail route is
not "reCAPTCHA-verified server-side".
default_headers() therefore sends no token header at all — sending a stale one
could only hurt.
🔀 Transport
The default is httpx: plain TLS, no extra dependency.
The upstream serves no Cloudflare challenge. Measured
(python3 scripts/transport_bench.py, 4 requests each): plain httpx,
curl_cffi, primp and wreq all returned 4/4 HTTP 200 with
cf-cache-status: DYNAMIC and no cf-mitigated.
| Transport | Result | Median | Note |
|---|---|---|---|
httpx (default) |
4/4 ok | 172 ms | no extra dependency |
curl_cffi |
4/4 ok | 175 ms | libcurl impersonation |
primp |
4/4 ok | 272 ms | sync-only → thread hop; not wired in |
wreq |
4/4 ok | 239 ms | BoringSSL, full JA3/JA4/HTTP2 |
So impersonation buys nothing here today. The others stay selectable for the day that changes:
PDDIKTIAW_TRANSPORT=curl_cffi uvicorn pddiktiaw:app --port 8000
PDDIKTIAW_TRANSPORT=wreq uvicorn pddiktiaw:app --port 8000
wreq is the strongest option (BoringSSL + full fingerprint control); its
response API is not httpx-shaped and is normalised inside _request().
📚 API docs & playground
| URL | What it is |
|---|---|
/ |
Service info (name, version, docs pointer) |
/docs |
Swagger UI — interactive playground. "Try it out" on any endpoint. |
/redoc |
ReDoc — readable reference docs |
/openapi.json |
Machine-readable OpenAPI spec (for codegen / clients) |
🚢 Deploy
pddiktiaw works with zero configuration — no required env var.
Option A — Fly.io
fly launch --no-deploy # reads fly.toml
fly deploy
fly.toml is preconfigured: port 8000, region sin, autoscaling, HTTPS forced,
health check on /api/v1/health.
Option B — Vercel
npx vercel
api/index.py + vercel.json are included. Caveats: cold starts (~2–5 s) and a
per-instance, ephemeral cache — fine for demos, not for latency-critical use.
Option C — Docker
docker build -t pddiktiaw .
docker run -p 8000:8000 pddiktiaw
Option D — Docker Compose
docker compose up -d
Prebuilt images are published to GHCR on every v* tag
(ghcr.io/pandamoon21/pddiktiaw, linux/amd64 + linux/arm64)
by .github/workflows/docker.yml.
🧪 Tests
Offline — no network needed:
python3 -m pytest tests/ -q # 52 passed, 1 skipped
python3 pddiktiaw.py self # fast offline sanity check
Covers: the Origin/Referer header contract, the no-token-header rule, route
mapping, envelope flattening (including the all-null no-match shape), cache TTL
and expiry, SQL change detection, error mapping, transport normalisation, the
route registry, the async client surface, and the CLI.
Every assertion here is a fact measured against the live upstream, pinned so a future upstream change fails loudly instead of silently returning wrong data.
Live probes (hit the real host, respect the 1 req/s throttle):
python3 scripts/doc_examples.py # re-run every documented example
python3 scripts/pencarian_probe.py # re-verify the enc-route behaviour
python3 scripts/transport_bench.py # re-verify all four HTTP transports
doc_examples.py exists because tests/ is offline and so cannot catch a README
example that has drifted — and one had: a pages=N snippet that silently
returned 100 records.
📦 Project layout
pddiktiaw/
├── pddiktiaw.py # the whole API: config, client, cache, routes, CLI
├── pyproject.toml # packaging (console script: pddiktiaw)
├── CLAUDE.md # AI-agent / dev instructions
├── CHANGELOG.md # version history
├── api/index.py # Vercel serverless entrypoint
├── vercel.json # Vercel config
├── requirements.txt # pinned for deploys (pyproject has >= bounds)
├── Dockerfile # python:3.12-slim, non-root
├── docker-compose.yml # one-command deploy, with healthcheck
├── fly.toml # Fly.io config
├── pytest.ini # pythonpath for tests
├── .env.example # every env var documented
├── scripts/
│ ├── doc_examples.py # live: re-run every example the README prints
│ ├── pencarian_probe.py # live: re-verify the enc-route quirk
│ ├── transport_bench.py # live: compare the four HTTP transports
│ └── scenarios.py # real-life scenario audit (offline / --live)
├── tests/
│ └── test_pddiktiaw.py # 52 offline checks
└── README.md
⚠️ Known limits
- ~100 records per call, sampled at random. See Search semantics. There is no page 2.
- No detail route except mahasiswa.
POST /api/detail/mhsis the only one; a dosen/pt/prodi id500s there. Usept_by_code()/dosen_by_nidn()/prodi_by_name()to resolve those to a single record. - No aggregate statistics endpoint. The "Statistik Pendidikan Tinggi" publications on the console are PDFs on external hosts, not API resources.
- Rate limit is hard. 60/min. The wrapper throttles to 1 req/s; don't remove it, and don't point heavy crawlers at this.
- Version drift. If the console's API changes, endpoints may start returning
403/404.route_for(),default_headers()and_upstream_error()are the constants to update — all in one file. - Unofficial and unaffiliated. For research and personal use.
🛠 Development
python3 -m pytest tests/ -q # offline
python3 pddiktiaw.py self # offline
python3 -m ruff check pddiktiaw.py tests/ scripts/
uvicorn pddiktiaw:app --reload --port 8000 # dev server
Conventions (full detail in CLAUDE.md):
- Single-file rule — new endpoints go in
pddiktiaw.py. Avoid new dependencies;fastapi+httpx+ stdlib cover almost everything, and the optional transports / DB drivers stay optional. - Never remove the throttle or the cache. 1 req/s is the upstream's own ceiling.
- Keep tests offline.
tests/must never touch the network. - Comments carry the measurement.
# measured: …/# ponytail: …— a claim without a measurement is how a wrapper like this goes wrong.
📜 License & disclaimer
MIT. Unofficial and not affiliated with Kementerian Pendidikan Tinggi, Sains dan Teknologi. This is a public government database containing personal academic records:
- Search endpoints are aggregate;
detail()resolves one person's full record (name, NIM, institution, programme, enrolment date, status). Use it for verification or research, not to build profiles of individuals. - The wrapper throttles to 1 req/s to stay inside the upstream's own 60/min limit. Do not remove it.
- No personal data is cached by default beyond the in-process TTL response cache.
Metadata
Release files for pddiktiaw 1.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| pddiktiaw-1.1.0.tar.gz | 36.2 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| pddiktiaw-1.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 64.0 kB
Release files / pddiktiaw-1.1.0.tar.gz
| Download URL | pddiktiaw-1.1.0.tar.gz |
|---|---|
| Size | 36.2 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
98d6f0a216121001c0d61bb74ede4a98dcdfef465d2249120fceefde118f64bf
|
|
BLAKE2b-256 checksum How to use checksums |
55be5cd6cdd5417eda01b46c17c021bbcc7f071d512192b980820fe89bccae52
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.12.10
|
Release files / pddiktiaw-1.1.0-py3-none-any.whl
| Download URL | pddiktiaw-1.1.0-py3-none-any.whl |
|---|---|
| Size | 27.8 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
e3bcf5241c4cf3a3f0a4360c88810572f700eda0784ed7b4e473be112a49dc53
|
|
BLAKE2b-256 checksum How to use checksums |
070f396b71497fa364f901ead550b41b7416ea7ed4b6785e2f9a027dda80ea9c
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.12.10
|