DocuLink Studio — Python SDK
Official Python SDK for the DocuLink Studio Customer API — OCR / LLM document processing and AI chat. It implements the full 32-endpoint contract with automatic authentication, token auto-refresh, typed models, multipart uploads, SSE chat streaming and Socket.IO real-time subscriptions.
- Distribution name (PyPI):
doculink-studio - Import package:
doculink_studio
⚠️ Breaking change in v0.4.0 — document status constants.
DocumentStatusnow matches the server's full lifecycle:"1"Uploaded,"2"/"3"OCR (processing/completed),"4"/"5"Schema,"6"/"7"Mapping,"9"Error —DocumentStatus.COMPLETEDchanged value from"4"to"7", andOCR/SCHEMA/MAPPINGwere replaced by*_PROCESSING/*_COMPLETED. See../CHANGELOG.md.
⚠️ Also in v0.4.0: the doc-scan
joinpayload is now the object{"room": ...}the server actually reads (bare-string joins were silently ignored — realtime never worked through 0.2.0–0.3.1);return_format_typetakes"1"/"2"(the return format'sType), not"JSON"|"XML"|"CSV"; newscan_documentpipeline helper +update_json_output.
⚠️ Breaking change in v0.3.0 — real-time subscription.
subscribe_document_scan(...)now takes the TaskUUID (the task id you use for upload), not the DocumentScanUUID. The server broadcastsdoc-scanevents to roomdoc-scan-{TaskUUID}; a0.2.0subscription that passed the scan UUID received nothing. Each event's payloadUUIDis still the DocumentScanUUID. See../CHANGELOG.md.
ℹ️ v0.3.1 —
subscribe_document_scan(...)now sends theauth={"token": ...}Socket.IO handshake the document server requires. Authenticate (or pre-set an access token) before subscribing; anonymous connects are rejected.
Install
pip install doculink-studio
Requires Python 3.9+. Runtime dependencies: requests, pydantic>=2,
python-socketio[client], websocket-client.
Quickstart
from doculink_studio import DoculinkClient
client = DoculinkClient(
provider_api_key="<35-char provider key>",
customer_api_key="<35-char customer key>",
email="you@example.com", # optional
# base_url / document_ws_url / chat_ws_url default to the test environment
)
# Authentication is automatic on the first call, but you can force it:
client.authenticate()
usage = client.get_usage()
print(usage.planCode, usage.quotaRemaining)
Configuration
| Kwarg | Default |
|---|---|
base_url |
https://test-api-provider.doculink.studio/api/v1 |
document_ws_url |
https://test-ws.doculink.studio |
chat_ws_url |
https://test-ws.doculink.studio:8000 |
timeout |
30 (SSE / sync chat use 300s automatically) |
access_token / refresh_token |
optional — pre-set to skip initial auth |
session |
optional requests.Session |
Authentication & auto-refresh
The client stores the access + refresh tokens after authenticate(). Every
authenticated request sets Authorization: Bearer <AccessToken>. On a 401 the
client transparently refreshes the access token and retries once; if refresh
fails it re-authenticates with the API keys and retries once.
Upload & process a document
import uuid
task_id = str(uuid.uuid4()) # TaskId = UUID v4 ที่คุณสร้างเอง (**v4 เท่านั้น** — เวอร์ชันอื่นถูกปฏิเสธ)
result = client.upload_file(
task_id=task_id,
file=b"...pdf bytes...", # bytes, a file path (str), or a stream
schema_uuid="<schema uuid>",
return_format_uuid="<return format uuid>",
return_format_type="1", # "1" public / "2" customer — ค่า Type ของ return format ที่เลือก (ไม่ใช่ชื่อ file format)
client_uuid=None, # optional
filename="invoice.pdf",
)
doc_id = result.DocumentScanUUID
# Drive the pipeline (async — wait for each stage's odd status via realtime):
client.ocr_process(task_id, doc_id) # → wait for status "3"
client.schema_process(task_id, doc_id) # → wait for status "5"
client.mapping_process(task_id, doc_id) # → wait for status "7"
output = client.get_json_output(task_id, doc_id)
print(output.JsonOutput)
Document status constants (even = stage running, odd = stage done — wait for
the odd status of each stage before calling the next endpoint):
DocumentStatus.UPLOADED ("1"), OCR_PROCESSING ("2"), OCR_COMPLETED
("3"), SCHEMA_PROCESSING ("4"), SCHEMA_COMPLETED ("5"),
MAPPING_PROCESSING ("6"), COMPLETED ("7"), ERROR ("9").
Real-time — document processing (Socket.IO)
from doculink_studio import DocScanUpdate
def on_update(u: DocScanUpdate):
print(u.Status, u.CurrentLog)
sub = client.subscribe_document_scan(task_id, on_update=on_update)
# ... later ...
sub.close()
The subscribe id is the TaskUUID — the same :taskid you upload to — not
the DocumentScanUUID. On connect the SDK emits join with the object
{"room": "doc-scan-<TaskUUID>"} (same shape as chat) and dispatches
doc-scan events parsed into DocScanUpdate. A task may hold several scans
through the one room; each update's UUID field is the DocumentScanUUID it
belongs to.
Chat
session_id = client.create_chat_session(model="gpt-4o-mini")
# Synchronous (blocks until the full reply is ready):
msg = client.send_chat_message_sync(session_id, "Summarise the invoice")
print(msg.Content)
# Streaming over SSE:
def on_event(evt):
if "chunk" in evt:
print(evt["chunk"], end="")
final = client.send_chat_message_stream(session_id, "Explain more", on_event)
print("\nFinal:", final.Content)
# History / listing
sessions = client.list_chat_sessions(status="active", page=1, limit=20)
history = client.get_chat_history(session_id, page=1, limit=50)
Chat over Socket.IO
sub = client.subscribe_chat_session(
session_id,
on_start=lambda e: print("start", e),
on_chunk=lambda chunk, e: print(chunk, end=""),
on_end=lambda msg, e: print("\ndone:", msg.Content),
on_error=lambda e: print("error", e),
)
client.send_chat_message(session_id, "hello") # async (202) — watch the socket
# ...
sub.close()
The chat join emits an object {"room": "chat-<sessionID>"} and passes
auth={"token": <access token>} on the handshake when a token is available.
Error handling
Every method raises ApiError on HTTP >= 400 or an envelope status: false
(including billing rejections that come back as HTTP 200 with status: false).
from doculink_studio import ApiError
try:
client.get_usage()
except ApiError as e:
print(e.status_code, e.message, e.status)
Development
python -m venv .venv
source .venv/Scripts/activate # Git Bash on Windows
pip install -e ".[dev]"
pytest -q
More docs
See the bundled API reference at ../docs/index.html and
the canonical contract in ../CONTRACT.md.
Release files for doculink-studio 0.4.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| doculink_studio-0.4.0.tar.gz | 28.6 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| doculink_studio-0.4.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size:48.2 kB
Release files / doculink_studio-0.4.0.tar.gz
| Download URL | doculink_studio-0.4.0.tar.gz |
|---|---|
| Size | 28.6 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
393329f7016bfeeb76976b2c1ca185ac4cad384fbec548c3b4673d4690ea0b57
|
|
BLAKE2b-256 checksum How to use checksums |
e68e91228414cdb17baf4db39706fb98828804c7af39743a90023b0e4d96357d
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.12.3
|
Release files / doculink_studio-0.4.0-py3-none-any.whl
| Download URL | doculink_studio-0.4.0-py3-none-any.whl |
|---|---|
| Size | 19.6 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
2669350980ae41fdf7ae2877dec0960f8ad67e62755c42fc666af238997e448a
|
|
BLAKE2b-256 checksum How to use checksums |
ec02034f18a47bf832f2826a0db29c4491f99dfb37d9addfb53306137403c00d
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.12.3
|