oruk — Python client for the oruk Speech API
Official Python SDK for oruk, the speech lab building audio models for transcription, multilabel emotion scores, speaking-style classification, and unified audio analysis.
This SDK calls the file API. With Resonance, send a prerecorded English audio file (WAV, FLAC, MP3, M4A, OGG, or WebM; up to 30 MB / 60 minutes), get structured results back. Resonance is oruk’s flagship speech recognition model. Plans include audio minutes, measured by the second with a one-second minimum. The separate Realtime preview supports 32 locales and phrase-level emotion scores over WebSocket; see the realtime reference.
Install
python -m pip install https://oruk.ai/sdk/oruk-0.2.14-py3-none-any.whl
Requires Python 3.10 or newer. Versioned wheel mirror.
Quickstart
Create an account at oruk.ai, choose a subscription plan, and create an API key in the developer portal. Standard self-serve plans begin with a 7-day trial (card required, $0 today).
import os
from oruk import Oruk
with Oruk(api_key=os.environ["ORUK_API_KEY"]) as client:
result = client.analyze("sample.wav", model="oruk-resonance")
print(result["text"]) # English transcript
print(result["emotions"]) # selected model scores; see interpretation below
print(result["styles"]) # selected speaking-style scores; can be empty
Spectra-2
Spectra-2 transcribes 25 languages and returns 15 emotion scores and 16 speaking-style scores in one request. Use mono 16 kHz WAV, finalized FLAC, PCM16, or float32 audio from 45 ms to 60 seconds, up to 4 MiB. The model reference lists the supported languages and outputs.
with Oruk(api_key=os.environ["ORUK_API_KEY"]) as client:
result = client.analyze("sample.flac", model="oruk-spectra-2")
print(result["text"], result["emotions"], result["styles"])
The SDK prepares available speech sessions during ordinary requests to reduce repeat-call latency. For an application that needs to prepare before its first recording, client.prepare_spectra2() returns whether a session is ready. Ordinary requests remain available when it returns False. Explicit request_id values keep their existing behavior. Spectra-2 results cover the whole clip and do not include word timestamps or speaker segments.
For smaller uploads, install soundfile>=0.13,<1 alongside this SDK and enable
compress_audio=True when creating the client. Eligible mono 16 kHz, 16-bit WAV
recordings are encoded as lossless FLAC only when the result is smaller. No audio
is resampled or reduced in precision. Reuse the same setting and original bytes
when retrying a request. Already encoded FLAC files pass through unchanged.
with Oruk(api_key=os.environ["ORUK_API_KEY"], compress_audio=True) as client:
client.prepare_spectra2() # Before recording or a latency-sensitive upload.
result = client.analyze("sample.wav", model="oruk-spectra-2")
Session capacity is tracked conservatively, including concurrent and uncertain requests. A recording that exceeds the remaining capacity uses ordinary admission directly, avoiding a predictably refused session request. The server still checks authorization, capacity, expiry and billing on every call.
Endpoints
| Method | API endpoint | Returns |
|---|---|---|
client.transcribe(file) |
POST /v1/audio/transcriptions |
English transcript |
client.emotions(file) |
POST /v1/audio/emotions |
Selected scores from 15 emotion labels, no transcription |
client.styles(file) |
POST /v1/audio/styles |
Selected scores from 16 speaking-style labels, no transcription |
client.affect(file) |
POST /v1/audio/affect |
emotion + style, no transcript |
client.analyze(file) |
POST /v1/audio/analysis |
transcript, labels, segments, tagged text |
client.proficiency(file, transcript=None) |
POST /v1/audio/proficiency |
Preview: CEFR band, 0–5 score, fluency, transcript |
Every method accepts a path, Path, or binary file object and an optional
request_id=. The first five endpoints use oruk-resonance by default;
oruk-fourier is another file model with its own output scope. Proficiency
uses oruk-proficiency-1, not Resonance or Fourier. A supplied proficiency
transcript= skips built-in transcription. See the endpoint reference
before changing models; their capabilities are not interchangeable.
On the first five endpoints, with model="oruk-resonance", pass diarize=True
(and optionally num_speakers= when you know the speaker count, 1–32) to label speakers: diarization locates the speaker turns,
then Resonance scores each speaker turn, so every segment carries a
speaker field. Other fields follow the endpoint: emotion-only output does
not gain a transcript. Speaker labels are local to the recording, not identities or roles. Diarization is included in plan minutes.
with Oruk(api_key=os.environ["ORUK_API_KEY"]) as client:
result = client.analyze("support-call.wav", model="oruk-resonance", diarize=True)
for seg in result["segments"]:
print(seg["speaker"], seg.get("text"), seg.get("emotions", []))
Emotion only: client.emotions(...) on Resonance runs the encoder and affect
head and never invokes the transcription decoder, so nothing is transcribed,
the result has no transcript. One audio minute uses one plan minute for either emotion-only or unified analysis; calling both separately processes the audio twice.
with Oruk(api_key=os.environ["ORUK_API_KEY"]) as client:
result = client.emotions("support-call.wav", model="oruk-resonance")
print(result["emotions"]) # actual selected labels and scores
print(result.get("text")) # None: no transcript is produced
The client preserves request IDs after ambiguous outcomes and across retries of HTTP 429, 500,
502, 503, and 504 with jittered backoff (two retries by default). Other HTTP
errors are not retried; network errors from httpx propagate to the caller. A first-attempt Spectra-2 session-capacity refusal can use an ordinary request within the configured retry limit, only after the server confirms that no inference was admitted.
A stable request ID helps tracing; it does not promise exactly-once processing. Errors raise OrukAPIError
with status, code, and request_id attributes.
Runnable file workflows
Set ORUK_API_KEY in your environment, download the complete
Python example, and use your own recording:
curl --fail -O https://oruk.ai/examples/analyze-file.py
python analyze-file.py sample.wav --task analysis > result.json
python analyze-file.py support-call.wav --task analysis --diarize > speakers.json
python analyze-file.py speaking-sample.wav --task proficiency > proficiency.json
Each command sends one logical request, with bounded HTTP retries if needed.
Running multiple commands processes the audio separately. --help describes
model selection, known speaker count, and an optional proficiency transcript
file. The program writes the complete API response to stdout and errors to
stderr. It does not fabricate missing labels or assume every proficiency
request was scored. For proficiency, use 30–60 seconds of spontaneous English
and inspect check.status; insufficient_audio does not establish a CEFR
level. A model estimate is not a language certificate.
Interpreting scores and usage
Spectra-2 returns all 15 emotion and 16 speaking-style scores. For Resonance, these labels are model vocabularies, not a guarantee that every response contains every label. Outputs are selected by model thresholds. If no emotion clears its threshold, the highest-scoring emotion is returned; styles can be empty. Several labels may be high and scores need not sum to one. These thresholds do not establish calibrated probabilities of a person's private feelings. Evaluate representative audio before choosing application thresholds. See label interpretation and the scope of the evaluations.
usage includes measured and billable audio duration and may carry reference
fields such as rate_per_minute_usd, estimated_cost_usd, or pricing_version.
Those reference estimates are not your subscription invoice. Actual charges
follow the plan allowance, overage terms, and billing records. One minute of
audio uses one plan minute per request; separate calls process and meter the
file separately. Keep API keys in server-side code. This SDK's file methods do
not implement the separate realtime WebSocket workflow.
Links
- Documentation and API reference: https://oruk.ai/docs
- Capabilities and scope: https://oruk.ai/capabilities
- Pricing: https://oruk.ai/pricing
- Benchmarks: https://oruk.ai/benchmarks/methodology
- Service status: https://oruk.ai/status
License
MIT
Orukeet
Use model oruk-orukeet for English transcription up to 60 seconds / 4 MiB. Every subscription includes an Orukeet allowance; see current plans and task rates. Optional emotion detection and speaker diarization draw from that same allowance. Extra usage shares the plan spending cap. See the Orukeet contract for availability, outputs, streaming, and limits.
Metadata
Release files for oruk 0.2.14
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| oruk-0.2.14-py3-none-any.whl | Python 3 | none | any | Details |
Release files / oruk-0.2.14-py3-none-any.whl
| Download URL | oruk-0.2.14-py3-none-any.whl |
|---|---|
| Size | 10.4 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
4b501aff199b452fffc2ed8b1586559efc7f2e56c1ec985967178e96ee006801
|
|
BLAKE2b-256 checksum How to use checksums |
0668ff2b0ea1dcf1e3eebea8c7acee3baaf6101f0cbfb53c03533f98746a2b05
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 3, 2026.
Transparency log