qalam
Correct, logical-order Arabic text extraction from digitally-born PDFs — without OCR.
Most PDF extractors hand back Arabic that is reversed, made of presentation forms, or
silently corrupted at every ligature. qalam fixes the reading order, folds the shaped
glyphs back to base letters, reconstructs multi-column reading order, and — when a page
has no usable text layer — says so instead of returning plausible-looking garbage.
import qalam
text = qalam.extract_text("guide.pdf")
doc = qalam.Document("guide.pdf")
print(doc.confidence) # 0.91
print(doc.pages_needing_ocr) # [2, 3, 42, 43]
for page in doc:
if page.needs_ocr:
print(page.number, page.reasons)
else:
print(page.text)
Document.text omits pages that need OCR, so a scanned page never masquerades as a
result. Read page.text directly if you want to see it anyway.
Structure, not just a string
A page is an ordered list of typed blocks. Each carries a kind, so you can branch on it
without reaching for isinstance.
for block in doc.page(6).blocks:
if block.kind == "table":
for row in block.to_rows(): # the shape csv.writer wants
print(row)
elif block.kind == "image":
block.save(block.file_name)
else:
print(block.text)
page.tables, page.images and page.lines are filtered views over the same blocks.
Tables know their reading order: on an Arabic page rows[0][0] is the rightmost cell,
where a reader starts.
t = doc.page(5).tables[0]
t.row_count, t.column_count, t.confidence
t.rows[1][0].text # 'المحافظات'
Images come back decoded — JPEG passed through untouched, everything else re-encoded as
PNG — or with an unsupported_reason saying why not. Never silently dropped.
Each line carries its geometry and styling — line.direction, line.size, line.color,
line.confidence — enough to rebuild a styled document.
page.tagged says whether the reading order came from the document's own structure tree
rather than from geometry.
HTML output
open("out.html", "w").write(doc.to_html())
Reading order, headings inferred from type size, colour, tables with dir="rtl", images
inlined. doc.heading_sizes() reports what the heading inference assumed.
For Markdown, convert the HTML with turndown or pandoc — Markdown cannot state text
direction, so an Arabic table comes out mirrored.
Built on a Rust core; the GIL is released during extraction and rendering.
Metadata
Release files for qalam 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| qalam-0.1.0.tar.gz | 184.0 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| qalam-0.1.0-cp39-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl | CPython 3.9 | abi3 | Linux glibc 2.17+ x86-64 | Details |
Total release size: 1.2 MB
Release files / qalam-0.1.0.tar.gz
| Download URL | qalam-0.1.0.tar.gz |
|---|---|
| Size | 184.0 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
ed7b78e5bc406870414f53f977ed38af476fbfb372772a7d78d1a8301a4dfc69
|
|
BLAKE2b-256 checksum How to use checksums |
e63eebde557987e34b17a83ffb23a1dd865d315cef916ad1740835944d855ffb
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.12.3
|
Release files / qalam-0.1.0-cp39-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl
| Download URL | qalam-0.1.0-cp39-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl |
|---|---|
| Size | 1.0 MB |
| Tags | CPython 3.9 Linux glibc 2.17+ x86-64 abi3 |
|
SHA-256 checksum How to use checksums |
d77fb7046b66ac6a668a75d2b2c96f38d3db518979f194b71bcbcc7de502ffd9
|
|
BLAKE2b-256 checksum How to use checksums |
e7ecb36899d035d9ab593c69c6976ebabeb6f3198d6ebbefde12ce6c59c2a955
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.12.3
|