Skip to main content

AllRead

Read any file as clean text in one line — local, private, encoding-safe.

import allread

text = allread.read("report.csv")   # markdown table
text = allread.read("notes.txt")    # plain text
text = allread.read("page.html")    # extracted text + tables

No config. No encoding crashes. Works the same on Windows, macOS and Linux.

Install

pip install allread            # core: txt, md, rst, log, csv, tsv, json, jsonl, xml, html
pip install "allread[pdf]"     # + PDF support (PyMuPDF)
pip install "allread[docx]"    # + Word documents
pip install "allread[xlsx]"    # + Excel workbooks
pip install "allread[all]"     # everything

Optional formats are lazy: allread.read("report.pdf") raises a clear ImportError telling you exactly which extra to install.

Why AllRead?

  • Encoding-safe by default — UTF-8, UTF-16/32, UTF-8-BOM, cp1252 and legacy encodings are handled automatically; Persian/Arabic/any non-ASCII content just works.
  • Zero required dependencies — the core is pure standard library. Heavy engines are optional extras, loaded lazily.
  • Structured files become Markdown — CSV/TSV/Excel come back as Markdown tables, ready for LLM prompts, notebooks or docs.
  • Same API everywhere — read() returns text, load() returns a Document with source and format metadata.

API

Function Returns Purpose
allread.read(path) str extracted text/markdown
allread.load(path) Document text + source + format
allread.supported_formats() dict formats and availability

Examples

doc = allread.load("sales.xlsx")
print(doc.format)        # "xlsx"
print(doc.text)          # markdown tables, one section per sheet

Roadmap

  • Core formats (txt, md, rst, log, csv, tsv, json, jsonl, xml, html)
  • Optional PDF / DOCX / XLSX engines with lazy imports
  • OCR for scanned PDFs and images (allread[ocr])
  • Audio transcription (allread[audio])
  • CLI: allread file.pdf > out.md

License

MIT

Release files for allread 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for allread 0.1.0
File Size Uploaded
allread-0.1.0.tar.gz 7.2 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for allread 0.1.0
File Interpreter ABI Platform
allread-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 16.1 kB

Release files / allread-0.1.0.tar.gz

Download URL allread-0.1.0.tar.gz
Size 7.2 kB
Tags Source
SHA-256 checksum
How to use checksums
2ce9dd1604357b4329596d95ec443cada84cc01372292c3808c09923f05de353
BLAKE2b-256 checksum
How to use checksums
7b2361e4e0a87a81bd4f1483e146637fb2ffc508f0b1244b034e2a4a9c625c3e
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.14.4

Release files / allread-0.1.0-py3-none-any.whl

Download URL allread-0.1.0-py3-none-any.whl
Size 8.9 kB
Tags Python 3
SHA-256 checksum
How to use checksums
a0045222ddae01626ebc09e66ff6581de179c8cbb296dd3d293c2a3011f3cf3e
BLAKE2b-256 checksum
How to use checksums
d28a3afb02762784772eabf9b3d0f2b838e5460068b59429b0f3950243fbe22c
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.14.4

Release history Release notifications | RSS feed

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page