Skip to main content

🚀 About Laiye ADP

ADP is Laiye's intelligent agent document processing product (Agentic Document Processing, referred to as ADP) , based on the general understanding ability of large models, without relying on rules and annotations, with the general understanding ability of multi-language, MultiModal Machine Learning, and multi-scene; autonomous planning and execution of intelligent agents, able to understand task goals, autonomous planning steps, invoke tools, and complete complex tasks; end-to-end business automation, from document input to business decision-making to human-machine collaboration, forming a complete closed loop.

agentic-doc-parse-and-extract is the official open-source CLI tool of ADP, supporting both manual terminal invocation and automatic invocation via AI Skill. With a single command, it can accomplish: structured document parsing + intelligent extraction of key fields, covering all scenarios including invoices, orders, certificates, bills, and general documents, outputting standard JSON, and seamlessly integrating with automation and AI workflows.


💡 Core Features

agentic-doc-parse-and-extract focuses on intelligent processing of the entire document workflow, taking into account both manual terminal calls and automatic calls by AI Agents. Its core functions cover all scenarios of parsing, extraction, and batch processing, requiring no complex configuration, and operations can be completed with a single command:

Function Name Function Description Optimal Scenario
Document Parsing Automatically recognize multi-format documents such as PDFs and images, convert messy unstructured content (e.g., scanned documents, handwritten text, complex layout documents) into standardized Structured Data, while preserving the original document hierarchy and key relationships Convert unstructured documents into Structured Data for LLM reading and subsequent extraction
Out Of The Box Document Extraction Based on the native AI capabilities of the ADP large model, it comes with built-in standardized extraction models for invoices, receipts, orders, commonly used certificates in China, etc. No need to configure rules or manual annotation, one-click extraction of key fields from various types of general documentation, outputting standard JSON Account Payable automation, expense management, procurement automation, quick entry of card and certificate information into the system
Custom Document Extraction Supports independent creation, editing, and management of personalized extraction applications, allowing configuration of exclusive extraction fields and recognition logic for enterprise-specific documentation and industry-customized forms Private extraction requirements for enterprise-specific documentation, industry-customized forms, and non-standardized documents
Task Query Supports asynchronous task submission and status query, enabling quick viewing of task execution progress, success/failure status, and final task processing results Batch task processing, asynchronous document processing, problem troubleshooting, and processing record tracing
Application Management Provides comprehensive application management capabilities, allowing users to view all available extraction applications (system-built + custom), query application details, and manage application tags Multi-scenario business switching, full lifecycle management of applications, and custom application management

Installation and Update

Install from Source Code

pip install -e .

Install from PyPI

pip install agentic_doc_parse_and_extract

Version Update

Note: This package is no longer maintained on PyPI. Please update via npm instead:

npm install -g @laiye-adp/agentic-doc-parse-and-extract-cli

Configure

Get an API key at https://adp-global.laiye.com/ (new users get 100 free credits per month).

adp config set --api-key <your-api-key>
adp config set --api-base-url https://adp-global.laiye.com
adp config get

Quick Examples

# List available apps
adp app-id list

# Parse a local document
adp parse local ./invoice.pdf --app-id <app-id>

# Extract key fields
adp extract local ./invoice.pdf --app-id <app-id>

# Parse a directory in async mode
adp parse local ./documents/ --app-id <app-id> --async

# Process a remote URL
adp extract url https://example.com/file.pdf --app-id <app-id>

# Query an async task
adp parse query <task-id>

# Two-phase async (submit + query separately, resumable)
adp extract local ./documents/ --app-id <app-id> --async --no-wait --export tasks.json
adp extract query --watch --file tasks.json

# Auto retry on failure (up to 2 retries)
adp parse local ./documents/ --app-id <app-id> --retry 2

# Check remaining credits
adp credit

Commands

AI agents should call adp schema for the machine-readable, authoritative command spec. The table below is a human-friendly summary.

Command Description
adp version Print version
adp config set Set API key / base URL
adp config get Show current config
adp config clear Clear config
adp app-id list List available apps
adp app-id cache Read app list from local cache
adp parse local <path> Parse local file/directory
adp parse url <url> Parse remote file (URL list file supported)
adp parse base64 <data> Parse Base64-encoded content
adp parse query <task-id...> Query async parse tasks (supports multiple IDs or --file)
adp extract local <path> Extract from local file/directory
adp extract url <url> Extract from remote file
adp extract base64 <data> Extract from Base64-encoded content
adp extract query <task-id...> Query async extract tasks (supports multiple IDs or --file)
adp custom-app create Create a custom extraction app
adp custom-app update Update custom app config
adp custom-app get-config Show app config
adp custom-app delete Delete a custom app
adp custom-app delete-version Delete a specific config version
adp custom-app ai-generate AI-recommend extraction fields
adp credit Show remaining credits
adp schema Output command schema (for AI agents)

Flags

Flag Description
--json Output JSON
--quiet Quiet mode, output result only
--lang <en|zh> Interface language
--app-id App ID (required for parse / extract)
--async Async mode
--no-wait Submit tasks only, do not wait for results (use with --async)
--export <path> Export result to file (single file) or directory (batch)
--timeout <seconds> Timeout (default 900s)
--concurrency <n> Concurrent workers (free: max 1, paid: max 2)
--retry <n> Retries for retryable errors (default 0)
--file <path> Read task IDs from JSON file (output of --no-wait, query only)

Async Workflow

For large files or batch jobs, submit with --async and the CLI returns a task-id. Poll for results with parse query / extract query:

adp parse local ./big.pdf --app-id <app-id> --async
# returns a task-id

adp parse query <task-id>

Two-Phase Async (--no-wait)

By default, --async submits and polls until completion — ideal for AI agents. For resumable workflows, use two-phase mode:

Phase 1: Submit tasks

adp extract local ./documents/ --app-id <app-id> --async --no-wait --export tasks.json

Output is a JSON array with task IDs:

[
  {"path": "invoice.pdf", "task_id": "task_abc123"},
  {"path": "contract.pdf", "task_id": "task_def456"}
]

Phase 2: Query results

adp extract query --watch --file tasks.json
adp extract query --watch --file tasks.json --export ./results/

Even if the CLI crashes mid-way, task IDs in tasks.json are preserved — resume anytime with query --file.

Batch Processing

When processing multiple files/URLs, the CLI writes each result to a separate file:

adp_results_20250417_153020/
├── _summary.json              # Summary (total, success, failed, per-file status)
├── invoice_01.pdf.json        # Successful result
├── contract_02.docx.json
└── report_03.pdf.error.json   # Error details
  • --export <dir> — specify output directory
  • Without --export — auto-creates adp_results_<timestamp>/
  • Single file — outputs to stdout or the --export file path

Exit Codes

Code Meaning
0 All success
1 All failed / system error
2 Parameter error
3 Resource not found
4 Permission denied
5 Conflict
6 Partial failure (some tasks failed in batch)

Environment Variables

Variable Description
ADP_API_KEY API key (overrides config file)
ADP_API_BASE_URL Service URL
ADP_LANG Interface language (en / zh)
ADP_LOG_LEVEL Log level (debug / info / warn / error)

Config Storage

  • Config dir: ~/.adp/
  • Config file: ~/.adp/config.json
  • Encrypted API key: ~/.adp/key.enc (AES-256-GCM)
  • App cache: ~/.adp/app_cache.json
  • Version check cache: ~/.adp/version_check.json (refreshed every 24h)

📜 License

We adopt a combined model of open-source tools + paid services: the CLI tool is completely free and open-source, making it easy for everyone to quickly integrate; while the core ADP intelligent parsing capability is a Public Cloud commercial service, billed based on actual usage, aiming to provide users with a highly accurate and stable document processing experience.

  • CLI Tool: Open source under the MIT License, freely available for use, modification, and distribution
  • ADP Service: AI document processing service based on Public Cloud, billed by usage, Billing Rules

Free Quota: New users can receive 100 free credits per month after registration, allowing them to experience full functionality

📞 Support and Contact

Release files for agentic-doc-parse-and-extract 1.10.3

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for agentic-doc-parse-and-extract 1.10.3
File Size Uploaded
agentic_doc_parse_and_extract-1.10.3.tar.gz 65.9 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for agentic-doc-parse-and-extract 1.10.3
File Interpreter ABI Platform
agentic_doc_parse_and_extract-1.10.3-py3-none-any.whl Python 3 none any Details

Total release size: 115.0 kB

Release files / agentic_doc_parse_and_extract-1.10.3.tar.gz

Download URL agentic_doc_parse_and_extract-1.10.3.tar.gz
Size 65.9 kB
Tags Source
SHA-256 checksum
How to use checksums
89095e8cc357f932813335f866fb3cadb3a01f76d2160726a817587d27bc02e2
BLAKE2b-256 checksum
How to use checksums
cb83a75cf48ab07ab8795a25a43fe0f10944f22fc283d28973f1ec59e9e28eea
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.1.0 CPython/3.8.18

Release files / agentic_doc_parse_and_extract-1.10.3-py3-none-any.whl

Download URL agentic_doc_parse_and_extract-1.10.3-py3-none-any.whl
Size 49.1 kB
Tags Python 3
SHA-256 checksum
How to use checksums
545c23b24b3cfa4538383fc6989f636c2a5dc4894b8fea4d2a89242ac4c7bc95
BLAKE2b-256 checksum
How to use checksums
fa05afc9fcc338b68d7c36c7b01d8fbd4eb01f07a5345e5fdb200b91792e2c75
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.1.0 CPython/3.8.18

Release history Release notifications | RSS feed

This release

1.10.3 This release

2 release files

1.10.1

2 release files

1.0.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page