Skip to main content

Docling

Docling MCP: making docling agentic

PyPI version PyPI - Python Version uv Ruff Pydantic v2 pre-commit License MIT PyPI Downloads LF AI & Data MCP Registry

A document processing service using the Docling-MCP library and MCP (Model Context Protocol) for tool integration.

Overview

Docling MCP is a service that provides tools for document conversion, processing and generation. It uses the Docling library to convert PDF documents into structured formats and provides a caching mechanism to improve performance. The service exposes functionality through a set of tools that can be called by client applications.

🆕 What's New in v2.0

Major Architecture Update: Docling MCP v2.0 introduces a hybrid architecture with support for both remote API and local conversion modes:

  • 🚀 90% Size Reduction: Base package is now ~50MB (down from ~500MB)
  • ⚡ Faster Installation: No model downloads required for default remote mode
  • 🌐 Remote API Support: Use Docling Serve for scalable cloud-based conversion
  • 💻 Local Mode Available: Install [local] extra for offline/local conversion
  • 🔄 Automatic Fallback: Optional fallback from remote to local mode
  • 🎯 Flexible Configuration: Choose the mode that fits your needs

Migration: Upgrading from v1.x? See MIGRATION_v2.md for detailed instructions.

Installation Options

Remote Mode (Recommended - Lightweight)

For users with access to Docling Serve API:

Getting Docling Serve: Visit docling-serve for installation guides. You can deploy it from published container images or look for managed Docling SaaS offerings.

pip install docling-mcp

Then configure your environment:

export DOCLING_MCP_SERVICE_URL=https://your-docling-service.example.com
export DOCLING_MCP_SERVICE_API_KEY=your-api-key-here
export DOCLING_MCP_CONVERSION_MODE=remote

Local Mode (Full Features)

For users who need local conversion or don't have Docling Serve access:

pip install docling-mcp[local]

Then configure your environment:

export DOCLING_MCP_CONVERSION_MODE=local

Hybrid Mode (Best of Both)

Install with local support and enable automatic fallback:

pip install docling-mcp[local]

Configure for remote with fallback:

export DOCLING_MCP_SERVICE_URL=https://your-docling-service.example.com
export DOCLING_MCP_CONVERSION_MODE=remote
export DOCLING_MCP_FALLBACK_TO_LOCAL=true

Features

  • Conversion tools:
  • Generation tools:
    • Document generation in DoclingDocument, which can be exported to multiple formats
  • Local document caching for improved performance
  • Support for local files and URLs as document sources
  • Memory management for handling large documents
  • Logging system for debugging and monitoring
  • RAG applications with Milvus upload and retrieval

Configuration

All settings use the DOCLING_MCP_ prefix and can be supplied as environment variables, in a .env file in the working directory, or via the env block of your MCP client config. Copy .env.example as a starting point.

Conversion mode

Variable Default Description
DOCLING_MCP_CONVERSION_MODE remote remote or local

Remote service (required when DOCLING_MCP_CONVERSION_MODE=remote)

Variable Default Description
DOCLING_MCP_SERVICE_URL URL of the Docling Serve instance
DOCLING_MCP_SERVICE_API_KEY API key for the service
DOCLING_MCP_SERVICE_TIMEOUT 300.0 Request timeout in seconds
DOCLING_MCP_SERVICE_MAX_RETRIES 3 Max retry attempts
DOCLING_MCP_FALLBACK_TO_LOCAL false Fall back to local if service is unreachable (requires docling-mcp[local])

Conversion pipeline (applies to both modes)

Variable Default Description
DOCLING_MCP_KEEP_IMAGES false Retain page images in output
DOCLING_MCP_IMAGES_SCALE 1.0 Image scale factor (increase to avoid tensor padding errors)
DOCLING_MCP_DO_OCR true Run OCR pipeline
DOCLING_MCP_DO_TABLE_STRUCTURE true Detect table structure

LlamaIndex RAG (--tools llama-index-rag)

Variable Default Description
DOCLING_MCP_LI_API_BASE http://127.0.0.1:1234/v1 OpenAI-compatible LLM endpoint
DOCLING_MCP_LI_API_KEY none API key for the LLM endpoint
DOCLING_MCP_LI_MODEL_ID ibm/granite-3.2-8b LLM model identifier
DOCLING_MCP_LI_EMBEDDING_MODEL BAAI/bge-base-en-v1.5 HuggingFace embedding model

LlamaStack (--tools llama-stack-rag / --tools llama-stack-ie)

Variable Default Description
DOCLING_MCP_LLS_URL http://localhost:8321 LlamaStack server URL
DOCLING_MCP_LLS_VDB_EMBEDDING all-MiniLM-L6-v2 Embedding model for vector DB
DOCLING_MCP_LLS_EXTRACTION_MODEL openai/gpt-oss-20b Model used for structured extraction

Setting variables in an MCP client config

{
  "mcpServers": {
    "docling": {
      "command": "uvx",
      "args": [
        "--from=docling-mcp",
        "docling-mcp-server"
      ],
      "env": {
        "DOCLING_MCP_CONVERSION_MODE": "remote",
        "DOCLING_MCP_SERVICE_URL": "https://your-docling-service.example.com",
        "DOCLING_MCP_SERVICE_API_KEY": "your-api-key-here"
      }
    }
  }
}

Getting started

The easiest way to install Docling MCP and connect it to your client is by launching it via uvx.

Depending on the transfer protocol required, specify the argument --transport, for example

  • stdio used e.g. in Claude for Desktop and LM Studio

    uvx --from docling-mcp docling-mcp-server --transport stdio
    
  • sse used e.g. in Llama Stack

    uvx --from docling-mcp docling-mcp-server --transport sse
    
  • streamable-http used e.g. in containers setup

    uvx --from docling-mcp docling-mcp-server --transport streamable-http
    

More options are available, e.g. the selection of which toolgroup to launch. Use the --help argument to inspect all the CLI options.

For developing the MCP tools further, please refer to the Developing section of CONTRIBUTING.md for instructions.

Integration with MCP clients

One of the easiest ways to experiment with the tools provided by Docling MCP is to leverage an AI desktop client with MCP support. Most of these clients use a common config interface. Adding Docling MCP in your favorite client is usually as simple as adding the following entry in the configuration file.

{
  "mcpServers": {
    "docling": {
      "command": "uvx",
      "args": [
        "--from=docling-mcp",
        "docling-mcp-server"
      ]
    }
  }
} 

When using Claude for Desktop, simply edit the config file claude_desktop_config.json with the snippet above or the example provided here.

In LM Studio, edit the mcp.json file with the appropriate section or simply click on the button below for a direct install.

Add MCP Server docling to LM Studio

Other integrations are described in the integrations page.

Examples

Converting documents

Example of prompt for converting PDF documents:

Convert the PDF document at <provide file-path> into DoclingDocument and return its document-key.

Generating documents

Example of prompt for generating new documents:

I want you to write a Docling document. To do this, you will create a document first by invoking `create_new_docling_document`. Next you can add a title (by invoking `add_title_to_docling_document`) and then iteratively add new section-headings and paragraphs. If you want to insert lists (or nested lists), you will first open a list (by invoking `open_list_in_docling_document`), next add the list_items (by invoking `add_listitem_to_list_in_docling_document`). After adding list-items, you must close the list (by invoking `close_list_in_docling_document`). Nested lists can be created in the same way, by opening and closing additional lists.

During the writing process, you can check what has been written already by calling the `export_docling_document_to_markdown` tool, which will return the currently written document. At the end of the writing, you must save the document and return me the filepath of the saved document.

The document should investigate the impact of tokenizers on the quality of LLMs.

Contributing

We welcome external contributions. See CONTRIBUTING.md for details on how to get started.

License

The Docling MCP codebase is under MIT license. For individual model usage, please refer to the model licenses found in the original packages.

LF AI & Data

Docling and Docling MCP is hosted as a project in the LF AI & Data Foundation.

IBM ❤️ Open Source AI: The project was started by the AI for knowledge team at IBM Research Zurich.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

docling_mcp-2.2.1.tar.gz (28.4 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

docling_mcp-2.2.1-py3-none-any.whl (37.7 kB view details)

Uploaded Python 3

File details

Details for the file docling_mcp-2.2.1.tar.gz.

File metadata

  • Download URL: docling_mcp-2.2.1.tar.gz
  • Upload date:
  • Size: 28.4 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for docling_mcp-2.2.1.tar.gz
Algorithm Hash digest
SHA256 a644a15ef591dcde66816d71815ab58d8023f03ad6319f531285fda2e3b6bbcc
MD5 a72e3d4a2bbd7989122e9e9013e4e544
BLAKE2b-256 573bf5e59ba8f2892de57d7faea926ee920673a1a4ec82eae80397417fc69c45

See more details on using hashes here.

File details

Details for the file docling_mcp-2.2.1-py3-none-any.whl.

File metadata

  • Download URL: docling_mcp-2.2.1-py3-none-any.whl
  • Upload date:
  • Size: 37.7 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for docling_mcp-2.2.1-py3-none-any.whl
Algorithm Hash digest
SHA256 3fb225528bd9b60264ec1e5787a742f83d98d32e057c627adef048dbc3d349f6
MD5 94f39ebbac78bcb8c19ba1074bc936ba
BLAKE2b-256 b9c5083170c8428afb82a558bab70c9c907991580b110c6b8d89b0ed6ffec00b

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page