Skip to main content

Tikara

Tikara Logo

Coverage Tests Security Scan PyPI GitHub License PyPI - Downloads GitHub issues GitHub pull requests GitHub stars

🚀 Overview

Tikara is a modern, type-hinted Python wrapper for Apache Tika, supporting over 1600 file formats for content extraction, metadata analysis, and language detection. It provides direct JNI integration through JPype for optimal performance.

from tikara import Tika

tika = Tika()
content, metadata = tika.parse("document.pdf")

⚡️ Key Features

  • Modern Python 3.12+ with complete type hints
  • Direct JVM integration via JPype (no HTTP server required)
  • Streaming support for large files
  • Recursive document unpacking
  • Language detection
  • MIME type detection
  • Custom parser and detector support
  • Comprehensive metadata extraction
  • Ships with embedded Tika JAR: works in air-gapped networks. No need to manage libraries.
  • Opinionated Pydantic wrapper over Tika's metadata model, with access to the raw metadata.

📦 Supported Formats

🌈 1682 supported media types and counting!

🛠️ Installation

pip install tikara

System Dependencies

Required Dependencies

  • Python 3.12+
  • Java Development Kit 11+ (OpenJDK recommended)

Optional Dependencies

Image and PDF OCR Enhancements (recommended)
  • Tesseract OCR (strongly recommended if you process images) (Reference ⇗)

    # Ubuntu
    apt-get install tesseract-ocr
    

    Additional language packs for Tesseract (optional):

    # Ubuntu
    apt-get install tesseract-ocr-deu tesseract-ocr-fra tesseract-ocr-ita tesseract-ocr-spa
    
  • ImageMagick for advanced image processing (Reference ⇗)

    # Ubuntu
    apt-get install imagemagick
    
Multimedia Enhancements (recommended)
  • FFMPEG for enhanced multimedia file support (Reference ⇗)

    # Ubuntu
    apt-get install ffmpeg
    
Enhanced PDF Support (recommended)

Enhanced PDF support with PDFBox Reference ⇗

Metadata Enhancements (recommended)
  • EXIFTool for metadata extraction from images Reference ⇗

    # Ubuntu
    apt-get install libimage-exiftool-perl
    
Geospatial Enhancements
  • GDAL for geospatial file support (Reference ⇗)

    # Ubuntu
    apt-get install gdal-bin
    
Additional Font Support (recommended)
  • MSCore Fonts for enhanced Office file handling (Reference ⇗)

    # Ubuntu
    apt-get install xfonts-utils fonts-freefont-ttf fonts-liberation ttf-mscorefonts-installer
    

For more OS dependency information including MSCore fonts setup and additional configuration, see the official Apache Tika Dockerfile.

📖 Usage

Example Jupyter Notebooks 📔

Basic Content Extraction

from tikara import Tika
from pathlib import Path

tika = Tika()

# Basic string output
content, metadata = tika.parse("document.pdf")

# Stream large files
stream, metadata = tika.parse(
    "large.pdf",
    output_stream=True,
    output_format="txt"
)

# Save to file
output_path, metadata = tika.parse(
    "input.docx",
    output_file=Path("output.txt"),
    output_format="txt"
)

Language Detection

from tikara import Tika

tika = Tika()
result = tika.detect_language("El rápido zorro marrón salta sobre el perro perezoso")
print(f"Language: {result.language}, Confidence: {result.confidence}")

MIME Type Detection

from tikara import Tika

tika = Tika()
mime_type = tika.detect_mime_type("unknown_file")
print(f"Detected type: {mime_type}")

Recursive Document Unpacking

from tikara import Tika
from pathlib import Path

tika = Tika()
results = tika.unpack(
    "container.docx",
    output_dir=Path("extracted"),
    max_depth=3
)

for item in results:
    print(f"Extracted {item.metadata['Content-Type']} to {item.file_path}")

🔧 Development

Environment Setup

  1. Ensure that you have the system dependencies installed

  2. Install uv:

    curl -LsSf https://astral.sh/uv/install.sh | sh
    
  3. Install python dependencies and create the Virtual Environment:

    make install
    

Common Tasks

Run make (or make help) to see all available targets. The most common ones:

# Setup
make install         # Install all dependencies (including dev)
make stubs           # Regenerate Java type stubs from the Tika JAR

# Lint & Format
make lint            # Run ruff linter (with auto-fix)
make format          # Run ruff formatter
make ruff            # Run linter and formatter together

# Test
make test            # Run tests with verbose output
make test-fast       # Run tests, skip slow benchmark/isolated markers
make test-coverage   # Run tests with coverage report (XML + terminal)

# Docs
make docs            # Build Sphinx HTML docs
make docs-open       # Build docs and open in browser

# Security
make safety          # Run safety dependency vulnerability scan

# Build & Release
make build           # Build sdist and wheel
make clean           # Remove build artifacts, caches, and generated reports

# CI / Pre-push
make ci              # Run full CI suite (lint → test → safety → docs)
make prepush         # Alias for ci — run before pushing

🤔 When to Use Tikara

Ideal Use Cases

  • Python applications needing document processing
  • Microservices and containerized environments
  • Data processing pipelines (Ray, Dask, Prefect)
  • Applications requiring direct Tika integration without HTTP overhead

Advanced Usage

For detailed documentation on:

  • Custom parser implementation
  • Custom detector creation
  • MIME type handling

See the Example Jupyter Notebooks 📔

🎯 Inspiration

Tikara builds on the shoulders of giants:

  • Apache Tika - The powerful content detection and extraction toolkit
  • tika-python - The original Python Tika wrapper using HTTP that inspired this project
  • JPype - The bridge between Python and Java

Considerations

  • Process isolation: Tika crashes will affect the host application
  • Memory management: Large documents require careful handling
  • JVM startup: Initial overhead for first operation
  • Custom implementations: Parser/detector development requires Java interface knowledge

📊 Performance Considerations

Memory Management

  • Use streaming for large files
  • Monitor JVM heap usage
  • Consider process isolation for critical applications

Optimization Tips

  • Reuse Tika instances
  • Use appropriate output formats
  • Implement custom parsers for specific needs
  • Configure JVM parameters for your use case

🔐 Security Considerations

  • Input validation
  • Resource limits
  • Secure file handling
  • Access control for extracted content
  • Careful handling of custom parsers

🤝 Contributing

Contributions welcome! The project uses Make for development tasks:

make prepush     # Run full CI suite (lint, test, coverage, safety, docs)

For developing custom parsers/detectors, Java stubs can be generated:

make stubs       # Generate Java stubs for Apache Tika interfaces

Note: Generated stubs are git-ignored but provide IDE support and type hints when implementing custom parsers/detectors.

Common Problems

  • Verify Java installation and JAVA_HOME environment variable
  • Ensure Tesseract and required language packs are installed
  • Check file permissions and paths
  • Monitor memory usage when processing large files
  • Use streaming output for large documents

📚 Reference

See API Documentation for complete details.

📄 License

Apache License 2.0 - See LICENSE for details.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

tikara-0.2.0.tar.gz (56.0 MB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

tikara-0.2.0-py3-none-any.whl (56.1 MB view details)

Uploaded Python 3

File details

Details for the file tikara-0.2.0.tar.gz.

File metadata

  • Download URL: tikara-0.2.0.tar.gz
  • Upload date:
  • Size: 56.0 MB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for tikara-0.2.0.tar.gz
Algorithm Hash digest
SHA256 2cbc09782a8babf8bcdb70e88ec4607d673e213b293281ce9f14cdc03947a0b3
MD5 21efdea2b6617c5f08cd5e2c119e938b
BLAKE2b-256 6457383e6201fe4f8077e5de4bb5f8b2bc88aa8bc366e4a2661f4df3181c4547

See more details on using hashes here.

Provenance

The following attestation bundles were made for tikara-0.2.0.tar.gz:

Publisher: release.yml on baughmann/tikara

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file tikara-0.2.0-py3-none-any.whl.

File metadata

  • Download URL: tikara-0.2.0-py3-none-any.whl
  • Upload date:
  • Size: 56.1 MB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for tikara-0.2.0-py3-none-any.whl
Algorithm Hash digest
SHA256 b1336f9d276fc20f9adc0c1dac19d539abcd8f1106d6e574ca25c97a3e81d0e5
MD5 7b4ac148f12021f8d48266796b382e89
BLAKE2b-256 4d90c8fdaef14e229c2bd036091e37e624af7bf8fdca2daa3eb5af5548f29030

See more details on using hashes here.

Provenance

The following attestation bundles were made for tikara-0.2.0-py3-none-any.whl:

Publisher: release.yml on baughmann/tikara

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

This release

0.2.0 This release

2 files

0.1.7

2 files

0.1.6

2 files

0.1.5.post1

2 files

0.1.5

2 files

0.1.4

2 files

0.1.3

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page