Extract structured information from documents using AI
Project description
Metaminer
A tool for extracting structured information from documents using AI.
Overview
Metaminer allows you to extract structured data from various document formats (PDF, DOCX, TXT, etc.) by asking natural language questions. It uses AI to analyze documents and return structured results in CSV or JSON format.
Installation
pip install metaminer
System Requirements
Metaminer requires pandoc to be installed on your system for document processing:
- Ubuntu/Debian:
sudo apt-get install pandoc - macOS:
brew install pandoc - Windows: Download from pandoc.org
Note: Metaminer uses pandoc for most document formats and PyMuPDF specifically for PDF processing to ensure optimal text extraction.
Usage
Command Line Interface
# Basic usage
metaminer questions.txt documents/
# Process single document
metaminer questions.txt document.pdf
# Save results to file
metaminer questions.txt documents/ --output results.csv
# JSON output format
metaminer questions.txt documents/ --format json --output results.json
# Custom API endpoint and model
metaminer questions.txt documents/ --base-url http://localhost:8000/api/v1 --model gpt-4
# Use specific AI model
metaminer questions.txt documents/ --model gpt-4
# Show normalized question structure with inferred types
metaminer questions.txt --show-questions --output questions_analysis.csv
# Verbose output for debugging
metaminer questions.txt documents/ --verbose
# Custom API key and model
metaminer questions.txt documents/ --api-key your-api-key --model gpt-4
Python Module
from metaminer import Inquiry, extract_metadata, Config
from metaminer import extract_text, extract_text_from_directory, get_supported_extensions
from metaminer import DataTypeInferrer, infer_question_types, setup_logging
import pandas as pd
# From question file
inquiry = Inquiry.from_file("questions.txt")
df = inquiry.process_documents("documents/")
# Direct questions
inquiry = Inquiry(questions=["Who is the author?", "What is the publication date?"])
df = inquiry.process_documents(["doc1.pdf", "doc2.docx"])
# Single document
result = inquiry.process_document("document.pdf")
# Process text directly (without files)
inquiry = Inquiry(questions=["Who is the author?", "What is the main topic?"])
result = inquiry.process_text("This is a research paper by Dr. Smith about machine learning.")
# Process multiple texts with concurrent processing
texts = ["Document 1 content...", "Document 2 content...", "Document 3 content..."]
results = inquiry.process_texts(texts) # Uses concurrent processing by default
# Pandas integration for seamless data processing
import pandas as pd
df = pd.DataFrame({'text': ["Doc 1 content", "Doc 2 content", "Doc 3 content"]})
# Method 1: Using apply
df['results'] = df['text'].apply(inquiry.process_text)
# Method 2: Using vectorized processing (better performance for large datasets)
results = inquiry.process_texts(df['text'].tolist())
df['results'] = results
# Extract text directly
text = extract_text("document.pdf")
# Extract text from directory
texts = extract_text_from_directory("documents/")
# Get supported file extensions
extensions = get_supported_extensions()
# Use configuration
config = Config()
print(f"Default API endpoint: {config.base_url}")
# Set up logging
logger = setup_logging(config)
# Infer data types for questions
questions = ["Who is the author?", "What is the publication date?", "How many pages?"]
type_suggestions = infer_question_types(questions)
for q, suggestion in type_suggestions.items():
print(f"{q}: {suggestion.suggested_type} - {suggestion.reasoning}")
# Use DataTypeInferrer directly
inferrer = DataTypeInferrer()
suggestion = inferrer.infer_single_type("What is the priority level?")
print(f"Suggested type: {suggestion.suggested_type}")
Question Formats
Text File (.txt)
One question per line:
Who is the author?
What is the publication date?
What is the main topic?
CSV File (.csv)
Structured format with optional field names, data types, and default values:
question,field_name,data_type,default
"Who is the author?",author,str,"Unknown"
"What is the publication date?",pub_date,date,
"How many pages?",page_count,int,0
"What is the document type?",doc_type,"enum(report,memo,letter)","report"
CSV Columns:
question(required): The question to ask about each documentfield_name(optional): Custom field name for the output (defaults to auto-generated)data_type(optional): Data type specification (defaults tostr)default(optional): Default value to use when extraction fails or returns empty
Supported data types:
str(default): Textint: Integer numbersfloat: Decimal numbersbool: True/False valuesdate: Date values (e.g., YYYY-MM-DD)datetime: Date and time values (e.g., YYYY-MM-DD HH:MM:SS)list(type): Arrays of values (e.g.,list(str),list(int),list(date))enum(val1,val2,val3): Single choice from discrete valuesmulti_enum(val1,val2,val3): Multiple choices from discrete values
Supported Document Formats
Thanks to pandoc integration and PyMuPDF, metaminer supports:
- PDF (.pdf)
- Microsoft Word (.docx, .doc)
- OpenDocument (.odt)
- Rich Text Format (.rtf)
- Plain text (.txt)
- Markdown (.md)
- HTML (.html)
- EPUB (.epub)
- LaTeX (.tex)
Concurrent Processing & Performance
Metaminer includes built-in concurrent processing capabilities for efficient batch processing of multiple texts while respecting API rate limits and system resources.
Performance Configuration
# For high-throughput processing
config = Config(
max_concurrent_requests=10, # More workers for faster processing
requests_per_minute=300, # Higher rate limit if your API supports it
batch_size=100 # Larger batches for better memory efficiency
)
# For rate-limited APIs
config = Config(
max_concurrent_requests=2, # Fewer workers to stay under limits
requests_per_minute=60, # Conservative rate limit
batch_size=20 # Smaller batches to reduce memory usage
)
Environment Variables
# Concurrent processing settings
export METAMINER_MAX_CONCURRENT_REQUESTS=5
export METAMINER_REQUESTS_PER_MINUTE=120
export METAMINER_BATCH_SIZE=50
Configuration
API Settings
By default, metaminer connects to a local AI server at http://localhost:5001/api/v1. You can customize this using environment variables or command-line options:
Environment Variables
# API Configuration
export OPENAI_API_KEY=your-api-key
export METAMINER_BASE_URL=http://your-api-server.com/api/v1
export METAMINER_MODEL=gpt-4
export METAMINER_TIMEOUT=60
export METAMINER_MAX_RETRIES=5
# Logging Configuration
export METAMINER_LOG_LEVEL=DEBUG
Command Line
metaminer questions.txt documents/ --base-url http://your-api-server.com/api/v1
Python
from metaminer import Inquiry, Config
# Create Config with explicit parameters
config = Config(
model="gpt-4",
base_url="https://api.openai.com/v1",
api_key="your-api-key"
)
inquiry = Inquiry.from_file("questions.txt", config=config)
# Or use individual parameters
config = Config(model="gpt-4")
inquiry = Inquiry.from_file("questions.txt", config=config)
# Or set environment variables before creating Config
import os
os.environ["METAMINER_BASE_URL"] = "http://your-api-server.com/api/v1"
os.environ["METAMINER_MODEL"] = "gpt-4"
config = Config() # Will use environment variables
inquiry = Inquiry.from_file("questions.txt", config=config)
Configuration Defaults
- Base URL:
http://localhost:5001/api/v1 - Model:
gpt-3.5-turbo - Timeout: 30 seconds
- Max Retries: 3
- Log Level: INFO
- Max File Size: 50MB
Output Format
Results include the extracted information plus metadata:
author,pub_date,page_count,_document_path,_document_name
"John Doe","2023-01-15",25,"/path/to/doc1.pdf","doc1.pdf"
"Jane Smith","2023-02-20",18,"/path/to/doc2.pdf","doc2.pdf"
Default Value Handling
When extraction fails or returns empty values, default values (if specified) are used:
author,doc_type,priority,_document_path,_document_name
"John Doe","report","high","/path/to/doc1.pdf","doc1.pdf"
"Unknown","report","medium","/path/to/doc2.pdf","doc2.pdf"
In this example, the second document had no extractable author, so the default "Unknown" was used.
Examples
Research Paper Analysis
# questions.txt
Who are the authors?
What is the title?
What journal was this published in?
What is the publication year?
What is the main research question?
What methodology was used?
Invoice Processing
question,field_name,data_type,default
"What is the invoice number?",invoice_number,str,"N/A"
"What is the total amount?",total_amount,float,0.0
"What is the invoice date?",invoice_date,date,
"Who is the vendor?",vendor_name,str,"Unknown Vendor"
"What is the due date?",due_date,date,
Legal Document Review
What type of document is this?
Who are the parties involved?
What is the effective date?
What is the termination date?
What are the key obligations?
Document Classification with Enums and Defaults
question,field_name,data_type,default
"What is the document type?",doc_type,"enum(report,memo,letter,invoice)","report"
"What topics are covered?",topics,"multi_enum(finance,hr,marketing,operations)","finance"
"What is the priority level?",priority,"enum(low,medium,high,urgent)","medium"
"What is the title?",title,str,"Untitled Document"
"Who is the author?",author,str,"Unknown"
Notes:
- When using enum types in CSV files, quote the entire type specification to prevent CSV parsing issues with commas
- Default values for enums must be valid enum options
- For multi-enum types, defaults can be single values or comma-separated lists
Data Type Inference
Metaminer includes an intelligent data type inference system that can automatically suggest appropriate data types for your questions. This feature uses AI to analyze question content and recommend the most suitable data types.
Using Type Inference
from metaminer import DataTypeInferrer, infer_question_types
# Infer types for multiple questions
questions = [
"Who is the author?",
"What is the publication date?",
"How many pages are there?",
"Is this document confidential?",
"What is the priority level?"
]
type_suggestions = infer_question_types(questions)
for question_id, suggestion in type_suggestions.items():
print(f"Question: {questions[int(question_id.split('_')[1])-1]}")
print(f"Suggested type: {suggestion.suggested_type}")
print(f"Reasoning: {suggestion.reasoning}")
print(f"Alternatives: {suggestion.alternatives}")
print()
CLI Type Analysis
You can also analyze your questions from the command line:
# Analyze questions and show suggested types
metaminer questions.txt --show-questions
# Save analysis to file
metaminer questions.txt --show-questions --output question_analysis.csv
This will output a structured analysis showing:
- Original questions
- Suggested data types
- Field names
- Reasoning for type suggestions
Type Inference Features
- Smart Analysis: Uses AI to understand question context and intent
- Fallback Logic: Provides sensible defaults when AI analysis fails
- Validation: Ensures all suggested types are valid metaminer data types
- Multiple Suggestions: Provides alternative type options
- Reasoning: Explains why each type was suggested
Development
Running Tests
pip install -e ".[dev]"
pytest
Test Coverage
The project includes comprehensive tests covering:
- Core functionality (document processing, question parsing)
- Data type validation and inference
- Error handling and edge cases
- CSV parsing with various formats
- Default value handling
- Date/datetime processing
- Enum type validation
- Text processing capabilities
Project Structure
metaminer/
├── __init__.py # Main exports
├── inquiry.py # Core Inquiry class
├── document_reader.py # Document text extraction
├── question_parser.py # Question file parsing
├── schema_builder.py # Pydantic schema generation
├── datatype_inferrer.py # AI-powered data type inference
├── extractor.py # Metadata extraction utilities
├── config.py # Configuration management
├── cli.py # Command-line interface
└── __main__.py # Module entry point
License
GNU Lesser General Public License v3.0 - see LICENSE file for details.
Contributing
Contributions are welcome! Please feel free to submit a Pull Request.
Project details
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file metaminer-0.3.6.tar.gz.
File metadata
- Download URL: metaminer-0.3.6.tar.gz
- Upload date:
- Size: 62.0 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.12.9
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
a6b61d31ad27a0b75f2cc9134a7e6be3a9bad11f9375d544299e70c35aba41bf
|
|
| MD5 |
30f5c18bcaa33e4ba8aacaf2847555d6
|
|
| BLAKE2b-256 |
fd9e5ec99057d4ee3d09d65c6abc89d5d35ca2f8a5b7fd60f1153b4a81f221da
|
Provenance
The following attestation bundles were made for metaminer-0.3.6.tar.gz:
Publisher:
publish-to-pypi.yml on travis4dams/metaminer
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
metaminer-0.3.6.tar.gz -
Subject digest:
a6b61d31ad27a0b75f2cc9134a7e6be3a9bad11f9375d544299e70c35aba41bf - Sigstore transparency entry: 230199712
- Sigstore integration time:
-
Permalink:
travis4dams/metaminer@ca57dc0cf7aac122baa9de3085a0844ea2f09e8c -
Branch / Tag:
refs/tags/v0.3.6 - Owner: https://github.com/travis4dams
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish-to-pypi.yml@ca57dc0cf7aac122baa9de3085a0844ea2f09e8c -
Trigger Event:
push
-
Statement type:
File details
Details for the file metaminer-0.3.6-py3-none-any.whl.
File metadata
- Download URL: metaminer-0.3.6-py3-none-any.whl
- Upload date:
- Size: 38.3 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.12.9
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
8c4decc9e3e778bdb8af9cdf803a01d86ba92b2e99c1e4794cd56f9ce565ea76
|
|
| MD5 |
66aae757cebfeb9e6302daaf23797865
|
|
| BLAKE2b-256 |
6bb2ad9d7b876efa32af4b87c758d5b3a0982db2faf7c4254ee902452ef75683
|
Provenance
The following attestation bundles were made for metaminer-0.3.6-py3-none-any.whl:
Publisher:
publish-to-pypi.yml on travis4dams/metaminer
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
metaminer-0.3.6-py3-none-any.whl -
Subject digest:
8c4decc9e3e778bdb8af9cdf803a01d86ba92b2e99c1e4794cd56f9ce565ea76 - Sigstore transparency entry: 230199715
- Sigstore integration time:
-
Permalink:
travis4dams/metaminer@ca57dc0cf7aac122baa9de3085a0844ea2f09e8c -
Branch / Tag:
refs/tags/v0.3.6 - Owner: https://github.com/travis4dams
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish-to-pypi.yml@ca57dc0cf7aac122baa9de3085a0844ea2f09e8c -
Trigger Event:
push
-
Statement type: