Convert India MCA annual return PDF forms MGT-7 and MGT-7A into structured JSON
Project description
mgt7-pdf-to-json
Convert India MCA annual return PDF forms MGT-7 and MGT-7A into structured JSON.
A Python CLI tool and library for parsing India Ministry of Corporate Affairs (MCA) annual return forms and converting them to structured JSON format with comprehensive logging and validation.
Quickstart
Installation
pip install -e .
Basic Usage
# Convert PDF to JSON (outputs to same directory)
mgt7pdf2json input.pdf
# Convert with custom output file
mgt7pdf2json input.pdf -o output.json
# Convert with processing statistics
mgt7pdf2json input.pdf -o output.json --include-stats
Python Library
from mgt7_pdf_to_json import Pipeline, Config
# Quick start with defaults
config = Config.default()
pipeline = Pipeline(config)
result = pipeline.process("input.pdf", output_path="output.json")
# Access results
print(f"Form Type: {result['meta']['form_type']}")
print(f"Company: {result['data']['company']['name']}")
What You Get
The tool converts PDF forms into structured JSON with:
- Company Information: CIN, name, registered address
- Financial Data: Turnover, net worth, financial year
- Meeting Records: Board meetings, AGM details
- Directors Information: Director details and changes
- Validation: Warnings and errors for data quality
See Usage section for more examples and advanced features.
Features
- ✅ CLI Tool: Easy-to-use command-line interface
- ✅ Python Library: Programmatic API for integration
- ✅ Multiple Mappers: Support for
default,minimal, anddboutput formats - ✅ Structured Logging: JSON and console logging with request_id tracking
- ✅ Artifacts: Optional intermediate file saving for debugging
- ✅ Validation: Built-in validation with warnings and errors
- ✅ Configurable: YAML-based configuration with CLI override support
- ✅ Production Ready: Comprehensive error handling and exit codes
Installation
Basic Installation
pip install -e .
Development Installation
pip install -e ".[dev]"
This installs additional development dependencies:
pytestandpytest-covfor testingruffandblackfor code formattingmypyfor type checking
Usage
CLI
Basic Usage
mgt7pdf2json examples/U17120DL2013PTC262515_mgt7.pdf
This creates U17120DL2013PTC262515_mgt7.json in the same directory.
With Output File
mgt7pdf2json examples/U17120DL2013PTC262515_mgt7.pdf -o output.json
With Custom Mapper
mgt7pdf2json examples/U17120DL2013PTC262515_mgt7.pdf --mapper minimal -o output.json
Available mappers:
default: Full JSON output with all parsed fieldsminimal: Minimal output with essential fields onlydb: Database-friendly format with flattened structure
With Configuration
mgt7pdf2json examples/U17120DL2013PTC262515_mgt7.pdf --config config.yml
Enable Debug Artifacts
mgt7pdf2json examples/U17120DL2013PTC262515_mgt7.pdf --debug-artifacts
This saves intermediate files (raw, normalized, parsed) in logs/artifacts/.
Strict Validation
mgt7pdf2json examples/U17120DL2013PTC262515_mgt7.pdf --strict
Fails if any required fields are missing.
Include Processing Statistics
mgt7pdf2json examples/U17120DL2013PTC262515_mgt7.pdf --include-stats -o output.json
Includes processing statistics (time, pages, tables, parsed fields) in the output JSON.
Output to Directory
mgt7pdf2json examples/U17120DL2013PTC262515_mgt7.pdf --outdir output/
Creates output/U17120DL2013PTC262515_mgt7.json.
Custom Logging
# Set log level
mgt7pdf2json examples/U17120DL2013PTC262515_mgt7.pdf --log-level DEBUG
# Use JSON logging format
mgt7pdf2json examples/U17120DL2013PTC262515_mgt7.pdf --log-format json
# Set custom log directory
mgt7pdf2json examples/U17120DL2013PTC262515_mgt7.pdf --log-dir custom_logs/
Fail on Warnings
mgt7pdf2json examples/U17120DL2013PTC262515_mgt7.pdf --fail-on-warnings
Exits with error code if any warnings are generated during processing.
Complete Example
mgt7pdf2json examples/U17120DL2013PTC262515_mgt7.pdf \
--output result.json \
--mapper minimal \
--config config.yml \
--log-level INFO \
--strict \
--include-stats
Python Library
from mgt7_pdf_to_json import Pipeline, Config
# Using default configuration
config = Config.default()
pipeline = Pipeline(config)
result = pipeline.process("input.pdf", output_path="output.json")
print(f"Form Type: {result['meta']['form_type']}")
print(f"Company CIN: {result['data']['company']['cin']}")
print(f"Warnings: {len(result['warnings'])}")
print(f"Errors: {len(result['errors'])}")
With Custom Configuration
from mgt7_pdf_to_json import Pipeline, Config
# Load from YAML file
config = Config.from_yaml("config.yml")
# Or override programmatically
config.logging.level = "DEBUG"
config.artifacts.enabled = True
config.pipeline.mapper = "minimal"
pipeline = Pipeline(config)
result = pipeline.process("input.pdf")
With Processing Statistics
from mgt7_pdf_to_json import Pipeline, Config
config = Config.default()
pipeline = Pipeline(config)
# Enable statistics collection
result = pipeline.process(
"input.pdf",
output_path="output.json",
include_stats=True
)
# Access statistics
if "statistics" in result.get("meta", {}):
stats = result["meta"]["statistics"]
print(f"Processing time: {stats['processing_total_duration_seconds']:.2f}s")
print(f"Pages: {stats['pages_count']}")
print(f"Tables: {stats['tables_count']}")
print(f"Parsed fields: {stats['parsed_fields_count']}")
Processing Without Output File
from mgt7_pdf_to_json import Pipeline, Config
config = Config.default()
pipeline = Pipeline(config)
# Process without saving to file (returns dict only)
result = pipeline.process("input.pdf")
# Access parsed data
form_type = result["meta"]["form_type"]
company_name = result["data"]["company"]["name"]
warnings = result["warnings"]
errors = result["errors"]
Error Handling
from mgt7_pdf_to_json import Pipeline, Config
from mgt7_pdf_to_json.exceptions import UnsupportedFormatError
config = Config.default()
pipeline = Pipeline(config)
try:
result = pipeline.process("input.pdf", output_path="output.json")
except FileNotFoundError:
print("Input file not found")
except ValueError as e:
if "scanned" in str(e).lower():
print("PDF appears to be scanned. OCR required.")
else:
print(f"Processing error: {e}")
except Exception as e:
print(f"Unexpected error: {e}")
Configuration
Configuration is managed via YAML files. See config.example.yml for a complete example.
Example Configuration
logging:
level: INFO
format: console
format_file: json
file: logs
date_format: "%d-%m-%Y"
artifacts:
enabled: false
dir: artifacts
save_raw: true
save_normalized: true
save_parsed: true
save_output: false
keep_days: 7
pipeline:
mapper: default
validation:
strict: false
required_fields:
- meta.form_type
- meta.financial_year.from
- meta.financial_year.to
- company.cin
- company.name
Output Format
Default Mapper
{
"meta": {
"request_id": "6a1d1c35-7f88-4e12-9e9f-8d3d4d1b6f5a",
"schema_version": "1.0",
"form_type": "MGT-7",
"financial_year": {
"from": "01/04/2024",
"to": "31/03/2025"
},
"source": {
"input_file": "example.pdf"
}
},
"data": {
"company": {
"cin": "U17120DL2013PTC262515",
"name": "TEGAN TEXOFAB PRIVATE LIMITED"
},
"turnover_and_net_worth": {
"turnover_inr": 891114630,
"net_worth_inr": 266771238
},
"meetings": {
"board_meetings": [
{
"date": "01/04/2024",
"directors_total": 2,
"directors_attended": 2
}
]
}
},
"warnings": [],
"errors": []
}
Exit Codes
The CLI uses standard exit codes:
0: Success1: Processing error (extraction/parsing/mapping/write error)2: Validation failed (in strict mode)3: Input file not found4: Unsupported format (cannot detect form type)5: Warnings as errors (--fail-on-warningsenabled)6: Configuration error
Development
Code Formatting
ruff format .
Linting
ruff check --fix .
Type Checking
mypy src/mgt7_pdf_to_json
Running Tests
# Run all tests
pytest
# Run with coverage
pytest --cov=src/mgt7_pdf_to_json --cov-report=term-missing
# Run specific test categories
pytest -m unit
pytest -m integration
pytest -m smoke
Project Structure
mgt7-pdf-to-json/
├── src/
│ └── mgt7_pdf_to_json/
│ ├── __init__.py
│ ├── cli.py # CLI interface
│ ├── config.py # Configuration management
│ ├── pipeline.py # Main pipeline orchestrator
│ ├── extractor.py # PDF extraction
│ ├── normalizer.py # Text normalization
│ ├── parser.py # Document parsing
│ ├── mappers.py # Output mappers
│ ├── validator.py # JSON validation
│ ├── artifacts.py # Artifact management
│ ├── logging_.py # Structured logging
│ ├── models.py # Data models
│ └── date_utils.py # Date parsing utilities
├── tests/ # Test suite
├── examples/ # Example PDF files
├── docs/ # Documentation
├── config.example.yml # Example configuration
└── pyproject.toml # Project configuration
License
MIT
Contributing
We welcome contributions! Please see CONTRIBUTING.md for detailed guidelines on:
- Development setup
- Coding standards
- Testing guidelines
- Commit message conventions
- Pull request process
Quick start:
- Fork the repository
- Create a feature branch
- Make your changes
- Run tests and linting
- Submit a pull request
Troubleshooting
Common Issues
"Input file not found" Error
Problem: The tool cannot find the specified PDF file.
Solutions:
- Check that the file path is correct
- Use absolute path if relative path doesn't work
- Ensure the file has
.pdfextension - Check file permissions (must be readable)
"Unsupported PDF format" Error
Problem: The PDF appears to be scanned or image-only.
Solutions:
- The PDF may be a scanned document requiring OCR
- Try using OCR tools to convert scanned PDFs to text-based PDFs
- Ensure the PDF contains extractable text (not just images)
"Validation failed" Error (in strict mode)
Problem: Required fields are missing in the parsed output.
Solutions:
- Check if the PDF is a valid MGT-7 or MGT-7A form
- Try processing without
--strictflag to see warnings instead - Enable
--debug-artifactsto inspect intermediate parsing results - Check the
errorsarray in the output JSON for details
Low Parsing Accuracy
Problem: Some fields are not parsed correctly.
Solutions:
- Enable
--debug-artifactsto inspect raw extracted text - Check the normalized text artifact to see how text was cleaned
- Review the parsed artifact to see what was extracted
- Some PDFs may have non-standard formatting
Memory Issues with Large PDFs
Problem: Processing fails or is slow with large PDF files.
Solutions:
- Ensure sufficient system memory
- Process files one at a time rather than in batch
- Consider splitting very large PDFs if possible
Getting Help
- Check the logs: Enable
--log-level DEBUGfor detailed information - Enable artifacts: Use
--debug-artifactsto inspect intermediate files - Review output: Check the
warningsanderrorsarrays in the JSON output - Open an issue: Provide the error message, PDF file type, and log output
FAQ
What PDF formats are supported?
Currently, the tool supports:
- MGT-7: Annual Return form for companies
- MGT-7A: Annual Return form for One Person Companies (OPC)
The PDF must contain extractable text (not scanned images).
Can I process multiple PDFs at once?
Currently, the CLI processes one PDF at a time. For batch processing, you can:
- Use a shell script to loop through files
- Use the Python library in a loop
- Process files in parallel using Python's
multiprocessing
How accurate is the parsing?
Parsing accuracy depends on:
- PDF quality and formatting
- Text extraction quality
- Form structure consistency
The tool includes validation to identify missing or incorrect fields. Use --strict mode for production to ensure all required fields are present.
Can I customize the output format?
Yes! You can:
- Use different mappers:
default,minimal, ordb - Create custom mappers by extending
BaseMapper - Process the output JSON programmatically to transform it
How do I handle warnings and errors?
- Warnings: Indicate missing optional fields or minor parsing issues
- Errors: Indicate missing required fields (in strict mode) or critical issues
- Use
--fail-on-warningsto treat warnings as errors - Check the
warningsanderrorsarrays in the output JSON
What are artifacts?
Artifacts are intermediate files saved during processing:
- Raw: Extracted text and metadata from PDF
- Normalized: Cleaned and normalized text
- Parsed: Structured parsed data
- Output: Final JSON output
Enable with --debug-artifacts for debugging parsing issues.
How do I contribute?
See CONTRIBUTING.md for detailed guidelines on:
- Development setup
- Code style and conventions
- Testing requirements
- Pull request process
Support
For issues and questions:
- GitHub Issues: Open an issue
- Security Issues: Report security vulnerability
- Discussions: Start a discussion
Project details
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file mgt7_pdf_to_json-0.1.9.tar.gz.
File metadata
- Download URL: mgt7_pdf_to_json-0.1.9.tar.gz
- Upload date:
- Size: 52.1 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.9.25
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
86bcbef8a63cfa948137770213b35f34eee13527c2e7690374dadcf18465bb55
|
|
| MD5 |
551a4ebfc62013170660c54981819ab6
|
|
| BLAKE2b-256 |
8ee5e3703c90958a777e62751845cb4071d109caf0527c74775dd075774905b4
|
File details
Details for the file mgt7_pdf_to_json-0.1.9-py3-none-any.whl.
File metadata
- Download URL: mgt7_pdf_to_json-0.1.9-py3-none-any.whl
- Upload date:
- Size: 34.2 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.9.25
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
a1f1ba98fc554ce19b3daf61bab81fcee0a8288635be48a6fe7e361e8a2a50ec
|
|
| MD5 |
cef2ef4979b6ce7822ecd8f773a4c3ec
|
|
| BLAKE2b-256 |
696674d78d38934a84cebdb404b14a81faa0a00eade977a976d9e9d3f72f38fe
|