Python script to parse and return records based on the record line. This script specifically handles microfile headers sequentially.
Project description
py-visual-cobol
│
├───py_visual_cobol
│ ├───record_extractor.py
│ ├───__init__.py
│ │
│ ├───constants
│ │ ├───params.py
│ │ └───__init__.py
│ │
│ └───utils
│ ├───bytes_converter.py
│ ├───segment_patterns_generator.py
│ └─── __init__.py
├───tests
│ │ test_bytes_converter.py
│ │ test_segment_patterns_generator.py
│ └───__init__.py
│
├───poetry.lock
├───pyproject.toml
└───README.md
from pyspark.sql.types import StructType, StructField, IntegerType, BinaryType
from py_visual_cobol.record_extractor import record_header_extractor
from pyspark.sql import SparkSession,functions as F
from pyspark import StorageLevel
import argparse
import mmap
def parse_args():
"""Parse command line arguments."""
parser = argparse.ArgumentParser(description="Process a microfile.")
parser.add_argument("file_path", type=str, help="Path to the microfile.")
return parser.parse_args()
if __name__ == "__main__":
# Parse command-line arguments
args = parse_args()
# Read the file content
with open(args.file_path, mode="rb") as file:
# Use mmap to map the file into memory
with mmap.mmap(file.fileno(), length=0, access=mmap.ACCESS_READ) as mmapped_file:
# Read data in a memory-efficient way
content = mmapped_file[:] # Only extract bytes as needed
records_length = [
126,
836,
96,
694,
302,
]
# Extract records using the segment patterns generated
records = record_header_extractor(content,records_length=records_length,debug=True)
print(len(records))
spark = SparkSession.builder \
.appName('app') \
.config("spark.executor.memory", "8g") \
.config("spark.driver.memory", "8g") \
.config("spark.executor.cores", "8") \
.config("spark.driver.maxResultSize", "0") \
.config("spark.dynamicAllocation.enabled", "true") \
.getOrCreate()
# Define schema for the DataFrame
schema = StructType([
StructField("rdw", IntegerType(), True),
StructField("value", BinaryType(), True)
])
df = spark.createDataFrame(records, schema=schema)
records.clear()
df.persist(StorageLevel.DISK_ONLY)
df.write.partitionBy("rdw").mode("overwrite").parquet("data", compression="snappy")
df.unpersist(True)
spark.stop()
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
py_visual_cobol-1.1.5.tar.gz
(9.4 kB
view details)
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file py_visual_cobol-1.1.5.tar.gz.
File metadata
- Download URL: py_visual_cobol-1.1.5.tar.gz
- Upload date:
- Size: 9.4 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: poetry/1.8.4 CPython/3.11.9 Windows/10
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
fd638292d77fc1df0072c197016d7070ecb8c2edeaef87af98acc1a10adedae7
|
|
| MD5 |
d0e6c2ecab54c60fcde2ebe298f4bf33
|
|
| BLAKE2b-256 |
2681cca9729d0f781c5b01b5f2dfd89e2bcde409ee0048f2e1da514bba930188
|
File details
Details for the file py_visual_cobol-1.1.5-py3-none-any.whl.
File metadata
- Download URL: py_visual_cobol-1.1.5-py3-none-any.whl
- Upload date:
- Size: 7.5 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: poetry/1.8.4 CPython/3.11.9 Windows/10
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
850a4feeee81539c21ec7b07367b8985787ae77b9e2b2388f881091ce8eeb982
|
|
| MD5 |
f956ecd05799b4f7239a7005ccafbaa7
|
|
| BLAKE2b-256 |
d5f01768838e5fbad37059c4c9686e77a86ed01b075f963520deb834510bcde9
|