Python script to parse and return records based on the record line. This script specifically handles microfile headers sequentially.
Project description
py-visual-cobol
│
├───py_visual_cobol
│ ├───record_extractor.py
│ ├───__init__.py
│ │
│ ├───constants
│ │ ├───params.py
│ │ └───__init__.py
│ │
│ └───utils
│ ├───bytes_converter.py
│ ├───segment_patterns_generator.py
│ └─── __init__.py
├───tests
│ │ test_bytes_converter.py
│ │ test_segment_patterns_generator.py
│ └───__init__.py
│
├───poetry.lock
├───pyproject.toml
└───README.md
from pyspark.sql.types import StructType, StructField, IntegerType, BinaryType
from py_visual_cobol.record_extractor import record_header_extractor
from pyspark.sql import SparkSession,functions as F
from pyspark import StorageLevel
import argparse
import mmap
def parse_args():
"""Parse command line arguments."""
parser = argparse.ArgumentParser(description="Process a microfile.")
parser.add_argument("file_path", type=str, help="Path to the microfile.")
return parser.parse_args()
if __name__ == "__main__":
# Parse command-line arguments
args = parse_args()
# Read the file content
with open(args.file_path, mode="rb") as file:
# Use mmap to map the file into memory
with mmap.mmap(file.fileno(), length=0, access=mmap.ACCESS_READ) as mmapped_file:
# Read data in a memory-efficient way
content = mmapped_file[:] # Only extract bytes as needed
records_length = [
126,
836,
96,
694,
302,
]
# Extract records using the segment patterns generated
records = record_header_extractor(content,records_length=records_length,debug=True)
print(len(records))
spark = SparkSession.builder \
.appName('app') \
.config("spark.executor.memory", "8g") \
.config("spark.driver.memory", "8g") \
.config("spark.executor.cores", "8") \
.config("spark.driver.maxResultSize", "0") \
.config("spark.dynamicAllocation.enabled", "true") \
.getOrCreate()
# Define schema for the DataFrame
schema = StructType([
StructField("rdw", IntegerType(), True),
StructField("value", BinaryType(), True)
])
df = spark.createDataFrame(records, schema=schema)
records.clear()
df.persist(StorageLevel.DISK_ONLY)
df.write.partitionBy("rdw").mode("overwrite").parquet("data", compression="snappy")
df.unpersist(True)
spark.stop()
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
py_visual_cobol-1.1.7.tar.gz
(9.5 kB
view details)
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file py_visual_cobol-1.1.7.tar.gz.
File metadata
- Download URL: py_visual_cobol-1.1.7.tar.gz
- Upload date:
- Size: 9.5 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: poetry/1.8.4 CPython/3.11.9 Windows/10
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
9b323fcf3e75a029036c86ae834fb34b382541146f35015e1083e6f7ad257861
|
|
| MD5 |
0a2e6a000d7f894d2128c50ca362133f
|
|
| BLAKE2b-256 |
0b014937393f891556510ea60754a130a2fe5db6a4cb7c06bbe750814e1c0425
|
File details
Details for the file py_visual_cobol-1.1.7-py3-none-any.whl.
File metadata
- Download URL: py_visual_cobol-1.1.7-py3-none-any.whl
- Upload date:
- Size: 7.5 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: poetry/1.8.4 CPython/3.11.9 Windows/10
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
0e1f52b38b09de80dad1bf145913efa1a19a3266fee3fdecf301ecd66c3bbb04
|
|
| MD5 |
da9b18f8e057e9a283f07b2d2723f80b
|
|
| BLAKE2b-256 |
96a6785b41248a81006a36fd32e5432fad06411550ba4960a932e87f95f49260
|