Python script to parse and return records based on the record line. This script specifically handles microfile headers sequentially.
Project description
py-visual-cobol
│
├───py_visual_cobol
│ ├───record_extractor.py
│ ├───__init__.py
│ │
│ ├───constants
│ │ ├───params.py
│ │ └───__init__.py
│ │
│ └───utils
│ ├───bytes_converter.py
│ ├───segment_patterns_generator.py
│ └─── __init__.py
├───tests
│ │ test_bytes_converter.py
│ │ test_segment_patterns_generator.py
│ └───__init__.py
│
├───poetry.lock
├───pyproject.toml
└───README.md
from pyspark.sql.types import StructType, StructField, IntegerType, BinaryType
from py_visual_cobol.record_extractor import record_header_extractor
from pyspark.sql import SparkSession,functions as F
from pyspark import StorageLevel
import argparse
import mmap
def parse_args():
"""Parse command line arguments."""
parser = argparse.ArgumentParser(description="Process a microfile.")
parser.add_argument("file_path", type=str, help="Path to the microfile.")
return parser.parse_args()
if __name__ == "__main__":
# Parse command-line arguments
args = parse_args()
# Read the file content
with open(args.file_path, mode="rb") as file:
# Use mmap to map the file into memory
with mmap.mmap(file.fileno(), length=0, access=mmap.ACCESS_READ) as mmapped_file:
# Read data in a memory-efficient way
content = mmapped_file[:] # Only extract bytes as needed
records_length = [
126,
836,
96,
694,
302,
160,
107,
76,
266,
68,
99,
78,
425,
]
# records_length = [
# 126,
# 79,
# 897,
# 242,
# 199,
# 180,
# 373,
# 186,
# 127,
# 163,
# 106,
# 88,
# 828,
# 219,
# 208,
# 67,
# 54,
# 175,
# 138,
# 214,
# 173,
# 291,
# 116,
# 271,
# 534,
# 172,
# 178,
# 198,
# 78 ,
# 527,
# 653,
# 132,
# 228,
# 159,
# 224,
# 74,
# 221,
# 65,
# 145,
# 982,
# 244,
# 82,
# 73,
# 112,
# 94,
# 142,
# 81,
# 87,
# 119,
# 59,
# 158,
# 226,
# 171,
# 114,
# 113,
# 72,
# 248,
# 64,
# 164,
# 91,
# 125,
# 63,
# 55,
# 89,
# 506,
# 315,
# 289,
# 1490,
# 1000,
# ]
# Extract records using the segment patterns generated
records = record_header_extractor(content,records_length=records_length,debug=True)
print(len(records))
spark = SparkSession.builder \
.appName('app') \
.config("spark.executor.memory", "8g") \
.config("spark.driver.memory", "8g") \
.config("spark.executor.cores", "8") \
.config("spark.driver.maxResultSize", "0") \
.config("spark.dynamicAllocation.enabled", "true") \
.getOrCreate()
# Define schema for the DataFrame
schema = StructType([
StructField("rdw", IntegerType(), True),
StructField("value", BinaryType(), True)
])
df = spark.createDataFrame(records, schema=schema)
records.clear()
df.persist(StorageLevel.DISK_ONLY)
df.write.partitionBy("rdw").mode("overwrite").parquet("data", compression="snappy")
df.unpersist(True)
spark.stop()
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
py_visual_cobol-1.1.3.tar.gz
(9.0 kB
view details)
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file py_visual_cobol-1.1.3.tar.gz.
File metadata
- Download URL: py_visual_cobol-1.1.3.tar.gz
- Upload date:
- Size: 9.0 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: poetry/1.8.4 CPython/3.11.9 Windows/10
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
40e031ad7687353283ed4cb78ada641e81496830369b29e34cab1d6830bfd4f5
|
|
| MD5 |
ba7d2356cef123422354710742e46ffb
|
|
| BLAKE2b-256 |
4128d7f849466d6adcfbeeea71dda9540f6ba898f2deb1730e95a57ce0d22457
|
File details
Details for the file py_visual_cobol-1.1.3-py3-none-any.whl.
File metadata
- Download URL: py_visual_cobol-1.1.3-py3-none-any.whl
- Upload date:
- Size: 6.8 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: poetry/1.8.4 CPython/3.11.9 Windows/10
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
778e17631987f1f5a276c7ee799bd305ef31cb3d8abcfbff88563117e57e4be5
|
|
| MD5 |
fa3d99b53814bbc3b4def1efcec24345
|
|
| BLAKE2b-256 |
b71900e5bfa6bdfb916f050225e18adf3d82f7f0ee6f338dae3066ab7b1e97a0
|