Dataset validator and Slack notifier for NeuralForgeAI
Project description
🛠️ NeuralForgeAI Data Prep (wyoloservice-data-prep)
A robust, enterprise-grade CLI utility and Python library dedicated to validating YOLO datasets and dispatching real-time notifications across the NeuralForgeAI ecosystem. Built by William Rodriguez (Wisrovi), this library acts as the foundational data gatekeeper before massive GPU compute resources are allocated for hyperparameter sweeps.
📖 Table of Contents
- Introduction
- Why Data Validation Matters
- Core Features
- Installation Guide
- CLI Usage & Examples
- Python API Usage
- Integration with CI/CD & Celery
- Slack Notifications
- Data Augmentation Roadmap
- Troubleshooting
- About the Author
🌟 Introduction
Training state-of-the-art computer vision models—whether for bounding box detection, semantic segmentation, or image classification—requires strictly formatted datasets. A missing val folder, an unmapped class in a .yaml file, or corrupted image streams can crash a distributed Optuna training sweep hours after it begins, wasting valuable GPU hours and engineering time.
The wyoloservice-data-prep library solves this by providing a lightweight, headless package that can be deployed anywhere—on raw Linux servers, inside Jenkins pipelines, or baked into Docker images. It validates datasets strictly according to Ultralytics YOLOv8 standards.
🎯 Why Data Validation Matters
In an automated MLOps environment (like NeuralForgeAI), training jobs are queued asynchronously. Without pre-validation:
- Jobs fail silently in remote celery workers.
- Logs are polluted with tracebacks from missing files.
- Teams lose track of which datasets are actually production-ready.
By using this tool, you enforce a strict "Shift-Left" data quality check. Datasets that pass validation immediately trigger a Slack notification to your team, signaling that they are clean and queued for NeuralForge.
✨ Core Features
- Deep YOLO Architecture Validation:
- Classification: Scans directories to ensure explicit
train/andval/subdirectories exist. Dynamically maps subdirectory names to class labels. - Detection & Segmentation: Parses
dataset.yamlfiles. Validates thattrain,val,nc(number of classes), andnamesfields exist. Verifies that the internal absolute/relative paths point to valid filesystem targets.
- Classification: Scans directories to ensure explicit
- Headless & CI/CD Ready: Designed to run without interactive prompts, making it perfect for GitHub Actions, GitLab CI, Jenkins, or cron jobs.
- Slack Integrations: Provides native, robust hooks to notify engineering teams across Slack channels. Custom payloads can alert managers of class imbalances or dataset readiness.
- Extensible Augmentation Hooks:
Includes structured stubs for balancing datasets dynamically using
albumentationsandopencv-python-headlessbefore the actual YOLO engine touches them.
⚙️ Installation Guide
The library is published on PyPI. It requires Python 3.10 or higher.
From PyPI (Recommended)
pip install wyoloservice-data-prep
From Source (Development)
git clone https://github.com/wisrovi/wyoloservice2_data_prep.git
cd wyoloservice2_data_prep
pip install -e .
Installing the package automatically registers the global CLI command wyolo-validate.
🚀 CLI Usage & Examples
The easiest way to use the library is through its built-in Command Line Interface.
Validating a Detection Dataset
If you have a standard YOLO detection dataset defined by a data.yaml file:
wyolo-validate --yaml /mnt/samba/datasets/arepo/cicatrices/data.yaml
Expected Output:
{
"valid": true,
"task": "detect",
"classes": 3,
"names": ["scratch", "dent", "rust"],
"train_path": "/mnt/samba/datasets/arepo/cicatrices/images/train"
}
Validating a Classification Dataset
For classification, YOLO expects a raw directory structure (without a yaml). Pass the root directory to the tool:
wyolo-validate --yaml /datasets/AIDIAGNOST/classification/ages_classification/
Expected Output:
{
"valid": true,
"task": "classify",
"classes": 4,
"names": ["ANCIANO", "ADULTO", "JOVEN", "NINO"],
"train_path": "/datasets/AIDIAGNOST/classification/ages_classification/train"
}
💻 Python API Usage
You can also import the logic directly into your own Python microservices or FastAPI endpoints.
from wyolo_data_prep.validator import check_yolo_dataset
# Validate detection
result_det = check_yolo_dataset("/path/to/detection/dataset.yaml", task_type="detect")
if result_det["valid"]:
print(f"Dataset is ready! Classes found: {result_det['names']}")
else:
print(f"Validation failed: {result_det['error']}")
# Validate classification
result_cls = check_yolo_dataset("/path/to/classification/dir", task_type="classify")
if result_cls["valid"]:
print(f"Found {result_cls['classes']} classes for classification.")
🔗 Integration with CI/CD & Celery
Within NeuralForge Celery Workers
This library is designed to be bundled inside your wisrovi/train_service worker images. Before the Celery task executes yolo train ..., it should call the python API to verify the mounted CIFS dataset isn't corrupted.
Notification Triggers
Combine validation with the Slack notifier to build a complete pipeline:
from wyolo_data_prep.validator import check_yolo_dataset
from wyolo_data_prep.notifier import send_slack_alert
dataset_path = "/path/to/dataset.yaml"
result = check_yolo_dataset(dataset_path)
if result["valid"]:
send_slack_alert(f"✅ Dataset {dataset_path} passed validation and is ready for training. Classes: {result['classes']}")
else:
send_slack_alert(f"❌ CRITICAL: Dataset validation failed for {dataset_path}. Error: {result['error']}")
📝 License
This project is open-sourced under the MIT License.
Author: William Rodriguez (Wisrovi) AI Leader & Solutions Architect
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file wyoloservice_data_prep-0.3.0.tar.gz.
File metadata
- Download URL: wyoloservice_data_prep-0.3.0.tar.gz
- Upload date:
- Size: 6.7 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.13.5
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
2a5e23f4dca525022d822bb00a8f06cb4f9e3c7eb870fce22dc0abe4e4874eaf
|
|
| MD5 |
a4be075f855051e64bb560571a948189
|
|
| BLAKE2b-256 |
ab4d2862d330d9f0d5b91f75deda959585ee5efae6259a3c33ebe03d32330963
|
File details
Details for the file wyoloservice_data_prep-0.3.0-py3-none-any.whl.
File metadata
- Download URL: wyoloservice_data_prep-0.3.0-py3-none-any.whl
- Upload date:
- Size: 7.5 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.13.5
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
c524008de7d5063c3b41d80f17f93569d114a6777fa6b2caa9f7add2fa937ae6
|
|
| MD5 |
76b4be317412ece5817c66eb504618c9
|
|
| BLAKE2b-256 |
4f4dca55bcc15f3bdb3cb815652ac41ec4a37287bd4ba5858fc14b5b4e682b31
|