Skip to main content

ModelScout 🤖

Intelligent ML Model Recommendation System

An automated machine learning tool that analyzes your dataset and recommends the best-fitting ML models. ModelScout uses Auto-sklearn to intelligently search through a vast hyperparameter space and identifies optimal models for your specific data.

🎯 Features

  • Automated Problem Detection: Automatically detects classification, regression, or clustering tasks
  • Smart Model Selection: Uses Auto-sklearn to find the best models for your data
  • Comprehensive Analysis: Provides detailed dataset analysis and insights
  • Multiple Formats: Generates reports in text, JSON, and table formats
  • REST API: Flask-based REST API for easy integration
  • Support for All ML Tasks: Classification, Regression, Time-series, and more

📋 Project Structure

ModelScout/
├── agent/                    # Core ML engine
│   ├── data_analyzer.py     # Dataset analysis module
│   ├── model_selector.py    # Model recommendation using Auto-sklearn
│   ├── reporter.py          # Report generation
│   ├── orchestrator.py      # Main pipeline orchestrator
│   └── __init__.py
├── api/                      # REST API
│   ├── main.py              # Flask API endpoints
│   └── __init__.py
├── data/                     # Sample datasets
├── models/                   # Trained models storage
├── outputs/                  # Generated reports
├── requirements.txt         # Python dependencies
├── demo.py                  # Demo script with examples
└── README.md

🚀 Quick Start

1. Installation

# Clone or navigate to the project directory
cd ModelScout

# Create virtual environment
python -m venv venv
source venv/bin/activate  # On Windows: venv\Scripts\activate

# Install dependencies
pip install -r requirements.txt

2. Basic Usage

from agent.orchestrator import ModelScout

# Initialize
scout = ModelScout(auto_train_time=300)

# Run complete pipeline
result = scout.run_full_pipeline(
    data_path='your_data.csv',
    target='target_column',
    report_path='outputs/report.txt'
)

# Access results
print(result['recommendations']['best_model_name'])
print(result['recommendations']['test_score'])
print(result['report'])

3. Step-by-Step Usage

from agent.orchestrator import ModelScout
import pandas as pd

scout = ModelScout()

# Load data
df = scout.load_data('data.csv')

# Analyze data
analysis = scout.analyze_data(df, target='label')
print(f"Problem Type: {analysis['target_analysis']['type']}")

# Get recommendations
recommendations = scout.recommend_models(df, 'label')
print(f"Best Model: {recommendations['best_model_name']}")
print(f"Test Score: {recommendations['test_score']}")

# Generate report
report = scout.generate_report(output_format='text', output_path='report.txt')

🔧 API Endpoints

Health Check

GET /health

Analyze Dataset

POST /api/analyze
Content-Type: application/json

{
    "file_path": "path/to/data.csv",
    "target": "target_column"
}

Get Recommendations

POST /api/recommend
Content-Type: application/json

{
    "file_path": "path/to/data.csv",
    "target": "target_column",
    "time_limit": 300
}

Generate Report

POST /api/report
Content-Type: application/json

{
    "file_path": "path/to/data.csv",
    "target": "target_column",
    "format": "text"
}

Full Pipeline

POST /api/pipeline
Content-Type: application/json

{
    "file_path": "path/to/data.csv",
    "target": "target_column",
    "time_limit": 300
}

🎮 Run Demo

python demo.py

The demo script:

  1. Creates sample datasets (Iris, Breast Cancer, Regression)
  2. Runs ModelScout on each dataset
  3. Generates comparison reports
  4. Demonstrates both classification and regression

📊 What ModelScout Analyzes

Data Characteristics

  • Dataset size and shape
  • Missing values and data quality
  • Feature types and counts
  • Memory usage

Target Variable

  • Problem type (Classification/Regression)
  • Class distribution (for classification)
  • Value range (for regression)
  • Class imbalance ratio

Feature Statistics

  • Numeric: mean, std, min, max, missing count
  • Categorical: unique values, missing count

🤖 How It Works

  1. Data Loading: Supports CSV, Excel, JSON formats
  2. Analysis: Comprehensive dataset profiling
  3. Problem Detection: Auto-detects ML task type
  4. Model Search: Auto-sklearn searches optimal models
  5. Evaluation: Train/test split and performance metrics
  6. Reporting: Generates detailed recommendations

📦 Dependencies

  • pandas: Data manipulation
  • scikit-learn: ML algorithms
  • auto-sklearn: Automated ML model selection
  • numpy: Numerical computing
  • matplotlib/seaborn: Visualization
  • flask: REST API
  • xgboost, lightgbm, catboost: Advanced models
  • imbalanced-learn: Class imbalance handling

🔍 Example Output

======================================================================
  ___  ___           _      _    ____  ___  _   _ ___
 |  \/  |          | |    | |  / ___ \/ _ \| | | |_  |
 | .  . | ___    __| | ___| | / /   \/ /_\ \ | | | / /
 | |\/| |/ _ \  / _` |/ _ \ | \ \   |  _  | | | |/ /
 | |  | | (_) || (_| |  __/ |  \ \__| | | | |_| / /
 |_|  |_|\___/  \__,_|\___|_|   \___/_| |_|\___/___/

======================================================================

DATA OVERVIEW
======================================================================
Dataset Shape: (150, 5) (rows, columns)
Memory Usage: 0.00 MB
Missing Values: 0 (0.00%)
Numeric Features: 4
Categorical Features: 0

TARGET VARIABLE ANALYSIS
----------------------------------------------------------------------
Problem Type: CLASSIFICATION
Unique Values: 3
Missing Values: 0
Class Imbalance Ratio: 1.00:1
Class Distribution:
  0: 50 (33.3%)
  1: 50 (33.3%)
  2: 50 (33.3%)

MODEL RECOMMENDATIONS
======================================================================
Best Model: RandomForestClassifier
Problem Type: CLASSIFICATION
Train Score: 1.0000
Test Score: 0.9333
Data Shape Used: (150, 4)
Number of Classes: 3

======================================================================

🛠️ Configuration

You can customize behavior by modifying parameters:

scout = ModelScout(
    auto_train_time=600  # Increase for more thorough search (seconds)
)

📝 License

This project is for educational and portfolio purposes.

🤝 Contributing

Feel free to extend ModelScout with:

  • Additional models
  • More data preprocessing options
  • Visualization enhancements
  • Performance optimizations

📞 Support

For issues or questions, refer to the demo.py script for usage examples.


Happy Model Scouting! 🎯

Metadata

Release files for modelscout-ai 0.1.2

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for modelscout-ai 0.1.2
File Size Uploaded
modelscout_ai-0.1.2.tar.gz 51.3 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for modelscout-ai 0.1.2
File Interpreter ABI Platform
modelscout_ai-0.1.2-py3-none-any.whl Python 3 none any Details

Total release size: 110.5 kB

Release files / modelscout_ai-0.1.2.tar.gz

Download URL modelscout_ai-0.1.2.tar.gz
Size 51.3 kB
Tags Source
SHA-256 checksum
How to use checksums
a4d4122af79c30bd5af8ea8e18440548e4f788d3732bf160d662c251fbfe856d
BLAKE2b-256 checksum
How to use checksums
d643e8b9b9d1a7d4f7df59373c2b055285899006c78f77097b75d174d995fb13
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.12.5

Release files / modelscout_ai-0.1.2-py3-none-any.whl

Download URL modelscout_ai-0.1.2-py3-none-any.whl
Size 59.1 kB
Tags Python 3
SHA-256 checksum
How to use checksums
a0ecb897716cab5625840523a992eae353ed6a79504f7248171b0fcc37072782
BLAKE2b-256 checksum
How to use checksums
d1d37e64de3bcc9696f3b4db823567e1c66574d640178c18f0423c4e6de9cce4
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.12.5

Release history Release notifications | RSS feed

This release

0.1.2 This release

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page