Skip to main content

AutoDataMind 🧠

Zero-Code Automated Data Science

PyPI version Python License

AutoDataMind is your Native Intelligence Layer for data science. No Pandas knowledge required. No ML expertise needed. Just simple function calls for complete automation.

🌍 Philosophy

Data science should be accessible to everyone, everywhere.

AutoDataMind democratizes data science for:

  • Emerging Markets: Africa, Asia, Latin America
  • Small Businesses: No data science budget
  • Students: Learning without complexity
  • Non-Technical Users: Business analysts, researchers

✨ Features

🎯 Zero-Code Automation

import autodatamind as adm

# Automatic data analysis - ONE line
adm.analyze("sales.csv")

# Automatic ML training - ONE line
model = adm.autotrain("sales.csv", target="revenue")

# Automatic dashboard - ONE line
adm.dashboard("sales.csv")

# Automatic deep learning - ONE line
dl_model = adm.auto_deep("data.csv", target="category")

# Automatic insights report - ONE line
report = adm.generate_insights("sales.csv")

🤖 6 Native Intelligence Agents

  1. DataAgent: Universal data loading (CSV, Excel, JSON, Parquet)
  2. ProfileAgent: Automatic data profiling and analysis
  3. VizAgent: Beautiful HTML dashboards
  4. MLAgent: Automatic machine learning
  5. DLAgent: Automatic deep learning (PyTorch)
  6. InsightAgent: Natural language insights

🚀 What AutoDataMind Does

  • Loads Data: CSV, Excel, JSON, Parquet - auto-detected
  • Cleans Data: Duplicates, missing values, outliers - automatic
  • Analyzes Data: Statistics, correlations, insights - comprehensive
  • Visualizes Data: HTML dashboards - professional
  • Trains ML Models: Regression, classification - auto-selected
  • Trains Deep Models: Neural networks - auto-built
  • Generates Reports: Narratives, recommendations - human-readable

📦 Installation

pip install autodatamind

🎓 Quick Start

Analyze Any Dataset

import autodatamind as adm

# Load and analyze - returns complete analysis
analysis = adm.analyze("your_data.csv")

# Access results
print(analysis['overview'])
print(analysis['statistics'])
print(analysis['insights'])

Train ML Model - Zero Code

# Automatic ML training
result = adm.autotrain("sales.csv", target="revenue")

# Get trained model
model = result['model']

# Get metrics
print(result['metrics'])
# {'rmse': 1234.56, 'mae': 987.65, 'r2': 0.89}

# Model saved automatically!

Create Dashboard - One Line

# Generate professional HTML dashboard
adm.dashboard("sales.csv")
# Opens in browser automatically!

Deep Learning - No PyTorch Knowledge

# Automatic deep learning
result = adm.auto_deep(
    "data.csv",
    target="category",
    epochs=50
)

# Get model and metrics
model = result['model']
print(result['metrics'])
# {'accuracy': 0.95}

Get Business Insights

# Generate narrative report
report = adm.generate_insights("sales.csv", target="revenue")

# Report includes:
# - Executive summary
# - Key findings
# - Statistical insights
# - Recommendations
# - Data quality assessment

📊 Complete Example

import autodatamind as adm

# 1. Load data (auto-detected format)
df = adm.read_data("sales.csv")

# 2. Clean data (automatic)
df_clean = adm.autoclean(df)

# 3. Analyze data
analysis = adm.analyze(df_clean)

# 4. Create dashboard
adm.dashboard(df_clean)

# 5. Train ML model
ml_result = adm.autotrain(df_clean, target="revenue")

# 6. Train deep learning model
dl_result = adm.auto_deep(df_clean, target="revenue", epochs=100)

# 7. Generate insights report
report = adm.generate_insights(df_clean, target="revenue")

# Done! 🎉

🎯 Use Cases

Business Analytics

# Analyze sales data
adm.analyze("sales_2024.csv")
adm.dashboard("sales_2024.csv")
adm.generate_insights("sales_2024.csv", target="total_sales")

Predictive Modeling

# Predict customer churn
result = adm.autotrain("customers.csv", target="churn")
print(f"Model accuracy: {result['metrics']['accuracy']:.2%}")

Data Exploration

# Explore new dataset
adm.analyze("new_data.csv")  # Get overview
adm.dashboard("new_data.csv")  # Visual exploration

Report Generation

# Generate executive report
report = adm.generate_insights(
    "quarterly_data.csv",
    target="profit",
    save_report=True
)

🏗️ Architecture

AutoDataMind uses 6 Native Intelligence Agents:

autodatamind/
├── core/               # Core functionality
│   ├── reader.py      # Universal data loader
│   ├── cleaner.py     # Automatic data cleaning
│   ├── utils.py       # Helper functions
│   └── validator.py   # Data validation
├── agents/            # Intelligence agents
│   ├── data_agent.py       # Data handling
│   ├── profile_agent.py    # Analysis & profiling
│   ├── viz_agent.py        # Visualization
│   ├── ml_agent.py         # Machine learning
│   ├── dl_agent.py         # Deep learning
│   └── insight_agent.py    # Narrative generation
└── models/            # ML/DL engines
    ├── auto_ml.py     # AutoML engine
    └── auto_dl.py     # AutoDL engine

💡 Philosophy: Native Intelligence Layer

Traditional Data Science:

# 40 lines of Pandas code
import pandas as pd
df = pd.read_csv("data.csv")
df = df.dropna()
df = df.drop_duplicates()
# ... 35 more lines ...

AutoDataMind:

# 1 line
adm.analyze("data.csv")

No Pandas knowledge required. No ML expertise needed.

🌍 Target Markets

Emerging Economies

  • Africa: Kenya, Nigeria, South Africa, Ghana
  • Asia: India, Bangladesh, Philippines, Vietnam
  • Latin America: Brazil, Mexico, Colombia

User Segments

  • Students: Learn data science without complexity
  • Small Businesses: No data science budget
  • Researchers: Focus on insights, not code
  • Analysts: Fast results without programming

🔬 Technical Details

Supported Data Formats

  • CSV: Auto-encoding detection (UTF-8, Latin-1, ISO-8859-1)
  • Excel: .xlsx, .xls
  • JSON: Multiple orientations
  • Parquet: High-performance columnar

Auto-Cleaning Features

  • Duplicate removal
  • Missing value handling (auto/drop/mean/median/mode)
  • Type fixing
  • Outlier removal
  • Data validation

ML Algorithms

  • Classification: RandomForest, GradientBoosting, Logistic Regression, KNN, Naive Bayes
  • Regression: RandomForest, GradientBoosting, Linear Regression, Ridge, Lasso

DL Architectures

  • MLP: Simple multilayer perceptron
  • Deep: Deep neural networks (128→64→32)
  • Wide: Wide networks (256→128→64)
  • Auto-selection based on data size

📈 Performance

Speed

  • Small datasets (<10K rows): <1 second
  • Medium datasets (10K-100K): <5 seconds
  • Large datasets (>100K): <30 seconds

Accuracy

  • AutoML: Competitive with manual tuning
  • AutoDL: State-of-the-art architectures
  • Auto-hyperparameter tuning: GridSearchCV optimization

🤝 Contributing

Contributions welcome! Areas of interest:

  • Additional data formats
  • More ML algorithms
  • Advanced DL architectures
  • New visualization types
  • Documentation improvements

📄 License

MIT License - see LICENSE file.

👤 Author

Idriss Olivier Bado

🙏 Acknowledgments

Built with:

  • pandas: Data manipulation
  • scikit-learn: Machine learning
  • PyTorch: Deep learning
  • matplotlib/seaborn: Visualization

🔗 Links


Made with ❤️ for the global data science community

Democratizing AI, one line of code at a time

Release files for autodatamind 0.1.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Built distribution (wheel)

Table of built distributions (wheels) for autodatamind 0.1.1
File Interpreter ABI Platform
autodatamind-0.1.1-py3-none-any.whl Python 3 none any Details

Release files / autodatamind-0.1.1-py3-none-any.whl

Download URL autodatamind-0.1.1-py3-none-any.whl
Size 36.7 kB
Tags Python 3
SHA-256 checksum
How to use checksums
8d67b4c7e39c101f79fe817753dbf23e94bd0fe8c5afc83d20a588eb2da76364
BLAKE2b-256 checksum
How to use checksums
b2bef6e161a0bdad4807ca11c6acb74373fd2e7cc67a7bf86d505140f1f2dbd6
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.12.0

Release history Release notifications | RSS feed

This release

0.1.1 This release

1 release file

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page