Skip to main content

🤖 DE Agentic - AI-Powered Data Engineering Assistant

▶ Start here

Thirty seconds, no install:

python3 de.py demo

Creates a sample database, profiles a file, runs a query, reads the schema and answers a question — standard library only. No API key, no pip install.

Prefer a de command on your PATH?

python3 -m pip install -e .   # then: de demo

The four commands:

python3 de.py profile customers.csv                      # profile a file
python3 de.py query  "SELECT * FROM customers LIMIT 5"   # read-only SQL
python3 de.py schema                                     # tables + relationships
python3 de.py ask   "how do I find null values?"         # model, or offline help

Plug in a real model (llama.cpp, Ollama, OpenAI or Anthropic) whenever you want:

python3 de.py doctor      # what is reachable?
python3 de.py setup       # detect one and write .env
python3 de.py demo --ai   # full demo, step 5 answered by your model
python3 de.py tc          # end-to-end test: the model writes SQL, we run it

llama.cpp, Ollama and OpenAI all speak the same OpenAI-compatible /v1/chat/completions, so one code path covers them. Precedence is environment > .env > config.yaml > defaults.

Cloud warehouses (ClickHouse Cloud and Snowflake)

The default remains a local database. To use a configured warehouse, put non-secret connection settings in config.yaml and keep the password in an environment variable:

database:
  type: clickhouse_cloud       # or: snowflake
  host: your-service.clickhouse.cloud
  port: 8443                   # Snowflake commonly uses 443
  database: analytics
  username: default
  password_env: CLICKHOUSE_PASSWORD
  secure: true
  # Snowflake also accepts warehouse, role and schema.

Then export the password and use --db configured:

export CLICKHOUSE_PASSWORD='...'
python3 -m pip install clickhouse-connect
python3 de.py profile --db configured
python3 de.py schema --db configured
python3 de.py query "SELECT count() FROM events" --db configured

For Snowflake, install snowflake-connector-python and set SNOWFLAKE_PASSWORD (or your chosen environment-variable name). Passwords are never stored in YAML; the YAML stores only the variable name.

Supported types are sqlite, duckdb, clickhouse_cloud and snowflake — there is no Postgres, MySQL or MongoDB support. Snowflake additionally accepts warehouse, role and schema.

Run these commands from the directory holding your config.yaml: the database: block is resolved relative to your working directory, so it will silently stop applying if you cd somewhere else.

📖 Full reference, including every key and the LLM side: docs/CONFIGURATION.md

📖 Full getting-started guide: QUICKSTART.md

Which docs are current.

Start with Quick Start and Usage above — they are verified against a real install.

Current reference docs: docs/CONFIGURATION.md (models and databases), docs/API.md (Python), docs/EXTENDING.md, and docs/TROUBLESHOOTING.md.

docs/COMMAND_REFERENCE.md, docs/TESTING.md, docs/QUICK_TEST_REFERENCE.md, docs/DEPLOYMENT.md, docs/MODEL_SELECTION.md, docs/AGENT_MODES_OPTIONS.md and docs/FIX_CLI_OPTIONS.md still document the legacy python -m src.cli run ... interface, which needs the optional legacy extra. They are kept for reference while that tree is migrated.

docs/MIGRATION.md is the guide for moving off them, and docs/QUALITY_REPORT.md / docs/USER_TESTING.md are dated records of the simplification work, not living instructions.


A comprehensive agentic system designed to assist data engineers with daily tasks including data ingestion, modeling, quality checks, warehousing, debugging, reverse engineering, and architecture simplification.

🆕 Now with Local LLM Support! Run completely offline with Ollama, or use OpenAI/Anthropic for cloud-based models. Switch between models per command for cost/quality optimization.

!Note: This is the part of README.md of Data-Agent Project, the project is covering for all topics related to data engineering. The project is still in development, but I would like to publish the documentation for reference and idea brainstorming into the to accelerating the working.

System Overview

System Overview

To understand how the project is going to resolve, run the demo below:

python ./examples/demo_local.py

python ./examples/demo_simple.py

🌟 Features

Task Categories

  • Data Ingestion: Automated data loading from various sources (APIs, databases, files)
  • Data Modeling: Schema design, ERD generation, normalization checks
  • Data Quality: Profiling, validation, anomaly detection
  • Warehousing: Pipeline generation, optimization, dbt assistance
  • Debugging: Error analysis, log parsing, performance profiling
  • Reverse Engineering: Schema extraction, lineage tracking, documentation generation
  • Architecture: Pattern detection, optimization recommendations, diagram generation

Execution Modes

  • 🚀 Core Modes (Free, Instant, No LLM Required):

    • query: Execute SQL queries on local database
    • reverse: Analyze database schema and generate documentation
    • profile: Profile data files with comprehensive statistics
    • interactive: Command-based system for database operations
  • 🤖 Agent Modes (LLM-Powered, Flexible Model Selection):

    • debug: Analyze errors and provide debugging guidance
    • quality: Data quality assessment with recommendations
    • model: Data modeling assistance
    • ingest: Data ingestion strategy and code generation
    • warehouse: Data warehouse architecture guidance
    • architect: System architecture analysis and optimization

Mental Model Principles

  • ReAct (Reasoning + Acting): Combines reasoning and action for better decision-making
  • Chain of Thought: Step-by-step reasoning for complex problems
  • Planning: Task decomposition and strategic planning
  • Reflection: Self-evaluation and continuous improvement
  • Memory: Context retention across interactions

LLM Flexibility

  • Multiple Providers: OpenAI (GPT-4, GPT-4o-mini), Ollama (llama2, mistral, codellama), Anthropic (Claude)
  • Per-Command Model Selection: Choose optimal model for each task using --model parameter
  • Cost Optimization: Use cheap models for daily work, premium for critical tasks
  • Local-First: Default to Ollama for free, offline operation

🚀 Quick Start

Prerequisites

  • Python 3.9 or newer
  • Optional: Ollama or another local model server
  • Optional: an API key for OpenAI or Anthropic
  • Optional: Docker, for the container image

Install

pip install de-agentic

That is the whole thing. The base install depends only on pyyaml and pydantic; profile, query, schema and init-db then run on the standard library alone. Two extras are available:

pip install "de-agentic[pretty]"   # rich tables (falls back to plain text)
pip install "de-agentic[data]"     # DuckDB files, Parquet/Excel profiling

de is now on your PATH. Check it:

de --version

Prefer not to install anything? python3 de.py demo runs the same walkthrough from a clone using only the standard library.

Create the sample database

The commands below read sample.sqlite, which is generated on demand:

de init-db      # customers, products, orders, order_items

Everyday commands

de profile customers.csv                       # column types, nulls, ranges
de query  "SELECT * FROM customers LIMIT 5"    # read-only SQL
de schema                                      # tables, columns, relationships
de ask    "how do I find null values?"         # answered by your model

ask works offline too: without a model it falls back to local guidance and tells you how to configure one.

Connect a model

de doctor      # which endpoints are reachable?
de setup       # detect one and write .env

Then re-run de ask, or try the full walkthrough with a real model:

de demo        # end to end
de demo --ai   # same, with step 5 answered by your model

llama.cpp, Ollama and OpenAI all speak the same OpenAI-compatible /v1/chat/completions, so one code path covers them. Precedence is environment > .env > config.yaml > defaults.

Per-command overrides are available when you want to switch providers without editing configuration:

de ask "..." --provider openai --model gpt-4o-mini

Using Docker

Quick Start with Docker Compose

# Build and start all services (app + postgres + ollama)
docker compose up -d

# Check status
docker compose ps

# View logs
docker compose logs -f de-agentic

Test the Deployment

This stack exists for the legacy agent interface, which needs the optional legacy extra that the base image does not install. If you only want the de commands, use the single-container form below instead.

# Interactive shell with the CLI (requires -it)
docker compose exec -it de-agentic de --help

# Create the sample database and inspect it
docker compose exec de-agentic de init-db
docker compose exec de-agentic de schema
docker compose exec de-agentic de query "SELECT * FROM customers LIMIT 5"

Single Container (No Dependencies)

The image installs this package from source, so the commands inside the container are the same de commands:

# Build the image
docker build -t de-agentic .

# Standard-library walkthrough, no model required
docker run --rm de-agentic de demo

# Or use the CLI directly
docker run --rm de-agentic de schema
docker run --rm de-agentic de query "SELECT 1 AS x"

# Profile a local file (mount volume)
docker run --rm -v "$PWD/customers.csv:/data/customers.csv" \
  de-agentic de profile /data/customers.csv

Note: the container's default command is de demo, not the legacy python -m src.cli, which needs optional dependencies the base image does not install.

📚 For complete deployment guide, see docs/DEPLOYMENT.md

📖 Usage

Command Line Interface

Core modes (free, instant, no model required)

These run on the standard library and need no model server:

de init-db                                       # create sample.sqlite
de query  "SELECT COUNT(*) FROM customers"       # read-only SQL
de schema                                        # tables, columns, relationships
de profile customers.csv                         # column types, nulls, ranges

Model-backed commands

ask and demo --ai route through whichever provider is configured. Use de doctor to see what is reachable and de setup to configure one.

de ask "which tables have no primary key?"       # answered by your model
de demo --ai                                     # full walkthrough, model answers
de tc                                            # end-to-end test of the SQL path

Select a provider per command without editing any config file:

de ask "..." --provider openai --model gpt-4o-mini
de ask "..." --provider ollama  --model llama3.2:1b

📖 docs/CONFIGURATION.md is the full reference: which file holds what, every LLM_* and database: key, ClickHouse Cloud and Snowflake, and how to check what resolved. Worth reading once — the database: block is read from your current directory while .env is read from the install directory, which surprises people.

Legacy CLI

The earlier agent-mode interface (run query, run debug, run reverse, ...) is still in the tree as src.cli, but it is mid-migration and needs the optional legacy extra:

pip install "de-agentic[legacy]"
de-agentic --help

Running de-agentic without that extra prints one line naming what is missing and exits with status 2 — it does not raise a traceback. de itself needs no extra and is unaffected.

It is not required for anything documented above. New code should target the de commands.

Legacy agent modes

Available only with the legacy extra, via de-agentic. Listed for reference; these are mid-migration and not covered by CI.

de-agentic run debug   --error="Connection timeout"
de-agentic run quality --file=data.csv --model=gpt-4o-mini
de-agentic run model   --description="E-commerce system"
de-agentic run ingest  --source="API" --target="postgres"
de-agentic run warehouse --requirements="Real-time analytics"
de-agentic run architect --description="Current pipeline"

Per-command --model works the same way here as with de.

Python API

from src.agents.de_agent import DEAgent
from src.tasks import DataQualityTask

# Initialize agent
agent = DEAgent()

# Run data quality check
task = DataQualityTask(
    file_path="customers.csv",
    profile=True,
    validate=True
)

result = agent.execute(task)
print(result)

Available Models

Provider Model Cost Best For
OpenAI gpt-4o-mini $0.15/1M tokens Daily work, fast responses
gpt-4 $30/1M tokens Critical issues, complex problems
gpt-4-turbo $10/1M tokens Balance of cost/quality
Ollama llama2 Free General purpose, offline
mistral Free Better code understanding
codellama Free Code generation, debugging
Anthropic claude-3.5-sonnet $3/1M tokens Long context, analysis
claude-3-opus $15/1M tokens Premium quality

🏗️ Architecture

de-agentic/
├── src/
│   ├── agents/          # Agent implementations
│   ├── tasks/           # Task definitions
│   ├── skills/          # Reusable skills
│   ├── tools/           # Integration tools
│   ├── workflows/       # Workflow orchestration
│   └── utils/           # Utilities
├── config/              # Configuration files
├── examples/            # Usage examples
├── tests/               # Test suite
└── docs/                # Documentation

🔧 Configuration

LLM Provider Configuration (.env)

# Default: Local Ollama (free, offline)
LLM_PROVIDER=ollama
OLLAMA_BASE_URL=http://localhost:11434
OLLAMA_MODEL=llama2

# OpenAI (requires API key)
LLM_PROVIDER=openai
OPENAI_API_KEY=sk-proj-...
OPENAI_MODEL=gpt-4o-mini

# Anthropic (requires API key)
LLM_PROVIDER=anthropic
ANTHROPIC_API_KEY=sk-ant-...
ANTHROPIC_MODEL=claude-3.5-sonnet

Agent Configuration (config/agent_config.yaml)

Edit config/agent_config.yaml to customize:

  • Enable/disable specific tasks and skills
  • Adjust mental model parameters
  • Configure database connections
  • Set execution limits
  • Fine-tune agent behavior

Database Setup

The demo.db DuckDB database includes:

  • customers table: 10 sample customers
  • orders table: 30 sample orders
  • products table: 10 sample products
  • Total revenue: $6,813.97

Recreate anytime with: de init-db --force

🤝 Contributing

Contributions are welcome! Please read our contributing guidelines first.

📚 Documentation

Comprehensive guides available in the docs/ directory:

🎯 Use Cases

Daily Data Engineering Tasks

  • Query databases without remembering SQL syntax
  • Profile new data files instantly
  • Debug pipeline errors with AI assistance
  • Generate data models from requirements

Cost Optimization Strategy

# Development/Testing (free, local)
--model=llama2

# Daily work (cheap, cloud)
--model=gpt-4o-mini

# Critical issues (premium, best quality)
--model=gpt-4

Example workflow

# 1. Profile a new file (free, instant, no model)
de profile new_data.csv

# 2. Read the schema
de schema

# 3. Run a check with a cheap model
de ask "does new_data.csv have quality issues?" --model gpt-4o-mini

# 4. Escalate to a stronger model for a hard question
de ask "how should I fix the encoding problem?" --model gpt-4

# 5. Verify with SQL (free, instant)
de query "SELECT count(*) FROM new_data"

🚀 Recent Updates

v2.1.2 - Bug fixes and configuration docs

  • ✅ de-agentic no longer dies with a ModuleNotFoundError traceback when the legacy extra is absent; it names the missing packages, prints the pip install line and exits 2
  • ✅ Table headers are no longer truncated with U+2026, which rendered as customer… on non-UTF-8 Windows consoles; values wrap instead
  • ✅ de init-db --db configured creates the configured file instead of one literally named configured
  • ✅ database.path now defaults to sample.sqlite, matching the CLI, so --db configured points at the file init-db creates
  • ✅ Removed a duplicated path/timeout declaration in DatabaseConfig
  • ✅ .env.example rewritten: it advertised Postgres, MySQL, MongoDB and HuggingFace, none of which are supported
  • ✅ New docs/CONFIGURATION.md covering both model and database setup, and it documents which config.yaml sections are read at runtime
  • ✅ test_config.py is isolated from the developer's shell; exported LLM_* variables no longer fail it

v2.1.1 - Documentation accuracy

  • ✅ Recent Updates now covers the 2.1.x line instead of stopping at v1.2.0
  • ✅ Added the MIT LICENSE file that pyproject.toml and the README referenced
  • ✅ Legacy reference docs are banner-marked and point at MIGRATION.md
  • ✅ Release workflow publishes only for strict vX.Y.Z tags

v2.1.0 - Packaging, CI and install docs

  • ✅ GitHub Actions: lint, a 3.9-3.12 test matrix, and a build that verifies the wheel
  • ✅ Tag-driven release: publish to PyPI via trusted publishing, plus a GitHub release
  • ✅ The wheel now declares its dependencies; import core.config works after install
  • ✅ harness data files and the src package ship correctly
  • ✅ de doctor no longer crashes when no model is configured
  • ✅ README and Dockerfile now match what actually installs and runs

v1.2.0 - Model Selection Feature

  • ✅ Added --model parameter for per-command model selection
  • ✅ Auto-provider detection (gpt* → openai, claude* → anthropic)
  • ✅ Cost optimization through flexible model selection
  • ✅ Comprehensive documentation (docs/MODEL_SELECTION.md)

v1.1.0 - Local Setup & CLI Enhancement

  • ✅ DuckDB local database with sample data
  • ✅ Ollama integration for local LLM support
  • ✅ Legacy execution modes (4 core + 6 agent, behind the legacy extra)
  • ✅ Enhanced CLI with rich terminal output
  • ✅ Upgraded to Typer 0.21.0 for better compatibility

v1.0.0 - Initial Release

  • ✅ 7 task categories with modular architecture
  • ✅ Mental model principles (ReAct, CoT, Planning, Reflection, Memory)
  • ✅ Multiple LLM provider support
  • ✅ Docker deployment support

📄 License

MIT License - see LICENSE file for details

🙏 Acknowledgments

Built with:

  • LangChain for LLM orchestration and agent framework
  • Ollama for local LLM deployment
  • OpenAI & Anthropic for cloud LLM options
  • DuckDB for in-memory analytics and local database
  • Typer for CLI framework
  • Rich for beautiful terminal output
  • Pandas & NumPy for data manipulation
  • SQLAlchemy for database connectivity
  • Great Expectations for data quality (optional)
  • SQLGlot for SQL parsing (optional)

💡 Pro Tips

  1. Start with Core Modes: Use query, reverse, profile, interactive for free, instant results
  2. Cost Control: Use --model=gpt-4o-mini for daily work, save --model=gpt-4 for critical issues
  3. Offline Mode: Configure LLM_PROVIDER=ollama in .env for completely offline operation
  4. Model Selection: See docs/MODEL_SELECTION.md for detailed guidance on choosing models
  5. Agent Modes: Read the first response from agent modes (it's usually complete) before any looping occurs

🐛 Troubleshooting

Ollama Issues

# Check if Ollama is running
ollama list

# Restart Ollama service
# Windows: Restart from system tray
# Linux: systemctl restart ollama

Model Not Found

# Pull the model first
ollama pull llama2
ollama pull mistral

Database Issues

# Recreate the sample database
de init-db --force

Query Syntax

# The SQL goes in quotes, after the subcommand
de query "SELECT * FROM customers"

For more troubleshooting help, see docs/MODEL_SELECTION.md and docs/COMMAND_REFERENCE.md.

📞 Support

  • 📖 Documentation: See docs/ directory
  • 🐛 Issues: GitHub Issues
  • 💬 Discussions: GitHub Discussions

Ready to get started?

pip install de-agentic
de init-db
de demo

llm-based-data-engineering-agents

Metadata

Release files for de-agentic 2.1.2

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for de-agentic 2.1.2
File Size Uploaded
de_agentic-2.1.2.tar.gz 164.3 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for de-agentic 2.1.2
File Interpreter ABI Platform
de_agentic-2.1.2-py3-none-any.whl Python 3 none any Details

Total release size: 319.5 kB

Release files / de_agentic-2.1.2.tar.gz

Download URL de_agentic-2.1.2.tar.gz
Size 164.3 kB
Tags Source
SHA-256 checksum
How to use checksums
7941650afffb8f10ecb06aa3a763f59665384329f65f5c36e26a85f358774d63
BLAKE2b-256 checksum
How to use checksums
0b996d67f22230f7076731c679df76d4c46e639f99dd3dea200a1949c822af90
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 30, 2026.

Transparency log

Release files / de_agentic-2.1.2-py3-none-any.whl

Download URL de_agentic-2.1.2-py3-none-any.whl
Size 155.2 kB
Tags Python 3
SHA-256 checksum
How to use checksums
18af6aae81192ec4df503e63c0a8cf561c5fc183d8dc8b2b8385dd9babf4dd35
BLAKE2b-256 checksum
How to use checksums
aef1bcb5a22a2515d8e367928ce489a3c3da97711f70d46aea2c99a07070383f
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 30, 2026.

Transparency log

Release history Release notifications | RSS feed

2.1.3

2 release files

This release

2.1.2 This release

2 release files

2.1.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page