🤖 DE Agentic - AI-Powered Data Engineering Assistant
▶ Start here
Thirty seconds, no install:
python3 de.py demo
Creates a sample database, profiles a file, runs a query, reads the schema and
answers a question — standard library only. No API key, no pip install.
Prefer a de command on your PATH?
python3 -m pip install -e . # then: de demo
The four commands:
python3 de.py profile customers.csv # profile a file
python3 de.py query "SELECT * FROM customers LIMIT 5" # read-only SQL
python3 de.py schema # tables + relationships
python3 de.py ask "how do I find null values?" # model, or offline help
Plug in a real model (llama.cpp, Ollama, OpenAI or Anthropic) whenever you want:
python3 de.py doctor # what is reachable?
python3 de.py setup # detect one and write .env
python3 de.py demo --ai # full demo, step 5 answered by your model
python3 de.py tc # end-to-end test: the model writes SQL, we run it
llama.cpp, Ollama and OpenAI all speak the same OpenAI-compatible
/v1/chat/completions, so one code path covers them. Precedence is
environment > .env > config.yaml > defaults.
Cloud warehouses (ClickHouse Cloud and Snowflake)
The default remains a local database. To use a configured warehouse, put non-secret
connection settings in config.yaml and keep the password in an environment variable:
database:
type: clickhouse_cloud # or: snowflake
host: your-service.clickhouse.cloud
port: 8443 # Snowflake commonly uses 443
database: analytics
username: default
password_env: CLICKHOUSE_PASSWORD
secure: true
# Snowflake also accepts warehouse, role and schema.
Then export the password and use --db configured:
export CLICKHOUSE_PASSWORD='...'
python3 -m pip install clickhouse-connect
python3 de.py profile --db configured
python3 de.py schema --db configured
python3 de.py query "SELECT count() FROM events" --db configured
For Snowflake, install snowflake-connector-python and set
SNOWFLAKE_PASSWORD (or your chosen environment-variable name). Passwords are never
stored in YAML; the YAML stores only the variable name.
📖 Full getting-started guide: QUICKSTART.md
Which docs are current.
Start with Quick Start and Usage above — they are verified against a real install.
docs/COMMAND_REFERENCE.md,docs/TESTING.md,docs/QUICK_TEST_REFERENCE.md,docs/DEPLOYMENT.md,docs/MODEL_SELECTION.md,docs/AGENT_MODES_OPTIONS.mdanddocs/FIX_CLI_OPTIONS.mdstill document the legacypython -m src.cli run ...interface, which needs the optionallegacyextra. They are kept for reference while that tree is migrated.
docs/MIGRATION.mdis the guide for moving off them, anddocs/QUALITY_REPORT.md/docs/USER_TESTING.mdare dated records of the simplification work, not living instructions.
A comprehensive agentic system designed to assist data engineers with daily tasks including data ingestion, modeling, quality checks, warehousing, debugging, reverse engineering, and architecture simplification.
🆕 Now with Local LLM Support! Run completely offline with Ollama, or use OpenAI/Anthropic for cloud-based models. Switch between models per command for cost/quality optimization.
!Note: This is the part of README.md of Data-Agent Project, the project is covering for all topics related to data engineering. The project is still in development, but I would like to publish the documentation for reference and idea brainstorming into the to accelerating the working.
System Overview
To understand how the project is going to resolve, run the demo below:
python ./examples/demo_local.py
python ./examples/demo_simple.py
🌟 Features
Task Categories
- Data Ingestion: Automated data loading from various sources (APIs, databases, files)
- Data Modeling: Schema design, ERD generation, normalization checks
- Data Quality: Profiling, validation, anomaly detection
- Warehousing: Pipeline generation, optimization, dbt assistance
- Debugging: Error analysis, log parsing, performance profiling
- Reverse Engineering: Schema extraction, lineage tracking, documentation generation
- Architecture: Pattern detection, optimization recommendations, diagram generation
Execution Modes
-
🚀 Core Modes (Free, Instant, No LLM Required):
query: Execute SQL queries on local databasereverse: Analyze database schema and generate documentationprofile: Profile data files with comprehensive statisticsinteractive: Command-based system for database operations
-
🤖 Agent Modes (LLM-Powered, Flexible Model Selection):
debug: Analyze errors and provide debugging guidancequality: Data quality assessment with recommendationsmodel: Data modeling assistanceingest: Data ingestion strategy and code generationwarehouse: Data warehouse architecture guidancearchitect: System architecture analysis and optimization
Mental Model Principles
- ReAct (Reasoning + Acting): Combines reasoning and action for better decision-making
- Chain of Thought: Step-by-step reasoning for complex problems
- Planning: Task decomposition and strategic planning
- Reflection: Self-evaluation and continuous improvement
- Memory: Context retention across interactions
LLM Flexibility
- Multiple Providers: OpenAI (GPT-4, GPT-4o-mini), Ollama (llama2, mistral, codellama), Anthropic (Claude)
- Per-Command Model Selection: Choose optimal model for each task using
--modelparameter - Cost Optimization: Use cheap models for daily work, premium for critical tasks
- Local-First: Default to Ollama for free, offline operation
🚀 Quick Start
Prerequisites
- Python 3.9 or newer
- Optional: Ollama or another local model server
- Optional: an API key for OpenAI or Anthropic
- Optional: Docker, for the container image
Install
pip install de-agentic
That is the whole thing. The base install depends only on pyyaml and
pydantic; profile, query, schema and init-db then run on the
standard library alone. Two extras are available:
pip install "de-agentic[pretty]" # rich tables (falls back to plain text)
pip install "de-agentic[data]" # DuckDB files, Parquet/Excel profiling
de is now on your PATH. Check it:
de --version
Prefer not to install anything? python3 de.py demo runs the same
walkthrough from a clone using only the standard library.
Create the sample database
The commands below read sample.sqlite, which is generated on demand:
de init-db # customers, products, orders, order_items
Everyday commands
de profile customers.csv # column types, nulls, ranges
de query "SELECT * FROM customers LIMIT 5" # read-only SQL
de schema # tables, columns, relationships
de ask "how do I find null values?" # answered by your model
ask works offline too: without a model it falls back to local guidance and
tells you how to configure one.
Connect a model
de doctor # which endpoints are reachable?
de setup # detect one and write .env
Then re-run de ask, or try the full walkthrough with a real model:
de demo # end to end
de demo --ai # same, with step 5 answered by your model
llama.cpp, Ollama and OpenAI all speak the same OpenAI-compatible
/v1/chat/completions, so one code path covers them. Precedence is
environment > .env > config.yaml > defaults.
Per-command overrides are available when you want to switch providers without editing configuration:
de ask "..." --provider openai --model gpt-4o-mini
Using Docker
Quick Start with Docker Compose
# Build and start all services (app + postgres + ollama)
docker compose up -d
# Check status
docker compose ps
# View logs
docker compose logs -f de-agentic
Test the Deployment
This stack exists for the legacy agent interface, which needs the optional
legacy extra that the base image does not install. If you only want the
de commands, use the single-container form below instead.
# Interactive shell with the CLI (requires -it)
docker compose exec -it de-agentic de --help
# Create the sample database and inspect it
docker compose exec de-agentic de init-db
docker compose exec de-agentic de schema
docker compose exec de-agentic de query "SELECT * FROM customers LIMIT 5"
Single Container (No Dependencies)
The image installs this package from source, so the commands inside the
container are the same de commands:
# Build the image
docker build -t de-agentic .
# Standard-library walkthrough, no model required
docker run --rm de-agentic de demo
# Or use the CLI directly
docker run --rm de-agentic de schema
docker run --rm de-agentic de query "SELECT 1 AS x"
# Profile a local file (mount volume)
docker run --rm -v "$PWD/customers.csv:/data/customers.csv" \
de-agentic de profile /data/customers.csv
Note: the container's default command is de demo, not the legacy
python -m src.cli, which needs optional dependencies the base image
does not install.
📚 For complete deployment guide, see docs/DEPLOYMENT.md
📖 Usage
Command Line Interface
Core modes (free, instant, no model required)
These run on the standard library and need no model server:
de init-db # create sample.sqlite
de query "SELECT COUNT(*) FROM customers" # read-only SQL
de schema # tables, columns, relationships
de profile customers.csv # column types, nulls, ranges
Model-backed commands
ask and demo --ai route through whichever provider is configured. Use
de doctor to see what is reachable and de setup to configure one.
de ask "which tables have no primary key?" # answered by your model
de demo --ai # full walkthrough, model answers
de tc # end-to-end test of the SQL path
Select a provider per command without editing any config file:
de ask "..." --provider openai --model gpt-4o-mini
de ask "..." --provider ollama --model llama3.2:1b
Legacy CLI
The earlier agent-mode interface (run query, run debug, run reverse,
...) is still in the tree as src.cli, but it is mid-migration and needs the
optional legacy extra:
pip install "de-agentic[legacy]"
de-agentic --help
It is not required for anything documented above. New code should target the
de commands.
Legacy agent modes
Available only with the legacy extra, via de-agentic. Listed for
reference; these are mid-migration and not covered by CI.
de-agentic run debug --error="Connection timeout"
de-agentic run quality --file=data.csv --model=gpt-4o-mini
de-agentic run model --description="E-commerce system"
de-agentic run ingest --source="API" --target="postgres"
de-agentic run warehouse --requirements="Real-time analytics"
de-agentic run architect --description="Current pipeline"
Per-command --model works the same way here as with de.
Python API
from src.agents.de_agent import DEAgent
from src.tasks import DataQualityTask
# Initialize agent
agent = DEAgent()
# Run data quality check
task = DataQualityTask(
file_path="customers.csv",
profile=True,
validate=True
)
result = agent.execute(task)
print(result)
Available Models
| Provider | Model | Cost | Best For |
|---|---|---|---|
| OpenAI | gpt-4o-mini | $0.15/1M tokens | Daily work, fast responses |
| gpt-4 | $30/1M tokens | Critical issues, complex problems | |
| gpt-4-turbo | $10/1M tokens | Balance of cost/quality | |
| Ollama | llama2 | Free | General purpose, offline |
| mistral | Free | Better code understanding | |
| codellama | Free | Code generation, debugging | |
| Anthropic | claude-3.5-sonnet | $3/1M tokens | Long context, analysis |
| claude-3-opus | $15/1M tokens | Premium quality |
🏗️ Architecture
de-agentic/
├── src/
│ ├── agents/ # Agent implementations
│ ├── tasks/ # Task definitions
│ ├── skills/ # Reusable skills
│ ├── tools/ # Integration tools
│ ├── workflows/ # Workflow orchestration
│ └── utils/ # Utilities
├── config/ # Configuration files
├── examples/ # Usage examples
├── tests/ # Test suite
└── docs/ # Documentation
🔧 Configuration
LLM Provider Configuration (.env)
# Default: Local Ollama (free, offline)
LLM_PROVIDER=ollama
OLLAMA_BASE_URL=http://localhost:11434
OLLAMA_MODEL=llama2
# OpenAI (requires API key)
LLM_PROVIDER=openai
OPENAI_API_KEY=sk-proj-...
OPENAI_MODEL=gpt-4o-mini
# Anthropic (requires API key)
LLM_PROVIDER=anthropic
ANTHROPIC_API_KEY=sk-ant-...
ANTHROPIC_MODEL=claude-3.5-sonnet
Agent Configuration (config/agent_config.yaml)
Edit config/agent_config.yaml to customize:
- Enable/disable specific tasks and skills
- Adjust mental model parameters
- Configure database connections
- Set execution limits
- Fine-tune agent behavior
Database Setup
The demo.db DuckDB database includes:
- customers table: 10 sample customers
- orders table: 30 sample orders
- products table: 10 sample products
- Total revenue: $6,813.97
Recreate anytime with: de init-db --force
🤝 Contributing
Contributions are welcome! Please read our contributing guidelines first.
📚 Documentation
Comprehensive guides available in the docs/ directory:
- QUICKSTART.md: 5-minute getting started guide
- DEPLOYMENT.md: Complete deployment and testing guide (Docker, remote, production)
- MODEL_SELECTION.md: Complete guide to choosing and using different LLM models
- COMMAND_REFERENCE.md: All CLI commands with examples
- AGENT_MODES_OPTIONS.md: Detailed comparison of all 10 execution modes
- DESIGN.md: Architecture and design principles
🎯 Use Cases
Daily Data Engineering Tasks
- Query databases without remembering SQL syntax
- Profile new data files instantly
- Debug pipeline errors with AI assistance
- Generate data models from requirements
Cost Optimization Strategy
# Development/Testing (free, local)
--model=llama2
# Daily work (cheap, cloud)
--model=gpt-4o-mini
# Critical issues (premium, best quality)
--model=gpt-4
Example workflow
# 1. Profile a new file (free, instant, no model)
de profile new_data.csv
# 2. Read the schema
de schema
# 3. Run a check with a cheap model
de ask "does new_data.csv have quality issues?" --model gpt-4o-mini
# 4. Escalate to a stronger model for a hard question
de ask "how should I fix the encoding problem?" --model gpt-4
# 5. Verify with SQL (free, instant)
de query "SELECT count(*) FROM new_data"
🚀 Recent Updates
v2.1.1 - Documentation accuracy
- ✅ Recent Updates now covers the 2.1.x line instead of stopping at v1.2.0
- ✅ Added the MIT
LICENSEfile thatpyproject.tomland the README referenced - ✅ Legacy reference docs are banner-marked and point at
MIGRATION.md - ✅ Release workflow publishes only for strict
vX.Y.Ztags
v2.1.0 - Packaging, CI and install docs
- ✅ GitHub Actions: lint, a 3.9-3.12 test matrix, and a build that verifies the wheel
- ✅ Tag-driven release: publish to PyPI via trusted publishing, plus a GitHub release
- ✅ The wheel now declares its dependencies;
import core.configworks after install - ✅
harnessdata files and thesrcpackage ship correctly - ✅
de doctorno longer crashes when no model is configured - ✅ README and Dockerfile now match what actually installs and runs
v1.2.0 - Model Selection Feature
- ✅ Added
--modelparameter for per-command model selection - ✅ Auto-provider detection (gpt* → openai, claude* → anthropic)
- ✅ Cost optimization through flexible model selection
- ✅ Comprehensive documentation (docs/MODEL_SELECTION.md)
v1.1.0 - Local Setup & CLI Enhancement
- ✅ DuckDB local database with sample data
- ✅ Ollama integration for local LLM support
- ✅ Legacy execution modes (4 core + 6 agent, behind the
legacyextra) - ✅ Enhanced CLI with rich terminal output
- ✅ Upgraded to Typer 0.21.0 for better compatibility
v1.0.0 - Initial Release
- ✅ 7 task categories with modular architecture
- ✅ Mental model principles (ReAct, CoT, Planning, Reflection, Memory)
- ✅ Multiple LLM provider support
- ✅ Docker deployment support
📄 License
MIT License - see LICENSE file for details
🙏 Acknowledgments
Built with:
- LangChain for LLM orchestration and agent framework
- Ollama for local LLM deployment
- OpenAI & Anthropic for cloud LLM options
- DuckDB for in-memory analytics and local database
- Typer for CLI framework
- Rich for beautiful terminal output
- Pandas & NumPy for data manipulation
- SQLAlchemy for database connectivity
- Great Expectations for data quality (optional)
- SQLGlot for SQL parsing (optional)
💡 Pro Tips
- Start with Core Modes: Use
query,reverse,profile,interactivefor free, instant results - Cost Control: Use
--model=gpt-4o-minifor daily work, save--model=gpt-4for critical issues - Offline Mode: Configure
LLM_PROVIDER=ollamain .env for completely offline operation - Model Selection: See
docs/MODEL_SELECTION.mdfor detailed guidance on choosing models - Agent Modes: Read the first response from agent modes (it's usually complete) before any looping occurs
🐛 Troubleshooting
Ollama Issues
# Check if Ollama is running
ollama list
# Restart Ollama service
# Windows: Restart from system tray
# Linux: systemctl restart ollama
Model Not Found
# Pull the model first
ollama pull llama2
ollama pull mistral
Database Issues
# Recreate the sample database
de init-db --force
Query Syntax
# The SQL goes in quotes, after the subcommand
de query "SELECT * FROM customers"
For more troubleshooting help, see docs/MODEL_SELECTION.md and docs/COMMAND_REFERENCE.md.
📞 Support
- 📖 Documentation: See
docs/directory - 🐛 Issues: GitHub Issues
- 💬 Discussions: GitHub Discussions
Ready to get started?
pip install de-agentic
de init-db
de demo
llm-based-data-engineering-agents
Metadata
Release files for de-agentic 2.1.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| de_agentic-2.1.1.tar.gz | 157.3 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| de_agentic-2.1.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 309.3 kB
Release files / de_agentic-2.1.1.tar.gz
| Download URL | de_agentic-2.1.1.tar.gz |
|---|---|
| Size | 157.3 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
55af42dbb94b34bacc727aa79692a9a78ca2b0920f843743e82f83d18c838f9e
|
|
BLAKE2b-256 checksum How to use checksums |
984e2faf08a2e3a28b31324fb525538afa0c66e72fc72998f069bc5c61a77bc3
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 30, 2026.
Transparency logRelease files / de_agentic-2.1.1-py3-none-any.whl
| Download URL | de_agentic-2.1.1-py3-none-any.whl |
|---|---|
| Size | 152.0 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
ce0e92b361497bf9ea757648054c7e4809ad31914357af078bb0007a73166adb
|
|
BLAKE2b-256 checksum How to use checksums |
c3bf1161382518856f4e497b63adfe562bb9c92bbc31b68bac8cd4f9decd6678
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 30, 2026.
Transparency log