✨ Highlights
- 🤖 Intelligent Agent Workflows: aide / automind / dsagent / data_interpreter / autokaggle / aflow, etc.
- 🔍 Discovery API: explore all available prompts and operators
- 📊 Data Management: unified data loading, task registry, and grading
- 🔧 Multi-model Support: OpenAI, GLM, DeepSeek, Qwen, and more
- 🧩 Extensible Architecture: custom tasks, workflows, and operators
- 📦 Smart Package Context: auto-detect installed packages to avoid incompatible code
- 🎯 Built-in Datasets: run demos without data preparation
- 📝 Full Traceability: logs, workspace, and artifacts saved automatically
🚀 Quick Start
1. Install
pip install dslighting python-dotenv
System requirements: Python 3.10+. Using a virtual environment is recommended.
🍎 macOS note (xgboost)
If you use xgboost, install the OpenMP runtime:
brew install libomp
Otherwise you may see XGBoostError: Library not loaded: libomp.dylib.
2. Configure environment variables
Create a .env file:
# .env
# Default model (required)
LLM_MODEL=glm-4
# Multi-model config (JSON)
LLM_MODEL_CONFIGS='{
"glm-4": {
"api_key": ["your-key-1", "your-key-2"],
"api_base": "https://open.bigmodel.cn/api/paas/v4",
"temperature": 0.7,
"provider": "openai"
},
"openai/deepseek-ai/DeepSeek-V3": {
"api_key": ["sk-siliconflow-key-1", "sk-siliconflow-key-2"],
"api_base": "https://api.siliconflow.cn/v1",
"temperature": 1.0
},
"gpt-4o": {
"api_key": "sk-your-openai-api-key",
"api_base": "https://api.openai.com/v1",
"temperature": 0.7
}
}'
Supported providers:
- OpenAI (GPT-4 / GPT-3.5)
- Zhipu AI (GLM-4)
- SiliconFlow (DeepSeek / Qwen / Kimi, etc.)
- Any OpenAI-compatible API
💡 Tip: call
load_dotenv()before importingdslighting.
🆕 Quick Experience
Option 1: Built-in dataset (zero setup)
from dotenv import load_dotenv
load_dotenv()
import dslighting
# No data prep required
result = dslighting.run_agent(task_id="bike-sharing-demand")
print(f"✅ Done! Score: {result.score}")
Built-in dataset example:
bike-sharing-demand(bike demand forecasting)
Option 2: Open-ended API (recommended for beginners)
import dslighting
# Analyze
result = dslighting.analyze(
data="./data/titanic",
description="Analyze passenger distribution",
model="gpt-4o"
)
# Process
result = dslighting.process(
data="./data/titanic",
description="Clean missing values and outliers",
model="gpt-4o"
)
# Model
result = dslighting.model(
data="./data/titanic",
description="Train a survival prediction model",
model="gpt-4o"
)
Option 3: Global config (recommended for multi-task)
from dotenv import load_dotenv
load_dotenv()
import dslighting
# Configure once, reuse everywhere
dslighting.setup(
data_parent_dir="/path/to/data/competitions",
registry_parent_dir="/path/to/registry"
)
agent = dslighting.Agent()
result = agent.run(task_id="bike-sharing-demand")
🌱 Beginner Usage
1. One-line demo (built-in dataset)
from dotenv import load_dotenv
load_dotenv()
import dslighting
result = dslighting.run_agent(task_id="bike-sharing-demand")
print(f"✅ Done! Score: {result.score}")
2. Open-ended API trio (Analyze / Process / Model)
import dslighting
# Analyze
_ = dslighting.analyze(
data="./data/titanic",
description="Analyze passenger distribution",
model="gpt-4o"
)
# Process
_ = dslighting.process(
data="./data/titanic",
description="Handle missing values and outliers",
model="gpt-4o"
)
# Model
_ = dslighting.model(
data="./data/titanic",
description="Train a survival prediction model",
model="gpt-4o"
)
3. Check results and workspace
print(result.workspace_path)
print(result.score)
🚀 Advanced Usage
1. Global config + reusable execution
import dslighting
# Configure once, reuse
dslighting.setup(
data_parent_dir="/path/to/data/competitions",
registry_parent_dir="/path/to/registry"
)
agent = dslighting.Agent(
workflow="aide",
model="gpt-4o",
max_iterations=5,
keep_workspace=True
)
result = agent.run(task_id="bike-sharing-demand")
2. Custom task registry (competition-style)
result = agent.run(
task_id="your-task-name",
data_dir="/path/to/data/competitions",
registry_dir="/path/to/registry"
)
3. Custom Agent (Operator / Workflow / Factory)
from dslighting.operators.custom import SimpleOperator
async def summarize(text: str) -> dict:
return {"summary": text[:200]}
summarize_op = SimpleOperator(func=summarize, name="Summarize")
class MyWorkflow:
def __init__(self, operators):
self.ops = operators
async def solve(self, description, io_instructions, data_dir, output_path):
_ = await self.ops["summarize"](text=description)
class MyWorkflowFactory:
def __init__(self, model="openai/gpt-4o"):
self.model = model
def create_agent(self):
return MyWorkflow({"summarize": summarize_op})
agent = MyWorkflowFactory().create_agent()
📦 Data Preparation
Method 1: MLE-Bench (recommended)
git clone https://github.com/openai/mle-bench.git
cd mle-bench
pip install -e .
python scripts/prepare.py --competition all
# Link data to DSLighting
ln -s ~/mle-bench/data/competitions /path/to/dslighting/data/competitions
Method 2: Custom dataset
data/competitions/
<competition-id>/
config.yaml
prepared/
public/
private/
More details:
🧭 Discovery API (Explore Components)
import dslighting
# List all prompts / operators
dslighting.explore()
List specific categories:
all_prompts = dslighting.list_prompts()
llm_ops = dslighting.list_operators(category="llm")
Get details:
from dslighting.prompts import get_prompt_info
from dslighting.operators import get_operator_info
print(get_prompt_info("create_improve_prompt"))
print(get_operator_info("PlanOperator"))
Full guide:
🧰 CLI Usage
After installation:
dslighting --help
Common subcommands:
dslighting help: help and quick guidedslighting workflows: list all workflowsdslighting example <workflow>: show workflow examplesdslighting quickstart: detailed quick startdslighting detect-packages: detect packages and write to config.yamldslighting show-packages: show detected packagesdslighting validate-config: validate configuration
🔧 Custom Tasks (Advanced)
your-project/
├── data/competitions/
│ └── your-task-name/
│ └── prepared/
│ ├── public/
│ └── private/
└── registry/
└── your-task-name/
├── config.yaml
├── description.md
└── grade.py
Example config.yaml:
id: your-task-name
name: Your Task Display Name
competition_type: simple
awards_medals: false
description: your-task-name/description.md
dataset:
answers: your-task-name/prepared/private/test_answer.csv
sample_submission: your-task-name/prepared/public/sampleSubmission.csv
grader:
name: rmsle # or accuracy, f1, mae, etc.
Run a custom task:
result = agent.run(
task_id="your-task-name",
data_dir="/path/to/data/competitions",
registry_dir="/path/to/registry"
)
📈 Checking Results
print(f"Workspace: {result.workspace_path}")
print(f"Score: {result.score}")
print(f"Cost: {result.cost}")
🧪 Web UI (Optional)
The Web UI requires the frontend/backend source. If you installed via pip, clone the repo:
git clone https://github.com/usail-hkust/dslighting.git
cd dslighting
Backend:
pip install -r web_ui/backend/requirements.txt
cd web_ui/backend
python main.py
Frontend:
cd web_ui/frontend
npm install
npm run dev
Open: http://localhost:3000
🎉 Latest Version: 2.7.9
Highlights:
- Comprehensive PyPI README with detailed documentation
- Enhanced installation guide with system requirements
- Multi-provider API setup examples (OpenAI, GLM, DeepSeek)
- Beginner and advanced usage examples
- Custom Agent tutorial for expert users
- Complete CLI and Web UI documentation
📚 Docs
- Quick Start: https://luckyfan-cs.github.io/dslighting-web/api/getting-started.html
- Data System: https://luckyfan-cs.github.io/dslighting-web/api/data-system.html
- GitHub: https://github.com/usail-hkust/dslighting
- PyPI: https://pypi.org/project/dslighting/
🤝 Contributing
Contributions are welcome!
- Fork the repo
- Create a feature branch (
git checkout -b feature/AmazingFeature) - Commit changes (
git commit -m 'Add some AmazingFeature') - Push to branch (
git push origin feature/AmazingFeature) - Open a Pull Request
📄 License
This project is licensed under AGPL-3.0.
If this project helps you, please give it a ⭐️
Made with ❤️ by USAIL Lab
Release files for dslighting 2.7.9
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| dslighting-2.7.9.tar.gz | 1.2 MB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| dslighting-2.7.9-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 3.6 MB
Release files / dslighting-2.7.9.tar.gz
| Download URL | dslighting-2.7.9.tar.gz |
|---|---|
| Size | 1.2 MB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
cd566b7061853b5d3102277bfa596627111856730fefb961ae9bd9cba39bf397
|
|
BLAKE2b-256 checksum How to use checksums |
dfba19c38960933d3751fa6b851242a3fbc68d5f56886a0c30a7ef639f499a46
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.10.19
|
Release files / dslighting-2.7.9-py3-none-any.whl
| Download URL | dslighting-2.7.9-py3-none-any.whl |
|---|---|
| Size | 2.4 MB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
c24da800423c5df80e4c8ab36ea5253684e4a9098bb6072f5d35fbc3af18ffc7
|
|
BLAKE2b-256 checksum How to use checksums |
07090f4b7fa2d06cadc89818aabb9e9d908109a827a2d92491d17183d23e044d
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.10.19
|