DataNova Python SDK
Official Python SDK for the DataNova Smart Analytics Platform REST API.
DataNova SDK provides a simple and production-ready Python interface for interacting with the DataNova Smart Analytics Platform.
It supports:
- Dataset upload and management
- Dataset preview and inspection
- Automated EDA
- Data cleaning
- Automatic visualizations
- Automated data analysis
- AI-powered questions
- Machine learning
- Asynchronous pipelines
- Pipeline status monitoring
- Developer API key management
- Synchronous and asynchronous API clients
Table of Contents
- Features
- Requirements
- Installation
- Configuration
- Quick Start
- Async Usage
- Dataset Operations
- Analytics
- AI Q&A
- Machine Learning
- Pipelines
- High-Level Convenience Methods
- API Key Management
- Error Handling
- Logging
- Supported API Endpoints
- Development
- Testing
- Project Structure
- Security
- License
Features
Core
- Synchronous client powered by
requests - Asynchronous client powered by
httpx - Automatic retry support
- Configurable request timeout
- SSL verification support
- Environment variable configuration
- Context manager support
- Safe JSON response parsing
- Structured API exceptions
- Request/correlation ID support
- Automatic MIME type detection
- Python
Pathand string file-path support
Data Analytics
- Dataset upload
- Dataset listing
- Dataset preview
- Exploratory Data Analysis
- Data cleaning
- Automatic visualizations
- Automated end-to-end analysis
- Analysis reports
- Generated analysis code
AI
- Natural-language questions about datasets
- Optional dataset-aware AI queries
- Combined automated analysis + AI question workflow
Machine Learning
- Regression
- Classification
Pipelines
- Execute asynchronous pipelines
- Check pipeline status
- Wait for pipeline completion automatically
Requirements
DataNova SDK requires:
- Python 3.8 or newer
- Internet access to the DataNova API
The SDK does not restrict itself to a specific future Python 3.x version.
For example:
Python 3.8+
Python 3.9+
Python 3.10+
Python 3.11+
Python 3.12+
Python 3.13+
Python 3.14+
Future Python 3.x versions
Actual compatibility with a future Python release may still depend on the underlying third-party dependencies used by the SDK.
Installation
Install from PyPI
The recommended installation method is:
pip install datanova-sdk
Verify the installation:
pip show datanova-sdk
Check the installed version:
python -c "import datanova; print(datanova.__version__)"
Install with Async Support
If you want to use AsyncDataNovaClient:
pip install "datanova-sdk[async]"
This installs the optional httpx dependency.
Install for Development
For SDK development and testing:
pip install "datanova-sdk[dev]"
Install from Local Source
If you cloned or downloaded the SDK source:
pip install ./datanova-sdk
For async support:
pip install "./datanova-sdk[async]"
For development:
pip install "./datanova-sdk[dev]"
Editable Development Installation
When actively modifying the SDK source:
pip install -e "./datanova-sdk[dev]"
Changes to the source code will then be reflected immediately without reinstalling the package.
Configuration
DataNova SDK accepts the API key directly or through an environment variable.
Option 1 — Pass API Key Directly
from datanova import DataNovaClient
client = DataNovaClient(
api_key="dn_live_your_api_key"
)
Option 2 — Environment Variable
Windows CMD
set DATANOVA_API_KEY=dn_live_your_api_key
Windows PowerShell
$env:DATANOVA_API_KEY="dn_live_your_api_key"
Linux / macOS
export DATANOVA_API_KEY="dn_live_your_api_key"
Then:
from datanova import DataNovaClient
client = DataNovaClient()
This is the recommended approach for production applications.
Custom API Base URL
The default API URL is:
https://datanova-fude.onrender.com
You can override it:
client = DataNovaClient(
api_key="dn_live_your_api_key",
base_url="https://your-api.example.com"
)
Or:
export DATANOVA_BASE_URL="https://your-api.example.com"
Quick Start
Synchronous Client
from datanova import DataNovaClient, DataNovaAPIError
try:
with DataNovaClient(
api_key="dn_live_your_api_key"
) as client:
# Check API health
health = client.health()
print("API Status:")
print(health)
except DataNovaAPIError as error:
print("DataNova API Error:")
print("Message:", error)
print("Status Code:", error.status_code)
print("Request ID:", error.request_id)
Complete Dataset Workflow
The following example demonstrates a typical DataNova workflow.
from datanova import DataNovaClient, DataNovaAPIError
try:
with DataNovaClient(
api_key="dn_live_your_api_key",
max_retries=3,
timeout=60
) as client:
# 1. Check API health
health = client.health()
print("Health:", health)
# 2. Upload dataset
upload = client.upload_dataset(
"sample_data.csv"
)
print("Upload response:")
print(upload)
dataset_id = (
upload.get("dataset_id")
or upload.get("id")
)
if dataset_id is None:
raise RuntimeError(
"Dataset ID was not returned by the API."
)
print("Dataset ID:", dataset_id)
# 3. Preview dataset
preview = client.preview_dataset(
dataset_id
)
print("Dataset Preview:")
print(preview)
# 4. Run EDA
eda = client.run_eda(
dataset_id
)
print("EDA:")
print(eda)
# 5. Generate visualizations
visualizations = client.generate_visualizations(
dataset_id
)
print("Visualizations:")
print(visualizations)
# 6. Run complete automated analysis
analysis = client.run_auto_analysis(
dataset_id
)
print("Automated Analysis:")
print(analysis)
except DataNovaAPIError as error:
print(
f"DataNova API Error "
f"({error.status_code}): {error}"
)
Dataset Operations
List Datasets
datasets = client.list_datasets()
print(datasets)
Function
client.list_datasets()
API
GET /api/v1/datasets
Upload Dataset
Supported files depend on the DataNova API configuration.
Example:
result = client.upload_dataset(
"sales.csv"
)
print(result)
Excel example:
result = client.upload_dataset(
"sales.xlsx"
)
Function
client.upload_dataset(file_path)
API
POST /api/v1/datasets/upload
The SDK automatically detects the MIME type from the filename.
Preview Dataset
preview = client.preview_dataset(
dataset_id=35
)
print(preview)
Function
client.preview_dataset(dataset_id)
API
GET /api/v1/datasets/{id}/preview
Analytics
Run EDA
result = client.run_eda(
dataset_id=35
)
print(result)
With a target column:
result = client.run_eda(
dataset_id=35,
target_column="Purchased"
)
Function
client.run_eda(
dataset_id,
target_column=None
)
API
POST /api/v1/analysis/eda
Clean Dataset
Basic:
result = client.clean_dataset(
dataset_id=35
)
With strategies:
result = client.clean_dataset(
dataset_id=35,
strategies={
"missing_values": "mean",
"duplicates": "remove"
}
)
Function
client.clean_dataset(
dataset_id,
strategies=None
)
API
POST /api/v1/analysis/clean
Generate Visualizations
Generate default visualizations:
result = client.generate_visualizations(
dataset_id=35
)
print(result)
Specify chart types:
result = client.generate_visualizations(
dataset_id=35,
chart_types=[
"bar",
"line",
"histogram",
"scatter"
]
)
Function
client.generate_visualizations(
dataset_id,
chart_types=None
)
API
POST /api/v1/analysis/visualizations
Run Automated Analysis
result = client.run_auto_analysis(
dataset_id=35
)
print(result)
Function
client.run_auto_analysis(dataset_id)
API
POST /api/v1/analysis/auto
Get Analysis Report
report = client.get_analysis_report(
analysis_id=10
)
print(report)
Function
client.get_analysis_report(analysis_id)
API
GET /api/v1/analysis/{id}/report
Get Generated Analysis Code
code = client.get_analysis_code(
analysis_id=10
)
print(code)
Function
client.get_analysis_code(analysis_id)
API
GET /api/v1/analysis/{id}/code
AI Q&A
Ask a general DataNova AI question:
answer = client.ask_question(
"What are the most important trends in my data?"
)
print(answer)
Ask a question about a specific dataset:
answer = client.ask_question(
query="Which product category has the highest sales?",
dataset_id=35
)
print(answer)
Function
client.ask_question(
query,
dataset_id=None
)
API
POST /api/v1/ask
Machine Learning
Regression
result = client.train_regression(
dataset_id=35,
target_column="Sales"
)
print(result)
Function
client.train_regression(
dataset_id,
target_column
)
API
POST /api/v1/ml/regression
Classification
result = client.train_classification(
dataset_id=35,
target_column="Purchased"
)
print(result)
Function
client.train_classification(
dataset_id,
target_column
)
API
POST /api/v1/ml/classification
Pipelines
Execute Pipeline
Define pipeline steps:
steps = [
{
"step": "clean"
},
{
"step": "eda"
},
{
"step": "visualizations"
}
]
result = client.execute_pipeline(
dataset_id=35,
steps=steps
)
print(result)
Function
client.execute_pipeline(
dataset_id,
steps
)
API
POST /api/v1/pipeline/execute
Check Pipeline Status
status = client.get_pipeline_status(
task_id="your_task_id"
)
print(status)
Function
client.get_pipeline_status(task_id)
API
GET /api/v1/pipeline/status/{task_id}
Wait for Pipeline Completion
Instead of manually polling the API:
result = client.wait_for_pipeline(
task_id="your_task_id"
)
print(result)
Customize polling:
result = client.wait_for_pipeline(
task_id="your_task_id",
poll_interval=3,
timeout=300
)
print(result)
Function
client.wait_for_pipeline(
task_id,
poll_interval=2.0,
timeout=300.0
)
The method automatically checks the pipeline status until it reaches a terminal state such as:
completed
success
succeeded
failed
error
cancelled
canceled
High-Level Convenience Methods
DataNova SDK also provides higher-level methods that combine existing API operations.
Analyze Dataset
Run automated analysis and optionally ask an AI question:
result = client.analyze_dataset(
dataset_id=35
)
print(result)
With an AI question:
result = client.analyze_dataset(
dataset_id=35,
question="What are the most important business insights?"
)
print(result)
Function
client.analyze_dataset(
dataset_id,
question=None
)
This internally uses existing DataNova API operations.
Upload and Analyze
Upload a dataset and automatically start analysis:
result = client.upload_and_analyze(
"sales.csv"
)
print(result)
With an AI question:
result = client.upload_and_analyze(
"sales.csv",
question="What are the main trends in this dataset?"
)
print(result)
Function
client.upload_and_analyze(
file_path,
question=None
)
The returned structure contains:
{
"upload": {...},
"analysis": {...}
}
Async Usage
Install async support:
pip install "datanova-sdk[async]"
Then:
import asyncio
from datanova import AsyncDataNovaClient
async def main():
async with AsyncDataNovaClient(
api_key="dn_live_your_api_key"
) as client:
health = await client.health()
print("Health:")
print(health)
datasets = await client.list_datasets()
print("Datasets:")
print(datasets)
if __name__ == "__main__":
asyncio.run(main())
Async Dataset Upload
async with AsyncDataNovaClient(
api_key="dn_live_your_api_key"
) as client:
result = await client.upload_dataset(
"sales.csv"
)
print(result)
Async Automated Analysis
async with AsyncDataNovaClient(
api_key="dn_live_your_api_key"
) as client:
result = await client.analyze_dataset(
dataset_id=35,
question="What are the major trends?"
)
print(result)
Async Pipeline Waiting
async with AsyncDataNovaClient(
api_key="dn_live_your_api_key"
) as client:
result = await client.wait_for_pipeline(
task_id="your_task_id",
poll_interval=2,
timeout=300
)
print(result)
API Key Management
Generate API Key
result = client.generate_api_key(
key_name="My Application",
environment="live"
)
print(result)
With scopes:
result = client.generate_api_key(
key_name="Analytics Application",
environment="live",
scopes=[
"datasets",
"analytics"
]
)
print(result)
Function
client.generate_api_key(
key_name,
environment="live",
scopes=None
)
API
POST /api/developer/api_key/generate
Revoke API Key
result = client.revoke_api_key(
key_id=123
)
print(result)
Function
client.revoke_api_key(key_id)
API
POST /api/developer/api_key/revoke
Error Handling
DataNova SDK provides the DataNovaAPIError exception.
from datanova import (
DataNovaClient,
DataNovaAPIError
)
try:
with DataNovaClient(
api_key="dn_live_your_api_key"
) as client:
result = client.health()
except DataNovaAPIError as error:
print("Message:", error)
print("Status:", error.status_code)
print("Payload:", error.payload)
print("Request ID:", error.request_id)
Available error information:
error.status_code
error.payload
error.request_id
For example:
try:
result = client.preview_dataset(999999)
except DataNovaAPIError as error:
if error.status_code == 404:
print("Dataset not found.")
elif error.status_code == 401:
print("Invalid API key.")
elif error.status_code == 429:
print("Rate limit exceeded.")
Logging
The SDK uses Python's standard logging module.
Enable logging:
import logging
logging.basicConfig(
level=logging.INFO
)
For detailed debugging:
logging.basicConfig(
level=logging.DEBUG
)
The synchronous client logger is:
datanova
The asynchronous client logger is:
datanova.async
Client Configuration
The synchronous client supports:
DataNovaClient(
api_key=None,
base_url=None,
timeout=60,
max_retries=3,
backoff_factor=0.5,
verify_ssl=True,
user_agent=None
)
Parameters
| Parameter | Description | Default |
|---|---|---|
api_key |
DataNova API key | Environment variable |
base_url |
DataNova API base URL | Official API |
timeout |
Request timeout in seconds | 60 |
max_retries |
Retry attempts | 3 |
backoff_factor |
Retry backoff | 0.5 |
verify_ssl |
Verify HTTPS certificates | True |
user_agent |
Custom User-Agent | SDK default |
Example:
client = DataNovaClient(
api_key="dn_live_your_api_key",
timeout=120,
max_retries=5,
backoff_factor=1.0,
verify_ssl=True
)
Automatic Retries
The SDK automatically retries supported transient HTTP failures.
Retryable status codes include:
429
502
503
504
Example:
client = DataNovaClient(
api_key="dn_live_your_api_key",
max_retries=3,
backoff_factor=0.5
)
Retry behavior is handled by the underlying HTTP session.
Context Manager
Using a context manager is recommended:
with DataNovaClient(
api_key="dn_live_your_api_key"
) as client:
print(client.health())
The HTTP session is automatically closed when the block finishes.
You can also manually close the client:
client = DataNovaClient(
api_key="dn_live_your_api_key"
)
try:
print(client.health())
finally:
client.close()
Supported API Endpoints
Complete Endpoint Reference
| Category | Python Function | HTTP | REST Endpoint |
|---|---|---|---|
| System | health() |
GET | /api/v1/health |
| Datasets | list_datasets() |
GET | /api/v1/datasets |
| Datasets | upload_dataset(file_path) |
POST | /api/v1/datasets/upload |
| Datasets | preview_dataset(id) |
GET | /api/v1/datasets/{id}/preview |
| Analytics | run_eda(id, target_column) |
POST | /api/v1/analysis/eda |
| Analytics | clean_dataset(id, strategies) |
POST | /api/v1/analysis/clean |
| Analytics | generate_visualizations(id) |
POST | /api/v1/analysis/visualizations |
| Analytics | run_auto_analysis(id) |
POST | /api/v1/analysis/auto |
| Analytics | get_analysis_report(id) |
GET | /api/v1/analysis/{id}/report |
| Analytics | get_analysis_code(id) |
GET | /api/v1/analysis/{id}/code |
| AI | ask_question(query, id) |
POST | /api/v1/ask |
| ML | train_regression(id, target) |
POST | /api/v1/ml/regression |
| ML | train_classification(id, target) |
POST | /api/v1/ml/classification |
| Pipelines | execute_pipeline(id, steps) |
POST | /api/v1/pipeline/execute |
| Pipelines | get_pipeline_status(task_id) |
GET | /api/v1/pipeline/status/{task_id} |
| Developer | generate_api_key(name) |
POST | /api/developer/api_key/generate |
| Developer | revoke_api_key(key_id) |
POST | /api/developer/api_key/revoke |
SDK Convenience Methods
The following methods combine existing API operations and do not represent additional REST endpoints:
| Method | Purpose |
|---|---|
wait_for_pipeline() |
Automatically poll pipeline status |
analyze_dataset() |
Run automated analysis + optional AI question |
upload_and_analyze() |
Upload dataset + run automated analysis |
Production Recommendations
For production applications:
1. Use environment variables
Recommended:
client = DataNovaClient()
with:
DATANOVA_API_KEY
configured in the deployment environment.
2. Do not hard-code API keys
Avoid:
api_key = "dn_live_real_secret_key"
Use:
import os
api_key = os.getenv("DATANOVA_API_KEY")
3. Use context managers
Recommended:
with DataNovaClient() as client:
result = client.health()
4. Configure retries
For production workloads:
client = DataNovaClient(
max_retries=3,
backoff_factor=0.5
)
5. Handle API errors
try:
result = client.health()
except DataNovaAPIError as error:
print(error)
Development
Clone the repository:
git clone https://github.com/joshichaitanya58/datanova.git
Move to the SDK directory:
cd datanova-sdk
Install development dependencies:
pip install -e ".[dev]"
Run Tests
Run the complete test suite:
pytest
Verbose output:
pytest -v
Run a specific test file:
pytest tests/test_client.py
Run async tests:
pytest tests/test_async_client.py
Build the Package
Install build tools:
python -m pip install --upgrade build twine
Build:
python -m build
This generates:
dist/
├── datanova_sdk-1.1.1-py3-none-any.whl
└── datanova_sdk-1.1.1.tar.gz
Validate the Package
Run:
python -m twine check dist/*
Expected result:
Checking dist/...
PASSED
Publish to PyPI
After validating the package:
python -m twine upload dist/*
PyPI will then allow users to install the SDK with:
pip install datanova-sdk
Project Structure
datanova-sdk/
│
├── datanova/
│ ├── __init__.py
│ ├── client.py
│ └── async_client.py
│
├── tests/
│ ├── test_client.py
│ └── test_async_client.py
│
├── README.md
├── LICENSE
├── pyproject.toml
└── setup.py
The tests/ directory is used for development and CI testing and is not
required for normal SDK runtime usage.
Security
Never commit real API keys to GitHub or other public repositories.
Use environment variables:
export DATANOVA_API_KEY="dn_live_your_api_key"
or configure the key through your cloud provider's environment-variable settings.
If an API key is accidentally exposed, revoke it immediately and generate a new one.
Version
Current SDK version:
1.1.1
Check it programmatically:
import datanova
print(datanova.__version__)
License
This project is licensed under the MIT License.
See LICENSE for details.
DataNova
Smart Analytics. Automated Insights. Data-Driven Decisions.
Official API:
https://datanova-fude.onrender.com
GitHub:
Metadata
Release files for datanova-sdk 1.1.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| datanova_sdk-1.1.1.tar.gz | 16.6 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| datanova_sdk-1.1.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 24.3 kB
Release files / datanova_sdk-1.1.1.tar.gz
| Download URL | datanova_sdk-1.1.1.tar.gz |
|---|---|
| Size | 16.6 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
2cc995311dc04c9b1e2fee60c11f6bf619b6c73976de30dbf26638e139b4e5f1
|
|
BLAKE2b-256 checksum How to use checksums |
217aa9def31ddd13d8cd690f8215b503d9e1263c6724358d2095f10c93162af1
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Release files / datanova_sdk-1.1.1-py3-none-any.whl
| Download URL | datanova_sdk-1.1.1-py3-none-any.whl |
|---|---|
| Size | 7.7 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
4990df0599de05126c9fbb3f5c031e28b034aba40e6b93c4b353f0d6b8b287fb
|
|
BLAKE2b-256 checksum How to use checksums |
78eaf52efb9b70acb4992bd84c167b074eb68729f17cfc5a35e2bae07126fa4c
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|