Skip to main content

DataNova Python SDK

Official Python SDK for the DataNova Smart Analytics Platform REST API.

PyPI Version Python Versions License

DataNova SDK provides a simple and production-ready Python interface for interacting with the DataNova Smart Analytics Platform.

It supports:

  • Dataset upload and management
  • Dataset preview and inspection
  • Automated EDA
  • Data cleaning
  • Automatic visualizations
  • Automated data analysis
  • AI-powered questions
  • Machine learning
  • Asynchronous pipelines
  • Pipeline status monitoring
  • Developer API key management
  • Synchronous and asynchronous API clients

Table of Contents


Features

Core

  • Synchronous client powered by requests
  • Asynchronous client powered by httpx
  • Automatic retry support
  • Configurable request timeout
  • SSL verification support
  • Environment variable configuration
  • Context manager support
  • Safe JSON response parsing
  • Structured API exceptions
  • Request/correlation ID support
  • Automatic MIME type detection
  • Python Path and string file-path support

Data Analytics

  • Dataset upload
  • Dataset listing
  • Dataset preview
  • Exploratory Data Analysis
  • Data cleaning
  • Automatic visualizations
  • Automated end-to-end analysis
  • Analysis reports
  • Generated analysis code

AI

  • Natural-language questions about datasets
  • Optional dataset-aware AI queries
  • Combined automated analysis + AI question workflow

Machine Learning

  • Regression
  • Classification

Pipelines

  • Execute asynchronous pipelines
  • Check pipeline status
  • Wait for pipeline completion automatically

Requirements

DataNova SDK requires:

  • Python 3.8 or newer
  • Internet access to the DataNova API

The SDK does not restrict itself to a specific future Python 3.x version.

For example:

Python 3.8+
Python 3.9+
Python 3.10+
Python 3.11+
Python 3.12+
Python 3.13+
Python 3.14+
Future Python 3.x versions

Actual compatibility with a future Python release may still depend on the underlying third-party dependencies used by the SDK.


Installation

Install from PyPI

The recommended installation method is:

pip install datanova-sdk

Verify the installation:

pip show datanova-sdk

Check the installed version:

python -c "import datanova; print(datanova.__version__)"

Install with Async Support

If you want to use AsyncDataNovaClient:

pip install "datanova-sdk[async]"

This installs the optional httpx dependency.


Install for Development

For SDK development and testing:

pip install "datanova-sdk[dev]"

Install from Local Source

If you cloned or downloaded the SDK source:

pip install ./datanova-sdk

For async support:

pip install "./datanova-sdk[async]"

For development:

pip install "./datanova-sdk[dev]"

Editable Development Installation

When actively modifying the SDK source:

pip install -e "./datanova-sdk[dev]"

Changes to the source code will then be reflected immediately without reinstalling the package.


Configuration

DataNova SDK accepts the API key directly or through an environment variable.

Option 1 — Pass API Key Directly

from datanova import DataNovaClient

client = DataNovaClient(
    api_key="dn_live_your_api_key"
)

Option 2 — Environment Variable

Windows CMD

set DATANOVA_API_KEY=dn_live_your_api_key

Windows PowerShell

$env:DATANOVA_API_KEY="dn_live_your_api_key"

Linux / macOS

export DATANOVA_API_KEY="dn_live_your_api_key"

Then:

from datanova import DataNovaClient

client = DataNovaClient()

This is the recommended approach for production applications.


Custom API Base URL

The default API URL is:

https://datanova-fude.onrender.com

You can override it:

client = DataNovaClient(
    api_key="dn_live_your_api_key",
    base_url="https://your-api.example.com"
)

Or:

export DATANOVA_BASE_URL="https://your-api.example.com"

Quick Start

Synchronous Client

from datanova import DataNovaClient, DataNovaAPIError

try:

    with DataNovaClient(
        api_key="dn_live_your_api_key"
    ) as client:

        # Check API health
        health = client.health()

        print("API Status:")
        print(health)

except DataNovaAPIError as error:

    print("DataNova API Error:")
    print("Message:", error)
    print("Status Code:", error.status_code)
    print("Request ID:", error.request_id)

Complete Dataset Workflow

The following example demonstrates a typical DataNova workflow.

from datanova import DataNovaClient, DataNovaAPIError

try:

    with DataNovaClient(
        api_key="dn_live_your_api_key",
        max_retries=3,
        timeout=60
    ) as client:

        # 1. Check API health
        health = client.health()

        print("Health:", health)

        # 2. Upload dataset
        upload = client.upload_dataset(
            "sample_data.csv"
        )

        print("Upload response:")
        print(upload)

        dataset_id = (
            upload.get("dataset_id")
            or upload.get("id")
        )

        if dataset_id is None:
            raise RuntimeError(
                "Dataset ID was not returned by the API."
            )

        print("Dataset ID:", dataset_id)

        # 3. Preview dataset
        preview = client.preview_dataset(
            dataset_id
        )

        print("Dataset Preview:")
        print(preview)

        # 4. Run EDA
        eda = client.run_eda(
            dataset_id
        )

        print("EDA:")
        print(eda)

        # 5. Generate visualizations
        visualizations = client.generate_visualizations(
            dataset_id
        )

        print("Visualizations:")
        print(visualizations)

        # 6. Run complete automated analysis
        analysis = client.run_auto_analysis(
            dataset_id
        )

        print("Automated Analysis:")
        print(analysis)

except DataNovaAPIError as error:

    print(
        f"DataNova API Error "
        f"({error.status_code}): {error}"
    )

Dataset Operations

List Datasets

datasets = client.list_datasets()

print(datasets)

Function

client.list_datasets()

API

GET /api/v1/datasets

Upload Dataset

Supported files depend on the DataNova API configuration.

Example:

result = client.upload_dataset(
    "sales.csv"
)

print(result)

Excel example:

result = client.upload_dataset(
    "sales.xlsx"
)

Function

client.upload_dataset(file_path)

API

POST /api/v1/datasets/upload

The SDK automatically detects the MIME type from the filename.


Preview Dataset

preview = client.preview_dataset(
    dataset_id=35
)

print(preview)

Function

client.preview_dataset(dataset_id)

API

GET /api/v1/datasets/{id}/preview

Analytics

Run EDA

result = client.run_eda(
    dataset_id=35
)

print(result)

With a target column:

result = client.run_eda(
    dataset_id=35,
    target_column="Purchased"
)

Function

client.run_eda(
    dataset_id,
    target_column=None
)

API

POST /api/v1/analysis/eda

Clean Dataset

Basic:

result = client.clean_dataset(
    dataset_id=35
)

With strategies:

result = client.clean_dataset(
    dataset_id=35,
    strategies={
        "missing_values": "mean",
        "duplicates": "remove"
    }
)

Function

client.clean_dataset(
    dataset_id,
    strategies=None
)

API

POST /api/v1/analysis/clean

Generate Visualizations

Generate default visualizations:

result = client.generate_visualizations(
    dataset_id=35
)

print(result)

Specify chart types:

result = client.generate_visualizations(
    dataset_id=35,
    chart_types=[
        "bar",
        "line",
        "histogram",
        "scatter"
    ]
)

Function

client.generate_visualizations(
    dataset_id,
    chart_types=None
)

API

POST /api/v1/analysis/visualizations

Run Automated Analysis

result = client.run_auto_analysis(
    dataset_id=35
)

print(result)

Function

client.run_auto_analysis(dataset_id)

API

POST /api/v1/analysis/auto

Get Analysis Report

report = client.get_analysis_report(
    analysis_id=10
)

print(report)

Function

client.get_analysis_report(analysis_id)

API

GET /api/v1/analysis/{id}/report

Get Generated Analysis Code

code = client.get_analysis_code(
    analysis_id=10
)

print(code)

Function

client.get_analysis_code(analysis_id)

API

GET /api/v1/analysis/{id}/code

AI Q&A

Ask a general DataNova AI question:

answer = client.ask_question(
    "What are the most important trends in my data?"
)

print(answer)

Ask a question about a specific dataset:

answer = client.ask_question(
    query="Which product category has the highest sales?",
    dataset_id=35
)

print(answer)

Function

client.ask_question(
    query,
    dataset_id=None
)

API

POST /api/v1/ask

Machine Learning

Regression

result = client.train_regression(
    dataset_id=35,
    target_column="Sales"
)

print(result)

Function

client.train_regression(
    dataset_id,
    target_column
)

API

POST /api/v1/ml/regression

Classification

result = client.train_classification(
    dataset_id=35,
    target_column="Purchased"
)

print(result)

Function

client.train_classification(
    dataset_id,
    target_column
)

API

POST /api/v1/ml/classification

Pipelines

Execute Pipeline

Define pipeline steps:

steps = [
    {
        "step": "clean"
    },
    {
        "step": "eda"
    },
    {
        "step": "visualizations"
    }
]

result = client.execute_pipeline(
    dataset_id=35,
    steps=steps
)

print(result)

Function

client.execute_pipeline(
    dataset_id,
    steps
)

API

POST /api/v1/pipeline/execute

Check Pipeline Status

status = client.get_pipeline_status(
    task_id="your_task_id"
)

print(status)

Function

client.get_pipeline_status(task_id)

API

GET /api/v1/pipeline/status/{task_id}

Wait for Pipeline Completion

Instead of manually polling the API:

result = client.wait_for_pipeline(
    task_id="your_task_id"
)

print(result)

Customize polling:

result = client.wait_for_pipeline(
    task_id="your_task_id",
    poll_interval=3,
    timeout=300
)

print(result)

Function

client.wait_for_pipeline(
    task_id,
    poll_interval=2.0,
    timeout=300.0
)

The method automatically checks the pipeline status until it reaches a terminal state such as:

completed
success
succeeded
failed
error
cancelled
canceled

High-Level Convenience Methods

DataNova SDK also provides higher-level methods that combine existing API operations.


Analyze Dataset

Run automated analysis and optionally ask an AI question:

result = client.analyze_dataset(
    dataset_id=35
)

print(result)

With an AI question:

result = client.analyze_dataset(
    dataset_id=35,
    question="What are the most important business insights?"
)

print(result)

Function

client.analyze_dataset(
    dataset_id,
    question=None
)

This internally uses existing DataNova API operations.


Upload and Analyze

Upload a dataset and automatically start analysis:

result = client.upload_and_analyze(
    "sales.csv"
)

print(result)

With an AI question:

result = client.upload_and_analyze(
    "sales.csv",
    question="What are the main trends in this dataset?"
)

print(result)

Function

client.upload_and_analyze(
    file_path,
    question=None
)

The returned structure contains:

{
    "upload": {...},
    "analysis": {...}
}

Async Usage

Install async support:

pip install "datanova-sdk[async]"

Then:

import asyncio

from datanova import AsyncDataNovaClient


async def main():

    async with AsyncDataNovaClient(
        api_key="dn_live_your_api_key"
    ) as client:

        health = await client.health()

        print("Health:")
        print(health)

        datasets = await client.list_datasets()

        print("Datasets:")
        print(datasets)


if __name__ == "__main__":
    asyncio.run(main())

Async Dataset Upload

async with AsyncDataNovaClient(
    api_key="dn_live_your_api_key"
) as client:

    result = await client.upload_dataset(
        "sales.csv"
    )

    print(result)

Async Automated Analysis

async with AsyncDataNovaClient(
    api_key="dn_live_your_api_key"
) as client:

    result = await client.analyze_dataset(
        dataset_id=35,
        question="What are the major trends?"
    )

    print(result)

Async Pipeline Waiting

async with AsyncDataNovaClient(
    api_key="dn_live_your_api_key"
) as client:

    result = await client.wait_for_pipeline(
        task_id="your_task_id",
        poll_interval=2,
        timeout=300
    )

    print(result)

API Key Management

Generate API Key

result = client.generate_api_key(
    key_name="My Application",
    environment="live"
)

print(result)

With scopes:

result = client.generate_api_key(
    key_name="Analytics Application",
    environment="live",
    scopes=[
        "datasets",
        "analytics"
    ]
)

print(result)

Function

client.generate_api_key(
    key_name,
    environment="live",
    scopes=None
)

API

POST /api/developer/api_key/generate

Revoke API Key

result = client.revoke_api_key(
    key_id=123
)

print(result)

Function

client.revoke_api_key(key_id)

API

POST /api/developer/api_key/revoke

Error Handling

DataNova SDK provides the DataNovaAPIError exception.

from datanova import (
    DataNovaClient,
    DataNovaAPIError
)

try:

    with DataNovaClient(
        api_key="dn_live_your_api_key"
    ) as client:

        result = client.health()

except DataNovaAPIError as error:

    print("Message:", error)
    print("Status:", error.status_code)
    print("Payload:", error.payload)
    print("Request ID:", error.request_id)

Available error information:

error.status_code
error.payload
error.request_id

For example:

try:
    result = client.preview_dataset(999999)

except DataNovaAPIError as error:

    if error.status_code == 404:
        print("Dataset not found.")

    elif error.status_code == 401:
        print("Invalid API key.")

    elif error.status_code == 429:
        print("Rate limit exceeded.")

Logging

The SDK uses Python's standard logging module.

Enable logging:

import logging

logging.basicConfig(
    level=logging.INFO
)

For detailed debugging:

logging.basicConfig(
    level=logging.DEBUG
)

The synchronous client logger is:

datanova

The asynchronous client logger is:

datanova.async

Client Configuration

The synchronous client supports:

DataNovaClient(
    api_key=None,
    base_url=None,
    timeout=60,
    max_retries=3,
    backoff_factor=0.5,
    verify_ssl=True,
    user_agent=None
)

Parameters

Parameter Description Default
api_key DataNova API key Environment variable
base_url DataNova API base URL Official API
timeout Request timeout in seconds 60
max_retries Retry attempts 3
backoff_factor Retry backoff 0.5
verify_ssl Verify HTTPS certificates True
user_agent Custom User-Agent SDK default

Example:

client = DataNovaClient(
    api_key="dn_live_your_api_key",
    timeout=120,
    max_retries=5,
    backoff_factor=1.0,
    verify_ssl=True
)

Automatic Retries

The SDK automatically retries supported transient HTTP failures.

Retryable status codes include:

429
502
503
504

Example:

client = DataNovaClient(
    api_key="dn_live_your_api_key",
    max_retries=3,
    backoff_factor=0.5
)

Retry behavior is handled by the underlying HTTP session.


Context Manager

Using a context manager is recommended:

with DataNovaClient(
    api_key="dn_live_your_api_key"
) as client:

    print(client.health())

The HTTP session is automatically closed when the block finishes.

You can also manually close the client:

client = DataNovaClient(
    api_key="dn_live_your_api_key"
)

try:
    print(client.health())

finally:
    client.close()

Supported API Endpoints

Complete Endpoint Reference

Category Python Function HTTP REST Endpoint
System health() GET /api/v1/health
Datasets list_datasets() GET /api/v1/datasets
Datasets upload_dataset(file_path) POST /api/v1/datasets/upload
Datasets preview_dataset(id) GET /api/v1/datasets/{id}/preview
Analytics run_eda(id, target_column) POST /api/v1/analysis/eda
Analytics clean_dataset(id, strategies) POST /api/v1/analysis/clean
Analytics generate_visualizations(id) POST /api/v1/analysis/visualizations
Analytics run_auto_analysis(id) POST /api/v1/analysis/auto
Analytics get_analysis_report(id) GET /api/v1/analysis/{id}/report
Analytics get_analysis_code(id) GET /api/v1/analysis/{id}/code
AI ask_question(query, id) POST /api/v1/ask
ML train_regression(id, target) POST /api/v1/ml/regression
ML train_classification(id, target) POST /api/v1/ml/classification
Pipelines execute_pipeline(id, steps) POST /api/v1/pipeline/execute
Pipelines get_pipeline_status(task_id) GET /api/v1/pipeline/status/{task_id}
Developer generate_api_key(name) POST /api/developer/api_key/generate
Developer revoke_api_key(key_id) POST /api/developer/api_key/revoke

SDK Convenience Methods

The following methods combine existing API operations and do not represent additional REST endpoints:

Method Purpose
wait_for_pipeline() Automatically poll pipeline status
analyze_dataset() Run automated analysis + optional AI question
upload_and_analyze() Upload dataset + run automated analysis

Production Recommendations

For production applications:

1. Use environment variables

Recommended:

client = DataNovaClient()

with:

DATANOVA_API_KEY

configured in the deployment environment.


2. Do not hard-code API keys

Avoid:

api_key = "dn_live_real_secret_key"

Use:

import os

api_key = os.getenv("DATANOVA_API_KEY")

3. Use context managers

Recommended:

with DataNovaClient() as client:
    result = client.health()

4. Configure retries

For production workloads:

client = DataNovaClient(
    max_retries=3,
    backoff_factor=0.5
)

5. Handle API errors

try:
    result = client.health()

except DataNovaAPIError as error:
    print(error)

Development

Clone the repository:

git clone https://github.com/joshichaitanya58/datanova.git

Move to the SDK directory:

cd datanova-sdk

Install development dependencies:

pip install -e ".[dev]"

Run Tests

Run the complete test suite:

pytest

Verbose output:

pytest -v

Run a specific test file:

pytest tests/test_client.py

Run async tests:

pytest tests/test_async_client.py

Build the Package

Install build tools:

python -m pip install --upgrade build twine

Build:

python -m build

This generates:

dist/
├── datanova_sdk-1.1.1-py3-none-any.whl
└── datanova_sdk-1.1.1.tar.gz

Validate the Package

Run:

python -m twine check dist/*

Expected result:

Checking dist/...
PASSED

Publish to PyPI

After validating the package:

python -m twine upload dist/*

PyPI will then allow users to install the SDK with:

pip install datanova-sdk

Project Structure

datanova-sdk/
│
├── datanova/
│   ├── __init__.py
│   ├── client.py
│   └── async_client.py
│
├── tests/
│   ├── test_client.py
│   └── test_async_client.py
│
├── README.md
├── LICENSE
├── pyproject.toml
└── setup.py

The tests/ directory is used for development and CI testing and is not required for normal SDK runtime usage.


Security

Never commit real API keys to GitHub or other public repositories.

Use environment variables:

export DATANOVA_API_KEY="dn_live_your_api_key"

or configure the key through your cloud provider's environment-variable settings.

If an API key is accidentally exposed, revoke it immediately and generate a new one.


Version

Current SDK version:

1.1.1

Check it programmatically:

import datanova

print(datanova.__version__)

License

This project is licensed under the MIT License.

See LICENSE for details.


DataNova

Smart Analytics. Automated Insights. Data-Driven Decisions.

Official API:

https://datanova-fude.onrender.com

GitHub:

https://github.com/joshichaitanya58/datanova

Metadata

Release files for datanova-sdk 1.1.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for datanova-sdk 1.1.1
File Size Uploaded
datanova_sdk-1.1.1.tar.gz 16.6 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for datanova-sdk 1.1.1
File Interpreter ABI Platform
datanova_sdk-1.1.1-py3-none-any.whl Python 3 none any Details

Total release size: 24.3 kB

Release files / datanova_sdk-1.1.1.tar.gz

Download URL datanova_sdk-1.1.1.tar.gz
Size 16.6 kB
Tags Source
SHA-256 checksum
How to use checksums
2cc995311dc04c9b1e2fee60c11f6bf619b6c73976de30dbf26638e139b4e5f1
BLAKE2b-256 checksum
How to use checksums
217aa9def31ddd13d8cd690f8215b503d9e1263c6724358d2095f10c93162af1
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.14

Release files / datanova_sdk-1.1.1-py3-none-any.whl

Download URL datanova_sdk-1.1.1-py3-none-any.whl
Size 7.7 kB
Tags Python 3
SHA-256 checksum
How to use checksums
4990df0599de05126c9fbb3f5c031e28b034aba40e6b93c4b353f0d6b8b287fb
BLAKE2b-256 checksum
How to use checksums
78eaf52efb9b70acb4992bd84c167b074eb68729f17cfc5a35e2bae07126fa4c
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.14

Release history Release notifications | RSS feed

This release

1.1.1 This release

2 release files

1.1.0

2 release files

1.0.1

2 release files

1.0.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page