Python SDK for DataSpace API
Project description
DataSpace Python SDK
A Python SDK for programmatic access to DataSpace resources including Datasets, AI Models, and Use Cases.
Installation
From PyPI (once published)
pip install dataspace-sdk
From Source
git clone https://github.com/CivicDataLab/DataExchange.git
cd DataExchange/DataExBackend
pip install -e .
For Development
pip install -e ".[dev]"
Quick Start
from dataspace_sdk import DataSpaceClient
# Initialize the client
client = DataSpaceClient(base_url="https://api.dataspace.example.com")
# Login with Keycloak token
user_info = client.login(keycloak_token="your_keycloak_token")
print(f"Logged in as: {user_info['user']['username']}")
# Search for datasets
datasets = client.datasets.search(
query="health data",
tags=["public-health"],
page=1,
page_size=10
)
# Get a specific dataset
dataset = client.datasets.get_by_id("dataset-uuid")
print(f"Dataset: {dataset['title']}")
# Get organization's resources
org_id = user_info['user']['organizations'][0]['id']
org_datasets = client.datasets.get_organization_datasets(org_id)
Features
- Authentication: Login with Keycloak tokens and automatic token refresh
- Datasets: Search, retrieve, and list datasets with filtering and pagination
- AI Models: Search, retrieve, and list AI models with filtering
- Use Cases: Search, retrieve, and list use cases with filtering
- Organization Resources: Get resources specific to your organizations
- GraphQL & REST: Supports both GraphQL and REST API endpoints
Authentication
Login with Keycloak
from dataspace_sdk import DataSpaceClient
client = DataSpaceClient(base_url="https://api.dataspace.example.com")
# Login with Keycloak token
response = client.login(keycloak_token="your_keycloak_token")
# Access user information
print(response['user']['username'])
print(response['user']['organizations'])
Token Refresh
# Refresh access token when it expires
new_token = client.refresh_token()
Check Authentication Status
if client.is_authenticated():
print("Authenticated!")
print(f"User: {client.user['username']}")
Working with Datasets
Search Datasets
# Basic search
results = client.datasets.search(query="education")
# Advanced search with filters
results = client.datasets.search(
query="health",
tags=["public-health", "covid-19"],
sectors=["Health"],
geographies=["India", "Karnataka"],
status="PUBLISHED",
access_type="OPEN",
sort="recent",
page=1,
page_size=20
)
# Access results
print(f"Total results: {results['total']}")
for dataset in results['results']:
print(f"- {dataset['title']}")
Get Dataset by ID
# Get detailed dataset information
dataset = client.datasets.get_by_id("550e8400-e29b-41d4-a716-446655440000")
print(f"Title: {dataset['title']}")
print(f"Description: {dataset['description']}")
print(f"Organization: {dataset['organization']['name']}")
print(f"Resources: {len(dataset['resources'])}")
List All Datasets
# List with pagination
datasets = client.datasets.list_all(
status="PUBLISHED",
limit=50,
offset=0
)
for dataset in datasets:
print(f"- {dataset['title']}")
Get Trending Datasets
trending = client.datasets.get_trending(limit=10)
for dataset in trending['results']:
print(f"- {dataset['title']} (views: {dataset['view_count']})")
Get Organization Datasets
# Get datasets for your organization
org_id = client.user['organizations'][0]['id']
org_datasets = client.datasets.get_organization_datasets(
organization_id=org_id,
limit=20,
offset=0
)
Working with AI Models
Search AI Models
# Basic search
results = client.aimodels.search(query="language model")
# Advanced search
results = client.aimodels.search(
query="llm",
tags=["nlp", "text-generation"],
model_type="LLM",
provider="OPENAI",
status="ACTIVE",
sort="recent",
page=1,
page_size=10
)
Get AI Model by ID
# Using REST endpoint
model = client.aimodels.get_by_id("model-uuid")
# Using GraphQL (more detailed)
model = client.aimodels.get_by_id_graphql("model-uuid")
print(f"Model: {model['displayName']}")
print(f"Type: {model['modelType']}")
print(f"Provider: {model['provider']}")
print(f"Endpoints: {len(model['endpoints'])}")
List All AI Models
models = client.aimodels.list_all(
status="ACTIVE",
model_type="LLM",
limit=20,
offset=0
)
Get Organization AI Models
org_id = client.user['organizations'][0]['id']
org_models = client.aimodels.get_organization_models(
organization_id=org_id,
limit=20,
offset=0
)
Working with Use Cases
Search Use Cases
# Basic search
results = client.usecases.search(query="health monitoring")
# Advanced search
results = client.usecases.search(
query="covid",
tags=["health", "monitoring"],
sectors=["Health"],
status="PUBLISHED",
running_status="COMPLETED",
sort="completed_on",
page=1,
page_size=10
)
Get Use Case by ID
usecase = client.usecases.get_by_id(123)
print(f"Title: {usecase['title']}")
print(f"Summary: {usecase['summary']}")
print(f"Status: {usecase['runningStatus']}")
print(f"Datasets used: {len(usecase['datasets'])}")
print(f"Organizations: {len(usecase['organizations'])}")
List All Use Cases
usecases = client.usecases.list_all(
status="PUBLISHED",
running_status="COMPLETED",
limit=20,
offset=0
)
Get Organization Use Cases
org_id = client.user['organizations'][0]['id']
org_usecases = client.usecases.get_organization_usecases(
organization_id=org_id,
limit=20,
offset=0
)
Error Handling
from dataspace_sdk import (
DataSpaceAPIError,
DataSpaceAuthError,
DataSpaceNotFoundError,
DataSpaceValidationError,
)
try:
dataset = client.datasets.get_by_id("invalid-uuid")
except DataSpaceNotFoundError as e:
print(f"Dataset not found: {e.message}")
except DataSpaceAuthError as e:
print(f"Authentication error: {e.message}")
# Try to refresh token
client.refresh_token()
except DataSpaceValidationError as e:
print(f"Validation error: {e.message}")
print(f"Details: {e.response}")
except DataSpaceAPIError as e:
print(f"API error: {e.message}")
print(f"Status code: {e.status_code}")
Advanced Usage
Pagination
# Manual pagination
page = 1
page_size = 20
all_datasets = []
while True:
results = client.datasets.search(
query="health",
page=page,
page_size=page_size
)
all_datasets.extend(results['results'])
if len(results['results']) < page_size:
break
page += 1
print(f"Total datasets fetched: {len(all_datasets)}")
Working with Multiple Organizations
# Get user's organizations
user_info = client.get_user_info()
for org in user_info['organizations']:
print(f"\nOrganization: {org['name']} (Role: {org['role']})")
# Get resources for each organization
datasets = client.datasets.get_organization_datasets(org['id'])
models = client.aimodels.get_organization_models(org['id'])
usecases = client.usecases.get_organization_usecases(org['id'])
print(f" Datasets: {len(datasets)}")
print(f" AI Models: {len(models)}")
print(f" Use Cases: {len(usecases)}")
Combining Search Results
# Search across all resource types
query = "health"
datasets = client.datasets.search(query=query, page_size=5)
models = client.aimodels.search(query=query, page_size=5)
usecases = client.usecases.search(query=query, page_size=5)
print(f"Found {datasets['total']} datasets")
print(f"Found {models['total']} AI models")
print(f"Found {usecases['total']} use cases")
API Reference
DataSpaceClient
Main client for interacting with DataSpace API.
Methods:
login(keycloak_token: str) -> dict: Login with Keycloak tokenrefresh_token() -> str: Refresh access tokenget_user_info() -> dict: Get current user informationis_authenticated() -> bool: Check authentication status
Properties:
datasets: DatasetClient instanceaimodels: AIModelClient instanceusecases: UseCaseClient instanceuser: Current user informationaccess_token: Current access token
DatasetClient
Client for dataset operations.
Methods:
search(...): Search datasets with filtersget_by_id(dataset_id: str): Get dataset by UUIDlist_all(...): List all datasets with paginationget_trending(limit: int): Get trending datasetsget_organization_datasets(organization_id: str, ...): Get organization's datasets
AIModelClient
Client for AI model operations.
Methods:
search(...): Search AI models with filtersget_by_id(model_id: str): Get AI model by UUID (REST)get_by_id_graphql(model_id: str): Get AI model by UUID (GraphQL)list_all(...): List all AI models with paginationget_organization_models(organization_id: str, ...): Get organization's AI models
UseCaseClient
Client for use case operations.
Methods:
search(...): Search use cases with filtersget_by_id(usecase_id: int): Get use case by IDlist_all(...): List all use cases with paginationget_organization_usecases(organization_id: str, ...): Get organization's use cases
Contributing
Contributions are welcome! Please feel free to submit a Pull Request.
License
AGPL-3.0 License
Support
For issues and questions, please open an issue on GitHub.
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file dataspace_sdk-0.3.0.tar.gz.
File metadata
- Download URL: dataspace_sdk-0.3.0.tar.gz
- Upload date:
- Size: 469.3 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.11.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
fa67836d302587860b353c8cb26bb974ecab57da4e5bdd616d1a6e96260c4957
|
|
| MD5 |
fc6f22d0391d4c74c4ed51cf402e1e4a
|
|
| BLAKE2b-256 |
69ca1601caab41683916c6566bb33d47c503fe80f38abde3ec0eb3d6247971cc
|
File details
Details for the file dataspace_sdk-0.3.0-py3-none-any.whl.
File metadata
- Download URL: dataspace_sdk-0.3.0-py3-none-any.whl
- Upload date:
- Size: 27.1 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.11.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
c9c1993025c9b0ac25e97f1f343168f8eea9a18a22e85a85756d50fa80858266
|
|
| MD5 |
2dbbaa9031db73f4a8ed79039082fae3
|
|
| BLAKE2b-256 |
43a2a29b5e52b250d6ac691091c03b448346218a1204f86ac4920dab2a6a2662
|