Python SDK for Nexus AI Platform - Complete text embedding & RAG capabilities
Project description
Nexus AI Python SDK
Official Python SDK for Nexus AI - A unified AI capabilities platform.
🎉 Latest Release v0.2.3
🚀 New: Standalone Text Embedding API - Complete implementation of embedding functionality
Production Ready - 95.2% test pass rate with 100% P0 core features passing.
Installation:
pip install keystone-ai
Quick Start:
from nexusai import NexusAIClient
client = NexusAIClient(api_key="your_api_key")
# Text generation
response = client.text.generate("Hello, AI!")
print(response.text)
# Text embedding (NEW in v0.2.3)
embedding = client.embeddings.create("你好世界")
vector = embedding.data[0].embedding # 768-dimensional vector
Features
- 🚀 Simple & Intuitive - Clean API design with sensible defaults
- 🔄 Multi-Model Support - 6 text models + 2 image models + 2 embedding models
- 📊 Text Embeddings - Standalone embedding API with BAAI/bge models
- 📡 Streaming - Real-time streaming for text generation
- 💬 Session Management - Stateful conversations with automatic context handling
- 🧠 Knowledge Bases - RAG capabilities with semantic search
- 🎨 Multi-Modal - Text, images, audio (ASR), and document processing
- 🔐 Type-Safe - Full type hints with Pydantic models
- 🌐 Production Ready - Defaults to production API at
https://nexus-ai.juncai-ai.com/api/v1
Installation
pip install keystone-ai
国内镜像加速:
# 清华镜像
pip install keystone-ai -i https://pypi.tuna.tsinghua.edu.cn/simple
# 阿里云镜像
pip install keystone-ai -i https://mirrors.aliyun.com/pypi/simple/
从源码安装:
git clone https://github.com/aidrshao/nexus-ai-sdk.git
cd nexus-ai-sdk
poetry install
Quick Start
1. Set up your API key
Create a .env file in your project root:
NEXUS_API_KEY=nxs_your_api_key_here
# SDK automatically uses production: https://nexus-ai.juncai-ai.com/api/v1
# For local development, set: NEXUS_BASE_URL=http://localhost:8000/api/v1
2. Initialize the client
from nexusai import NexusAIClient
# Simple - uses production API automatically
client = NexusAIClient(api_key="nxs_your_api_key")
# Or read from environment variables
client = NexusAIClient()
# For local development
client = NexusAIClient(
api_key="nxs_your_api_key",
base_url="http://localhost:8000/api/v1"
)
3. Generate text
# Simple mode (省心模式) - uses default model
response = client.text.generate("写一首关于春天的诗")
print(response.text)
# With model selection
response = client.text.generate(
prompt="Explain quantum computing",
model="gpt-5-mini", # Recommended: fast and cost-effective
temperature=0.7,
max_tokens=500
)
print(response.text)
print(f"Tokens used: {response.usage.total_tokens}")
# Available models (三档体系):
# 🥇 高端: "gpt-5" (fastest premium), "gemini-2.5-pro" (strongest reasoning)
# 🥈 中端: "gpt-5-mini" (recommended), "gpt-4o-mini" (alternative)
# 🥉 经济: "deepseek-v3.2-exp" (cheapest)
4. Stream text generation
for chunk in client.text.stream("Tell me a story"):
if "delta" in chunk:
print(chunk["delta"].get("content", ""), end="", flush=True)
print()
5. Work with sessions (conversations)
# Create a session
session = client.sessions.create(
name="My Chat",
agent_config={
"model": "gpt-5-mini", # Recommended model for conversations
"temperature": 0.7
}
)
# Have a conversation
response = session.invoke("My name is Alice")
print(response.response.content)
response = session.invoke("What's my name?")
print(response.response.content) # Remembers "Alice"
# Get conversation history
history = session.history()
for message in history:
print(f"{message.role}: {message.content}")
6. Generate images
# Simple mode
image = client.images.generate("A futuristic city")
print(image.image_url)
# With options
image = client.images.generate(
prompt="A sunset over mountains, digital art",
model="doubao-seedream-4-0-250828", # Default recommended model (ByteDance Doubao)
aspect_ratio="16:9", # Use ratio instead of pixel size
num_images=1
)
print(f"Image: {image.image_url}")
# Supported aspect ratios: "1:1", "16:9", "9:16", "4:3", "3:4", "21:9"
# Image models: "doubao-seedream-4-0-250828" (default), "gemini-2.5-flash-image" (alternative)
7. Speech-to-Text (ASR)
# Upload audio file
file_meta = client.files.upload("meeting.mp3")
# Transcribe
transcription = client.audio.transcribe(
file_id=file_meta.file_id,
language="zh"
)
print(transcription.text)
8. Text Embeddings (文本向量化) 🆕
Transform text into high-dimensional vectors for semantic search, similarity computation, and machine learning applications.
Available Models
- BAAI/bge-base-zh-v1.5 (768 dimensions) - High-quality Chinese embedding model, supports mixed Chinese-English text
- BAAI/bge-large-zh-v1.5 (1024 dimensions) - Large-scale Chinese embedding model with higher precision
Single Text Embedding
# Create embedding for a single text
response = client.embeddings.create(
input="人工智能正在改变世界",
model="BAAI/bge-base-zh-v1.5" # Default model
)
vector = response.data[0].embedding # 768-dimensional vector
print(f"Vector dimensions: {len(vector)}")
print(f"Token usage: {response.usage.total_tokens}")
Batch Text Embedding
# Process multiple texts at once
texts = [
"人工智能技术发展迅速",
"机器学习是AI的核心",
"深度学习模型性能优异",
"自然语言处理理解人类语言"
]
response = client.embeddings.create(
input=texts,
model="BAAI/bge-base-zh-v1.5"
)
# Extract all vectors
vectors = [item.embedding for item in response.data]
print(f"Generated {len(vectors)} embeddings")
Optimized Large-Scale Processing
# Efficient processing for large datasets
large_texts = [f"文档内容 {i}" for i in range(1000)]
response = client.embeddings.create_batch(
texts=large_texts,
model="BAAI/bge-base-zh-v1.5",
batch_size=50 # Process 50 texts per batch
)
print(f"Processed {response.batch_info.total_texts} texts")
print(f"Processing time: {response.batch_info.processing_time:.2f}s")
Similarity Calculation
import numpy as np
def cosine_similarity(vec1, vec2):
"""Calculate cosine similarity between two vectors."""
vec1, vec2 = np.array(vec1), np.array(vec2)
return np.dot(vec1, vec2) / (np.linalg.norm(vec1) * np.linalg.norm(vec2))
# Generate embeddings for comparison
texts = ["AI技术发展", "人工智能进步", "今天天气很好"]
response = client.embeddings.create(input=texts)
vectors = [item.embedding for item in response.data]
# Calculate similarities
sim1 = cosine_similarity(vectors[0], vectors[1]) # Similar texts
sim2 = cosine_similarity(vectors[0], vectors[2]) # Different texts
print(f"Similar texts similarity: {sim1:.4f}") # ~0.85+
print(f"Different texts similarity: {sim2:.4f}") # ~0.3-
Model Information & Health Check
# List available models
models = client.embeddings.list_models()
for model in models.data:
print(f"Model: {model.id}")
print(f"Dimensions: {model.dimensions}")
print(f"Description: {model.description}")
# Check service health
health = client.embeddings.health_check()
print(f"Service status: {health.status}")
Performance Tips
- Batch Processing: Use
create_batch()for >10 texts for better performance - Optimal Batch Size: 20-50 texts per batch for best throughput
- Token Efficiency: Batch processing reduces per-text overhead
- Model Selection: Use
bge-basefor speed,bge-largefor accuracy
9. Knowledge Base & RAG (检索增强生成)
Step 1: Create Knowledge Base with Custom Configuration
# Create knowledge base with custom chunking and embedding settings
kb = client.knowledge_bases.create(
name="Company Docs",
description="Internal documentation",
embedding_model="BAAI/bge-base-zh-v1.5", # 向量化模型 (default)
chunk_size=1000, # 文档切片大小(字符数)
chunk_overlap=200 # 切片重叠大小(字符数)
)
print(f"Knowledge Base ID: {kb.kb_id}")
Configuration Parameters:
embedding_model: Embedding model for vectorization (default:BAAI/bge-base-zh-v1.5)chunk_size: Document chunk size in characters (default: 1000)chunk_overlap: Overlap between chunks in characters (default: 200)
Step 2: Upload Documents (Asynchronous Processing)
import time
# Upload document
task = client.knowledge_bases.upload_document(
kb_id=kb.kb_id,
file="company_policy.pdf"
)
# ⚠️ Important: Document processing is asynchronous (10-60 seconds)
# The task object contains task_id but NOT doc_id
print(f"Document submitted. Task ID: {task.task_id}")
# Wait for processing to complete
timeout = 60
start_time = time.time()
while time.time() - start_time < timeout:
status = client._internal_client.request("GET", f"/tasks/{task.task_id}")
if status["status"] == "completed":
doc_id = status["output"]["doc_id"]
chunk_count = status["output"]["chunk_count"]
print(f"✅ Document processed successfully!")
print(f" Document ID: {doc_id}")
print(f" Chunks created: {chunk_count}")
break
elif status["status"] == "failed":
print(f"❌ Processing failed: {status['error']['message']}")
break
print(f"⏳ Status: {status['status']}...")
time.sleep(2)
Alternative: File Reuse Across Knowledge Bases
# Upload file once
file_meta = client.files.upload("shared_policy.pdf")
# Add to multiple knowledge bases
task1 = client.knowledge_bases.add_document(kb_sales.kb_id, file_meta.file_id)
task2 = client.knowledge_bases.add_document(kb_support.kb_id, file_meta.file_id)
task3 = client.knowledge_bases.add_document(kb_hr.kb_id, file_meta.file_id)
Step 3: Semantic Search
# Search for relevant content
results = client.knowledge_bases.search(
query="What is the vacation policy?",
knowledge_base_ids=[kb.kb_id],
top_k=3, # Return top 3 most relevant chunks
similarity_threshold=0.7 # Minimum similarity score (0-1)
)
# Display search results
print(f"Found {results.total_results} results:")
for result in results.results:
print(f"\n📄 Score: {result.similarity_score:.2f}")
print(f" Content: {result.content[:100]}...")
print(f" Source: {result.metadata.get('filename', 'Unknown')}")
Search Parameters:
query: Search query textknowledge_base_ids: List of KB IDs to searchtop_k: Number of results to return (default: 5)similarity_threshold: Minimum similarity score, 0-1 (default: 0.7)
Step 4: RAG Generation (Retrieval + Generation)
# Combine retrieved context with text generation
context = "\n\n".join([r.content for r in results.results])
# Generate answer based on retrieved context
answer = client.text.generate(
prompt=f"""Based on the following context, answer the question accurately.
Context:
{context}
Question: What is the vacation policy?
Answer:""",
model="gpt-5-mini", # Recommended for RAG
temperature=0.3 # Lower temperature for more accurate answers
)
print(f"\n🤖 Answer:\n{answer.text}")
📖 Complete Guide: See Knowledge Base Async Guide for detailed async processing patterns, error handling, and production best practices.
Configuration
The SDK can be configured via environment variables or constructor parameters:
| Environment Variable | Default | Description |
|---|---|---|
NEXUS_API_KEY |
(required) | Your API key |
NEXUS_BASE_URL |
https://nexus-ai.juncai-ai.com/api/v1 |
API base URL |
NEXUS_TIMEOUT |
30 |
Request timeout (seconds) |
NEXUS_MAX_RETRIES |
3 |
Maximum retry attempts |
NEXUS_POLL_INTERVAL |
2 |
Task polling interval (seconds) |
NEXUS_POLL_TIMEOUT |
300 |
Task polling timeout (seconds) |
Production vs Development Mode
Production Mode (Default):
# Uses production API by default - zero configuration needed!
client = NexusAIClient(api_key="nxs_your_api_key")
# → Connects to https://nexus-ai.juncai-ai.com/api/v1
Local Development Mode:
# Set environment variable
export NEXUS_BASE_URL=http://localhost:8000/api/v1
Or in code:
client = NexusAIClient(
api_key="nxs_dev_key",
base_url="http://localhost:8000/api/v1"
)
Error Handling
The SDK provides specific exception types for different error scenarios:
from nexusai import NexusAIClient
from nexusai.error import (
AuthenticationError,
RateLimitError,
NotFoundError,
APITimeoutError,
)
client = NexusAIClient()
try:
response = client.text.generate("Hello")
except AuthenticationError:
print("Invalid API key")
except RateLimitError as e:
print(f"Rate limited. Retry after {e.retry_after}s")
except NotFoundError:
print("Resource not found")
except APITimeoutError:
print("Request timed out")
except Exception as e:
print(f"Unexpected error: {e}")
Context Manager
The client supports context manager for automatic cleanup:
with NexusAIClient() as client:
response = client.text.generate("Hello")
print(response.text)
# Client automatically closed
📚 Complete Documentation
📖 Complete Documentation Index - All documentation in one place
Quick Links
- Quick Start Guide - 5-minute tutorial to get started
- API Reference for Developers - Complete API documentation
- Knowledge Base Async Guide 🆕 - Complete guide for async document processing
- Application Developer FAQ - Common questions answered
- Error Handling Guide - Best practices for error handling
- Model Selection Guide - Choosing the right model
- Changelog - Version history and updates
Code Examples
Check out the examples/ directory:
- basic_usage.py - Core features demonstration
- error_handling.py - Error handling patterns
Model Documentation
Text Models (6 available):
- 🥇 Premium:
gpt-5,gemini-2.5-pro - 🥈 Standard:
gpt-5-mini - 🥉 Budget:
deepseek-v3.2-exp(default),gpt-4o-mini
Image Models (2 available):
doubao-seedream-4-0-250828(default)gemini-2.5-flash-image
Requirements
- Python 3.8+
- httpx >= 0.25.0
- pydantic >= 2.5.0
- python-dotenv >= 1.0.0
Contributing
Contributions are welcome! Please feel free to submit a Pull Request.
License
This project is licensed under the MIT License - see the LICENSE file for details.
Support & Community
- PyPI: https://pypi.org/project/keystone-ai/
- Documentation: https://nexus-ai.juncai-ai.com/docs
- GitHub: https://github.com/aidrshao/nexus-ai-sdk
- Issues: https://github.com/aidrshao/nexus-ai-sdk/issues
- Email: support@nexus-ai.com
Star History
If you find this project helpful, please consider giving it a ⭐ on GitHub!
Changelog
v0.2.2 (2025-10-07) - Documentation Enhancement
Focus: Comprehensive documentation for async knowledge base processing
- 📚 Added Knowledge Base Async Guide - 13,000+ word comprehensive guide
- 3 async processing implementation methods
- Task status flow and error handling
- Performance optimization best practices
- Production-ready patterns
- 📖 Enhanced API documentation with detailed async workflow
- 🐛 Fixed documentation inconsistencies and version numbers
- ✅ 100% Knowledge Base RAG tests passing (6/6)
v0.2.1 (2025-10-06) - Stable Release
First Production-Ready Release - 95.2% test pass rate
- 🎉 First stable release on PyPI as
keystone-ai - ✅ All P0 core features passing (100%)
- 🔄 6 text models + 2 image models supported
- 🧠 Full RAG capabilities with semantic search
- 📡 Real-time streaming support
- 💬 Session management with context
- 🎨 Multi-modal support (text, images, audio)
v0.2.0 (2025-10-04) - Messages Format Support
- ✨ Messages format for multi-turn conversations
- ✨ File list API with pagination
- 🐛 Fixed streaming functionality
- 🐛 UTF-8 encoding improvements
v0.1.0 (2025-10-04) - Initial Release
- Initial alpha release
- Text generation (sync, async, streaming)
- Image generation with task polling
- Session management
- Audio processing (ASR)
- Knowledge base management
- File upload system
- Full type hints with Pydantic
For detailed changelog, see CHANGELOG.md
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file keystone_ai-0.2.3.tar.gz.
File metadata
- Download URL: keystone_ai-0.2.3.tar.gz
- Upload date:
- Size: 31.2 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.13.3
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
2384211cc89fd9d52549139ba0943b3a8b6a6f98d2a15ee339b9b93768e3cf5f
|
|
| MD5 |
bb384826bfd034c3c123a318aeffe7b8
|
|
| BLAKE2b-256 |
35078123b082f4bcbce853d9245f52052ffc64ceef69e72f438dd7b63961fa73
|
File details
Details for the file keystone_ai-0.2.3-py3-none-any.whl.
File metadata
- Download URL: keystone_ai-0.2.3-py3-none-any.whl
- Upload date:
- Size: 41.8 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.13.3
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
3a4153b442dde37a3b7907d9c0308d0f3e9bbfebac744936be9035d86aa5c1b4
|
|
| MD5 |
94e28bbabec02d5a0d4a617080d780c9
|
|
| BLAKE2b-256 |
1b68a4b23d18a0fd295397d3b7585e0250ef4d943481f9f327eccd35b427ee87
|