🔥 SAPPHIRE: High-Performance Compute for Apple Silicon 🔥
SAPPHIRE is a complete CUDA replacement that extracts 1.6 TFLOPS from Apple Silicon's AMX accelerator. Train and run AI models on Mac Mini for 50x less cost and 23x less power than NVIDIA hardware.
🚀 Performance
| Operation | SAPPHIRE | NVIDIA H100* |
|---|---|---|
| SGEMM | 1.56 TFLOPS | 60 TFLOPS |
| Flash Attention | 943 GFLOPS | ~20 TFLOPS |
| Conv2D | 1.57 TFLOPS | ~30 TFLOPS |
| INT8 Quantize | 6.3 B elem/s | ~50 B elem/s |
H100 costs $30,000 and uses 700W. Mac Mini costs $599 and uses 30W.
Price/Performance: SAPPHIRE wins by 50x!
📦 Installation
pip install sapphire-compute
🔥 Quick Start
import sapphire
import numpy as np
# Matrix multiplication at 1.6 TFLOPS
A = np.random.randn(4096, 4096).astype(np.float32)
B = np.random.randn(4096, 4096).astype(np.float32)
C = sapphire.matmul(A, B) # Uses AMX!
# Flash Attention V5
Q = np.random.randn(2, 16, 512, 64).astype(np.float32)
K = np.random.randn(2, 16, 512, 64).astype(np.float32)
V = np.random.randn(2, 16, 512, 64).astype(np.float32)
out = sapphire.flash_attention(Q, K, V)
# CUDA compatibility (drop-in replacement!)
cuda = sapphire.cuda
cuda.is_available() # True on Mac!
🧠 LLM Inference
from sapphire.llm import LlamaInference
# Load and run Llama on Mac Mini
model = LlamaInference("meta-llama/Llama-2-7b")
output = model.generate("The future of AI is", max_tokens=100)
print(output)
🔗 S-Fabric Clustering
Connect multiple Mac Minis for distributed compute:
from sapphire.sfabric import Cluster
# Create cluster over Thunderbolt 5
cluster = Cluster(["mac1:9999", "mac2:9999", "mac3:9999"])
cluster.connect()
# Distributed training
cluster.allreduce(gradients)
🏗️ Architecture
SAPPHIRE Stack
├── Python API (numpy-compatible)
├── Native Library (159 C functions)
│ ├── SGEMM (cblas → AMX)
│ ├── Flash Attention V5
│ ├── Conv2D (cuDNN replacement)
│ ├── Quantization (INT8/INT4)
│ └── cuSOLVER (LU, QR, SVD, Cholesky)
├── Lariat Transpiler (CUDA → Sapphire)
└── S-Fabric RDMA (Multi-Mac clustering)
📊 Benchmarks
Run the full benchmark suite:
python -m sapphire.benchmark
🎯 Key Features
- 159 Native Functions: Complete ML/AI operation coverage
- Flash Attention V5: Memory-efficient attention at 943 GFLOPS
- Zero-Copy UMA: Unified Memory Architecture exploitation
- Lariat CUDA Transpiler: Run CUDA code unchanged
- S-Fabric RDMA: Thunderbolt 5 multi-Mac clustering
- INT8 Quantization: 6.3 billion elements/second
🆚 NVIDIA Comparison
| Metric | Mac Mini + Sapphire | NVIDIA H100 |
|---|---|---|
| Cost | $599 | $30,000 |
| Power | 30W | 700W |
| TFLOPS/$ | 0.0026 | 0.002 |
| TFLOPS/W | 0.052 | 0.086 |
Conclusion: For most AI workloads, Sapphire on Mac Mini is the most cost-effective solution.
📄 License
MIT License - Use freely, no NVIDIA required!
🙏 Credits
Built by Svector Corporation - Making AI accessible to everyone.
🔥 NVIDIA's monopoly is over. The future runs on Apple Silicon. 🔥
Metadata
Release files for sapphire-compute 1.0.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| sapphire_compute-1.0.1-py3-none-any.whl | Python 3 | none | any | Details |
Release files / sapphire_compute-1.0.1-py3-none-any.whl
| Download URL | sapphire_compute-1.0.1-py3-none-any.whl |
|---|---|
| Size | 212.9 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
4c3eb9b6c905bab063b0ab0e5bdd124f9f853135910926c74aca7ef1bd5e622e
|
|
BLAKE2b-256 checksum How to use checksums |
60304dba17925aaa9cefe2f88e2a85313b230b4f1762bc6d15b35f6ff1d3ab57
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.12.11
|