2 projects
inferhost
Self-hosted, multi-modal AI server for your own GPU: chat/vision LLMs (Qwen, Llama, Gemma, DeepSeek), text-to-speech (Kokoro, OuteTTS), and image generation (SDXL, Flux, Z-Image, Qwen-Image) behind one OpenAI-compatible endpoint. Wraps llama.cpp and stable-diffusion.cpp, auto-downloads GGUF models from Hugging Face, hot-swaps VRAM, speculative decoding — no compiling. Type `inferhost` and you're done.
cadence-core
Neural model for next clinical event prediction from EHR sequences using the Narrative Velocity framework