2 projects
innards
See inside open-weight LLMs per message: KV cache, context growth and memory, predicted vs measured, across MHA, GQA, MLA, sliding-window and hybrid models.
tai-aitutor
Plain-Python RAG building blocks — LLMs, embeddings, chunking, retrieval and evaluation — as a migration path off LlamaIndex.