13 projects
retainkit
Context and memory policies for LLM agents, scored by the evidence that survives the token budget
judgekit
Bias probes for LLM judges, each with a bootstrap confidence interval: position, verbosity, self-preference, calibration and stability.
raterkit
Reliability, rogue-rater, drift and leakage diagnostics for labelled data
arenakit
Bradley-Terry leaderboards with simultaneous confidence control
abkit
Sample-ratio, peeking, multiple-testing and winner's-curse checks for experiment readouts
ppofolio
PPO ensemble for multi-asset crypto portfolio allocation (BTC/ETH/SOL) with a pre-registered evaluation protocol, walk-forward backtests, and a paper-trading safety chain
dynamic-pricing-lab
Dynamic pricing simulator: Thompson-sampling demand learning, forward-looking customers, and advertising
spark-search-ranking
Counterfactual learning-to-rank for marketplace search logs in PySpark
rankkit
Ranking evaluation with error bars: NDCG, MRR and MAP with confidence intervals, plus position-bias correction for click logs
judgepanel
Estimate LLM-judge accuracy without gold labels and aggregate judge panels: Dawid-Skene EM, agreement statistics, bootstrap uncertainty
calikit
Calibration auditing for probabilistic predictions: reliability diagrams, ECE, Brier decomposition, and temperature scaling
abeval
A/B-test statistics for LLM evals: error bars, paired comparisons, and sample-size planning
trajectory-judge
Measuring what LLM judges miss when an agent reaches the right answer the wrong way