ncaa_bbStats
ncaa_bbStats is an open-source Python package for retrieving, parsing, and analyzing college baseball data: NCAA Division I, II, and III team statistics (2002–2026), player statistics (2021–2026), MLB Draft history (1965–2025), draft detail with signing bonuses (2021–2026), RPI and schedule strength, program finances, and a draft-prediction model with scouting reports.
Built for analysts, developers, and fans. Everything is cached locally, so it works offline; scraping is opt-in.
The draft model in this package powers a public site where you can browse a board, look up a player, or score a stat line of your own: https://codemateo15-ncaa-draft-app.share.connect.posit.cloud/
That site lives in a separate repository, which is a distinct project and is not covered by this package's MIT licence — it carries no licence, so all rights are reserved. This package is MIT; the app is not.
Note This project is under active development.
Documentation
Documentation: ncaa_bbStats on ReadTheDocs
PyPI: ncaa-bbStats
Data sources and terms: DATA_PROVENANCE.md
Install
pip install ncaa_bbStats # everything except predictions
pip install "ncaa_bbStats[model]" # + draft predictions
pip install "ncaa_bbStats[explain]" # + SHAP explanations
pip install "ncaa_bbStats[scrape]" # + re-scraping the sources yourself
Requires Python 3.10 or later.
A tour
from ncaa_bbStats import (
team_profile, leaderboard, scouting_report, resolve_team, luckiest_teams,
)
# Everything about one program in one season, across every dataset
p = team_profile("Tennessee", 2024)
p["record"] # 60-13, .822
p["rpi"]["rpi_rank"] # 1
p["draft"]["picks"] # 8
p["pythagorean"] # expected .807 against an actual .822
# Leaderboards that sort the right way round
leaderboard("era", stat_type="pitching", year=2025, min_ip=60, n=10)
leaderboard("cwrc+", year=2025, conference="SEC", n=10)
leaderboard("hr", per="career", qualifier="noMin", n=5)
# Every source spells schools differently; one id resolves them all
resolve_team("Eastern Ill.") == resolve_team("EIU") == resolve_team("Eastern Illinois")
# Who won more than their run differential deserved?
luckiest_teams(2025, n=5)
# A scouting report
print(scouting_report("Kade Anderson", 2025))
What's in it
| Dataset | Coverage |
|---|---|
| NCAA team statistics | 2002–2026, Divisions I–III |
| Player statistics | 2021–2026, Division I |
| MLB Draft history | 1965–2025, 69,169 picks |
| MLB Draft detail (bonuses, slots, biography) | 2021–2026, 3,685 picks |
| RPI, strength of schedule, quadrant records | 2021–2026, Division I |
| Program finances (EADA) | 2021–2025, carried forward to 2026 |
| Draft prospect rankings | 2021–2026 |
| Team registry | 1,023 programs |
| Player registry | 27,283 players |
Modules
Team stats
get_team_stat, display_team_stats, display_specific_team_stat,
list_all_teams, plot_team_stat_over_years, average_all_team_stats,
average_team_stat_str, average_team_stat_float
Team registry
One canonical team_id per program, so datasets that spell schools differently
can be joined. Keyed on the federal IPEDS unitid where known, which survives
rebrands — Dixie State and Utah Tech share an id. Division is a per-season
attribute, not part of identity.
resolve_team, resolve_team_verbose, team_info, team_aliases,
team_seasons, team_division, team_conference, list_teams,
list_conferences, crosswalk
Player stats
list_players, list_batters, list_pitchers, player_seasons,
batting_stat, pitching_stat, get_player_rows, load_player_frame,
list_available_years
The cache stores counting statistics only. Every rate and advanced statistic is computed when you read it, from those counts plus league constants this package derives from its own NCAA team data — so they can never fall out of step.
Advanced stats
cwoba, cwraa, cwrc, cwrc_plus, cwsb, cspd, cfip, clob_pct,
league_constants, seasons_with_constants
College-calibrated analogues of the familiar sabermetric statistics, built the same way but with league constants regressed from NCAA play rather than borrowed from elsewhere. See DATA_PROVENANCE.md for the method and measured correlations.
Leaderboards
leaderboard, stat_direction, qualification_rules
Takes the sort direction from the statistic, so a top-ERA list contains good pitchers. Supports playing-time floors, team and conference filters, and career aggregation that rebuilds rates from summed components.
Draft
parse_mlb_draft, get_drafted_players_mlb, get_drafted_players_college,
print_draft_picks_mlb, print_draft_picks_college (1965–2025)
draft_pick, draft_class, draft_history, slot_value, signing_bonus,
bonus_vs_slot, overslot_picks, biggest_bonuses, draft_demographics,
conference_draft_counts, state_pipeline (2021–2026, with bonuses and slots)
prospect_rank, prospect_board, prospect_vs_actual, biggest_draft_risers,
biggest_draft_fallers
RPI and program finances
rpi_rank, strength_of_schedule, rpi_table, rpi_record,
quadrant_record, home_road_neutral, nonconference_profile,
rpi_over_years, best_wins
program_finance, budget_percentile, roster_size, coaching_staff_size,
richest_programs, conference_spending, finance_vs_rpi
Pythagorean expectation
get_pythagorean_expectation, compare_pythagorean_expectation,
luck_rating, luckiest_teams, unluckiest_teams, pythagorean_exponent,
conference_exponents
Cross-dataset
team_profile, player_profile, draft_yield, dollars_per_draft_pick,
conference_report, pipeline, compare_teams
Scouting and draft prediction
scouting_report, predict_draft_probability, predict_draft_order,
draft_board, explain_prediction, predict_from_stats, is_draft_eligible,
model_card
Two models: whether a player-season leads to being drafted (PR-AUC 0.703, ROC-AUC 0.957 on a held-out 2026) and where a drafted player falls in their class (Spearman 0.647). The held-out season is always the most recent one with complete draft labels, so it moves forward each release. Explanations come from SHAP where installed, with a gain-based fallback that says which it used.
Read model_card() before quoting any of it — it carries the limitations,
including that Stage 1 precision depends on the base rate you apply it to, that
eligibility is inferred rather than looked up, and that order predictions have a
mean absolute error of 78 places, so they separate tiers rather than picks.
from ncaa_bbStats import predict_from_stats
result = predict_from_stats(
"pitcher", age=21,
stats={"era": 2.40, "so": 130, "bb": 25, "ip": 95.0},
team="LSU", season=2025,
)
print(result["report"])
result["confidence"] # 'low' -- reports how much had to be imputed
Reference
- Team stat abbreviations
- Player stat abbreviations
- Team registry — how team names resolve
- Data provenance — sources, terms, and known limitations
Examples
Runnable notebooks covering every public function, with outputs saved so they
read without executing anything, live in notebooks/:
pip install -e ".[all]" jupyter
jupyter lab notebooks/
Regenerating the data
Builders live in tools/ and the *_store modules; none of them ship in the
wheel. See tools/README.md.
python -m ncaa_bbStats.team_store --years 2026 # scrape NCAA team stats
python tools/build_league_constants.py # refit the run values
python tools/build_team_registry.py # rebuild the registry
python -m ncaa_bbStats.model_store # retrain the draft models
python -m pytest tests/ -q
Planned
- Player statistics re-sourced from NCAA's own published data — partly done.
src/data/player_stats_cache_ncaa/covers 2021–2025 with no third-party export anywhere in the chain, and reproduces the default cache at r ≥ 0.998 on every column including this package's derived metrics. Read it withload_player_frame(..., source="ncaa"). It is not the default yet because NCAA publishes no date of birth (soageis empty) and the upstream mirrors stopped updating mid-2026. Promoting it needs a retrain, a registry rebuild, and another route to 2026 — see DATA_PROVENANCE.md - IPEDS identifiers backfilled for Division II and III programs
- Team game results with win-loss tracking
- Park factors, which currently limit
cwrc_plus
Found a bug or want a feature? Open an issue.
Support
Star this repo and share to help support!
Contact
Mateo Biggs, mateojohn2024@gmail.com
Release files for ncaa-bbStats 1.3.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| ncaa_bbstats-1.3.1.tar.gz | 20.1 MB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| ncaa_bbstats-1.3.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 40.1 MB
Release files / ncaa_bbstats-1.3.1.tar.gz
| Download URL | ncaa_bbstats-1.3.1.tar.gz |
|---|---|
| Size | 20.1 MB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
c1ee540cc457adc6c9121efa5608acbccf9c4c6d728512157696ff9399678b09
|
|
BLAKE2b-256 checksum How to use checksums |
f8a171f1bd192d110ebd390004d07b7756f71fd5c858682e5c6212dc8580ad68
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.1.0 CPython/3.12.7
|
Release files / ncaa_bbstats-1.3.1-py3-none-any.whl
| Download URL | ncaa_bbstats-1.3.1-py3-none-any.whl |
|---|---|
| Size | 20.0 MB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
ef779890f3a0038923e0190f25e3682e39322d59ef489619601cdef61429e214
|
|
BLAKE2b-256 checksum How to use checksums |
57a185963585ac186608360eb4ae5488f4771919f60d0b619045801070068961
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.1.0 CPython/3.12.7
|