association
A local-first NBA stats pipeline: fetch from ESPN's stats APIs into Parquet, build a DuckDB analytics warehouse from it, and ask questions about it in plain English — answered by a local LLM via Ollama, with no cloud API calls anywhere.
Ask it something
uv tool install 'association-py[web] @ git+https://github.com/jeffknupp/association@v5.0.0'
association web
# association is serving at http://127.0.0.1:40525 (ctrl-c to stop)
No default port — it binds a free one and prints the URL for your terminal to linkify. Common question shapes are rendered from structured data: leaderboards and game logs as tables, a multi-season history as a sparkline, a team's record as a card. Anything without a renderer still answers, in the same text the CLI prints — and every rendered answer keeps that text one click away.
Every answer says whether a template produced it or the question was refused for want of a reading, because that is the most useful thing you can know about how far to trust it. Each message is a new question; there is no conversation memory.
Shot charts and NetPoints fingerprints draw in the conversation, served from the same directory the CLI writes to — so a chart made at the terminal opens in the browser, and the CLI still writes the identical standalone file:
The page follows your system theme, and so do the charts — both screenshots above are the same build, one light and one dark.
Or from the command line
association query "who led the league in assists this season?"
# Nikola Jokic led the league in assists per game in the 2026 regular season
# (minimum 20 games), at 10.7.
association query "how many times did the 76ers play the Celtics this season?"
# The Philadelphia 76ers and the Boston Celtics met 4 times in the 2026 regular
# season, splitting them 2-2.
association query "Luka Doncic vs Shai Gilgeous-Alexander this season"
# Luka Doncic vs Shai Gilgeous-Alexander, 2026 regular season:
# Luka Doncic Shai Gilgeous-Alexander
# games 64 68
# points 33.5 31.1
# rebounds 7.7 4.3
# assists 8.3 6.6
# ...
# net pts/100 +6.55 +9.91
association query "top 5 rebounders on the Lakers in the playoffs"
# Deandre Ayton led the Los Angeles Lakers in rebounds per game in the 2026
# postseason (minimum 5 games), at 9.6. Next: LeBron James (6.7), Austin Reaves (4.0), Rui Hachimura (4.0), Luke Kennard (3.5).
association query "Steph Curry's 3pt percentage over the past 4 seasons"
# Stephen Curry, 3PT% by regular season, 2023-2026 (most recent first):
# season G 3PT% 3PM 3PA
# 2026 43 39.3 190 484
# 2025 70 39.7 311 784
# 2024 74 40.8 357 876
# 2023 56 42.7 273 639
association query "who had the most assists in a single game this season?"
# Ryan Nembhard had the most assists in a single game in the 2026 regular season:
# 23, on 2026-04-12 vs CHI. Next: Isaiah Collier (22), Josh Giddey (19).
Questions like these are answered in about a second. Anything outside that set is refused, and the refusal says what is missing - an intent nothing answers, a narrowing the data cannot honor, a season a table does not reach. No model writes SQL here: the one that used to, for questions no template covered, answered one in 23 and is gone.
The templates cover the shapes real NBA stat questions take, measured against StatMuse's live query feed. For example:
- Against one opponent: "jaylen brown last 8 games vs pistons", "evan mobley avg against bucks" (the averages, then the meetings behind them)
- Games kept under or over a line, one game of a series, a season by its place in a career: "Sga games with under 14 fta", "paul reed gamelog with 25 minutes", "maxey's stats for game 4 against the knicks", "how many 40+ point games does lebron have in his 18th season"
- Splits: "Nikola Jokic home and away splits", "Joe Ingles stats when starting vs coming off the bench"
- With or without a teammate: "Celtics record without Tatum"
- A record under a condition: "Sixers record when Embiid scores 30 points"
- Two players' meetings: "lebron vs kawhi head to head"
- A quarter or half: "How many points did Jokic score in the 3rd quarter?", "76ers 4th quarter scoring against Boston"
- Streaks: "Lakers longest winning streak this season"
- Careers: "career points leaders", "Jokic career averages"
- Team rankings, lines and outlook: "which team scores the most points per game", "Knicks home record", "what are the celtics playoff odds"
A question that narrows to something no template can honor, such as "on back-to-backs", "in the Finals" or "since returning from injury", is not answered for everything instead: it is refused, naming the narrowing. A question about a season a table does not reach, such as a 1996 shot chart, is refused, and the answer says why.
Some questions render a chart instead of text:
association query "plot Stephen Curry's shot chart from his last game this season"
# Rendered shot chart for Stephen Curry (7/14 made, 50.0%) to query_output/shotchart_stephen_curry_401811054.html
That needs play-by-play data pulled first (--include-pbp), and writes a
self-contained, theme-aware HTML/SVG file — open it in a browser.
A player's NetPoints "fingerprint" — how they add value, across 20 play-type skills — renders the same way, and needs no play-by-play:
association query "plot Shai Gilgeous-Alexander's fingerprint for 2025"
# Rendered NetPoints fingerprint (total) for Shai Gilgeous-Alexander (2025 season, percentile scale) to query_output/fingerprint_shai_gilgeous_alexander_2025_total_percentile.html
Naming two players draws both on the same axes and shades each skill to whoever leads it.
Asking about one game ("steph curry's fingerprint from his last game") draws
that game in its own net points rather than a per-100 rate. It needs the
per-game NetPoints files, which are an opt-in pull
(--include-net-points-daily).
Features
- Resumable, rate-limited fetch from ESPN's stats APIs — checkpointed per game and season, safe to interrupt, cheap to re-run
- A local DuckDB warehouse built from the Parquet on disk, rebuildable any time without touching the network
- Data coverage auditing — cross-check what's on disk against what ESPN reports, optionally live
- Natural-language queries with no cloud calls — everything runs against a local Ollama model
- Fast, deterministic answers for common question shapes (rankings by any stat or over a career, player and team lines, splits, with/without a teammate, streaks, head-to-head, comparisons, game logs, shot charts, fingerprints, multi-season history, …), and a refusal naming what is missing for anything else
- Computed advanced stats ESPN's API doesn't expose directly — true shooting %, effective FG%, usage rate, game score
- NetPoints ratings from ESPN Analytics — player/team ratings plus a per-play-type "fingerprint" breakdown, as numbers or as a radar plot
- A local web interface (
association web) — the same answers in a chat-shaped page, rendered as tables, sparklines and cards per question shape, with progress streamed while a slow question runs - A full trace of every query — command, tool calls, timing, and answer — written to disk regardless of verbosity
- Shell completion for bash, zsh, and fish
Setup
uv tool install git+https://github.com/jeffknupp/association@v5.0.0
brew install ollama # or see https://ollama.com/download
ollama serve &
ollama pull qwen2.5:3b # the normalizer - required, ~1.9GB
One model: it copies the names and the stat out of the question for the parser, which reads everything else from the words.
Note —
pip install association-pyis the install command from the next release on: the distribution is namedassociation-pyon PyPI, since PyPI refuses the bare name though nobody holds it (the import package, theassociationcommand and this repository keep it). Nothing is on PyPI yet, so install from the tag as above for now. Releases also attach their wheel and sdist to the releases page. This note goes away once the first version is up.
To work on association itself, clone the repo and sync it instead (see
Development).
Hardware
Recommended: 8GB of RAM. The normalizer's model uses about 2GB once loaded, leaving headroom for DuckDB and the OS. No GPU is required.
Shell completion
Tab-complete subcommands, options, and --log-level's choices — generated
directly from the CLI's own command definitions, so there's nothing to keep in
sync by hand.
# bash
echo 'source /path/to/association/completions/association.bash' >> ~/.bashrc
# zsh
echo 'source /path/to/association/completions/association.zsh' >> ~/.zshrc
# fish
cp completions/association.fish ~/.config/fish/completions/
Or generate it fresh, which picks up any future CLI changes automatically:
eval "$(_ASSOCIATION_COMPLETE=bash_source association)" # bash, in ~/.bashrc
eval "$(_ASSOCIATION_COMPLETE=zsh_source association)" # zsh, in ~/.zshrc
_ASSOCIATION_COMPLETE=fish_source association | source # fish, in ~/.config/fish/config.fish
Documentation
Full documentation — architecture, a command reference, usage recipes, data source notes, and the complete API — is at association.readthedocs.io, or build it locally:
uv sync --extra docs
scripts/build_docs.sh # docs/_build/html/index.html
Data
Fetches teams, games/box scores, standings, player and team season stats, ESPN's Basketball Power Index, and optionally play-by-play, shot charts, and win probability. It also pulls NetPoints — ESPN Analytics' advanced player/team rating — from a separate, unauthenticated source. See Data sources for exactly what's fetched from where, and Architecture for how the warehouse and query engine are built from it.
Project layout
src/association/
cli/ the `association` command (commands.py) and the default warehouse/data paths (paths.py)
nba/ what fetch and query both need to know: seasons and Eastern dates, franchise names by season, each table's coverage floor, NetPoints categories
fetch/ client, endpoints, parse, storage, pipeline, warehouse
repairs/ load-time repairs of ESPN's faults, and the filtered game list and rebuilt box lines
check/ data coverage report, cross-checked live against ESPN
query/ the reader (parse.py, what the model is told in normalizer.py, and the stages in router.py), query templates, entity resolution, leaderboard, conditions (splits, with/without, streaks), team metrics, shot chart, fingerprint, prompt/knowledge base, tools, court and radar renderers, agent loop
web/ the local web interface: HTTP API, one-at-a-time runner, single-page app
scripts/
backfill_markers.sh re-derive completion markers for data fetched before they existed
check_coverage.py verify each table's coverage floor against a built warehouse
check_nicknames.py verify the player nickname table against a built warehouse
check_net_points_games.py verify which game each per-game NetPoints row lands on
completions/ generated bash/zsh/fish shell completion scripts
tests/ pytest, one file per source module
.history/ one command/trace/timing log per question, from query and web (gitignored)
Development
uv sync --extra dev --extra docs --extra web
pre-commit install # one-time, wires the git hook
pytest -q
All three extras are needed even just to run the tests and gates: the docs
build is one of the pre-commit hooks, and tests/web/ imports fastapi
directly with no skip guard, so it fails to collect without the web extra
installed.
ruff and mypy run as pre-commit hooks, along with docstring coverage and
type completeness checks on src/. The same checks run in CI on push and pull
request, alongside the docs build. Any commit touching src/ must also update
CHANGES.md,
enforced by a pre-commit hook.
Known limitations
- ESPN's stats API is undocumented and unofficial — endpoints or shapes can change without notice.
- ESPN's own Real Plus-Minus (RPM) isn't available at a stable JSON endpoint — NetPoints (see above) is its successor and is included. PER, Win Shares, BPM, and VORP are not included; see the computed advanced stats notes in the docs for why.
- NetPoints only has data back to the 2018-19 season — nothing earlier exists on their side.
net_points_team(season-level) only ever reflects the current season;net_points_team_game(per-game,--include-net-points-daily) has full history instead.net_points_player_gameandnet_points_player_game_fingerprint(both per-game,--include-net-points-daily) match players by exact display-name text, so spelling differences between sources can leave a real player's game rows unmatched rather than wrongly matched.- Tables start in different seasons. Standings from 1987-88; playoffs from 1989; box scores and regular-season games from 1993-94; play-by-play from 2001-02 (about half of that season); shot charts from 2001-02 (only part of 2001-02 and 2002-03); ESPN's power index from 2016-17; win probability from 2017-18; NetPoints from 2018-19. These are ESPN's gaps, and no pull fills them. A question below a table's floor is refused with the reason.
- The 2001 playoffs are missing ten games, Games 1-4 of the Lakers-76ers Final among them, and ESPN's archive has them nowhere. A 2001 playoff answer carries a note saying which series are short.
- Careers. Career answers cover only careers that reached 1993-94, because players are found through box scores: Kareem Abdul-Jabbar, Larry Bird and Julius Erving are not in the warehouse, and a career list says it is not all-time.
- Empty box scores. Every game Chicago or New Orleans played from 2012-13
to 2017-18, playoffs included, has an empty box score for both teams in
ESPN's data, apart from two games. Per-game and per-condition answers (game
logs, single-game highs, threshold counts, splits, streaks, with/without,
head-to-head) rebuild the missing line from play-by-play for points,
rebounds, assists, steals, blocks, and made field goals and free throws, and
say the figure is rebuilt; turnovers and fouls are refused rather than
guessed. Season totals summed directly from the box scores are still short
(about 87% of ESPN's own totals for those team-seasons) since a rebuilt
season aggregate is right only about half the time -
player_season_stats, a separate ESPN endpoint, is unaffected. - No round or conference data. Nothing records a playoff round or a team's conference.
data check --livecross-checks are opt-in and can be slow for seasons without a local completion marker yet —pullfirst to build those up.
Metadata
Release files for association-py 5.0.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| association_py-5.0.0.tar.gz | 1.6 MB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| association_py-5.0.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 2.6 MB
Release files / association_py-5.0.0.tar.gz
| Download URL | association_py-5.0.0.tar.gz |
|---|---|
| Size | 1.6 MB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
d15bd894936b709c957681aa067655e76d20c289aaa364a6d108505db6a17fce
|
|
BLAKE2b-256 checksum How to use checksums |
f11e53c9837e93f4dbe792593ddead77dd6de777e26174cc008b21254c19a417
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 9, 2026.
Transparency logRelease files / association_py-5.0.0-py3-none-any.whl
| Download URL | association_py-5.0.0-py3-none-any.whl |
|---|---|
| Size | 975.1 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
2d3522787b121d8622b6eb9bc2b5c653e66361e568062657eb6262d7a405d60a
|
|
BLAKE2b-256 checksum How to use checksums |
538ec5b2e8aefa2612d256477147a65b564863ba034fcb434795da795e0932f6
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 9, 2026.
Transparency log