Assumptions¶
Every constant the models use, generated from the registry TOML files in
src/athletevalue/assumptions/data when the site is built. athletevalue
assumptions prints the same list, and a governance test fails on a bare number
in modelling code. Status: ● reported, ◆ derived, ▲ estimated, ○ scenario.
Override any entry with --assumptions my.toml.
Program economics¶
economics.bid_calibration_max_gap¶
Largest allowed gap between predicted and observed bid rates in any decile of predicted probability.
Value 0.05 (probability) · basis modelling_choice · status ▲ estimated
Rationale. Scored on the fitting rows (about 490 team-seasons per decile), so a gap this size is a misspecified logistic form rather than sampling noise. Through 2026 the largest gap is 0.025.
economics.bid_ridge_lambda¶
Ridge penalty on the bid model's coefficients, the intercept excluded.
Value 0.01 (penalty) · basis modelling_choice · status ▲ estimated
Rationale. Present for numerical stability, not shrinkage: leaving one season out can come close to separating the data, where an unpenalised fit runs coefficients off to large values. The coefficients are on the natural win-percentage scale and are near 15, so a penalty of 1 already distorts calibration; at 0.01 no coefficient moves materially.
economics.bootstrap_draws¶
School-cluster bootstrap replicates for the revenue model.
Value 400 (count) · basis modelling_choice · status ▲ estimated
Rationale. Enough to estimate an 80% interval of a coefficient to within a few percent.
economics.conference_unit_share_multiplier¶
Multiplier on an equal split of tournament-unit money among conference members. 1.0 is an equal split; conferences that reward the earning school use more.
Value 1 (multiplier) · basis user_input · status ○ scenario
economics.excluded_seasons¶
Seasons left out of the revenue and bid fits.
Value 2,020, 2,021 (seasons) · basis modelling_choice · status ▲ estimated
Rationale. The 2020 NCAA tournament was cancelled, so no team has a bid that season. Attendance limits in 2020-21 cut ticket revenue for reasons unrelated to winning.
economics.house_revenue_share_cap_2026¶
Per-school cap on direct revenue sharing with athletes across all sports under the House v. NCAA settlement, 2025-26. Context for roster budgets, which also include collective money; not used in any calculation.
Value 20,500,000 (USD) · basis published_third_party · status ▲ estimated
Source: ESPN: Judge OK's $2.8B settlement, paving way for colleges to pay athletes
The annual cap is expected to start at roughly $20.5 million per school in 2025-26 and increase every year during the decade-long deal.
economics.max_prob_nonpositive_win_effect¶
Largest allowed share of bootstrap draws with a current-season win effect at or below zero.
Value 0.1 (probability) · basis modelling_choice · status ▲ estimated
Rationale. Program value rests on wins raising revenue. If more than one draw in ten says a win does not, the dollar figures would present noise as value. Through 2026 no draw is at or below zero.
economics.panel_first_season¶
First EADA season used in the revenue panel (needs one lagged season).
Value 2,012 (season) · basis modelling_choice · status ▲ estimated
Rationale. Possession data start in 2011, so 2012 is the first season with a lagged outcome.
economics.panel_min_seasons¶
Schools need this many non-allocated seasons to enter the revenue fit.
Value 4 (count) · basis modelling_choice · status ▲ estimated
Rationale. School fixed effects with two lags are not identified from fewer seasons.
economics.revenue_reference_seasons¶
Recent seasons averaged to set a school's revenue base for dollars per win.
Value 3 (count) · basis modelling_choice · status ▲ estimated
Rationale. Smooths one-off accounting changes in EADA without reaching back past a coaching era.
economics.tournament_unit_value¶
Total value of one NCAA men's tournament unit, paid to the conference in annual installments. The model discounts those installments to present value; see economics.unit_payout_installments and economics.unit_discount_rate.
Value 2,000,000 (USD) · basis published_third_party · status ▲ estimated
Source: Deseret News, via Yahoo Sports: Here's how NCAA tournament schools make money in March Madness
Each of the now 135 units up for grabs in recent years have been worth about $2 million.
economics.unit_discount_rate¶
Annual rate used to bring later unit installments to present value, and to discount next season's revenue carry-over by one year.
Value 0.05 (rate per year) · basis user_input · status ○ scenario
economics.unit_payout_first_year¶
Years between the tournament and the first installment. Units earned in one tournament are distributed the following April, so the first dollar arrives in the next fiscal year, not the season being valued.
Value 1 (years) · basis modelling_choice · status ▲ estimated
Rationale. The NCAA Division I revenue distribution plan pays the basketball performance fund each April for units earned in previous tournaments, which falls in the fiscal year after the season. Stated as a modelling choice because no page quotable under the citation rules pins the lag itself.
economics.unit_payout_installments¶
Annual installments a tournament unit is paid in. The unit value is the total of these, so the model discounts them rather than counting the whole sum at once.
Value 6 (years) · basis published_third_party · status ▲ estimated
Source: Deseret News, via Yahoo Sports: Here's how NCAA tournament schools make money in March Madness
Conferences distribute the money, roughly $350,000 a year over six years, to their schools as they see fit.
economics.unit_revenue_overlap¶
How the tournament-unit component and the EADA bid effect are kept from counting the same money. Units earned in one tournament are paid from the next April, so the same-season bid effect cannot contain them and the annual figure never overlaps; the first installment can sit inside next season's reported revenue. One of: "next_season_installment" subtracts that installment from the two-season figure; "none" assumes the filings exclude unit money; "exclude_units" drops the unit component from the headline figures and reports it alongside.
Value next_season_installment (policy) · basis user_input · status ○ scenario
Roster market¶
market.budget_range_level¶
Probability mass assigned to a cited (low, high) budget range.
Value 0.8 (probability) · basis modelling_choice · status ▲ estimated
Rationale. The study reports averages "between" two figures without a stated interval. Reading the range as covering most programs, not all, leaves room for outliers in both directions without inventing a distribution the source does not give.
market.min_possession_share¶
Players on the floor for less than this share of team possessions are treated as unpaid walk-ons and receive no allocation.
Value 0.05 (share) · basis user_input · status ○ scenario
market.model.label_deal_types¶
Deal types counted as roster pay when building training labels.
Value revenue_share, collective (deal types) · basis modelling_choice · status ▲ estimated
Rationale. Revenue-share and collective payments are what a roster budget buys. Endorsements and appearances are commercial NIL, a different quantity this model does not price.
market.model.min_labels¶
Labeled player-seasons needed before the fitted model can be used.
Value 40 (count) · basis modelling_choice · status ▲ estimated
Rationale. About three observations per feature. Below this, leave-one-school-out intervals are too coarse to beat the allocation in any meaningful way.
market.model.min_schools¶
Distinct schools needed among labels; cross-validation holds out one school at a time.
Value 10 (count) · basis modelling_choice · status ▲ estimated
Rationale. Pay is set school by school. With fewer schools, held-out error mostly measures which school was left out, not how well the model generalizes.
market.model.require_beats_allocation¶
Use the fitted price only if its held-out log error beats the allocation's and a tier median.
Value True (flag) · basis modelling_choice · status ▲ estimated
Rationale. A model that cannot beat the transparent allocation out of sample adds complexity without adding accuracy.
market.model.ridge_grid¶
Ridge penalties tried on standardized features, chosen by leave-one-school-out error.
Value 0.1, 1, 10, 100 (dimensionless) · basis modelling_choice · status ▲ estimated
Rationale. From nearly unpenalized to nearly intercept-only on a few dozen rows; a decade apart because error is flat within a factor of three.
market.performance_exponent¶
Exponent on a player's impact above replacement when splitting pay within a role. 0 pays every starter the same; 1 pays in proportion to impact.
Value 1 (dimensionless) · basis user_input · status ○ scenario
market.performance_floor¶
Impact above replacement assigned to a player whose estimate is at or below replacement, so a paid roster spot never allocates zero or negative dollars.
Value 0.5 (points per 100 possessions) · basis user_input · status ○ scenario
market.power_conferences¶
Conferences in the top roster-spending tier, as named by the cited study.
Value ACC, Big 12, Big East, Big Ten, SEC (conference names) · basis published_third_party · status ▲ estimated
The power conferences – the Big Ten, Big 12, ACC, SEC and Big East – spend between $7 million and $10 million on a roster, on average, according to Opendorse.
market.role_weight_bench¶
Pay of a bench player relative to a starter (low, high).
Value 0.45, 0.7 (ratio) · basis published_third_party · status ▲ estimated
Rotational players also earn roughly 15-25% less than starters while bench players make 30-55% less
market.role_weight_rotation¶
Pay of a rotation player relative to a starter (low, high).
Value 0.75, 0.85 (ratio) · basis published_third_party · status ▲ estimated
Rotational players also earn roughly 15-25% less than starters while bench players make 30-55% less
market.roster_budget_power_2026¶
Average men's basketball roster budget, power tier, 2025-26 (low, high).
Value 7,000,000, 10,000,000 (USD) · basis published_third_party · status ▲ estimated
spend between $7 million and $10 million on a roster, on average
market.roster_budget_tier_b_2026¶
Men's basketball roster budget, B tier, 2025-26 (low, high).
Value 2,000,000, 5,000,000 (USD) · basis published_third_party · status ▲ estimated
From there, the "B Tier" rosters in men's basketball cost between $2 million and $5 million.
market.roster_budget_tier_c_2026¶
Men's basketball roster budget, C tier, 2025-26 (low, high).
Value 1,000,000, 2,000,000 (USD) · basis published_third_party · status ▲ estimated
Other low- and mid-major leagues have rosters between $1 million and $2 million in the "C Tier."
market.rotation_size¶
Players ranked 6th and below by playing time who count as rotation rather than bench.
Value 3 (count) · basis user_input · status ○ scenario
market.tier_b_conferences¶
Conferences the cited study places in its middle spending tier, in this package's conference names. The article names the American, Atlantic 10, Mountain West, Pac-12 and Horizon League; the Pac-12 has no members in the 2025-26 data.
Value AAC, Atlantic 10, MWC, Horizon (conference names) · basis published_third_party · status ▲ estimated
Conferences in that category include the American, Atlantic 10, Mountain West, Pac-12 and Horizon League.
Player impact¶
mbb.impact.cv_folds¶
Folds for cross-validation, grouped by game so no game is split.
Value 5 (count) · basis modelling_choice · status ▲ estimated
Rationale. Standard k for ~6,000 games; each fold still holds every team.
mbb.impact.drop_garbage_time¶
Exclude possessions the source flags as garbage time.
Value True (flag) · basis modelling_choice · status ▲ estimated
Rationale. Lineups in decided games do not try to win the possession, which biases coefficients of deep-bench players. The flag comes from the source data.
mbb.impact.home_court_band¶
Plausible home-court advantage, points per 100 possessions; a sign/scale gate.
Value 1, 4 (points per 100 possessions) · basis published_third_party · status ▲ estimated
Source: SportsDataverse ncaa-mbb-hoops-data: RAPM model card, gate table
home-court advantage band | [1.0, 4.0]
mbb.impact.intercept_band¶
Plausible fitted intercept (points per 100 possessions); a scale-bug gate.
Value 95, 112 (points per 100 possessions) · basis published_third_party · status ▲ estimated
Source: SportsDataverse ncaa-mbb-hoops-data: RAPM model card, gate table
intercept era band (scale-bug catcher) | [95, 112]
mbb.impact.lambda_grid¶
Penalties tried when the lambda is chosen by grouped cross-validation.
Value 100, 300, 1,000, 3,000, 10,000 (dimensionless) · basis modelling_choice · status ▲ estimated
Rationale. Half-decade spacing around the published value. Held-out error is flat near its minimum on this scale, so a finer grid changes the choice without changing the error.
mbb.impact.min_possessions¶
Players below this many offensive-plus-defensive possessions / 2 share a single pooled column instead of getting their own coefficient.
Value 100 (possessions) · basis modelling_choice · status ▲ estimated
Rationale. About three full games. Below it the ridge estimate is almost entirely prior; pooling keeps those possessions in the fit and yields an empirical low-minute baseline to compare with the replacement-level convention.
mbb.impact.reference_min_possessions¶
Offensive possessions required for a player to enter the reference comparison.
Value 500 (possessions) · basis modelling_choice · status ▲ estimated
Rationale. Roughly a third of a starter's season; below it both fits are prior-dominated.
mbb.impact.reference_spearman_min¶
Minimum Spearman correlation between this package's net RAPM and the SportsDataverse league-wide RAPM, for players above the comparison threshold.
Value 0.9 (correlation) · basis modelling_choice · status ▲ estimated
Rationale. Both fits use the same possessions and penalty but differ in pooling, garbage-time handling and neutral sites, so identical rankings are not expected. Below 0.9 a design-matrix bug is more likely than a modelling difference.
mbb.impact.ridge_lambda¶
Ridge penalty on player coefficients.
Matches the published SportsDataverse league-wide model so results can be regression-tested against it. The penalty implies a prior SD of sigma/sqrt(lambda) per coefficient. Cross-validation by game is available and reported alongside.
Value 1,000 (dimensionless) · basis published_third_party · status ▲ estimated
Source: SportsDataverse ncaa-mbb-hoops-data: RAPM model card
Ridge regression (λ = 1000) over the possession-level on/off
mbb.impact.ridge_lambda_with_prior¶
Ridge penalty when shrinking toward the team-adjusted box-score prior.
Value 3,000 (dimensionless) · basis modelling_choice · status ▲ estimated
Rationale. Leak-free grouped cross-validation (box totals and team adjustment rebuilt per fold), lambda 1000/3000/10000/30000: 2026 held-out error 5,182.4 / 5,177.6 / 5,178.1 / 5,178.7; 2025 5,052.8 / 5,046.8 / 5,046.5 / 5,046.8. The curve is flat from 3000 up; the smallest value on the flat part keeps the most weight on each player's own possessions.
mbb.impact.sigma2_band¶
Plausible residual variance per possession, times 100^2; an SE scale gate.
Value 11,000, 15,000 ((points per 100 possessions)^2) · basis published_third_party · status ▲ estimated
Source: SportsDataverse ncaa-mbb-hoops-data: RAPM model card, gate table
σ̂² era band (SE scale-bug catcher) | [11000, 15000]
mbb.impact.torvik_min_teams¶
Teams that must be matched to Torvik for the team-rating comparison to count.
Value 250 (count) · basis published_third_party · status ▲ estimated
Source: SportsDataverse ncaa-mbb-hoops-data: RAPM model card, gate table
Torvik external (league-wide only) | ≥ 250 joined teams AND Spearman(team_net, adjem) ≥ 0.93
mbb.impact.torvik_spearman_min¶
Minimum Spearman correlation of team net rating with Torvik AdjOE minus AdjDE.
Value 0.93 (correlation) · basis published_third_party · status ▲ estimated
Source: SportsDataverse ncaa-mbb-hoops-data: RAPM model card, gate table
Torvik external (league-wide only) | ≥ 250 joined teams AND Spearman(team_net, adjem) ≥ 0.93
Box-score prior¶
mbb.prior.box_ridge¶
Ridge penalty on standardized box-model coefficients.
Value 1 (dimensionless) · basis modelling_choice · status ▲ estimated
Rationale. Negligible against ~13,000 possession-weighted rows; it only stabilizes the solve when shot-location rates are collinear.
mbb.prior.max_cv_error_ratio¶
Largest allowed ratio of held-out error with the prior to error without it.
Value 1 (ratio) · basis modelling_choice · status ▲ estimated
Rationale. A prior that raises held-out error should not be used; this gate reports it.
mbb.prior.rate_pseudo_possessions¶
League-average possessions added to each player's box rates before prediction.
Value 100 (possessions) · basis modelling_choice · status ▲ estimated
Rationale. Matches the RAPM pooling threshold. A player with 30 possessions keeps about a quarter of his raw rate; a starter with 2,000 keeps 95%.
mbb.prior.returning_min_gain¶
Smallest allowed gain in correlation with next season's no-prior ratings, ratings with the prior minus ratings without it, for returning players.
Value 0 (correlation) · basis modelling_choice · status ▲ estimated
Rationale. The prior exists to predict a player better than that player's own noisy season does. If it does not raise the correlation with the next season's independent fit, it is not earning its place. 2025 ratings give 0.524 with the prior and 0.400 without.
mbb.prior.returning_min_possessions_earlier¶
Offensive possessions a player needs in the earlier season to enter the returning-player check.
Value 200 (possessions) · basis modelling_choice · status ▲ estimated
Rationale. Low enough to include the part-time players the prior matters most for, high enough that the rating is not the pooled coefficient.
mbb.prior.returning_min_possessions_later¶
Offensive possessions a player needs in the later season to enter the returning-player check.
Value 500 (possessions) · basis modelling_choice · status ▲ estimated
Rationale. The later no-prior rating is the target, so it needs enough possessions to be more signal than noise; 500 matches the reference-comparison threshold.
mbb.prior.team_adjustment¶
Shift box predictions so each team's players sum to its no-prior team rating.
Value True (flag) · basis modelling_choice · status ▲ estimated
Rationale. Box rates ignore opponent strength. Without the adjustment the 2026 team ratings match Torvik at Spearman 0.925 (0.963 with no prior at all); with it, 0.971. The adjustment is what restores the scale of team strength that shrinkage toward zero compresses: it, not the box scores, accounts for most of the held-out gain.
mbb.prior.training_seasons¶
Most recent seasons before the target season used to fit the box model.
Value 4 (count) · basis modelling_choice · status ▲ estimated
Rationale. Four seasons give about 13,000 player-seasons, enough for 14 coefficients, while staying within one rules and pace era. Only earlier seasons are used so the prior never sees the season it is applied to or any later one.
mbb.prior.use_box_prior¶
Shrink player ratings toward a box-score prediction instead of toward zero.
Value True (flag) · basis modelling_choice · status ▲ estimated
Rationale. In cross-validation that holds out whole games and rebuilds both the box totals and the team adjustment from training games only, held-out error at lambda 3000 falls from 5,186.6 to 5,177.6 in 2026 and from 5,055.5 to 5,046.8 in 2025. Most of that gain comes from the team adjustment: a prior with zero box information and the same adjustment reaches 5,182.1 in 2026. The prior_cv_error_ratio gate rechecks the whole prior for every season validated.
Wins¶
mbb.wins.bench_rank_range¶
Playing-time ranks within a team that define the bench for bench_median.
Value 9, 12 (rank) · basis modelling_choice · status ▲ estimated
Rationale. Five starters and three rotation players are the paid core in the market allocation; the next four are who plays when one of them cannot. Ranks below 12 are walk-ons and injuries.
mbb.wins.exponent_search_bounds¶
Interval searched when fitting the Pythagorean exponent.
Value 2, 30 (dimensionless) · basis modelling_choice · status ▲ estimated
Rationale. Brackets every published basketball exponent (about 8 to 17) with room on both sides.
mbb.wins.game_margin_sd¶
Standard deviation of a game's final margin around its expected margin, fitted each season from team ratings and actual scores.
Value derived from data (points) · basis derived_from_data · status ◆ derived
mbb.wins.interval_level¶
Probability mass of reported intervals.
Value 0.8 (probability) · basis modelling_choice · status ▲ estimated
Rationale. An 80% interval still shows most of the noise in one college season, and stays narrow enough to separate a starter from a bench player.
mbb.wins.margin_sd_band¶
Validation band for the fitted game-margin spread, points.
Value 10, 12.5 (points) · basis modelling_choice · status ▲ estimated
Rationale. The D1-only fit gives 11.15 (2025) and 10.87 (2026). A spread outside this band means ratings, venues or scores disagree with the games, for example after a game-matching failure, and every WAR figure scales with its inverse.
mbb.wins.margin_sd_d1_only¶
Fit the game-margin spread on games between two D1 teams only.
Value True (flag) · basis modelling_choice · status ▲ estimated
Rationale. Non-D1 opponents share one coefficient, so their margins are not something the ratings can predict; including them raised the spread from 10.9 to 11.6 points in 2026.
mbb.wins.non_d1_opponent_floor¶
Rate a non-D1 opponent no worse than the worst D1 team in the win model.
Value True (flag) · basis modelling_choice · status ▲ estimated
Rationale. The shared non-D1 coefficient (about -55 per 100 in 2026) is pulled down by the blowouts it absorbs; no D1 team rates below about -35. Either value makes those games near-certain wins, so the floor changes little and avoids an implausible number.
mbb.wins.pythag_exponent¶
Exponent in expected win% = ORtg^x / (ORtg^x + DRtg^x), fitted each season by least squares on team win percentage.
Value derived from data (dimensionless) · basis derived_from_data · status ◆ derived
mbb.wins.pythag_exponent_reference¶
Published college exponents, used only to check the fitted exponent.
Value 10.25, 11.5 (dimensionless) · basis published_third_party · status ▲ estimated
Source: kenpom.com blog: Ratings Explanation (archived)
I am using 11.5 as the exponent.
mbb.wins.replacement_definition¶
Which replacement level WAR and pay use: nba_convention (the Box Plus/Minus
constant below), pooled (this season's fitted coefficient for players under the modelling
threshold) or bench_median (the possession-weighted median rating of players ranked in the
bench range by playing time on each team).
Value bench_median (choice) · basis user_input · status ○ scenario
mbb.wins.replacement_level¶
Net rating of a replacement player under the NBA Box Plus/Minus convention.
Kept as one of three definitions. It was borrowed from another league; in these data the pooled low-minute coefficient sits far below it (about -12 in 2026), and the choice roughly doubles every WAR and dollar figure, so it is a scenario input, not a fact.
Value -2 (points per 100 possessions) · basis user_input · status ○ scenario
Source: Basketball-Reference glossary: Value Over Replacement Player
the points per 100 TEAM possessions that a player contributed above a replacement-level (-2.0) player
mbb.wins.simulation_draws¶
Monte Carlo draws used to propagate rating uncertainty into wins and dollars.
Value 4,000 (count) · basis modelling_choice · status ▲ estimated
Rationale. Monte Carlo error on an 80% interval endpoint is under 1% of its width at 4,000 draws.
mbb.wins.team_war_min_correlation¶
Smallest allowed correlation between a team's summed player WAR and its wins.
Value 0.7 (correlation) · basis modelling_choice · status ▲ estimated
Rationale. Summed linear WAR correlates 0.79 (2025) and 0.80 (2026) with team wins. It cannot reach 1: WAR is linear in net rating, wins also depend on close-game luck, and garbage time is excluded. A drop below 0.7 would mean the shares, schedule slopes or ratings no longer add up to team results.