No look-ahead
Ratings and ensemble weights are frozen at their pre-match values before the result updates them.
Duna replays beach volleyball history in order. Each model must publish its probability before it sees the result, then earn its place on accuracy, calibration, Brier score, and log loss.
Ratings and ensemble weights are frozen at their pre-match values before the result updates them.
Simple baselines, Elo variants, Duna ablations, and an online adaptive ensemble compete on the same matches.
A challenger does not reach production because it sounds advanced. It must improve walk-forward error and remain explainable.
Duna models team strength with extra weight on the weaker partner, reflecting targeting and side-out pressure. Updates consider score margin, uncertainty, evidence quality, and repeat-opponent decay. Weekly display gains are capped, while losses remain uncapped.
A player’s public 1.00–8.00 number is a readable projection of an internal strength and uncertainty state. World ranking points remain a separate official signal; they are never relabeled as Sand Rating.
The adaptive ensemble is an online machine-learning model: after each result, it shifts weight toward component models with lower prior log loss. The current lab does not use reinforcement learning, and Duna will not claim that it does unless an actual RL policy is trained, evaluated, and documented.
AI agents may propose new features or challenger models. They cannot silently change the live rating. Every proposal must run through the same chronological evaluation, data-integrity review, versioned configuration, and human promotion gate.
“Champion” means the lowest Brier score in this run, with log loss as the tie-breaker. It is not an automatic production deployment.
The diagonal is ideal calibration. Each marker shows predicted probability against the observed win rate; its number is the match count in that bucket.

