ProbGoal Analysis · Model audit

ProbGoal passed its release gate. Here is what that actually means

The production model beat a historical-frequency baseline on untouched matches in both proper scoring rules — while calibration stayed inside the safety limit.

A validation badge is useful only if the standard behind it is visible. ProbGoal's current production model was promoted after a chronological holdout test, baseline comparison, calibration checks and paired resampling all passed the configured release gate.

The model was judged on probabilities, not just picks

A model that calls the most likely result correctly can still be badly calibrated or dangerously overconfident. ProbGoal therefore uses multiclass Log Loss as its primary score and multiclass Brier as its secondary score. Lower is better for both.

The comparison was made on the end of the timeline

The holdout contains 186 Champions League target matches that were not used to fit the final decision. The model's top-result hit rate was 56.5% versus 49.5% for the baseline, but that hit-rate advantage is supporting context rather than the release criterion.

The release gate is deliberately stricter than 'better once'

Both probabilistic scores beat baseline, and the paired date-block bootstrap kept the candidate-minus-baseline upper confidence bound below zero for both. Calibration ECE was 4.5%, below the 12% release ceiling. The historical-frequency baseline itself has lower ECE, which is why ProbGoal does not claim to dominate every diagnostic.

Passing this gate means the model earned the right to be used in production under its defined validation protocol. It does not mean every favourite wins, every probability is exact, or future performance is guaranteed.

Sources
Method note

ProbGoal editorial text is generated from a locked factual payload and is blocked from publication when its percentages, scorelines or supporting fact references fail validation.