Accuracy

186 untouched holdout matches · 10/10 release checksModel promoted Release gate passed
2026/27 public forecast scoring starts with Matchday 2.Matchday 1 is observed competition evidence, not an accuracy sample: ProbGoal had no frozen public pre-match snapshot before those 18 games. No retrospective forecast is reconstructed.
RELEASE GATE · PASSED

The model beat its baseline where probability quality matters.

On 186 matches the final model never used for fitting, both proper scoring rules improved against a simple historical-frequency benchmark.

10/10release checks passed
LOG LOSSPRIMARY
0.9555vs 1.0211

6.4% lower than baseline. Lower is better: confident wrong predictions are penalised heavily.

BRIER SCORESECONDARY
0.5646vs 0.6166

8.4% lower than baseline. Lower means the full home/draw/away probability vector was closer to reality.

TOP-RESULT HIT RATECONTEXT
56.5%vs 49.5%

+7.0 pp versus baseline. Useful context, but not the metric used to select a probabilistic model.

01
Unseen means unseen

The holdout sits at the end of the timeline. Those matches were not used to fit the final decision.

02
The advantage survived resampling

The paired date-block bootstrap kept candidate-minus-baseline below zero for both Log Loss and Brier at the configured 95% confidence level.

03
Calibration passed its safety gate

Model ECE is 4.5% against a release ceiling of 12.0%. Lower is better.

So, does it work?

It passed the release standard. That means it outperformed the defined baseline on untouched historical matches under the pre-registered gates. It does not mean every favourite wins, every probability is exact, or future performance is guaranteed.