Methodology

How ProbGoal builds and validates forecastsTransparent model pipeline
5 seasonshistorical depthRolling OOSvalidationCalibratedprobabilities2M runsofficial release

Five seasons stored. Recent football weighted more.

ProbGoal retains five completed seasons plus the current season where available. Candidate models use continuous exponential time decay so influence decreases smoothly with match age rather than changing abruptly at season boundaries.

Attack and defence are separate

The primary goals family estimates club attacking strength and defensive strength separately, together with home advantage. This allows clubs with similar results but different scoring and conceding profiles to remain statistically distinct.

European matches bridge domestic leagues

Domestic competitions estimate within-league club strength. Cross-league UEFA matches connect those domestic systems through an explicit league-strength bridge instead of assuming that equal domestic positions imply equal European strength.

Rolling out-of-sample testing

Every validation fold trains only on earlier matches and is scored on later matches. Test windows do not overlap. Because the first product is the Champions League, model selection is evaluated specifically on historical Champions League outcomes.

An untouched final holdout

The last validation folds are kept out of candidate selection. The winning model and its holdout calibrator are determined without seeing those results, then the complete pipeline is judged on that final temporal holdout.

Proper scoring and calibration

Log Loss is the primary selection metric and multiclass Brier Score is secondary. Reliability is measured across home, draw and away probabilities. The selected model is calibrated with multiclass temperature scaling fitted only on prior out-of-sample predictions.

A baseline must be beaten

The final pipeline is compared with a fold-specific historical home, draw and away frequency baseline. Date-block bootstrap intervals are used to test whether apparent gains are robust rather than just a favorable point estimate.

Regulation-time integrity

Extra-time matches are excluded from the current 90-minute fitting layer when the historical source does not provide a guaranteed regulation-time score split. Penalty-shootout tallies are not counted as match goals.

Tournament simulation

The competition engine uses the verified 36-club, 144-fixture UEFA draw, applies the official league-phase ranking order and propagates the resulting positions through the prescribed knockout bracket. A public release requires exactly 2,000,000 seeded Monte Carlo runs.

Goals, discipline and penalties are explicit modelling layers

The validated match target is home/draw/away probability, so the tournament engine projects that calibrated vector into a bounded score layer for goal-based tie-breaks. Future disciplinary cards cannot be known ex ante; only if teams remain tied through the preceding official criteria, a seeded exchangeable proxy occupies that tie-break position. Penalty shoot-outs remain 50/50 until a separate penalty model is validated.

No forced publication

A completed backtest is not automatically a successful model. If the independent release gate fails, ProbGoal keeps public probabilities blocked and preserves the failure as a diagnostic result.

PG Score is descriptive — and deliberately separate

PG Score is a 0–100 observed-performance index, not a tournament probability and not an input to the validated match model. Its pre-season version was selected from 100 historical-prior candidates using 836 temporal validation pairs. The selected specification normalizes six observed-result signals inside each competition-season, uses a 540-day exponential half-life and weights context volume by the square root of matches. The richer Attack, Defence, Dominance, Efficiency and Discipline live layer remains blocked until homogeneous boxscore coverage and robustness evidence are sufficient.