From the source
“If the lottery is an intensification of chance, a periodic infusion of chaos into the cosmos, would it not be desirable for chance to intervene at all stages of the lottery and not merely in the drawing?” — Jorge Luis Borges, The Lottery in Babylon O pen any recent image-generation paper and the central claim usually rests on a single number — the Fréchet Inception Distance (FID).
FID is the closest thing image generation has to an arbiter: a half-unit shift reorders the leaderboard; a decade of recipes have been justified by single-Inception-unit gains; and budgets in the low millions of GPU-hours hinge on which architecture lands a few decimals lower.
But behind every reported FID number sits a chain of pseudo-random draws (parameter initialisation, minibatch order, per-step Gaussian noise injected by the training loss, hardware stochasticity, and the initial noise drawn at sampling time), any of which could have produced a potentially different score had the seed been different.
Conventional wisdom considers this variance in FID to be negligible, especially for well-trained models.
In this paper we show that the FID reproducibility gap is real and is a serious concern .
…