This spectral energy distance is a proper scoring rule with respect to the distribution over magnitude-spectrograms of the generated waveform audio and offers statistical consistency guarantees.
While such an update can be accomplished by re-training on the complete data, the process is inefficient and prevents real-time and on-device learning.