Representing probability distributions by the gradient of their density functions has proven effective in modeling a wide range of continuous data modalities.
Since the data are high-dimensional or the network is large-scale, communication load can be a bottleneck for the efficiency of distributed algorithms.