Technology
[R1/R2] Infinite width assumption: the infinite width assumption is needed due to the technical detail that the norm
We thank reviewers for their valuable comments. We respond to the main concerns below. Similar to that in Zhang et al. [31], we chose 10k block ResNet to stress the We will rephrase L243 to better express this. Derivative of weights depend on this term due to the chain rule. We will make this explicit in the revised manuscript.
e4dd5528f7596dcdf871aa55cfccc53c-AuthorFeedback.pdf
We show that predicting bias from images alone is more challenging than prediction from12 text. The features are then applied on visual-only data,23 e.g. We tried a variant of this approach of predicting latent text topics from images24 and obtained 0.681 on our full dataset (much lower than our method; compare to Tab. 1 in main). The model achieved 0.626 acc. Further balancing of loss hyperparameters could potentially improve the result.