Goto

Collaborating Authors

 Optimization





NaturalCounterfactualsWithNecessaryBacktracking

Neural Information Processing Systems

Ourmethodologyincorporates a certain amount of backtracking when needed, allowing changes in causally preceding variables tominimize deviations from realistic scenarios. Specifically, we introduce a novel optimization framework that permits but also controls the extent of backtracking with a "naturalness" criterion. Empirical experiments demonstrate the effectiveness of our method.




4b3cc0d1c897ebcf71aca92a4a26ac83-Paper-Conference.pdf

Neural Information Processing Systems

More specifically,for the output features ofthe penultimate layer, for each class the within-class features converge to their means, and the means of different classes exhibit a certain tight frame structure, which is also aligned withthelastlayer'sclassifier.


Catastrophic Goodhart: regularizing RLHF with KL divergence does not mitigate heavy-tailed reward misspecification

Neural Information Processing Systems

However, if error is heavy-tailed, some policies obtain arbitrarily high reward despite achieving no more utility than the base model-a phenomenon we call catastrophic Goodhart. We adapt a discrete optimization method to measure the tails of reward models, finding that they are consistent with light-tailed error.