A fundamental problem in deep neural networks is toverify or certify that a trained network is robust, i.e. not susceptible to adversarial attacks [11, 29, 39].
Instead, we seek to learn a fair distribution overthearms. Drawing onalong lineofresearch ineconomics and computer science, we use theNash social welfareas our notion of fairness.
BLUE (Biomedical Language Understanding Evaluation) is abenchmarkfor 10 datasetsrepresenting 5 tasks [34]. BLURB (Biomedical Language Understanding and Reasoning Benchmark) includes 13 datasetsand 7 tasks [19].