Country
AD-DROP: Attribution-DrivenDropoutforRobust LanguageModelFine-Tuning
Pre-training large language models (PrLMs) on massive unlabeled corpora and fine-tuning them on downstream tasks has become a new paradigm [1-3]. Their success can be partly attributed to the self-attention mechanism [4], yet these self-attention networks are often redundant [5, 6] and tend to cause overfitting when fine-tuned on downstream tasks due to the mismatch between their overparameterization and the limited annotated data [7-13]. To address this issue, various regularization techniques such as data augmentation [14, 15], adversarial training [16, 17]), and dropout-based methods [11,13,18]have been developed.
SupplementaryMaterials: BiologicalCredit AssignmentthroughDynamicInversion ofFeedforwardNetworks
Notethattheaccuracyof δl 1 isnotmeasureddirectlyforReLU because it does not have an explicit inversion. This precludes stability forα = 0 and dl > dl 1 (expanding layer), as the matrix productWlBl will be singular. Forexample,inthenonlinearregression experiment shown in the main text, we initialize the SLDI feedback asB = B1B2, whereB1 and B2 are the feedback matrices for sequential DI. Once again, when the controller has no leak, this will produce the same steady state assequential dynamicinversion. We study a simple case here as an illustration, and leave a more thorough analysis for futurework.
On Measuring Fairness in Generative Models Supplementary Material
These were not included in the main paper due to space limitations. In Sec 4.1 of main paper, we have proposed a statistical model for the sensitive attribute classifier Generators are not completely biased. Given that a generator is trained on a reliable dataset with the availability of all classes of a given sensitive attribute, coupled with the advancement in generator's architecture, it is a fair assumption that the generator would learn some representation of each class in the sensitive attribute and not be completely Here, we provide more information on the necessary assumptions and the expanded forms of the equations. A.2, we will similarly provide more information on MLE value of Population Mean. A.1, we can equate the sample mean to the expanded theoretical model: µ Now given that the classifier's accuracy Fairness in generative models is defined as Equal Representation meaning that the generator is supposed to generate an equal number of samples for each element of an attribute, e.g., an equal number In the main paper Sec.3, we discussed that there could be considerable error in the fairness measurement, In our extended experiments in Sec.