Goto

Collaborating Authors

 self-correct bias


Language models might be able to self-correct biases--if you ask them

MIT Technology Review

The second test used a data set designed to check how likely a model is to assume the gender of someone in a particular profession, and the third tested for how much race affected the chances of a would-be applicant's acceptance to a law school if a language model was asked to do the selection--something that, thankfully, doesn't happen in the real world. The team found that just prompting a model to make sure its answers didn't rely on stereotyping had a dramatically positive effect on its output, particularly in those that had completed enough rounds of RLHF and had more than 22 billion parameters, the variables in an AI system that get tweaked during training. GPT-3 has around 175 million parameters.) In some cases, the model even started to engage in positive discrimination in its output. Crucially, as with much deep-learning work, the researchers don't really know exactly why the models are able to do this, although they have some hunches.