A Human Evaluation Details A.1 Unlearning Toxicity Human Eval Details

Neural Information Processing Systems 

In total we have 1200 comparisons, and each comparison is rated by 3 raters. In total we have 2400 comparisons, and each comparison is rated by 3 raters. These were: 1. Coherence: Is the system's generation aligned in meaning and topic with the prompt? We sampled 100 prompts randomly from the corpus, and then evaluated 19 different algorithms. HITs was 2.2K, and the total number of ratings was 6.6K.

Similar Docs  Excel Report  more

TitleSimilaritySource
None found