Implicit biases from word embeddings: methods – RealThinks
Friends, this post is an extension of my previous work at analyzing language in reduced space. Here's a quick refresher in case you don't want to go back and read the whole thing: Using word2vec, I took a several blocks of text and projected them into a 100-dimensional1The default dimensions vector space. These vectors are called word embeddings, and we can perform algebraic manipulations on them, including getting comparisons like Havana:Cuba:: Berlin:Germany. Inspired by this article on implicit sexism in word embeddings2link to arvix publication, I'm trying to both reproduce their work and extend it to biases beyond sexism. Just like the authors, I used a word embedding model trained on Google News.
Dec-25-2017, 18:21:45 GMT
- Country:
- North America > Cuba
- La Habana Province > Havana (0.25)
- Europe > Germany
- Berlin (0.25)
- North America > Cuba
- Technology: