How to measure using word vectors

I'm attempting to understand how bias can be measured using word embeddings. Reading the article https://towardsdatascience.com/gender-bias-word-embeddings-76d9806a0e17

What is the bias being identified in the above statement ? Is the bias here that a woman cannot be seen as a doctor when a man is involved ?

Is a neutral bias for a either a man or woman being identified is where there is a small difference between woman,doctor man,doctor , represented a vector : $woman + doctor \approx man + doctor$ ?

Solution

You would expect that

woman + doctor = man + doctor

Or rewritten:

woman + doctor - man = doctor

But since it is 'nurse' in that word embedding space, that is an indicator for bias towards women in healthcare to be percieved as nurses. Doctors are associated more with men in the corpus from which the embeddings were trained, so it can be concluded that the corpus (and the learned word embedding) has a gender bias.

Training a Keras model to identify leap years
Ideas for Extracting Blade Tip Coordinates from masked Wind Turbine Image
Macro VS Micro VS Weighted VS Samples F1 Score
Doing PyWavelets calculation on GPU
Training loss increases instead of decrease with epochs
How to save a Dataset in multiple shards using `tf.data.Dataset.save`
why explain logit as 'unscaled log probabililty' in sotfmax_cross_entropy_with_logits?
What is the loss function used in Trainer from the Transformers library of Hugging Face?
Using features extracted using a pretrained CNN as new features for an CNN/NN
InvalidArgumentError: No DNN in stream executor while training a TensorFlow RetinaNet model on Google Colab
ALS (Alternating Least Square) algorithm in multiple rankings for a user
How does one set the pad token correctly (not to eos) during fine-tuning to avoid model not predicting EOS?
How to create image of confusion matrix in Python
Cross-validation with nb method
GPU utilization almost always 0 during training Hugging Face Transformer
The “Forward/Backward Passage Size” is too large for the pytorch model (Yolov3)
How many images(minimum) should be there in each classes for training YOLO?
Why do neural networks work so well?
Will larger batch size make computation time less in machine learning?
Why KL divergence is negative in Pytorch?
Creating a voice identification system using machine learning
Forward pass with all samples
override pytorch Dataset efficiently
fit method in sklearn
Calibrating Probabilities in lightgbm or XGBoost
Implementation of F1-score, IOU and Dice Score
Is it ok to have the training history very similar to the validation history?
How to understand Shapley value for binary classification problem?
Stochastic Gradient Descent for Logistic Regression always returns a cost of Inf and weight vector never gets any closer
Data format for Libsvm SVR training in Matlab