Pretty cool so hello everyone happy New Year and welcome back to uh our NL seminar Series today it is our great pleasure to have Arjun Supra morning as our guest speaker they are computer science PhD computer science students at the University of California Los Angeles the research focuses on inclusive graphs machine learning and natural language processing including fairness bias Ethics and integrating queer perspectives they are there are further a core organizer of queer in uh AI today they will give a presentation on bias and Power in NLP as a reminder there will be a QA
session at the end of the talk if we still have room for that of course but feel free to ask your questions anytime you might have them throughout the talk and as usual for online participants please keep using your mics Unless if you have a question or comments to ensure the clarity of the talk please join me in welcoming Arjun and I hope you enjoy this time thank you hi everyone uh Miriam gave a great introduction but yeah the title of my talk is bias and Power in NLP and I want to say that these
slides do contain examples of some stereotypes and associations that could be offensive um and triggering all right so I guess the way that I'm Going to do this talk is not like here is my project here's the introduction here at this methodology and everything but I'm going to kind of integrate it into the broader landscape of how I perceive bias um an LP and kind of connecting that to harms and maybe a more Justice learned approach going into the future so to begin with everybody uses an NLP angle are familiar with these things things like
writing suggestions on the Grammarly machine translation toxicity detection um we've had this kind of question pop up a lot in recent years by one of these applications have biases uh kind of motivated by a lot of empirical investigations into this let's take that machine translation example from the previous slide where we want to translate my friend is a doctor into the same thing in Spanish so if you look at Google translate and maybe like 2016 or Something you would translate my friend's doctor into mi amigo and this kind of is uh the Assumption here is
that you're translating doctor into masculine you know you might say well maybe that's just the generic um within Spanish and so it's okay to do that however we also would just want to have representation of both as well as Doctor we're going to do these translations and so we see now Google translate it's kind of shifted And now provides both of these translations but in some senses this is one form of a bias that could manifest right like we're kind of implicitly centering the masculine masculine more female or in general maybe you're assuming that all
doctors is our men and I was thinking one example of bias but you might ask like bias seems like a very overloaded term you might have heard of it from statistics uh you might have heard an implicit bias before There's a lot of different ways in which bias is used so you know we might ask what exactly is bias so there are a whole lot of definitions of bias that I can think of at least in a context of NLP there's kind of an umbrella definition that I'm providing here um that's my my definition a
lot of different definitions I kind of consider it unjust unfair or presidential treatment by a model of people who face Discrimination and marginalization within Society and some examples of this include representation biases maybe you have stereotypes of different groups and under representation or over representation at different groups in your training data as well as the predictions that your Market is making there's also quality of service biases um so if suppose that for a lot of marginalized users your model just does Not perform as well for them then that degrades their experience with the model um
and that degrades our quality of service compute um I know Katie was just complaining about access to compute and so there's a lot of like lack of access to compute throughout the world especially in the global South and not quantities and other bias and languages we claim to make a lot of multi-lingual like lingual language models that you Know comprise English French German and Spanish and like pretty much no other language in the world and so we tend to have this we tend to Center Global languages from the global north um and you'll notice that
all these different definitions of bias overlap quite a bit like you know some of them may seem as causes of other biases but they're all in some ways unjust or unfair on uh the kind of prejudicial treatment of some group of people Um I do want to say here that many papers do lack a very clear conceptualization or biases we see papers coming out that are like yeah we're tackling bias in like this model but I have no idea what they mean it's like because they're like stereotypes with respect to gender and occupations like can
you please be more specific about what you mean so I kind of want to make it a goal that people are more explicit about how they conceptualize Bias when talking about it yes so I was wondering about like uh the kind of angle and bias of like I'm uh absence absence explanatory information are making a decision on the classes on some kind of classification that like you know is using um you know a prior that you know where where like um uh things are not not strong enough to understand like you know for example Like
a heteronormative bias of like well you know oh I um hello male presenting person I assume you're looking for a girlfriend or something like that um uh ours would you say that that is covered in some of your definitions here or that it's specifically excluded or that it's something else to consider or how does that a factor that I I'm not sure if that factors into one of those angles or not it probably factors into Like all of them okay and and the dot dot dot as well so it's kind of a very like a
lot of fluid categories you know there's stereotypes about uh heteronormativity that kind of factor in forms we have about it clearly if I you know um receive something like like do you are looking for a girlfriend my quality of service would be degraded and so it kind of fits in there as well yeah and I don't have compute so I wouldn't I would not have made that model to begin with um so yeah I'm going to go through some more examples what I was talking about previously so this is um very similar paper by kaiway
Chung um from a few years ago about learning headings and how they can be dreadfully sexist so the task here is to use word embeddings to complete the analogy he is to Something in the left column uh she is to Something in the right column and The first few seem okay right like he is the uncle uh she is Dan this dance he is the line as she is the Lioness um and then we're like well he is the surgeon as she is to nurse um and he is super faster as she is to associate
professor like all these things um are very like they capture a lot of the stereotypes that are present uh within the data and also just the kind of Co-occurrences of different phrases that you know the models the word embeddings are picking up um from our training Corpus um this is an example of representation bias we're really here capturing a lot of the stereotypes about binary gender that are being uh that are you know in our in our training data um there's also this paper on adversarial triggers and you might have seen this where you could
kind of throw In a prompt I believe this is GPT gpk2 you can I'm not going to read out the prompt but you can throw in a prompt and it'll start outputting a whole lot of text that's very offensive um and you know it's not even just like adversarial triggers like I don't have to throw in like phrases that are maybe kind of not intelligible so I did this with chat gbt like a few days ago to write an intense Political ad about homosexuality and I got a bunch of stuff about we must protect the
traditional family and defend the sanctity of marriage preserve the moral fabric of our society you know these are important things that you want to include in a political ad and yeah and these kinds of things are like you can throw in certain prompts and get models to respond in ways that cause the models to throw on stereotypes about different groups to degrade my Experiences as a user and yeah that's kind of some more biases there so this is from a paper that we wrote in 2021 on um gender exclusivity in language Technologies in particular how
our treatment of gender within NLP is very binary so this you know is an example of a quality of service bias so let's just take the top sentence here that it's a co-reference resolution task In the sentence Alan and Amy went to the market and Amy asked Allen if he wants anything and so the goal here is to match all the expressions in this sentence that refer to the same entity um so we see at the top that I got Alan Alan he you know that refers to the same entity and and Amy and Amy
refer to the same entity here you know we don't actually know what pronouns Alan or Amy use but you know we're just going to go with the fact that Allen Uses he pronouns up at the top and then if we tried the same thing in the second sentence but we change the key to they we're assuming that Allen uses their pronouns um we see that they is just not like connected to any of the entities there um and this is definitely a quality of service bias you can also attribute to a whole lot of other
things like representation advice maybe there's just not a whole lot of people who use their Pronouns within your training data which actually we did do analyzes on and I'll present them in a few seconds about how they singular thing is highly underrepresented in like uh Wikipedia data okay so before I jump into that I kind of want to talk about some sources of these biases that I just mentioned and again connecting it back to some of the work that I've done recently so before we kind of jump into these Things I really want to emphasize
that NLP models are associate Technical Systems we are kind of there's a pipeline by which we build NLP models there is infrastructure generated around NLP models and everything along this pipeline contributes to buys it's not just the training data right so it's kind of like these historical biases a term that I'm not you know particularly happy with because there's you know they're kind of Still there today they're not like past inequities but just in the ways in which data is generated whether it be hiring data at different companies um like this captures a lot of
things like sexism racism homophobia and then the ways in return sampling on this data we don't know which kind of groups of people we're excluding when we kind of curating our data set how are we doing annotation how are we measuring people's identities in particular we Know that with the U.S census uh race is defined as per five different categories which are like like I don't understand how they came with those categories they're very reflective again historical biases but yet we still do measurement with them um and that's kind of problematic and then there's also
learning biases right like again it's not just in training data we're optimizing for uh overall utility in a lot of cases we're optimizing not for Fairness but we're optimizing for just kind of overall accuracy and without thinking about the constitution of our evaluation sets and also deployment biases like who is this model being deployed on like you know we have these things and I'll talk about this more when it comes to justice but we can make our recidivism prediction models there but if they're being disproportionately deployed and on black individuals it doesn't really matter if
they're Fair We're kind of like already choosing a certain group of people on whom we're deploying these models and then subjecting them to uh cyclical incarceration Etc I guess so to dive into some of these parts of the pipeline especially with respect to data curation and learning biases skewed samples end up causing a lot of biases so this is an example from the paper that we did on gender exclusivity We looked at Wikipedia text from March of 2021 alone it's about 4.5 billion tokens and we looked at the composition of these different pronouns and and
in this case today is both singular um and plural but he and is 15 million has 15 million occurrences within the data she has 4.8 billion in current system and data and then we get to like Neo pronouns like Z and Z we get 4.5 000 and actually most of these are corporations like most of these are like The instances are like corporation uh abbreviations like ax media and polish words so it's not even like occurrences of Neo pronouns it's very expensive it's supposed to be for those of us they're supposed to be just like
uh pronouns and individuals users like Neo pronouns like uh I guess within the last two years people have been opting for more gender-neutral terms um besides he she and they that fit their identity yeah so people are using These terms but they're definitely not reflected in like Wikipedia text owl so for example um z z e goes like Z and then the year a time over and then the genitive is here's h-i-r-s and um zxc just works like like say grammatically the same to say then but with an X yeah so yeah that's that's great
explanation so like Z is at the market a sentence like Z is at the market I Told Sam this right it's just like kind of taking these pronouns particular and then adopting it language evolves language is cool um but yeah we kind of see that with this um kind of disparities and the representation of different pronouns this leads to a lot of issues because if we don't see things like singular day and we don't see things like singular um Z we can start assuming the people's Pronouns language models for hallucinating things uh we can start
perpetrating this gender and you can start erasing the existence of non-binary people and trains people entirely because the models are not capable of generating sentences or recognizing entities uh that use these pronouns and then the outputs of models as we see especially with authentic automation bias and everything are viewed as sources of Truth or scientific Knowledge you know Galactica right like with Facebook and everything people kind of tend to think that what a model output is correct especially if they're not doing NLP research or not doing AI research and then this causes human authors and
other language models right because we have now the situation where language models are taking in a text generated by other language models um and stereotypically portraying non-binary individuals who are ignore Them and then guess what kind of goes back into the Erasure Cycles is this amplification of Erasure over and over again so actually I do have an example of language models taking input from other language models a lot of the I I'm not going to say where this was but a lot of like machine translation models nowadays that train on unrepresented languages and corporate get
more language you know from scraping The internet and a lot of these sites that they're getting the information from are actually porn sites that were automatically translated so languages like baths have a lot of toxic translations simply because they all most of the data is straight from vast porn sites yeah yeah that's really that um tainted examples does that mean that like these are the only places that this kind of data can be found because it's Hard to find baskets for porn sites um so that I don't know if I don't know enough about um
I I would assume that like it forms a large portion of the data on the internet but I mean I don't actually know how much mass a text exists on the internet I'm not too well versed on like how Spain and uh like subjected fast people to like you know Decades of not being able to speak Fast and how that like reflects and online forums or other places where you can access Tech I don't know about masks specifically but the the prevailing estimate is that like 50 of the total traffic and content on the web
is born okay so just in general everything I guess some questions about the language distribution um foreign it's a really small section of Spain yes I mean I'm not going to get into it even further but um so when I was in like the Basque Country right like there's a lot of issues where people uh don't speed fast simply because like their parents kind of lost that during the Franco dictatorship and then when one person in the group doesn't speak fast and everyone switches to Spanish and then that becomes a thing where like like Spanish
is reinforced so it's very it's language language based cultural Assimilation it's like so easy to perpetuate in that way the cycle right like it works the same way yeah well yeah it said human authors and language yeah yeah there's actually a really good I should I should have put that here but if you're interested my friend was on The Basque radio and talked about large language models in in bask and like the intersections of those two I can send you a link um if you're curious about That okay tainted examples is another another thing here
so this is kind of based on when the supposedly ground truth data that we have with that we're training our models on reflect a lot of stereotypes and historical inequities so this was a great paper by Martin um on looking at the racial biases within toxicity detection so because um a lot of the annotators that we have and deciding whether things are toxic or Not or even just content moderators you know for real on Twitter um decide what whether something is toxic or not they are overwhelmingly white men um and they're not very well versed
African-American English um and so they kind of overwhelmingly decide that African-American vernacular English is inherently toxic I mean this is definitely reflected in things like perspective API which is a very common toxicity classifier we see that just Very common uh phrases used by black individuals on Twitter are labeled as like highly toxic uh that's definitely not the same for a lot of texts that comes from White individuals another example of examples is Amazon's resume screening system so you just have a system where they would take in resumes and they would look at whether the person
could potentially be a good candidate and they found out for some reason for some reason women were Just not costing the resume screening at all uh man it's because well the people who are using training data from the past 10 years of hiring in Amazon and Amazon's not particularly the most like gender inclusive place in the world um and there's a lot of sexism reflected in whether someone was decided to be a good candidate and the associations of that with the content of the resume in Fact they found their resumes that were being rejected the
most are the ones that contained any terms related to women so Society women Engineers scripts like these words were like just instantly causing resumes to get rejected from the screening system um some more uh examples here simplified disparities so I talked about um this earlier from the paper that we had a couple of years ago It feels weird to say it like a couple of years ago like 2021 but the ways in which these kind the ways in which these disparities in the training data set Translate into the learning algorithm is is very like pronounced
so if we were to train glove embeddings on the Wikipedia data that I mentioned about I mentioned earlier if we take the word pronoun and then look at the closest neighbors that were here in like the embedding space That we have plus we see like you know his man things are expected um similarly with she right like these things seem okay because again we're including it's really hard to inbiguate between singular and plural but as soon as we get to like the New York pronouns like what are these like these are polish words and corporations
right like we can see that it picked up but those were the Primary instances um of these of these two terms um and similarly for evaluation so this is from the data set mimic which is a clinical notes data set and it kind of like uh has some notes about diagnoses based on like patients uh based on nurses and doctors notes about the patient and we see that like for example Asian individuals comprise 1.9 the test data set so there's some evaluation biases There right like you can get 98.1 accuracy on that training set on
the test set rather I'd be like I did great and it just like does not work for anybody who's who they do um and I guess the other thing I want to you know make clear here is that a lot of people when they see this they're like just collect more training data right like just just do it like get more training data first of course it's very hard but on top of that we need to be Careful about predatory inclusion here uh we can kind of go out to different places um use people's labor
very exploitatively collect more data about them but if we don't think about the deployment biases where these models are being deployed and it doesn't positively impact these individuals on who you're collecting more data you know whose benefit is that like it's during predatory to do something like that And the last thing I want to talk about here is proxies so this is also from our paper um we have sentence John 8 mask sandwich they were hungry so this is where the kind of the model start hallucinating pronouns right like we get the very explicit indication
here that John uses they pronouns however if we look at the kind of scores predicted by the mass language model demo on lnlp and we see that we got Johnny his sandwich that's the top prediction 60 in fact I don't think any of these are there so it's like John ate a sandwich Johnny another sandwich right like Johnny another Santa which is over John ate their sandwich like amongst all the things that could possibly fill that Moss token um and the reason for this is because the model is clearly using John um as a proxy
for masculinity and kind of filling in terms that are masculine In order to kind of conform to the the co-occurrence distributions that we see within our training data it's interesting about the giant acre sandwich case presumably John is not going to have her as a pronoun and that's reinforced that they were hungry thing right yeah but the the sandwich could belong to somebody else right and so and that's true for the first sentence too and so it's kind of hesitance Banks it's it's using the mass So it's got to like okay there's like I I
have a fire that John and dies are linked right and also like it could be somebody else a sandwich right and so I'm gonna um ownership of the sandwich right like I think that generally like people don't collectively own sandwiches so I kind of would expect that no there's not a sandwich cabal or whatever but I feel like you know just that alone like the fact that there's no they are here Like you know um uh you know the deli down the corner I really like their sandwich right actually that would be a great case
a singular thing right because like that would I mean foreign but I I agree that you would probably want to see value above me at least somewhere at least somewhere or even just like if you don't know their pronouns at all maybe you Johnny John's Sandwich or like Johnny their sandwich because I feel like we still don't have enough information about like what pronouns they may use from like the sentence alone yeah I mean I think if if we're trying to shy away from like a student pronouns like Ada sandwich or the sandwich or fairly
acceptable um yeah and for sure the period's not causing there to not be any attention paid to uh Boundaries then that would make sense but right so I guess this is something that we had done um in the paper I don't have all the results here but like we generated some templates or effectively we like kind of swapped out the pronouns in a second sentence like he she they and they've used like the top most gender neutral names like in the US uh and we found that the second sentence pronoun was like consistently used Um
predicted for the mask as long as it was like he or she all right so we talked about bias but not kind of wanted to shift the conversation a little bit from bias to harness um so bias is great I think like it's like an interesting tool for looking at kind of the outputs of models or training data and thinking about it maybe even statistically um but I think all of us here really Care about making the world a better place right and like non-harming individuals so we want to see how these biases translate into
Harms in the real world when these models are actually deployed so there's representational harms that we might expect these include the stereotypes that we mentioned earlier negative generalizations misrepresentation of different groups um you know these are awful because you kind of see that If you have a model like chat GPT that's just like during the stereotypes I'm not having a good time but then there's also like allocative hearts that result from so that's why we really see an unfair distribution of resources because of the kinds of biases that these models contain a lot of unfair
consequence maybe someone doesn't get a loan uh maybe someone doesn't get hired Amazon because like their resume contains the word women in it right like these are All kind of very serious harms that arise for models having biases and particularly I want to like drill in ask me which bottles are becoming increasingly prevalent and you know these critical decision-making systems which impact people uh we want to make sure that we don't make worse the harms that are already posed to already marginalized communities within our society we don't want to basically be reinforcing the inequality and
power Relation that exist with them you know our contacts at least so yeah I think like this is kind of motivation for some of the paper that I talked about on gender exclusivity um like we surveyed a bunch of non-binary individuals in particular ask them about Farms that they've experienced or that they would expect to see or you know they just played around ballot and it'll be demo Um with different prompts and like looked at what they got what kind of representational allocational harms uh with the experience for these different tasks game entity recognition co-reference
resolution and machine translation um and so right like if we see a lot of things in representational harms for ndr where we're systematically mistagging your pronouns and singular they as non-person entities we saw it right like We need to see that it's just not working for these individuals it's not representing them in the outputs we can also see that this translates into real world issues because if we have the same system that's based on this ner or co-raft building block um to look at the instances of discrimination in historical documents against non-binary and trans individuals
towards counting instances of these discrimination to like you know enact More anti-discrimination policies we see it's going to under count instances of discrimination against trans and non-binary folks simply because it doesn't recognize them as entities and that's a serious problem because it's going to delay the kind of legislation that we need for the sake of keep you know keeping Transit non-binary folks safer um so how do we go about measuring biases um evaluating biases such that in A way that's kind of like thinking about their arms um so the way that people tend to do
this um it's often based on contrasting sentence fairs or some kind of toxicity detection in the case of language generation so on the left hand side we have this nli-based bias evaluation where we had a premise and a hypothesis so the premise is the doctor is striving and the Hypothesis is he is driving and we see that the uh the nli model here is like yeah there's a high impalement there that the that he's driving because we said the doctors driving but if you look at the contrasting case here where the premise of the doctor
is driving and the hypothesis now she is driving we see that it thinks that there's a much lower entailment there of the hypothesis being true so now you might ask how reflective is this of like The real world use cases of such a model so we don't really know actually right because I don't know how people type in the doctor is driving he is driving on a daily basis I feel like that's not the primary use case of both the allylons that are kind of coming out but this is one way that people do it
it's not super grounded in any real world use case and that actually might be a pitfall that I'll talk about in a second with regards to looking at free form Generation Um take that same example from earlier where we're looking at defending our moral Fabric and we put this through a machine learning based toxicity classifier to look at biases that it may have but if you're a call from the tainted example slide a few slides back these toxicity classifiers are not not foolproof right they actually do have they actually don't under count instances of toxicity
they misrepresent Different uh different kinds of language example African-American vernacular English as toxic language even when it's not so this is also not foolproof so some pitfalls here did you try the doctor as a name essentially the sandwich sentence uh oh lucky doctor a mask sandwich he he was hungry versus uh she was hungry and just at least see if it's able to make that consistency I don't remember if we tried that But that would be interesting that's a really good point but like the doctor ate their sandwich so there's a challenge too but I'm
just telling doctor she just tried to Dr Aiden asked her if she was hungry and she was hungry right uh if if it's going to resolve his sandwich and her sandwich right I think so that one I definitely have tried we don't have data like you know kind of aggregated for that but most definitely It sounds like Dr ate his sandwich she was hungry oh but the second thing was she was hungry in that case it definitely actually would say she is so it's at least able to yeah Johnny for example yeah the doctor if
he does kick he does get past that it's kind of like when that signal was explicitly there that the person is you know she then it would it would put in her [Music] but yeah So there are a whole lot of pitfalls that I kind of want so most of this talk is kind of like focusing on a pitfall so like biased evaluation the way you currently do it so this is a table from a paper by Sula and Blodgett on an inventory that falls in fairness Benchmark data set that's the best team ever it's
called stereotyping or region salmon um and you'll see why if you if you read the paper but one there's some examples Here of different ways in which different issues with the ways in which we're kind of measuring these biases we might have groups logical failures my favorite is like the unmarkedness issue which is the second to last row in the table which is where you're looking at which one is a language model prefer probabilistically like what is the probability of generating that sequence Of words it's like straight man Drew is done and fired or the
game Andrew was kind of fired I'm like who said the street I feel like that's just assumed right like like you're you're straight until someone says that you're gay and so like it's that kind of like contrast there um we don't really know what that entails for the bias evaluation because clearly the straight manager has gotten fired is not it's not something that Most people would say some other pitfalls um and this comes from a paper that we wrote in early 2022 they're very anglo-centric so a lot of these gender revised evaluations focus on Western
professions things like doctor receptionist nurse um and all things like priest um or other professions that you might find in other contexts also a lot of biases are very grounded in the global North I think one that I've been very interested in uh in the past year is like cast bias in Hindi and there's an excellent paper on this called social aware biased measurements for Hindi language representations there's a lot of ways in which people's names and the kinds of jobs that they hold are related to people's cast and that kind of manifests when we're
looking at biases in in the language presentations so we need to more critically examine Social context right like we're coming at this overwhelmingly from a U.S and a global North perspective and we do need to integrate integrating some more voices um from the global South and specifically like thinking about this with the global contest in mind to realize that like reflexively we are not the world right like a lot of the biases that we uncover are not going to be here Even relevant to a lot of other portions um of the world we must also
I think co-design these culturally aware more culturally aware biased measurements you know but those whose languages and cultures are actively being excluded other things we're kind of focused on Prestige forms of English who are kind of uh reincarnating a lot of the toxicity toxicity classifier problems where the people who are coming with These templates uh they're they speak a very specific you know academic standard North American English dialect and that definitely also influences the way in which we evaluation it's also a lot of false claims of external validity if you've worked these templates and biased
evaluation we'll see that there's a lot of poor syntactic diversity in the evaluations that are being run they're very contextual right like with the contrasting sentence pair I don't know What Downstream application that's going to be used in but I'm running the bias evaluation anyways Upstream but again that I don't know what the implications about student evaluation are for for any kind of like Downstream task that I'm going to be using that modeling also we also really need to go beyond binary gender this is one of those like thinking traps that a lot of computer
scientists uh you know myself included falling all the papers seemed like Overwhelmingly focus on binary gender uh as kind of like this um example of like two groups you know with respect to which you can balance some kind of prediction but like how do we know that it you know expands to race or like uh queerness or anything like that it really doesn't actually that's why I really like Katie's work kind of looking at like the these data centers specifically with Respect to anti-glare biases because the ways in which um toxic and stereotype toxic language
and stereotypes manifest is so different for different identities and different groups um it's not enough to be like yeah this works for men and women so right like you know we're good we can do this for any any two groups in society I think we also need to go beyond binary gender when we do these evaluations so I haven't made that case Clear enough from like the past two slides but also not incrementally I think there's the idea that gender is a starting point race is like a second step and then like we'll kind of
slowly get there and cover all identities but that's like not the case there's no reason that binary genders on a path towards looking at harms and biased against non-binary trans people or any group for that matter like everyone's this is kind of like relationality the Concept relationality from intersectionality like there are so many different ways in which groups are different of course we all face some sort of sheer depression but we do need to be looking we can't be identity agnostic we can't erase people's identity towards making a general framework for bias evaluation and with
all these false claims of external validity we may actually be causing further epistemic violence onto Marginalized people by creating a veneer of fairness we run this evaluation we say I got zero percent bias in my model and therefore it's ready to be deployed no please don't do that I feel like from for most for the most part it depends on how you run your bias evaluation I think bias evaluation should be a good indication of how your model can go wrong and that you need to continually monitor the model afterwards because you can never Say
a model is biased rate especially given how high stakes are like that statement is without continual evaluation we also use a lot of measurements of identity that are very unreliable or problematic right for example a lot of these template-based data sets for binary gender bias or based on pronouns or like gender names and they use that to infer people's gender or as a proxy for the gender Yeah that's not exactly great it's like reinforcing biases in other ways towards like a lot of Transit non-binary people and things like sexuality and disability are often unobservable from
text alone um like you know it's just because of the ways in which certain things are considered taboo or confidential or the Norms around um ability um and heteronormativity kind of lead us to act within this context uh we don't Often communicate that in you know the same explicit ways we might with other kinds of identities in language and ultimately we have the tendency to assume that all these identities are somewhat known and like measurable and discreet and immutable and not intersecting it really does not capture the full fluidity um of which people's identities kind
of occur and finally there's a lot of like There's a lot of faces in parity I think like we have one conceptualization of biased measures that we kind of rely upon overwhelmingly within our community it's like if we have the same amount of bias for both groups we're go but this kind of push towards neutrality definitely reinforces a lot of existing power relations within Society right like this great example from the intersectionality book by Patricia Collins Where it says like you can't they talk about the FIFA World Cup and how you have like a team
from Morocco on one side they don't talk about that because that just happened on the other side and in terms of resources and ability uh in terms of resources and ability to play like kinds of racial discrimination better faced the playing field is like not level right like we have friends all the way up here we have Morocco on the other Side and the referee kind of is this performative like individuals like yeah we're gonna make this game fair make sure everyone uh gets the same advantages like keeps you know enforce their rules correctly so
you're basically just letting Morocco continue to slide at the rate that they're already sliding like down this playing field and that's kind of what that pushed toward neutrality does and we don't consider other forms of Justice here right distributive or representational Justice there's so many other ways that we could think of justice and ready the baby has done great work on this topic so I really encourage improving biased evaluation through a lot more reflexive documentation um I think transparency is a very good first step towards kind of talking these problems like how do you define
bias and how does this definition of bias align With definitions of harm are you using proxies to infer people's identities are there any other data sets already exist that could do this measurement how does your benchmark kind of address issues at that other Benchmark is not addressing or what are some problems with those data sets um I think we need a lot more documentation for the day that we do not just in bias evaluation of course like just across the board but I think you Know it's especially critical in the space when we're talking about
harms and biases um so this is some ongoing work with uh some people at Microsoft research okay um and this is just general evaluation in NLP but I kind of would like people to be more specific about what kind of capabilities do they want their models to have and what is that Benchmark assessing by testing a model um with you know a data set at metric Like what are you hoping to see that the model uh possesses as a result of it doing well on your benchmark and how make correct outputs like for example toxicity
labels be contested um you know in your benchmark how many acceptable methods of solving the task or unacceptable methods be allowable or not um and I'm not going to read through all this stuff here but I think like one question that I really want people to Take away and like you need to start writing the limitations section of their paper maybe not in the limitation section I don't like the idea of a limitation section but like you know limitations throughout the paper um what does it mean for your for your model if it achieves a
perfect score on your on your task or your benchmark or your like bias evaluation Data set like what does that actually mean for society like what could you take away from that what is the implication of achieving such a number and in the process of proposing that Benchmark are you limiting progress to only working on one conceptualization of bias are you saying that charity-based bias is the only way forward or are you holding space for others to propose Alternatives this has very real consequences right like the ways and if I do anything that's not like
parity-based bias um you know with a few exceptions I don't get apply one get funding right like there's this very kind of attractive nature to being able to measure bias in this like these two quantities are equal and now I'm able to solve that kind of way like that's what band of stuff kind of really likes um there are also some ways for mitigating bias in contextual Representation so the reason I'm talking about this is this is the work that I did when I was a research engineer at Allen NLP uh build player fairness library
for the LLP um variance toolkit for the island NLP Library I'm very sad apparently they're shutting it down so use it while you can I I guess I guess there's not maintaining it anymore as well yeah there's some techniques out is When you use contextualization right like you're not just doing projections on embeddings you're not just removing things from static word embeddings like the context in which a word first in a sentence affects the kinds of biased information that the representation for that word captures uh like you know when you have a whole sentence things
like she you know gets a lot of information probably from words like that So there's some techniques here I think for the sake of time I'm gonna skip this for now uh but you have my slides so we kind of I kind of implemented some of these techniques um at lmlp and saw some improvements not a whole lot and again I also don't know what the implications of these tests actually mean for any JavaScript applications the reason I put this up here is I feel like I've kind of changed my perspective A lot in like
the last two years since like March 2021 um when I did this and I kind of like to see like how my view on my fairness and bias has changed throughout the last how much ever amount of time so there are again of course you know meat headphones I like talking about pitfalls pitfalls and bias mitigation uh biased mitigation has to go hand in hand with real world auditing um the problem is it's however like very Often post-hoc like it's kind of weird that we train a model that's like supposed to be good and then
we de-bias it afterwards right like like what is why is that not like the first intent you know from the beginning why is it always that the model is something and then we strip away the bad parts of it because that inevitably reinforces a lot of the issues that exist from the model again all things external ability like before and I think I think that's most Important to me is to happen and like the us we don't tell the model anything about slavery we don't have a model anything about um Stonewall or whatever right we
we kind of just we are like here's two groups any two groups like just make sure that the buy is the same for both of them and By ignoring the historical and social context we can't really accommodate any sort of prepared interventions remedy And Equity you don't have the context to actually identify certain stereotypes and know about like where what is the origin about how harmful is that stereotype right we just treat everything as the same and then we also going back to the whole Morocco friends thing like we should not just be like letting
people fly the same rate down a field like we should be turning the fuel pump so that it's like slightly More Level and that's what it would be by Reparative interventions and so this is again the paper from early that was on from early 2022 um and kind of the concluding note of this shifting from bias and horns to social justice so discussions on large language models cannot be divorced from The Wider power structures within which they exist there's a whole lot of things here kind of going back to the compute bias like language bias
large language models are Expensive who is making large language models like primarily Google and Facebook they're a very centralized way you know you can access their models sometimes on like uh from like you know their servers and I don't think that one model that's super scalable and supposedly inclusive and sits on the server in Menlo Park is going to be applicable to anybody who's living in like Nairobi that's there's definitely a mismatch there scale is Definitely an emesis to inclusivity a lot of the reasons these biases pop up is because we're attempting to make these
models at scale and claiming that they're inclusive where in reality we have to privilege some group over another group inherently if they're making these models really big the thing may work for everybody um language is Multicultural right but marginal models are not we can feed them language but there's no kind of input About race or ethnicity not everybody who speaks French speaks French the same way uh people speak French and former colonies are France right that's super different than like anything that um anyone like from split speak but we know that the culture there actually
matters uh when they're speaking a language and also the fact that only a few people can make large language models allows a lot of powerful actors to control Research right I don't know about you but I'm not particularly I don't know who sponsors isi I don't I'm not particularly comfortable with redacted you know in charge of like all my large language models um especially given other awful things that they've done in the past this is a very interesting table that I liked that we kind of threw together in the in the paper looking at the
author location and languages of different Large language models that have been released so we see primarily the locations like U.S and China multinational mostly just means U.S like maybe some like one person from Australia or something and the language is sort of like English Chinese that's pretty much it actually and then multilingual as a Cadillac Spanish German French like the usual like romantic languages uh romance languages rather from like the from the global North so that's kind of the impetus here to start thinking about a more Justice granted approach to NLP and bias and harms
in general I think our goal here for anyone who's working on bias or just I mean everyone should be thinking about the ethics of their technology but we care about minimizing social inequality I hope on stream cares about reducing social inequality and I think we need to be Careful think of like be more explicit Even in our papers of how are we grounding the work that we're doing in social inequality I really like the gender change paper in some ways just because they talk about how like the fact that the reason they're focusing on gender
recognition is uh commercial Technologies because those things actually get used by law enforcement right they're integrating that social Context um into their like their motivation for for making that for doing that project and their goal is to reduce social inequality because law enforcement primarily targets um black individuals especially black men um there's also a lot of power relations that we think about how is your NLP system like centralizing power how is it decentralizing power how is it participating in existing power Relations um and contact is key right going back to the stuff about how one
of our stuff is very anglo-centric what about the rest of the world like we are not the world like you know us here in this sixth floor conference room overlooking the marinas not the world and co-creation how can we better create sustained um because sustain like sustained Partnerships with different communities Uh towards making the technology that works for all them better yeah how come we used to equip different communities to make their own models so that it better represents their context um and not kind of rely on you know our understanding of like the the
US and Los Angeles and the other things um there's another thing that's like not here that I meant to add but it's infrastructures I think whenever I talk about this a lot of folks assume that It's like on them as individuals to fix this but that's so hard and like the ways in which Academia operate and which like conferences operate make it nearly impossible to address any of these things if you're submitting to ecl and you have eight pages who's going to talk about power relations right like you're like looking to squeeze that the tech
table into like the seventh page um and so we do need to have these discussions with a broader ACL community And other communities that everyone's part of um towards kind of providing more space for people to be more reflexive and critical of their own work and thinking about its role in society who's using it who's it being used on I really like the broader impact and limitation sections that we're doing now and I think they're actually very good steps response NLP checklist that Oliver can be filling out That's a very good step towards kind of
being more flexive in this way some more readings in the papers that I touched upon that I um was co-author on um yeah and that's it thank you so much for your time I'm happy to take any questions in the last six minutes that would be four minutes thank you so much Andrew for the great and excitable talk and now I will open the floor through more questions from The audience yes um so I know this is just like an example you're reading but I think it's uh it could be with like there's a useful
Network I would have to get it uh sort of higher level abstract um questions so the the Lego World Cup example for example um uh one is like would you so like would you actually make the argument that like the more just action is for the referee To be harsher to the French team and then also um I was thinking maybe part of the reason why we uh he talks about how like people say like oh the you know the bias your like performance is the same for these two groups so therefore our model is
like good I think like besides there being like issues like certain social issues of why we do that I think there's also maybe like an evaluation issue or like part of why we do that could be This because it's very easy it's an easy thing to evaluate like does this work equally for this whereas if you're trying to include like the the offset first and then evaluate for that they've got to be typical so it also asks like to be thought of how you evaluate yeah so with the with the football question I'm not gonna
answer that because it's being recorded um but I think I'm more confident making a stronger assertion isn't like NLP Um but I I do think like this paper called algorithmic reparation and I do align a lot with like their views which basically state that because uh redlining and like the discriminatory like um practice of not giving loans to Black individuals we do need to be over kind of like doing the over granted loans to Black communities um yeah so I think I think algorithms are so widespread and like very commonly Used in like housing markets
and like financial markets and so it's good that they also reply reflects the repair it in some are not blind to identities but rather try to remediate like past um inequity and the second question which is on I do think that I think that's actually part of the infrastructure is right like I do think part of it is it's very easy to just compare two numbers And also the reason that I can give you a better answer is because like the way in which we do research in this community is like we have to cite
somebody else who did it first and like we need to kind of build off of other people's work um and this is a kind of research that got cited a lot and the kind of research that gets funded a lot and so that's what I tend to think about overwhelmingly because that's what you Know my infrastructure allows me to think about so I don't have I don't have a better answer for you but I think overall we just need to collectively spend more time thinking about other ways and it's hard like it's really hard it's
not something that comes about um isolated sure thanks I think Katie has a question I think the other other person yeah so let me start with giving yellow cards To the Argentinian players so are we going back that is my bigger question so all these things that the model has learned is from a data set which is exactly a reflection of a society right so RV or the researchers the good chance of coming in and saying let's reduce the bias but then this will it reflect in a society that's my bigger problem like I would
love it on society but are we the right judge for it like roughly who did for who he Decided like okay never mind this is what happened eventually that they interviewed him and then Messi came up saying yeah they felt like he was yeah yeah so to your first question I would kind of contest the idea that the data we traded on this exactly reflected Society especially because the mask of porn sites that I was talking about earlier um I don't think most of the data we Trade contains like I also like to talk about
Wikipedia the data down probably like there's no existence of like new opponency it's just not the distribution is not the same as reality most people don't interact with automatically translated about supporting sites right like there's a whole lot of stuff going on there but with respect to um your second question about who is the judge if not I am obvious that we are the judge of that as researchers because who else will view that I think like we can't just keep on pushing accountability down the road there's just like you know saying kicking the can
down the road uh when I think I was very shocked when I went to the not by Europeans but by the discussions differences between Europeans and Americans about regulation and who is responsible for making these decisions a lot of Europeans have a lot Of faith especially in France have a lot of faith and like the central government and their ability to regulate things um but I don't think that actually ends up like getting done like really like I don't think the government has the knowledge all the time to even like know about the kind of
stuff that I'm talking about here um and then I'm I grew up in I get my contacts where I just like I grew up in the in the Bay Area and so I don't Really I don't I think I trust the government very much just because I think the American enemy doesn't want to trust like the federal government to regulate like a lot of Technology properly um and also just because of Oppression and things like that and like past legislation so I just I I think it starts with us right like I don't know which
body would otherwise um and it can seem like a survival thing For a lot of marginalized communities right like it look if I don't do it like that just means that more trans people are going to get harmed for example uh until someone decides that it's appropriate you know like some some white dude from Texas decides time that Transit book can like stop dying um so my question was did we get up on the wrong foot just by going with the co-operance concept in the first place It definitely helped the machine learning models learn a
lot of things but it's also learning all these things [Music] so that's just an open person I have to ask but we started with you know cochrance with you know 2014 training terms for paper glow and then it went on to work to back memories but yeah it's helpful but is that the real way I don't know I think they're a really great tool Um yeah I just think there's that going back to the pipeline image from earlier there's a whole lot of things that factor into the harms that we have from NLP systems that
are not just like co-occurrences although they do they do play their role too let's conceptualize so you know linguistically historically Association fixes were running out of time you can Take one last question from both yes so very interesting talk thank you uh so you have one uh great example there the gender bias horrible gen advice in the resume scanning so now you're in charge to fix this you have three to six months uh what would you do uh you know without creating new biases and the perspective of the question is you know are there shortest
firm solutions that are feasible for us to work on within three or six Months or is the goal really to change society over the next 10 to 20 50 years and you know just great awareness uh so there's you know two different timelines yeah so what's your set specifically for this kind of this might this might not be an example this might not be the answer you'd like but I would throw this system away and not improve it no because I think like first off as soon as you recognize that It's doing something like that
just just trash it like that's not something you want to keep on letting happen for three to six months but also like job security is a very real thing like Financial Financial disenfranchisement like um an income they're all like super important I know all the grad students there and we're like yeah like money money's good um but I think like if something like That is very high stakes and really should be considered more carefully and thought thought through that yeah I think like you should not really have a machine doing single-handedly I do think that
there is some potentially some ways in which we can kind of have human AI collaboration uh when we're doing these resume screenings because humans are also not exactly the most like Fair individuals either like we're not a lot Of the recruiters are not always thinking about uh reparation they're always thinking about justice but I also wonder if like these systems have a way of like highlighting uh different parts of people's resume that they uh like recruiters might have missed right they can't just do it single-handedly and also like definitely not that system not the one
that Amazon used because it was just highlighting things like Scripps College like I don't know if like that's Something that will basically work in a collaboration like you have to make um the system not make the final decisions you maybe wouldn't want to make a little bit more Bare Bones not like I'll put a single value like good Canada bad candidate but rather like highlight things in the text that explain like what what would contribute to someone's candidacy and then have someone review that as part of the process Cool so to be respectful of everyone's
time um if you have any more questions please feel free to reach out to Arjun offline and I would like to thank you all for joining us today and see you next time