Bias and Fairness in Data Science and Machine Learning

Bias and Fairness in Data Science and Machine Learning

Where does bias come from?

Bias in data science and machine learning may come from the source data, algorithmic or system bias, and cognitive bias. Imagine that you are analyzing criminal records for two districts. The records include 10,000 residents from district A and 1,000 residents from district B. 100 district A residents and 50 district B residents have committed crimes in the past year. Will you conclude that people from district A are more likely to be criminals than people from district B? If simply comparing the number of criminals in the past year, you are very likely to reach this conclusion. But if you look at the criminal rate, you will find that district A’s criminal rate is 1% which is less than district B. Based on this analysis, the previous conclusion is biased for district A residents. This type of bias is generated due to the analyzing method, thus we call it algorithmic bias or system bias. Does the criminal based analysis guarantee an unbiased conclusion? The answer is no. It could be possible that both districts have a population of 10,000. This indicates that the criminal records have the complete statistics of district A, yet only partial statistics of district B. Depending on how the reports data is collected, 5% may or may not be the true criminal rate for district B. As a consequence, we may still arrive at a biased conclusion. This type of bias is inherent in the data we are examining, thus we call it data bias. The third type of bias is cognitive bias, which arises from our perception of the presented data. An example is that you are given the conclusions from two criminal analysis agencies. You tend to believe one over another because the former has a higher reputation, even though the former may have the biased conclusion. Read a real world case of machine learning algorithms being racially biased on recidivism here: https://www.nytimes.com/2017/10/26/opinion/algorithm-compas-sentencing-bias.html.

Bias is everywhere

With the explosion of data and technologies, we are immersed in all kinds of data applications. Think of the news you read everyday on the internet, the music you listen to through service providers, the ads displayed while you are browsing webpages, the products recommended to you when shopping online, the information you found through search engines, etc., bias can be present everywhere without people’s awareness. Like “you are what you eat”, the data you consume is so powerful that it can in fact shape your views, preferences, judgements, and even decisions in many aspects of your life. Say you want to know whether some food is good or bad for health. A search engine returns 10 pages of results. The first result and most of the results on the first page are stating that the food is healthy. To what extend do you believe the search results? After glancing at the results on the first page, will you conclude that the food is beneficial or at least the benefits outweigh the harm? How likely will you continue to check results on the second page? Are you aware that the second page may contain results of the harm of the food so that results on the first page results are biased? As a data scientist, it is important to be careful to avoid biased outcomes. But as a human being who lives in the world of data, it is more important to be aware of the bias that may exist in your daily data consumption.

Bias v.s. Fairness

It is possible that bias leads to unfairness, but can it be biased but also fair? The answer is yes. Think bias as the skewed view of the protected groups, fairness is the subjective measurement of the data or the way data is handled. In other words, bias and fairness are not necessarily contradictory to each other. Consider the employee diversity in a US company. All but one employees are US citizens. Is the employment structure biased toward US citizens? Yes, if this is a result of the US citizens being favored during the hiring process. Is it a fair structure? Yes and No. According to the Rooney Rule, this is fair since the company hired at least one minority. While according to statistical parity, this is unfair since the number of US citizens and noncitizens are not equal. In general, bias is easy and direct to measure, yet fairness is subtler due to the various subjective concerns. There are just so many different fairness definitions to choose from, let alone some of which are contradictory to each other. Check out this tutorial https://www.youtube.com/watch?v=jIXIuYdnyyk for some examples and helpful insights of fairness definitions from the perspective of a computer scientist.­­­

InfoSeekers attend ECIR 2019!

InfoSeekers attend ECIR 2019!

InfoSeeker Souvick Ghosh attended the 41st Annual European Conference on Information Retrieval in Cologne, German. Souvick presented “Exploring Result Presentation in Conversational IR using a Wizard-of-Oz Study” at ECIR as part of the Doctoral Consortium.

Souvick Ghosh (center) Doctoral Consortium group photo at the 41st Annual European Conference on Information Retrieval (ECIR 2019.)

Souvick Ghosh presented his work that reflects on recent researches in conversational IR that have explored problems related to context enhancement, question-answering, and query reformulations. His work focused on result presentation over audio channels. The linear and transient nature of speech makes it cognitively challenging for the user to process a large amount of information. Presenting the search results (from SERP) is equally challenging, as it is not feasible to read out the list of results. He proposes a study to evaluate the users’ preference of modalities when using conversational search systems. The study aims to understand how results should be presented in a conversational search system. Through observation of how users search using audio queries, interact with the intermediary, and process the results presented, insight can be developed on how to present results more efficiently in a conversational search setting. Additionally, there are plans to explore the effectiveness and consistency of different media in a conversational search setting. Observations in this work will inform future designs and help to create a better understanding of such systems. 

Souvick had a few words to reflect on his experience at ECIR 2019: “I was lucky to have Dr. Udo Kruschwitz as my mentor and we had some great discussions about my dissertation ideas, research in general, and the life of a Ph.D. student. It also gave me the opportunity to catch up with some old friends in Europe and make some new ones.”

InfoSeekers attend the New Jersey Big Data Alliance Symposium!

InfoSeekers attend the New Jersey Big Data Alliance Symposium!

InfoSeeker Matthew Mitsui attended the 6th Annual New Jersey Big Data Alliance Symposium. The title of the symposium this year was The Future of Big Data: Artificial Intelligence and Machine Learning, and it was hosted at New Jersey City University.

Matthew Mitsui presented “Multi-Faceted Information Seeking Leveraging Big Data” at the symposium. It was co-authored by some of our other InfoSeekers: Souvick Ghosh, Ruoyuan Gao, and Chirag Shah.

Matthew Mitsui presenting at the 6th Annual New Jersey Big Data Alliance Symposium.

Their presentation addressed the complexities of the search process and the multitude of obstacles and issues an information seeker can encounter; viz., information task and resource limitations; information quality; information bias. They identified that those obstacles are often presented to the user through the tools employed during the search process, and their aim was to explore how search tools can be improved in order to foster collaboration with the user and surmount these obstacles. They addressed the need for search tools that can assist the user through three primary approaches: search task assistance, assessing information quality, and counterbalancing bias.

InfoSeekers attend CHIIR 2019 Conference!

InfoSeekers attend CHIIR 2019 Conference!

This month, some of our InfoSeekers attended the 2019 annual CHIIR conference in Glasgow, Scotland. Here are some of the highlights!

Rutgers University InfoSeeking students, Jiqun Liu, Souvick Ghosh, and Diana Soltari attended CHIIR, along with InfoSeekers Chirag Shah and Matthew Mitsui.

Infoseekers at the Glasgow City Chambers for CHIIR 2019.

InfoSeeker, Diana Soltari presented Coagmento 3.0, which is an interactive web application that allows researchers to prototype web-search behavior studies through a GUI. The demonstration presented the front-end administrative functionality of Coagmento, including, but not limited to, stage and questionnaire creation. 

Diana Soltari presenting Coagmento 3.0 at CHIIR 2019.

InfoSeeker, Jiqun Liu presented several papers this year at CHIIR! He presented one full paper, one short paper, and one doctoral consortium paper. Jiqun’s papers reported user studies on the interactions between task, information seeking intentions, and user search behavior in information seeking episodes.

Jiqun Liu presenting a short paper at CHIIR 2019.
Jiqun Liu presenting a paper at CHIIR 2019.
Creating social impact through research

Creating social impact through research

Over the years, our lab has done some really groundbreaking work in the fields of information seeking, interactive information retrieval, social and collaborative search, social media, and human-computer interaction. Almost all of it had been geared toward scholarly communities. It makes sense. After all, we are operating in an academic research setting.

But lately at least I have been pondering about how what we do could and should benefit the society. And I don’t mean it in subtle, indirect, or some hypothetical ways. Sure, everything we do has a positive impact on people, starting with people doing that work. It earns them class credits, diplomas, and salaries. It also helps educate students and train professionals in certain skills. But that’s still a very small sample of population. Beyond that, some of our research and technologies developed through that work have impacted various government, educational, research, for-profit, and non-profit organizations in furthering their agendas.

And yes, from time to time we have helped out the United Nations (UN) and a few other organizations more directly with their data problems.

That is still not enough. There are many important issues in the world to address and those of us in privileged positions should do more.

And that is where we launched a new effort called Science for Social Good (S4SG). Under this umbrella, we started rethinking some of our existing works and how they could help address one of the issues of societal importance. Since we already had ties with the UN, and I regularly participate in some of their activities, it made logical sense to start with what the UN considers as a set of important issues. As it happens, the UN has a list of 17 Sustainable Development Goals (SDGs), which they hope will be met by year 2030. And we decided to be a part of the solution.

The UN’s list seemed comprehensive enough, and so, we started from that list and first identified a few organizations who aim to address at least some of those SDGs. And then, we looked inward — to see what activities that we do could help with these SDGs. The result was a pleasant surprise. Several of our projects do actually directly connect to one or more of these SDGs. In other words, by solving those research and development problems, we are directly or indirectly helping the UN (and the world) meet those SDGs. Some of the most common SDGs that our projects are addressing include Good Health and Well-being (SDG-3), Education (SDG-4), and Reduced Inequality (SDG-10).

More importantly, creating the S4SG platform has allowed us to rethink some of our future research activities and see if we could better align them with the societal impact in mind. This is not always easy, but it’s almost always worth doing.

Visit S4SG.org to learn more.

2018: Year in Review!

2018: Year in Review!

As we close out the fall semester and rapidly approach the end of 2018, we must pause to reflect on everything we accomplished this year in the InfoSeeking Lab. We had two students successfully defend their dissertations; six students passed their qualifying exams; and, two students defended their dissertation proposals. The InfoSeeking Lab hosted the 2018 CHIIR conference, as well as attended the ASIS&T 2018 and CSCW 2018 conferences. There were over a half a dozen publications and some of our InfoSeekers were recognized for their contributions to research in information science. Of course, we also made time to run in the Big Chill and socialize as a group. Here’s to a great year of hard work and a hunger to top it all in 2019!

Information Retrieval (IR) Fairness: What is it and what can we do?

Information Retrieval (IR) Fairness: What is it and what can we do?

When you search for information in a search engine such as Google, a list of results is displayed for you to explore further. This process is called Information Retrieval. The contents of the search results are collected based on certain criteria that are catered to you such as: past search history to match your interests, geographic location to relate to what is relevant based on your physical location, and advertising that has been targeted to match your interests and geographic location. These criteria are coded into algorithms to automate the information retrieval process catered to your needs, or what you would potentially consider to be relevant.

For example, let’s say you are searching for information about the healthiness of coffee and you search for “is coffee good for your health.” You may be looking for information that confirms your belief about the benefits of coffee, or you may be simply asking the question “whether coffee is good or bad for your health.” If you asked this question to a human who was an expert in facts about coffee you would likely get an answer that weighs the benefits and harms of coffee. Ideally, when you enter this same question in a search engine, it should return both the goodness and badness about coffee.

Unfortunately, searching for “is coffee good for your health” and “is coffee good for your health” will return a different set of information that is catered to your needs, and not the question as a whole. As a result, catering specifically to the user can create bias. If you are only seeing information that relates to what you are already interested in, or what is geographically near you, there are other perspectives that are intentionally filtered out of the results list.

So, how can we improve the algorithms that are used in the information retrieval process to incorporate more perspectives to reduce bias?

InfoSeeker Ruoyuan Gao is currently working on addressing the presence of bias found in search engine results. Currently, she is exploring several strategies to investigate the relationship between information usefulness and fairness within search engines such as Google. She proposes developing tools to measure the degree of bias in order to create a more balanced list of search results that includes many relevant perspectives for a search topic.

Big Chill 2018!

Big Chill 2018!

This weekend was the annual Big Chill, a charity 5k. As always, some of our InfoSeekers joined in on the fun, exercise, and the good cause! The sun came out just long enough to break from the regular rain we’ve been experiencing to make the run enjoyable.

While some of the runners were getting ready for the race to begin, InfoSeeker Manasa Rath was able to get a shot of the big crowd.

Here’s our very own InfoSeeker, Matt Mitsui getting ready to make his way to the finish line:

And, check out this aerial shot of the race from the InfoSeeking Lab!

This marks the tail end of the Fall 2018 semester. What a great way to energize the start of finals!

InfoSeekers Attend ASIS&T and CSCW November Conferences!

InfoSeekers Attend ASIS&T and CSCW November Conferences!

This month, some of our InfoSeekers attended the 2018 annual ASIS&T conference in Vancouver and the annual CSCW in Jersey City, NJ. Here are some highlights!

Yiwei Wang, presenting the team’s poster at CSCW 2018.

Rutgers University InfoSeeking students, Jiquin Liu, Soumick Mandal, and Yiwei Wang attended CSCW

(Conference on Computer-Supported Cooperative Work and Social Computing) in Jersey City, NJ on Nov 3 – 7. The team presented their poster titled, “Persuasion by Peer or Expert for Web Search.” They presented their preliminary findings on the persuasiveness of two sources of search advice, cognitive authority and peer advice, and their influences on search behaviors.

 

Next, our InfoSeekers attended ASIS&T in Vancouver on Nov 10 – 14. Souvick Ghosh and Manasa Rath participated in SIGInfoLearn workshop. They discussed relevant research that supports searching as learning.

Manasa Rath, center, and Souvick Ghosh, far right, at the SIGInfoLearn Workshop at ASIS&T 2018.

Yiwei Wang also attended ASIS&T and co-organized this year’s SIG-USE Symposium with Annie Chen (University of Washington), Melissa Ocepek (University of Illinois Urbana Champaign), and Devendra Potnis (University of Tennessee at Knoxville). The SIG-USE Symposium is an annual workshop held by ASIS&T Special Interest Group on Information Needs, Seeking, and Use and it focuses on the behavioral and cognitive activities of users, and their affective states as they interact with information. The theme this year was Moving Toward the Future of Information Behavior Research and Practice. It was an engaging and inspiring event, and they had 42 participants this year!

Manasa Rath, far right, receiving the ASIS&T New Leader Award at ASIS&T 2018.

Finally, we are very excited to announce that Manasa Rath received the New Leader Award at ASIS&T. She was one of six students to receive the award. Additionally, Manasa will be working with ASIST Board of Directors in the Professional Development committee. Congratulations, Manasa!

Manasa Rath and her ASIS&T New Leader Award at ASIS&T 2018.