Jump to Content

Defining the technology of today and tomorrow.

Philosophy

We strive to create an environment conducive to many different types of research across many different time scales and levels of risk.
Learn more about our Philosophy

Philosophy

People

Our researchers drive advancements in computer science through both fundamental and applied research.
Learn more about our People

People
Foundational ML & Algorithms

Algorithms & theory new

Data Management

Data Mining & Modeling

Information Retrieval & the Web

Machine Intelligence

Machine Perception

Machine Translation

Natural Language Processing

Speech Processing

Algorithms & theory new

Data Management

Data Mining & Modeling

Information Retrieval & the Web

Machine Intelligence

Machine Perception

Machine Translation

Natural Language Processing

Speech Processing

Computing Systems & Quantum AI

Distributed Systems & Parallel Computing

Hardware & Architecture

Mobile Systems

Networking

Quantum Computing

Robotics

Security, Privacy, & Abuse Prevention

Software Engineering

Software Systems

Distributed Systems & Parallel Computing

Hardware & Architecture

Mobile Systems

Networking

Quantum Computing

Robotics

Security, Privacy, & Abuse Prevention

Software Engineering

Software Systems

Science, AI & Society

Climate & Sustainability

Economics & Electronic Commerce

Education Innovation

General Science

Health & Bioscience

Human-Computer interaction & Visualization

Responsible AI

Climate & Sustainability

Economics & Electronic Commerce

Education Innovation

General Science

Health & Bioscience

Human-Computer interaction & Visualization

Responsible AI

Explore all research areas
Projects

We regularly open-source projects with the broader research community and apply our developments to Google products.
Learn more about our Projects

Projects

Publications

Publishing our work allows us to share ideas and work collaboratively to advance the field of computer science.
Learn more about our Publications

Publications

Resources

We make products, tools, and datasets available to everyone with the goal of building a more collaborative ecosystem.
Learn more about our Resources

Resources
Shaping the future, together.
Collaborate with us

Student programs

Supporting the next generation of researchers through a wide range of programming.
Learn more about our Student programs

Student programs

Faculty programs

Participating in the academic research community through meaningful engagement with university faculty.
Learn more about our Faculty programs

Faculty programs

Conferences & events

Connecting with the broader research community through events is essential for creating progress in every aspect of our work.
Learn more about our Conferences & events

Conferences & events

Collaborate with us
Careers
Blog

Empirical methodology for crowdsourcing ground truth

Anca Dumitrache

Benjamin Timmermans

Chris Welty

Lora Mois Aroyo

Oana Inel

Semantic Web Journal, 12:3; 2021 (2021)

Google Scholar

Abstract

The process of gathering ground truth data through human annotation is a major bottleneck in the use of information
extraction methods for populating the Semantic Web. Crowdsourcing-based approaches are gaining popularity in the attempt to
solve the issues related to volume of data and lack of annotators. Typically these practices use inter-annotator agreement as a
measure of quality. However, in many domains, such as event detection, there is ambiguity in the data, as well as a multitude
of perspectives of the information examples. We present an empirically derived methodology for efficiently gathering of ground
truth data in a diverse set of use cases covering a variety of domains and annotation tasks. Central to our approach is the use of
CrowdTruth metrics that capture inter-annotator disagreement. We show that measuring disagreement is essential for acquiring
a high quality ground truth. We achieve this by comparing the quality of the data aggregated with CrowdTruth metrics with
majority vote, over a set of diverse crowdsourcing tasks: Medical Relation Extraction, Twitter Event Identification, News Event
Extraction and Sound Interpretation. We also show that an increased number of crowd workers leads to growth and stabilization
in the quality of annotations, going against the usual practice of employing a small number of annotators.

Research Areas