🤖 AI & Beyond

Effective Strategies for Enhancing Data Quality

Building a Responsible Approach to Data Collection

with the Partnership on AI

At DeepMind, our goal is to ensure that everything we do adheres to the highest standards of safety and ethics, in alignment with our Operating Principles. A key focus for us is how we collect our data. Over the past year, we’ve collaborated with the Partnership on AI (PAI) to carefully assess these challenges and co-develop standardized best practices and processes for responsible human data collection.

Human Data Collection

More than three years ago, we established our Human Behavioural Research Ethics Committee (HuBREC), a governance group modeled on academic institutional review boards (IRBs), such as those found in hospitals and universities. The aim is to protect the dignity, rights, and welfare of the human participants involved in our studies. This committee oversees behavioral research that involves experiments with humans as the subject of study, including investigations into how humans interact with artificial intelligence (AI) systems during decision-making processes.

In addition to projects involving behavioral research, the AI community has increasingly focused on ‘data enrichment’—tasks carried out by humans to train and validate machine learning models, such as data labeling and model evaluation. While behavioral research typically utilizes voluntary participants, data enrichment involves individuals being compensated for tasks that enhance AI models.

These tasks are generally conducted through crowdsourcing platforms, often raising ethical considerations related to worker pay, welfare, and equity, which may lack the necessary guidance or governance systems to meet sufficient standards. As research labs progress towards developing increasingly sophisticated models, reliance on data enrichment practices will likely escalate, emphasizing the need for stronger guidance.

The Best Practices

In accordance with our Operating Principles, we are dedicated to upholding and contributing to best practices in AI safety and ethics, encompassing fairness and privacy, to prevent unintended outcomes that could lead to risks of harm.

Following PAI’s recent white paper on Responsible Sourcing of Data Enrichment Services, we have collaborated to develop our practices and processes for data enrichment. This includes five steps that AI practitioners can follow to enhance the working conditions for individuals involved in data enrichment tasks:

  • Select an appropriate payment model and ensure all workers are paid above the local living wage.
  • Design and conduct a pilot before launching a data enrichment project.
  • Identify suitable workers for the desired task.
  • Provide verified instructions and/or training materials for workers to follow.
  • Establish clear and regular communication mechanisms with workers.

Together, we created the necessary policies and resources, gathering multiple rounds of feedback from our internal legal, data, security, ethics, and research teams. We then piloted them on a small scale before rolling them out organization-wide. These documents offer greater clarity on how to effectively set up data enrichment tasks at DeepMind, enhancing our researchers’ confidence in study design and execution. This initiative has increased the efficiency of our approval and launch processes and significantly improved the experience of individuals involved in data enrichment tasks.

Further information on responsible data enrichment practices and their integration into our existing processes is documented in a recent case study, illustrating the implementation of such practices at DeepMind. We also provide helpful resources and supporting materials for AI practitioners and organizations looking to develop similar processes.

Looking Forward

While these best practices form the foundation of our work, we believe that they should not be solely relied upon to ensure the highest standards of participant or worker welfare and safety in research. Each project at DeepMind is unique, which is why we have a dedicated human data review process to continually engage with research teams and identify and mitigate risks on a case-by-case basis.

This effort aims to serve as a resource for other organizations interested in enhancing their data enrichment sourcing practices. We aspire to foster cross-sector conversations that may further develop these guidelines and resources for teams and partners. Through this collaboration, we hope to ignite broader discussions on how the AI community can continue to advance norms of responsible data collection and collectively establish better industry standards.

Learn more about our Operating Principles.