🤖 AI & Beyond

Engaging with AI: Enhancing Language Model Development

New Research on Aligning Conversational Agents with Human Values

Language is an essential human trait and the primary means by which we communicate information including thoughts, intentions, and feelings. Recent breakthroughs in AI research have led to the creation of conversational agents that are capable of nuanced communication with humans. These agents utilize large language models—computational systems trained on extensive text corpora to predict and produce text through advanced statistical techniques.

While models like InstructGPT, Gopher, and LaMDA have demonstrated exceptional performance across various tasks such as translation, question-answering, and reading comprehension, they also present a number of potential risks and failure modes. These concerns include the generation of toxic or discriminatory language and the dissemination of false or misleading information. Such shortcomings hinder the effective application of conversational agents and highlight the ways in which they fail to meet certain communicative ideals.

Most existing approaches to aligning conversational agents have primarily concentrated on anticipating and mitigating harm. However, our new paper, In Conversation with AI: Aligning Language Models with Human Values, takes a different perspective by exploring the nature of successful communication between humans and artificial conversational agents. It examines the values that should guide these interactions across various conversational domains.

Insights from Pragmatics

To address the aforementioned challenges, the paper leverages insights from pragmatics, a field in linguistics and philosophy that emphasizes the significance of conversation purpose, context, and related norms in effective communication. Paul Grice, a notable linguist and philosopher, modeled conversation as a cooperative effort among participants, suggesting that they should:

  • Speak informatively
  • Tell the truth
  • Provide relevant information
  • Avoid obscure or ambiguous statements

Our paper illustrates that a refinement of these maxims is necessary before they can be effectively applied to evaluate conversational agents, given the variations in goals and values inherent within different conversational contexts.

Discursive Ideals

For instance, scientific investigation and communication primarily aim to understand or predict empirical phenomena. In this context, a conversational agent supporting scientific inquiry should ideally restrict itself to making statements backed by sufficient empirical evidence, or qualify its positions according to relevant confidence intervals. For example, an agent stating, “At a distance of 4.246 light years, Proxima Centauri is the closest star to Earth,” should do this only after confirming its accuracy.

Conversely, a conversational agent serving as a moderator in public political discourse must exhibit different virtues. Here, the focus is on managing differences and fostering productive cooperation within a community, necessitating the promotion of democratic values such as toleration, civility, and respect. This underscores why the generation of toxic or prejudicial language by language models is particularly concerning; such language undermines the fundamental value of equal respect within the conversation context. Moreover, the strengths associated with scientific inquiry, such as the thorough presentation of empirical data, may be less critical in public deliberation.

Additionally, in the realm of creative storytelling, communicative exchanges focus on novelty and originality, values that differ significantly from those mentioned earlier. In this setting, greater flexibility regarding imaginative content may be acceptable; however, safeguarding against harmful content disguised as ‘creative uses’ remains vital.

Paths Ahead

This research carries significant implications for the development of aligned conversational AI agents. First, these agents must embody diverse traits based on the contexts of their deployment, indicating that there is no universal model for language alignment. Instead, the suitable mode and evaluative criteria for an agent—including standards of truthfulness—will differ depending on the context and objective of a conversational exchange.

Furthermore, conversational agents hold the potential to foster more respectful and constructive conversations over time through a process known as context construction and elucidation. Even if individuals are unaware of the values guiding a particular conversational practice, an agent can assist them in understanding these values by integrating them into the conversation, enriching the communication experience for the human speaker.