Anthropic Launches $5 Million Grant Program for Evaluating AI's Impact on Wellbeing

·

A major player in the development of artificial intelligence (AI) systems has announced a significant investment in understanding their impact on users’ wellbeing. Anthropic, the company behind the popular conversational AI model Claude, is launching a $5 million grant program to fund independent research into how AI affects those who use it.

According to the company, AI systems have become an integral part of modern life, with many people relying on them for work, learning, and problem-solving. However, as these systems become more conversational and emotionally supportive, there is a growing need for clear standards on their behavior in sensitive conversations.

The grant program aims to address this gap by providing direct funding, access to Anthropic’s models, and technical support to researchers building open-source evaluations of AI’s impact on wellbeing. These evaluations will help the industry measure how its models affect users and develop more effective safeguards against potential harms.

One of the key challenges in evaluating AI’s impact on wellbeing is assessing context-dependent behaviors. Unlike simple accuracy metrics, which can be applied to individual answers or tasks, wellbeing requires a deeper understanding of complex conversations and nuanced user interactions.

For instance, a user may not immediately share thoughts of self-harm during an initial conversation with an AI model. However, as the conversation progresses and more context is revealed, it becomes clear that the user needs a cautious response from the model to ensure their safety. Similarly, a response that might be reasonable in one context could be harmful or inappropriate in another.

Anthropic acknowledges these complexities and has been working on developing safeguards to identify sensitive conversations and provide appropriate responses. However, the company recognizes that this is an ongoing process requiring continuous research and improvement.

The grant program seeks to invite more experts from various fields, including clinicians, psychologists, methodologists, and others, to contribute their expertise to this emerging field. By funding independent evaluations and benchmarks of user wellbeing, Anthropic hopes to create a comprehensive understanding of AI’s impact on users’ lives.

As part of the program, Anthropic is sharing guidance from its Safeguards team on what makes a wellbeing evaluation rigorous enough to build upon. The company emphasizes that successful evaluations should clearly define what they are measuring and involve clinical and subject-matter experts in their design and validation.

Additionally, grantees will be expected to test both precautions and harms, reflecting how users actually use AI systems in real-world scenarios. This includes constructing multi-turn conversations where risk escalates and context shifts over the course of a long conversation. Finally, evaluations should validate their graders against real subject-matter experts.

Anthropic is encouraging researchers from around the world to apply for the grant program by September 21. Applicants who are selected to submit full proposals will be notified by October 5. The company’s efforts in this area demonstrate its commitment to developing more responsible and user-centric AI systems.