Artificial intelligence

Anthropic Allocates $5 Million to Fund Independent Evaluations of AI’s Impact on User Well-Being

Anthropic has launched a $5 million grants program to fund independent, open-source research measuring the impact of AI systems on user well-being. The program focuses on multi-turn conversations, mental health crises, and the risks of over-response or over-refusal.

2026-09-27
3 min read
13 views
certi.news Editorial Team
Anthropic Allocates $5 Million to Fund Independent Evaluations of AI’s Impact on User Well-Being

Anthropic announced on August 25, 2026, the launch of a $5 million grants program to fund independent research on AI’s impact on user well-being. The program will provide direct funding, access to the company’s models, and technical support to researchers developing open-source evaluation tools that any developer can use.

The company says intelligent systems have become part of work, learning, and problem-solving, while some people also use them as conversation partners or as a source of emotional support during difficult times. However, the sector still lacks clear standards defining how models should behave when a user seeks companionship or tries to cope with a mental health crisis.

Why Is User Well-Being Difficult to Measure?

Anthropic believes this aspect cannot be reduced to examining a single response. A user may not disclose thoughts of self-harm at the beginning of a conversation, while the need for a more cautious response may emerge after several messages. Similarly, advice that is appropriate in one context may be harmful in another.

The company gives the example of advice about a balanced diet or exercise to a user asking about weight loss; such advice could become inappropriate or actually harmful if the user reveals a history of eating disorders. Anthropic says it is working to develop safeguards that detect these patterns and is publishing research on the types of conversations users have with Claude to improve those safeguards and their evaluation methods.

What Will the Grants Research?

Anthropic’s Safeguards department has published guidance for evaluations and benchmarks that it considers suitable foundations for further development. The core criteria include:

  • Clearly defining what is being measured, what constitutes success or failure, and why it matters.
  • Involving clinical experts and specialists in designing and validating the evaluation.
  • Testing precautions and harms together, including the risks of over-compliance and over-refusal.
  • Simulating how people use AI, particularly through multi-turn conversations in which the context changes and the risk may escalate.
  • Verifying the accuracy of automated evaluation tools by comparing them with the judgments of real experts.

What Is Changing in Practice?

Anthropic is not announcing a ready-made unified benchmark; instead, it is funding the development of independent, open evaluations that could help the sector develop such a foundation. The importance of the initiative lies in shifting the discussion from testing the safety of individual responses to examining model behavior across an entire conversation, while accounting for differences in context and the possibility of indirect harm.

Recipients will work with complete independence and will publish their work as open-source projects. The call is aimed at specialists including physicians, mental health professionals, methodology experts, and others. The deadline for applications is September 21, while invited organizations will be notified to submit full proposals by October 5.

The program’s effectiveness will remain tied to researchers’ ability to turn complex concepts such as well-being and psychological harm into measurable and verifiable criteria. Funding evaluations alone also does not determine how their results will be used to design protective barriers or handle sensitive cases.

News source
Anthropic Newsroom
Open original source ↗
c
Author

certi.news Editorial Team

In the same category

You may also like

View all news