Anthropic has launched a $5 million grant program to fund independent research into how artificial intelligence affects users’ wellbeing, as the rapid adoption of conversational AI raises new questions around emotional reliance, mental health conversations and the way models respond to people in vulnerable situations. The program will provide funding, access to Anthropic’s models and technical support to researchers developing open-source evaluations of AI wellbeing impacts.
The initiative is designed to create independent benchmarks that can help the broader AI industry measure how models behave in sensitive, long-running conversations. Anthropic says selected researchers will work independently and publish their evaluations as open-source projects that other developers can use. Applications are due September 21, with applicants selected to submit full proposals expected to be notified by October 5.
Anthropic Wants Independent Tests For AI Wellbeing
Anthropic’s new program focuses on a problem that is more difficult to measure than conventional AI performance. A model’s answer can often be assessed for accuracy or relevance using a single response, but wellbeing-related interactions may require understanding an entire conversation and how a user’s circumstances change over time.
For example, a user might initially ask a seemingly ordinary question and only later reveal that they are experiencing a crisis. Similarly, advice that is appropriate in one situation could become harmful when additional context emerges.
Anthropic says this makes wellbeing evaluations particularly challenging because researchers need to assess not only individual responses but also how AI behavior evolves over multiple interactions.
What The $5 Million Program Will Provide
| Program Component | Details |
|---|---|
| Total funding | $5 million |
| Primary focus | Independent AI wellbeing research |
| Model access | Access to Anthropic models |
| Technical support | Provided to selected researchers |
| Research output | Open-source evaluations |
| Researchers targeted | Clinicians, psychologists, methodologists and other experts |
| Application deadline | September 21, 2026 |
| Full-proposal notifications | October 5, 2026 |
The structure is intended to broaden participation in AI wellbeing research rather than keeping evaluation work entirely within AI companies. Anthropic says it wants outside expertise to help develop more rigorous standards for evaluating model behavior.
Why AI Wellbeing Is Becoming A Bigger Issue
AI assistants are increasingly being used for more than information retrieval, coding and productivity. People also turn to conversational models for personal advice, emotional support and difficult decisions.
Anthropic’s own April 2026 research found that roughly 6% of Claude conversations in its sampled dataset involved personal guidance, including questions related to relationships, careers, finances, health and wellness, parenting, ethics and spirituality. The study analyzed approximately 1 million Claude.ai conversations from March and April and identified about 38,000 conversations involving personal guidance.
Personal Guidance Conversations With Claude
| Measure | Anthropic Research |
|---|---|
| Conversations sampled | 1 million |
| Unique-user conversations analyzed | ~639,000 |
| Personal-guidance conversations identified | ~38,000 |
| Share involving personal guidance | ~6% |
| Domains examined | 9 |
More than 75% of the personal-guidance conversations in that study fell into four broad categories: health and wellness, professional and career issues, relationships, and financial matters.
The findings help explain why AI wellbeing has become a distinct research area. The more frequently people use conversational systems for personal guidance, the more important it becomes to understand whether those systems provide appropriate responses as situations become more complex.
Anthropic Wants Multi-Turn Evaluations
One of the central principles of the new grant program is that researchers should evaluate AI behavior across realistic, multi-turn interactions.
A single prompt can provide an incomplete picture of risk. In a longer conversation, users may disclose additional information, change their requests or move from a low-risk situation into a significantly more sensitive one.
Anthropic’s Safeguards team therefore recommends that evaluations simulate conversations in which risk escalates and context changes over time.
What Anthropic Says Strong Evaluations Should Measure
Researchers applying to the program are encouraged to:
- Clearly define what constitutes a pass or fail and explain why.
- Include clinical and subject-matter experts in designing and validating evaluations.
- Test both precautions and potential harms.
- Evaluate overcompliance as well as overrefusal.
- Use multi-turn scenarios that reflect how people actually interact with AI.
- Validate automated graders against assessments from real subject-matter experts.
The emphasis on both overcompliance and overrefusal is important. A model that complies too readily with a harmful request can create risk, but a system that refuses too broadly may also fail users who need appropriate assistance.
The Challenge Of Measuring Harm
Wellbeing cannot always be reduced to a simple accuracy score.
A response could appear supportive in isolation but have a different effect when viewed within a longer interaction. Conversely, an AI assistant may need to challenge or redirect a user rather than simply agreeing with them.
Anthropic has previously studied this issue through research into sycophancy and personal guidance. Its analysis of Claude conversations examined excessive validation or praise and how those behaviors varied across different types of personal advice.
This creates a need for evaluations that consider context, user vulnerability and the cumulative effects of model behavior.
Examples Of Evaluation Questions
| Question | Why It Matters |
|---|---|
| Does the model recognize escalating risk? | Risk may emerge gradually |
| Does it avoid excessive validation? | Agreement can be inappropriate in sensitive situations |
| Does it refuse appropriately? | Excessive refusal can prevent useful assistance |
| Does it adapt to new context? | User circumstances can change during a conversation |
| Are automated graders reliable? | Poor grading can produce misleading safety scores |
| Are experts involved? | Wellbeing judgments require subject expertise |
The program is therefore aimed at improving the methodology used to measure AI wellbeing rather than simply producing another set of model benchmarks.
Open-Source Research Could Create Common Standards
Anthropic says selected grantees will publish their work as open-source projects. This could allow researchers, AI developers and other organizations to use the resulting evaluations rather than creating separate measurement systems from scratch.
That approach could be particularly useful because there is currently no single universally accepted standard for measuring the wellbeing effects of conversational AI.
Different companies may use different definitions, test scenarios and grading systems. Open-source benchmarks could provide a common framework for comparing models and identifying weaknesses.
Potential Benefits Of Open Benchmarks
- Greater comparability: Models could be evaluated against shared criteria.
- Independent scrutiny: Researchers outside AI companies could challenge industry assumptions.
- Faster iteration: Developers could use publicly available tests to improve safeguards.
- Expert participation: Clinicians and psychologists could contribute specialized knowledge.
- Long-term tracking: Evaluations could be repeated as models and user behavior change.
The effectiveness of the program will depend on whether the resulting benchmarks are sufficiently rigorous and widely adopted.
Anthropic Is Expanding Its External Research Programs
The $5 million wellbeing initiative is part of a broader pattern in which Anthropic is funding research outside the company.
Anthropic launched its Economic Futures Program in 2025 to support research into AI’s effects on labor, productivity and the economy. In July 2026, the company announced a $200 million Economic Futures Research Fund to support larger external research projects and policy experiments focused on preparing for AI-driven economic change.
The company has also created programs and partnerships around AI’s societal impact, scientific research and workforce development. Its public transparency materials list external research and public-benefit initiatives covering areas including global health, education, economic mobility and AI’s effects on work.
The new wellbeing grants therefore fit into a larger strategy of encouraging outside researchers to investigate consequences of increasingly capable AI systems.
AI Companies Face Growing Pressure To Demonstrate Safety
The initiative comes as AI systems become more deeply embedded in everyday life.
People increasingly use conversational assistants for work, education, research and personal guidance. As models become more capable and conversations become longer, the distinction between an information tool and an interactive digital companion can become less clear.
This creates new responsibilities for AI developers. Companies must determine not only whether their models can complete tasks, but also whether they behave appropriately when users are vulnerable or when seemingly harmless conversations develop into high-risk situations.
Independent evaluations could help provide an external check on those systems.
The Bigger Picture
Anthropic’s $5 million wellbeing research program reflects a broader shift in AI safety research from evaluating isolated model responses to examining how AI behaves throughout real-world interactions. The focus on independent researchers, clinical expertise and open-source benchmarks suggests that the company sees wellbeing measurement as an area where the wider research community needs to develop stronger standards.
The initiative also highlights a central challenge for conversational AI: the same flexibility that makes these systems useful for personal guidance can make their potential effects difficult to measure. As people increasingly turn to AI for emotional support and sensitive advice, robust evaluations will become increasingly important for determining whether models respond appropriately across changing and complex circumstances.
Looking Ahead
Anthropic’s immediate goal will be to identify researchers capable of developing rigorous, reproducible wellbeing evaluations. The September 21 application deadline and subsequent full-proposal process will determine which research teams receive support. If the resulting benchmarks are genuinely independent, technically sound and broadly usable, they could become useful tools for evaluating multiple AI systems rather than only Anthropic’s models.
The larger test will be whether the AI industry adopts common approaches to measuring wellbeing. Open-source evaluations could help move the field toward shared standards, but their value will ultimately depend on expert validation, realistic multi-turn testing and continued updating as models and user behavior evolve. Anthropic’s funding puts additional resources behind that effort at a time when conversational AI is becoming an increasingly important part of how people seek information, advice and support.
Get the day’s top stories in your inbox
One concise email. No spam, unsubscribe anytime.



