Anthropic has launched a $5 million grant program to fund independent research into how artificial intelligence affects users’ wellbeing, as the rapid adoption of conversational AI raises new questions around emotional reliance, mental health conversations and the way models respond to people in vulnerable situations. The program will provide funding, access to Anthropic’s models and technical support to researchers developing open-source evaluations of AI wellbeing impacts.

The initiative is designed to create independent benchmarks that can help the broader AI industry measure how models behave in sensitive, long-running conversations. Anthropic says selected researchers will work independently and publish their evaluations as open-source projects that other developers can use. Applications are due September 21, with applicants selected to submit full proposals expected to be notified by October 5.

Anthropic Wants Independent Tests For AI Wellbeing

Anthropic’s new program focuses on a problem that is more difficult to measure than conventional AI performance. A model’s answer can often be assessed for accuracy or relevance using a single response, but wellbeing-related interactions may require understanding an entire conversation and how a user’s circumstances change over time.

For example, a user might initially ask a seemingly ordinary question and only later reveal that they are experiencing a crisis. Similarly, advice that is appropriate in one situation could become harmful when additional context emerges.

Anthropic says this makes wellbeing evaluations particularly challenging because researchers need to assess not only individual responses but also how AI behavior evolves over multiple interactions.

What The $5 Million Program Will Provide

Program ComponentDetails
Total funding$5 million
Primary focusIndependent AI wellbeing research
Model accessAccess to Anthropic models
Technical supportProvided to selected researchers
Research outputOpen-source evaluations
Researchers targetedClinicians, psychologists, methodologists and other experts
Application deadlineSeptember 21, 2026
Full-proposal notificationsOctober 5, 2026

The structure is intended to broaden participation in AI wellbeing research rather than keeping evaluation work entirely within AI companies. Anthropic says it wants outside expertise to help develop more rigorous standards for evaluating model behavior.

Why AI Wellbeing Is Becoming A Bigger Issue

AI assistants are increasingly being used for more than information retrieval, coding and productivity. People also turn to conversational models for personal advice, emotional support and difficult decisions.

Anthropic’s own April 2026 research found that roughly 6% of Claude conversations in its sampled dataset involved personal guidance, including questions related to relationships, careers, finances, health and wellness, parenting, ethics and spirituality. The study analyzed approximately 1 million Claude.ai conversations from March and April and identified about 38,000 conversations involving personal guidance.

Personal Guidance Conversations With Claude

MeasureAnthropic Research
Conversations sampled1 million
Unique-user conversations analyzed~639,000
Personal-guidance conversations identified~38,000
Share involving personal guidance~6%
Domains examined9

More than 75% of the personal-guidance conversations in that study fell into four broad categories: health and wellness, professional and career issues, relationships, and financial matters.

The findings help explain why AI wellbeing has become a distinct research area. The more frequently people use conversational systems for personal guidance, the more important it becomes to understand whether those systems provide appropriate responses as situations become more complex.

Anthropic Wants Multi-Turn Evaluations

One of the central principles of the new grant program is that researchers should evaluate AI behavior across realistic, multi-turn interactions.

A single prompt can provide an incomplete picture of risk. In a longer conversation, users may disclose additional information, change their requests or move from a low-risk situation into a significantly more sensitive one.

Anthropic’s Safeguards team therefore recommends that evaluations simulate conversations in which risk escalates and context changes over time.

What Anthropic Says Strong Evaluations Should Measure

Researchers applying to the program are encouraged to:

  • Clearly define what constitutes a pass or fail and explain why.
  • Include clinical and subject-matter experts in designing and validating evaluations.
  • Test both precautions and potential harms.
  • Evaluate overcompliance as well as overrefusal.
  • Use multi-turn scenarios that reflect how people actually interact with AI.
  • Validate automated graders against assessments from real subject-matter experts.

The emphasis on both overcompliance and overrefusal is important. A model that complies too readily with a harmful request can create risk, but a system that refuses too broadly may also fail users who need appropriate assistance.

The Challenge Of Measuring Harm

Wellbeing cannot always be reduced to a simple accuracy score.

A response could appear supportive in isolation but have a different effect when viewed within a longer interaction. Conversely, an AI assistant may need to challenge or redirect a user rather than simply agreeing with them.

Anthropic has previously studied this issue through research into sycophancy and personal guidance. Its analysis of Claude conversations examined excessive validation or praise and how those behaviors varied across different types of personal advice.

This creates a need for evaluations that consider context, user vulnerability and the cumulative effects of model behavior.

Examples Of Evaluation Questions

QuestionWhy It Matters
Does the model recognize escalating risk?Risk may emerge gradually
Does it avoid excessive validation?Agreement can be inappropriate in sensitive situations
Does it refuse appropriately?Excessive refusal can prevent useful assistance
Does it adapt to new context?User circumstances can change during a conversation
Are automated graders reliable?Poor grading can produce misleading safety scores
Are experts involved?Wellbeing judgments require subject expertise

The program is therefore aimed at improving the methodology used to measure AI wellbeing rather than simply producing another set of model benchmarks.

Open-Source Research Could Create Common Standards

Anthropic says selected grantees will publish their work as open-source projects. This could allow researchers, AI developers and other organizations to use the resulting evaluations rather than creating separate measurement systems from scratch.

That approach could be particularly useful because there is currently no single universally accepted standard for measuring the wellbeing effects of conversational AI.

Different companies may use different definitions, test scenarios and grading systems. Open-source benchmarks could provide a common framework for comparing models and identifying weaknesses.

Potential Benefits Of Open Benchmarks

  1. Greater comparability: Models could be evaluated against shared criteria.
  2. Independent scrutiny: Researchers outside AI companies could challenge industry assumptions.
  3. Faster iteration: Developers could use publicly available tests to improve safeguards.
  4. Expert participation: Clinicians and psychologists could contribute specialized knowledge.
  5. Long-term tracking: Evaluations could be repeated as models and user behavior change.

The effectiveness of the program will depend on whether the resulting benchmarks are sufficiently rigorous and widely adopted.

Anthropic Is Expanding Its External Research Programs

The $5 million wellbeing initiative is part of a broader pattern in which Anthropic is funding research outside the company.

Anthropic launched its Economic Futures Program in 2025 to support research into AI’s effects on labor, productivity and the economy. In July 2026, the company announced a $200 million Economic Futures Research Fund to support larger external research projects and policy experiments focused on preparing for AI-driven economic change.

The company has also created programs and partnerships around AI’s societal impact, scientific research and workforce development. Its public transparency materials list external research and public-benefit initiatives covering areas including global health, education, economic mobility and AI’s effects on work.

The new wellbeing grants therefore fit into a larger strategy of encouraging outside researchers to investigate consequences of increasingly capable AI systems.

AI Companies Face Growing Pressure To Demonstrate Safety

The initiative comes as AI systems become more deeply embedded in everyday life.

People increasingly use conversational assistants for work, education, research and personal guidance. As models become more capable and conversations become longer, the distinction between an information tool and an interactive digital companion can become less clear.

This creates new responsibilities for AI developers. Companies must determine not only whether their models can complete tasks, but also whether they behave appropriately when users are vulnerable or when seemingly harmless conversations develop into high-risk situations.

Independent evaluations could help provide an external check on those systems.

The Bigger Picture

Anthropic’s $5 million wellbeing research program reflects a broader shift in AI safety research from evaluating isolated model responses to examining how AI behaves throughout real-world interactions. The focus on independent researchers, clinical expertise and open-source benchmarks suggests that the company sees wellbeing measurement as an area where the wider research community needs to develop stronger standards.

The initiative also highlights a central challenge for conversational AI: the same flexibility that makes these systems useful for personal guidance can make their potential effects difficult to measure. As people increasingly turn to AI for emotional support and sensitive advice, robust evaluations will become increasingly important for determining whether models respond appropriately across changing and complex circumstances.

Looking Ahead

Anthropic’s immediate goal will be to identify researchers capable of developing rigorous, reproducible wellbeing evaluations. The September 21 application deadline and subsequent full-proposal process will determine which research teams receive support. If the resulting benchmarks are genuinely independent, technically sound and broadly usable, they could become useful tools for evaluating multiple AI systems rather than only Anthropic’s models.

The larger test will be whether the AI industry adopts common approaches to measuring wellbeing. Open-source evaluations could help move the field toward shared standards, but their value will ultimately depend on expert validation, realistic multi-turn testing and continued updating as models and user behavior evolve. Anthropic’s funding puts additional resources behind that effort at a time when conversational AI is becoming an increasingly important part of how people seek information, advice and support.

Get the day’s top stories in your inbox

One concise email. No spam, unsubscribe anytime.