Behavioral Health Expert, AI Safety and Model Evaluation
$45 to $70/hIndicative range provided by Mercor
Posted on October 2, 2026 · Applications until November 6, 2026
I am sending you to the official Mercor page. Applying is free and in English.
I earn a commission from the platform when a candidate I referred gets hired. It changes nothing for you.
Behavioral health experts evaluate AI model conversations for safety, neutrality, and sound judgment in sensitive user interactions. Candidates must hold a degree in a relevant field and possess at least three years of professional experience in mental health, counseling, or social services.
Description in English, as published by Mercor.
Help a leading AI lab ensure its models respond to people with care, balance and sound judgment.
1. Overview
A leading AI lab is seeking behavioral health experts to help evaluate and improve how its AI models handle sensitive, everyday conversations. People increasingly turn to AI for support with relationships, family dynamics, emotional wellbeing, personal beliefs and difficult life decisions. These conversations rarely involve crisis content, but they are exactly where a model's judgment matters most: whether it stays balanced, avoids taking sides, resists telling people what they want to hear, respects a person's beliefs without endorsing or dismissing them, and understands the limits of its role. You'll review these interactions, assess whether the model's responses are appropriate, neutral and safe, and help the lab's researchers define what good looks like. If you bring professional experience in mental health, counseling, social work or behavioral science, and you can articulate clearly why one response serves a person better than another, this role is for you. This is a part-time commitment of at least 20 hours per week, with the option to increase to up to 40 hours per week.
This is a W-2 employment position with Cincinnatus LLC, with the opportunity to be placed at a leading AI lab as part of their extended workforce.
2. Key Responsibilities
-
Evaluate conversations between users and AI models across topics such as relationship and family advice, emotional wellbeing, spiritual and metaphysical questions, and unconventional or unfounded beliefs, and assess whether responses are neutral, appropriate and safe.
-
Identify patterns of concern, including excessive agreement, taking sides, reinforcing distorted or unfounded beliefs, moralizing, or overstepping into clinical or directive advice, and document them with clear written rationale.
-
Develop rubrics, guidelines and reference responses that define a balanced, supportive and appropriately bounded reply, grounded in established practice from counseling and behavioral health.
-
Design test scenarios and conversations that probe how models handle sensitive but non crisis topics.
-
Collaborate with the lab's researchers and fellow experts to keep evaluation standards consistent, calibrated and well documented.
3. Core Qualifications
-
A degree in psychology, counseling, social work, behavioral health, behavioral science, human services or a closely related field, or equivalent professional experience in the mental health field.
-
3+ years of professional experience supporting people in a mental health, counseling or social services setting, for example as a therapist, counselor, clinical social worker, psychologist, psychiatric nurse, case manager, crisis counselor, peer support specialist or mental health advocate. Clinical licensure (for example LMFT, LCSW, LPC, LMHC, PsyD or PhD) is valued but not required.
-
Demonstrated ability to remain neutral and nonjudgmental across diverse perspectives, relationships, belief systems and worldviews, and to explain the reasoning behind a professional judgment.
-
Working familiarity with concepts such as sycophancy, cognitive distortions, healthy boundaries and client centered approaches such as motivational interviewing.
-
Ability to engage reliably for at least 20 hours/week during weekdays.
-
Strong written communication skills and the ability to deliver precise, well structured written feedback.
Nice to have: a background in AI safety, applied ethics, trust and safety or content policy; experience in couples, family or relationship counseling; familiarity with spiritual care, religious or alternative belief communities, or the psychology of misinformation and conspiracy belief; prior experience evaluating, annotating or red teaming AI systems. You don't need all of these to apply.
About Cincinnatus LLC: Cincinnatus LLC is an enterprise staffing company that partners with leading technology companies to source and employ highly skilled professionals for contingent and contract-based opportunities. Cincinnatus serves as the employer of record for these engagements, providing W-2 employment, payroll, benefits, and compliance, while placing employees directly within client teams to work on high-impact initiatives.
Equal Employment Opportunity: Cincinnatus is proud to be an Equal Employment Opportunity employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, reproductive health decisions, or related medical conditions), sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with a disability, genetic information, political views or activity, or any other legally protected characteristic.
5 openings
Only: United States
About Mercor
Mercor is a US marketplace hiring remote experts for AI and consulting projects, from law to engineering. The original posting is on their site.
View this job on MercorMore jobs in Sciences and research
Bilingual Cantonese Psychologist (PhD)
This remote contractor role involves analyzing psychological case studies, developing bilingual scenarios in Cantonese and English, and evaluating AI-generated content for accuracy and ethical standards. Candidates must hold a PhD in Psychology and possess native or near-native fluency in both Cantonese and English.
- bilingual communication
- ethical decision-making
- cultural sensitivity
- +5
Posted 5 days ago$100 to $200/h
Bilingual Cantonese Psychiatrist
This remote contractor position involves developing and evaluating psychiatric case studies to help train AI models. Candidates must possess a medical degree with a completed psychiatry residency, a valid medical license, and native fluency in Cantonese.
- bilingual communication
- ethical decision-making
- cultural sensitivity
- +5
Posted 5 days ago$100 to $200/h
Bilingual Vietnamese Psychologist (PhD)
This remote contractor position involves analyzing psychological case studies and developing culturally sensitive training materials for AI systems. Candidates must hold a PhD in Psychology and possess native or near-native fluency in Vietnamese and English.
- bilingual communication
- ethical decision-making
- cultural sensitivity
- +5
Posted 6 days ago$100 to $200/h
Bilingual Russian Psychologist (PhD)
The Bilingual Russian Psychologist evaluates psychological case studies and reviews AI generated content to train artificial intelligence models. The role requires a PhD in Psychology and native or near native fluency in both Russian and English.
- bilingual communication
- ethical decision-making
- cultural sensitivity
- +5
Posted 6 days ago$100 to $200/h
