Applied Computer Science Benchmark Specialist
$66 to $84/hIndicative range provided by Mercor
Posted on August 13, 2026 · Applications until November 5, 2026
I am sending you to the official Mercor page. Applying is free and in English.
I earn a commission from the platform when a candidate I referred gets hired. It changes nothing for you.
The Applied Computer Science Benchmark Specialist authors and reviews rigorous academic assessment content to advance AI capabilities. Applicants must hold a PhD or be a doctoral candidate in Computer Science or a related field, and possess a strong command of graduate-level computer science theory.
Description in English, as published by Mercor.
Role Overview
We are seeking expert computer scientists to author and review high-quality academic assessment content for an AI research initiative. You will write and verify rigorous multiple-choice questions across core computer science domains, evaluate solution quality, and help establish gold-standard benchmarks used to advance AI capabilities.
You will be assigned one of two task types:
-
Question Authoring, Create original, challenging multiple-choice questions in your area of computer science expertise, rate their difficulty, and submit them for review.
-
Question Verification, Review pre-written questions for accuracy, clarity, and rigor. Edit where needed, rate difficulty, and document any changes made.
Computer Science Domains Covered
Accelerator / GPU Kernel Engineering, Formal Methods & Automated Reasoning, Computer Architecture & Accelerators, Distributed Systems, DevOps & Site Reliability, Data Engineering & Databases, Cloud & Infrastructure, OS & Systems Kernel, Machine Learning Engineering, Web & API Development, Embedded Systems Engineering, Computer Graphics & Game Development, Mobile Engineering.
Key Responsibilities
-
Author original computer science questions that test deep conceptual understanding, not surface-level recall
-
Ensure questions are unambiguous, self-contained, and precisely defined, all necessary information must be in the problem statement
-
Rate each question's difficulty: Medium (intro undergraduate), Hard (advanced undergraduate), or Expert (post-graduate and above)
-
Provide 1 correct answer and 9 plausible but subtly incorrect alternatives that challenge expert-level solvers
-
Write step-by-step Chain-of-Thought solutions with clear, concise intermediate steps in markdown format
-
Supply 1, 5 academic references per question from reputable sources (peer-reviewed journals, university repositories)
-
For verification tasks: flag issues with clarity, completeness, precision, or solvability and justify any edits made
Ideal Qualifications
-
PhD or doctoral candidate in Computer Science, Electrical Engineering, or a closely related field
-
Master's degree considered for candidates with exceptional depth in a specific subdomain
-
Strong command of graduate-level CS theory, algorithms, systems design, and/or machine learning
-
Research publications, industry experience at top tech companies, or competitive programming background is a strong plus
-
Excellent written English and ability to express complex ideas clearly and concisely
More About the Opportunity
-
Expected commitment: 10+ hours/week
-
Asynchronous, fully remote work
Open worldwide
About Mercor
Mercor is a US marketplace hiring remote experts for AI and consulting projects, from law to engineering. The original posting is on their site.
View this job on MercorMore jobs in Sciences and research
Bilingual Cantonese Psychologist (PhD)
This remote contractor role involves analyzing psychological case studies, developing bilingual scenarios in Cantonese and English, and evaluating AI-generated content for accuracy and ethical standards. Candidates must hold a PhD in Psychology and possess native or near-native fluency in both Cantonese and English.
- bilingual communication
- ethical decision-making
- cultural sensitivity
- +5
Posted 5 days ago$100 to $200/h
Bilingual Cantonese Psychiatrist
This remote contractor position involves developing and evaluating psychiatric case studies to help train AI models. Candidates must possess a medical degree with a completed psychiatry residency, a valid medical license, and native fluency in Cantonese.
- bilingual communication
- ethical decision-making
- cultural sensitivity
- +5
Posted 5 days ago$100 to $200/h
Behavioral Health Expert, AI Safety and Model Evaluation
Behavioral health experts evaluate AI model conversations for safety, neutrality, and sound judgment in sensitive user interactions. Candidates must hold a degree in a relevant field and possess at least three years of professional experience in mental health, counseling, or social services.
Posted 5 days ago$45 to $70/h
Bilingual Vietnamese Psychologist (PhD)
This remote contractor position involves analyzing psychological case studies and developing culturally sensitive training materials for AI systems. Candidates must hold a PhD in Psychology and possess native or near-native fluency in Vietnamese and English.
- bilingual communication
- ethical decision-making
- cultural sensitivity
- +5
Posted 6 days ago$100 to $200/h
