STEM & Technical Specialist, Rubric Design

SME Careers
Sciences and research
Remote
United States only
Contract

$100/hIndicative range provided by SME Careers

Posted on September 29, 2026 · Applications until November 6, 2026

Apply on SME Careers

I am sending you to the official SME Careers page. Applying is free and in English.

I earn a commission from the platform when a candidate I referred gets hired. It changes nothing for you.

Tailor my resume to this job

The STEM and Technical Specialist designs advanced evaluation rubrics, creates challenging technical questions, and assesses AI-generated content to improve frontier models. Candidates must hold a bachelor's degree in a technical discipline and possess at least three years of domain experience.

Description in English, as published by SME Careers.

In this hourly, remote contractor role, you will work as a STEM & Technical Subject Matter Expert (SME) to create challenging expert-level questions, evaluate AI-generated technical content, and design detailed rubrics that define what a correct, rigorous, and expert-quality answer should contain.

You will develop problems within your area of expertise that test advanced reasoning rather than simple factual recall. You will identify weaknesses in frontier AI models, including incorrect assumptions, incomplete reasoning, mathematical or scientific errors, missed edge cases, and answers that appear plausible but fail under expert scrutiny.

A major focus of this role is rubric design and writing. You will translate complex technical judgment into clear, specific, and measurable evaluation criteria, including required concepts, reasoning steps, acceptable alternative approaches, critical errors, and partial-credit considerations. You may also critique and improve rubrics developed by other experts.

This role is with SME Careers, a fast-growing AI Data Services company and subsidiary of SuperAnnotate, delivering training data for many of the world’s largest AI companies and foundation-model labs. Your technical expertise directly helps improve the world’s premier AI models by making their reasoning more accurate, rigorous, and reliable.

Responsibilities

  • Create expert-level technical questions: Develop difficult questions that test genuine technical reasoning and expose weaknesses in advanced AI systems.
  • Design evaluation rubrics: Write structured criteria defining what a correct, rigorous, and complete response must contain.
  • Define partial-credit standards: Identify essential reasoning steps, acceptable alternatives, minor errors, and critical failures.
  • Evaluate AI-generated solutions: Assess outputs for correctness, reasoning quality, completeness, methodology, and technical precision.
  • Identify hidden reasoning errors: Detect cases where an answer reaches the correct conclusion using invalid assumptions or flawed logic.
  • Critique peer rubrics: Review other experts’ criteria for technical accuracy, clarity, completeness, and scoring consistency.
  • Develop reference content: Produce expert answers, explanations, critiques, and other gold-standard content.
  • Support AI training: Create questions, rubrics, evaluations, preference judgments, and reference-answer pairs suitable for RL and SFT workflows.
  • Maintain technical rigor: Ensure content reflects appropriate professional or academic standards within the relevant discipline.

Requirements

  • Bachelor’s degree or higher in Mathematics, Physics, Chemistry, Engineering, Computer Science, Statistics, Applied Science, or another technical discipline.
  • Strong professional proficiency in English, minimum C1, with the ability to write precise technical explanations and evaluation criteria.
  • 3+ years of professional, academic, research, or industry experience in your stated technical domain.
  • Demonstrated advanced expertise in a clearly defined technical field or specialization.
  • Ability to create challenging technical questions that require multi-step reasoning, domain expertise, or professional judgment.
  • Strong ability to design detailed evaluation rubrics defining required reasoning, correct methodology, acceptable alternatives, partial-credit criteria, and critical errors.
  • Ability to distinguish between a correct final answer and an answer supported by valid versus flawed reasoning.
  • Comfortable reviewing and critiquing other experts’ rubrics for ambiguity, missing criteria, redundancy, technical inaccuracies, or poor scoring design.
  • High attention to detail when evaluating formulas, assumptions, units, methodology, edge cases, logical consistency, and technical terminology.
  • Experience with exam writing, academic grading, peer review, research review, technical QA, standards development, or assessment design is strongly preferred.
  • Prior experience with AI evaluation, RLHF, SFT, benchmarking, data annotation, prompt design, or LLM evaluation is preferred.
  • Reliable, self-directed, and able to deliver consistent quality in an hourly, remote contractor workflow.

Required skills

  • STEM
  • Mathematics
  • Physics
  • Chemistry
  • Engineering
  • Computer Science
  • Technical Reasoning
  • AI Evaluation
  • Rubric Design
  • rubric writing
  • Benchmarking
  • Model Evaluation
  • Question Writing
  • Expert Review
  • English
  • Assessment Design
  • Technical Writing
  • Academic Grading
  • Exam Item Writing
  • Peer Review
  • Research Review
  • Quality Assurance (QA)
  • Standards Development
  • Scoring Methodology
  • Partial Credit Scoring
  • Error Analysis
  • Critical Thinking
  • Multi-step Reasoning
  • Statistical Analysis
  • Applied Mathematics

Only: United States

About SME Careers

SME Careers is the expert platform of SuperAnnotate, hiring remote specialists to train and evaluate AI models, from languages to law. The original posting is on their site.

View this job on SME Careers

More jobs in Sciences and research

Bilingual Cantonese Psychologist (PhD)

micro1
Sciences and research
Expert
Remote
Contract

This remote contractor role involves analyzing psychological case studies, developing bilingual scenarios in Cantonese and English, and evaluating AI-generated content for accuracy and ethical standards. Candidates must hold a PhD in Psychology and possess native or near-native fluency in both Cantonese and English.

  • bilingual communication
  • ethical decision-making
  • cultural sensitivity
  • +5

Posted 5 days ago$100 to $200/h

Bilingual Cantonese Psychiatrist

micro1
Sciences and research
Expert
Remote
Contract

This remote contractor position involves developing and evaluating psychiatric case studies to help train AI models. Candidates must possess a medical degree with a completed psychiatry residency, a valid medical license, and native fluency in Cantonese.

  • bilingual communication
  • ethical decision-making
  • cultural sensitivity
  • +5

Posted 5 days ago$100 to $200/h

Behavioral Health Expert, AI Safety and Model Evaluation

Mercor
Sciences and research
Remote
United States only
Part time

Behavioral health experts evaluate AI model conversations for safety, neutrality, and sound judgment in sensitive user interactions. Candidates must hold a degree in a relevant field and possess at least three years of professional experience in mental health, counseling, or social services.

Posted 5 days ago$45 to $70/h

Bilingual Vietnamese Psychologist (PhD)

micro1
Sciences and research
Expert
Remote
Contract

This remote contractor position involves analyzing psychological case studies and developing culturally sensitive training materials for AI systems. Candidates must hold a PhD in Psychology and possess native or near-native fluency in Vietnamese and English.

  • bilingual communication
  • ethical decision-making
  • cultural sensitivity
  • +5

Posted 6 days ago$100 to $200/h