Senior Software Engineer - AI Evaluation / Coding Agents

Turing
Software engineering
Remote
Contract

Posted on September 9, 2026 · Applications until October 18, 2026

Apply on Turing

I am sending you to the official Turing page. Applying is free and in English.

I earn a commission from the platform when a candidate I referred gets hired. It changes nothing for you.

The Senior Software Engineer evaluates and improves AI coding models by reviewing generated code, identifying failure modes, and creating evaluation signals. Candidates need at least five years of hands-on software engineering experience, strong proficiency in a major production language, and solid code review skills.

Description in English, as published by Turing.

Freelance · Remote · North America, LATAM, or India

About Turing

Turing is one of the world’s leading AGI infrastructure companies, working with frontier AI labs to accelerate model development through high-quality training data, evaluations, and engineering talent.

About the Role

We’re looking for experienced, hands-on software engineers to help evaluate and improve AI coding models.

Rather than primarily building production applications, you’ll work with coding agents across real-world repositories and assess the quality of their work. You’ll review generated code and agent behavior, determine whether solutions are technically correct, identify failure modes, and create the evaluation signals and feedback used to improve model performance.

Think of the coding agent as another engineer whose work you’re reviewing: Can it understand the task? Did it choose the right approach? Is the resulting code correct, robust, and maintainable? Can you explain precisely where it succeeded or failed?

What You’ll Do

  • Evaluate AI-generated code and solutions across real-world software repositories

  • Review agent behavior, tool usage, and code changes for correctness and quality

  • Identify technical errors, weak approaches, and recurring model failure modes

  • Compare model outputs and explain why one solution is better than another

  • Create and refine rubrics and evaluation criteria for coding tasks

  • Produce high-quality evaluation and preference data used to improve coding models

  • Build and maintain pipelines and infrastructure supporting data generation, collection, and evaluation workflows

  • Synthesize findings from data work into clear write-ups, updates, and recommendations for the team

  • Collaborate closely with researchers and engineers to translate qualitative judgment into scalable processes

  • Share clear, actionable findings with AI researchers and engineers

What We’re Looking For

  • 5+ years of hands-on software engineering experience

  • Strong proficiency in Python, TypeScript/JavaScript, Go, or another major production language

  • Experience working in substantial real-world codebases

  • Strong code-review skills and technical judgment

  • Ability to clearly explain why an implementation is correct, incorrect, or could be improved

  • Strong written communication

  • Experience using modern LLMs or AI coding tools

Experience with LLM evaluation, coding agents, RLHF, preference data, rubric design, or post-training is a plus, but not required.

Engagement Details

  • Compensation: Market rate; please provide a specific hourly rate expectation

  • Availability: 40 hours/week preferred, with at least 6 hours of Pacific Time overlap

  • Type: Independent contractor

  • Duration: Approximately 3 months

  • Start: As soon as possible

  • Location: North America, LATAM, or India

Evaluation Process

  • AI interview (~25 minutes)

  • Practical code/AI evaluation exercise (~30 minutes)

  • Hiring manager interview (~20 minutes)

The practical exercise focuses on your ability to review and evaluate AI-generated code, not competitive programming or algorithm puzzles.

Required skills

  • Python
  • Typescript
  • JavaScript
  • Go
  • Java
  • GitHub
  • CI/CD
  • Open Source
  • LLM
  • Code Reviews
  • Code Analysis

About Turing

Turing hires remote experts for AI and engineering projects, from software to medicine. The original posting is on their site.

View this job on Turing

More jobs in Software engineering

Python Engineering Manager, LLM Training & Evaluation

Turing
Software engineering
Remote
Full time

This remote freelance role involves leading large teams of Python engineers and data scientists to support foundational LLM training and evaluation workflows. Candidates must have five or more years of software engineering experience with strong Python proficiency and proven management skills.

  • Python
  • Software Development

Posted yesterday

Domain Expert - Engineering

Turing
Software engineering
Remote
Full time

This remote role involves building scalable front end features using React and JavaScript. Candidates need at least three years of experience as a front end engineer and a degree in computer science or equivalent.

  • Data Engineering

Posted 2 days ago

Open Source Contributor (GitHub)

micro1
Software engineering
Expert
Remote
Contract

The Open Source Contributor explores, debugs, and modifies front-end and back-end components for an AI training project. Applicants need a verifiable GitHub profile with meaningful open-source contributions and strong software engineering experience across multiple programming languages.

  • Python3
  • JAVA
  • Rust
  • +3

Posted 2 days ago$100 to $150/h

Sr. Full-Stack Software Engineer

micro1
Software engineering
Expert
Remote
Contract

The Senior Full Stack Software Engineer builds, debugs, and evaluates front end and back end application components for an AI training project. The role requires strong professional experience developing production full stack applications and hands on expertise across multiple programming languages.

  • Python
  • React
  • JavaScript
  • +5

Posted 2 days ago$50 to $100/h