Ingénieur logiciel senior - Évaluation d'IA / Agents de codage
Publié le 9 septembre 2026 · Candidatures jusqu'au 18 octobre 2026
Je vous redirige vers la page officielle de Turing. La candidature est gratuite et se fait en anglais.
Je touche une commission de la plateforme si un candidat que j'ai orienté est recruté. Cela ne change rien pour vous.
L'ingénieur logiciel senior évalue et améliore les modèles de codage par IA en examinant le code généré, en identifiant les modes de défaillance et en créant des signaux d'évaluation. Les candidats ont besoin d'au moins cinq ans d'expérience pratique en génie logiciel, d'une solide maîtrise d'un grand langage de production et de compétences avérées en révision de code.
Description en anglais, telle que publiée par Turing.
Freelance · Remote · North America, LATAM, or India
About Turing
Turing is one of the world’s leading AGI infrastructure companies, working with frontier AI labs to accelerate model development through high-quality training data, evaluations, and engineering talent.
About the Role
We’re looking for experienced, hands-on software engineers to help evaluate and improve AI coding models.
Rather than primarily building production applications, you’ll work with coding agents across real-world repositories and assess the quality of their work. You’ll review generated code and agent behavior, determine whether solutions are technically correct, identify failure modes, and create the evaluation signals and feedback used to improve model performance.
Think of the coding agent as another engineer whose work you’re reviewing: Can it understand the task? Did it choose the right approach? Is the resulting code correct, robust, and maintainable? Can you explain precisely where it succeeded or failed?
What You’ll Do
Evaluate AI-generated code and solutions across real-world software repositories
Review agent behavior, tool usage, and code changes for correctness and quality
Identify technical errors, weak approaches, and recurring model failure modes
Compare model outputs and explain why one solution is better than another
Create and refine rubrics and evaluation criteria for coding tasks
Produce high-quality evaluation and preference data used to improve coding models
Build and maintain pipelines and infrastructure supporting data generation, collection, and evaluation workflows
Synthesize findings from data work into clear write-ups, updates, and recommendations for the team
Collaborate closely with researchers and engineers to translate qualitative judgment into scalable processes
Share clear, actionable findings with AI researchers and engineers
What We’re Looking For
5+ years of hands-on software engineering experience
Strong proficiency in Python, TypeScript/JavaScript, Go, or another major production language
Experience working in substantial real-world codebases
Strong code-review skills and technical judgment
Ability to clearly explain why an implementation is correct, incorrect, or could be improved
Strong written communication
Experience using modern LLMs or AI coding tools
Experience with LLM evaluation, coding agents, RLHF, preference data, rubric design, or post-training is a plus, but not required.
Engagement Details
Compensation: Market rate; please provide a specific hourly rate expectation
Availability: 40 hours/week preferred, with at least 6 hours of Pacific Time overlap
Type: Independent contractor
Duration: Approximately 3 months
Start: As soon as possible
Location: North America, LATAM, or India
Evaluation Process
AI interview (~25 minutes)
Practical code/AI evaluation exercise (~30 minutes)
Hiring manager interview (~20 minutes)
The practical exercise focuses on your ability to review and evaluate AI-generated code, not competitive programming or algorithm puzzles.
Compétences recherchées
- Python
- Typescript
- JavaScript
- Go
- Java
- GitHub
- CI/CD
- Open Source
- LLM
- Code Reviews
- Code Analysis
À propos de Turing
Turing recrute des experts à distance pour des projets d'IA et d'ingénierie, du développement à la médecine. L'offre originale est consultable sur leur site.
Voir l'offre sur TuringAutres postes en Développement logiciel
Responsable d'ingénierie Python, entraînement et évaluation de LLM
Ce poste indépendant à distance consiste à diriger de grandes équipes d'ingénieurs Python et de data scientists pour soutenir l'entraînement et l'évaluation des LLM fondamentaux. Les candidats doivent justifier d'au moins cinq ans d'expérience en génie logiciel, d'une solide maîtrise de Python et de compétences avérées en gestion.
- Python
- Software Development
Publié hier
Expert sectoriel - Ingenierie
Ce role distant consiste a batir des fonctionnalites front-end evolutives avec React et JavaScript. Les candidats ont besoin d'au moins trois annees d'experience en tant qu'ingenieur front-end et d'un diplome en informatique ou equivalent.
- Data Engineering
Publié il y a 2 j
Contributeur open source (GitHub)
Le contributeur open source explore, débogue et modifie des composants front-end et back-end pour un projet d'entraînement en IA. Les candidats ont besoin d'un profil GitHub vérifiable avec des contributions open source significatives et d'une solide expérience en génie logiciel dans plusieurs langages de programmation.
- Python3
- JAVA
- Rust
- +3
Publié il y a 2 j100 $US à 150 $US/h
Ingénieur logiciel full stack senior
L'ingénieur logiciel full stack senior conçoit, débugue et évalue des composants d'application front end et back end pour un projet de formation en IA. Le poste requiert une solide expérience professionnelle dans le développement d'applications full stack destinées à la production ainsi qu'une expertise pratique sur plusieurs langages de programmation.
- Python
- React
- JavaScript
- +5
Publié il y a 2 j50 $US à 100 $US/h