MLE Bench, ML Engineers
Posted on January 27, 2026 · Applications until October 18, 2026
I am sending you to the official Turing page. Applying is free and in English.
I earn a commission from the platform when a candidate I referred gets hired. It changes nothing for you.
Machine Learning Engineers build, run, and modify model training, evaluation, and inference pipelines for benchmark-driven AI systems. Candidates need at least three years of professional machine learning experience and strong proficiency in Python.
Description in English, as published by Turing.
About Turing
Based in San Francisco, California, Turing is the world’s leading research accelerator for frontier AI labs and a trusted partner for global enterprises deploying advanced AI systems. Turing supports customers in two ways: first, by accelerating frontier research with high-quality data, advanced training pipelines, plus top AI researchers who specialize in coding, reasoning, STEM, multilinguality, multimodality, and agents; and second, by applying that expertise to help enterprises transform AI from proof of concept into proprietary intelligence with systems that perform reliably, deliver measurable impact, and drive lasting results on the P&L
Role Overview
We are looking for experienced Machine Learning Engineers (MLE Bench) to contribute to benchmark-driven evaluation projects focused on real-world machine learning systems. This role involves hands-on work with production-grade ML codebases, model training and evaluation pipelines, and deployment-oriented workflows to help assess and improve the capabilities of advanced AI systems.
The ideal candidate is comfortable bridging research and engineering, working deeply with models, data, and infrastructure in realistic ML environments.
What does day-to-day life look like?
- Work with real-world ML codebases to support MLE Bench, style evaluation tasks.
- Build, run, and modify model training, evaluation, and inference pipelines.
- Prepare datasets, features, and metrics for ML benchmarking and validation.
- Debug, refactor, and improve production-like ML systems for correctness and performance.
- Evaluate model behavior, failure modes, and edge cases relevant to benchmark tasks.
- Write clean, reproducible, and well-documented Python code for ML workflows.
- Participate in code reviews to ensure high standards of engineering quality.
- Collaborate with researchers and engineers to design challenging, real-world ML engineering tasks for AI system evaluation.
Requirements
- Minimum 3+ years of overall experience as a Machine Learning Engineer or Software Engineer (ML-focused).
- Strong proficiency in Python for machine learning and data workflows.
- Hands-on experience with model training, evaluation, and inference pipelines.
- Solid understanding of machine learning fundamentals (supervised/unsupervised learning, evaluation metrics, optimization).
- Experience working with ML frameworks (e.g., PyTorch, TensorFlow, JAX, or similar).
- Ability to understand, navigate, and modify complex, real-world ML codebases.
- Experience writing readable, reusable, and maintainable production-quality code.
- Strong problem-solving and debugging skills.
- Excellent spoken and written English communication skills.
Perks of Freelancing With Turing
- Work in a fully remote environment.
- Opportunity to work on cutting-edge AI projects with leading LLM companies.
Offer Details
- Commitments Required: At least 4 hours per day and minimum 20 hours per week with overlap of 4 hours with PST.
- Engagement Type: Contractor assignment (no medical/paid leave)
- Duration of Contract: 3 months (adjustable based on engagement)
- Location: India, Pakistan, Nigeria, Kenya, Egypt, Ghana, Bangladesh, Turkey, Brazil, Mexico
Evaluation Process
- Technical Interview with live coding challege (60 mins)
Required skills
- Machine Learning
- Python
About Turing
Turing hires remote experts for AI and engineering projects, from software to medicine. The original posting is on their site.
View this job on TuringMore jobs in AI and machine learning
QC Engineer - DataOS Management
This remote contractor role involves overseeing quality control pipelines, evaluating technical project outputs, and ensuring high standards for AI data workflows. Candidates must possess strong software engineering fundamentals and previous experience as a quality control expert or technical project lead.
- QA
- Quality Management
Posted 3 days ago
LLM Trainer - Agent Function call
The LLM Trainer designs multi-turn conversations and simulates tool use to help foundational AI companies train and evaluate their models. The role requires strong technical reasoning skills, experience with APIs and data formats, and three years of professional technical experience.
- Python
- Java
- JavaScript
Posted 4 days ago
Member of Technical Staff, Enterprise AI
This role embeds within enterprise AI systems to diagnose failures, design evaluation datasets, and run experiments. Candidates need a Master's degree in Computer Science or Machine Learning and strong experience in designing ML evaluation frameworks.
- Research Signal Judgment
- ML-Oriented Data Design
- Ops-to-Research Translation
- +1
Posted 4 days ago$300,000 to $700,000/yr
Forward Deployed Engineer
The Forward Deployed Engineer collaborates with leading AI labs and enterprises to build ML pipelines, data intelligence systems, and agentic workflows. Applicants must be strong Python engineers with professional experience building production systems and working with large language models.
- Python
- LLM Systems
- ML Infrastructure
- +1
Posted 4 days ago$300,000 to $650,000/yr