Mercor logo
Mercor
Posted 52 days agoVerified live 1d ago

ML Engineer - Model Evaluation Expert

Brief overview

Remote
MastersOr in progress
$60–$90/hrStated range
1+ yrsMinimum
Machine Learning Model TrainingMachine Learning Model EvaluationEnd-to-End ExperimentationPythonGitLarge Language Model EvaluationReinforcement LearningBenchmark and Task Authoring

About the company

Mercor is an AI-based hiring platform that improves the recruitment process by matching talent with opportunities tailored to their skills.

Visa sponsorship history

2 years sponsoring, last filed FY2026

Data powered by U.S. Department of Labor. This does not guarantee sponsorship for this specific role.
$161,637median wage / yr
LCA Certified ApplicationsAn early filing step, not a visa approval: it signals intent, not confirmed sponsorship.
20256
202611
Top sponsored roles
Software EngineerSenior Strategic Project LeadStrategic Project LeadSenior Product ManagerMachine Learning Engineer

Job description

Summary

Mercor connects elite creative and technical talent with leading AI research labs. The company is seeking a Machine Learning Engineer to design research-based tasks, run and analyze training experiments, evaluate frontier models, and collaborate with researchers on rigorous and fair task development.

Responsibilities

  • Design tasks by transforming real ML research ideas into well-defined, multi-step tasks
  • Run experiments by implementing changes, executing training experiments, and analyzing results to define correct solutions
  • Explore reinforcement learning concepts such as reward functions and training behavior in task development
  • Evaluate frontier models' performance on tasks and identify areas of improvement
  • Collaborate with researchers to ensure tasks are consistent, rigorous, and fair

Skills

  • • MSc or PhD in machine learning, computer science, or a related STEM field
  • • 1+ years in a research or research-engineering role
  • • Experience in training and evaluating ML models and conducting end-to-end experiments
  • • Proficiency in Python and Git
  • • Strong understanding of large language models and their evaluation
  • • Basic knowledge of reinforcement learning
  • • Experience in AI training, model evaluation, or benchmark/task authoring

Qualifications

Must Haves

  • • MSc or PhD in machine learning, computer science, or a related STEM field
  • • 1+ years in a research or research-engineering role
  • • Experience in training and evaluating ML models and conducting end-to-end experiments
  • • Proficiency in Python and Git
  • • Strong understanding of large language models and their evaluation

Nice to Haves

  • • Basic knowledge of reinforcement learning
  • • Experience in AI training, model evaluation, or benchmark/task authoring

Benefits

  • Remote work
  • Paid weekly

More jobs like this