Weekday AI (YC W21) logo
Weekday AI (YC W21)
Posted 66 days agoVerified live 12h ago

Machine Learning Engineer - Model Evaluation & Experimentation

Brief overview

Remote
MastersOr in progress
$60–$90/hrStated range
1+ yrsMinimum
Machine LearningLarge Language Models (LLMs)PythonGitReinforcement LearningMachine Learning ExperimentationBenchmark DevelopmentAI EvaluationModel TrainingModel OptimizationResearch Engineering

About the company

Weekday AI (YC W21) logo
Weekday AI (YC W21)jobs.weekday.works

We are a YC-backed recruitment startup. Find select jobs posted by premium YC as well as VC backed startups here. Hand-curated by Weekday team.

Job description

Summary

Weekday AI is a pioneering company focused on building next-generation evaluation benchmarks for frontier AI models. They are seeking experienced Machine Learning Engineers and Researchers to design and implement sophisticated machine learning challenges that establish high-quality evaluation benchmarks for advanced AI systems.

Responsibilities

  • Design realistic machine learning benchmark tasks based on research workflows, including model implementation, experimentation, training, evaluation, and performance analysis
  • Translate open-ended research concepts into structured, reproducible evaluation tasks with clearly defined success criteria
  • Implement machine learning solutions using Python, execute experiments, and produce reference implementations that demonstrate correct methodology and expected outcomes
  • Develop benchmark tasks involving reinforcement learning concepts such as reward functions, policy optimization, training dynamics, and model behavior where applicable
  • Evaluate AI-generated solutions by identifying implementation errors, experimental flaws, incorrect reasoning, and unsupported conclusions
  • Collaborate with AI researchers and fellow subject matter experts to continuously improve benchmark quality, technical rigor, and evaluation consistency

Skills

  • Master's degree, PhD, or equivalent practical experience in Machine Learning, Computer Science, Artificial Intelligence, Data Science, or another quantitative STEM discipline
  • Minimum 1 year of professional experience in machine learning research, research engineering, applied AI, or another research-intensive technical role
  • Strong hands-on experience designing, training, evaluating, and optimizing machine learning models through complete experimental workflows
  • Practical experience conducting machine learning experiments, including experiment setup, hyperparameter tuning, execution, validation, and analysis
  • Strong understanding of modern Large Language Models (LLMs), their capabilities, limitations, and evaluation methodologies
  • Proficiency in Python and Git, with experience working in both script-based and notebook-based development environments
  • Familiarity with reinforcement learning concepts—including reward functions, policy optimization, and training behavior
  • Excellent analytical thinking, creativity, attention to detail, and the ability to solve complex, open-ended technical problems independently
  • Strong written communication skills for documenting experimental methodologies and technical findings
  • Ability to commit approximately 35 hours per week on a consistent basis
  • Experience developing or evaluating large language models, foundation models, or generative AI systems
  • Background in reinforcement learning, deep learning, distributed training, or model optimization
  • Familiarity with benchmark design, AI safety evaluations, or research-quality experimentation
  • Experience contributing to research publications, open-source machine learning projects, or advanced AI systems

Qualifications

Must Haves

  • Master's degree, PhD, or equivalent practical experience in Machine Learning, Computer Science, Artificial Intelligence, Data Science, or another quantitative STEM discipline
  • Minimum 1 year of professional experience in machine learning research, research engineering, applied AI, or another research-intensive technical role
  • Strong hands-on experience designing, training, evaluating, and optimizing machine learning models through complete experimental workflows
  • Practical experience conducting machine learning experiments, including experiment setup, hyperparameter tuning, execution, validation, and analysis
  • Strong understanding of modern Large Language Models (LLMs), their capabilities, limitations, and evaluation methodologies
  • Proficiency in Python and Git, with experience working in both script-based and notebook-based development environments
  • Familiarity with reinforcement learning concepts—including reward functions, policy optimization, and training behavior
  • Excellent analytical thinking, creativity, attention to detail, and the ability to solve complex, open-ended technical problems independently
  • Strong written communication skills for documenting experimental methodologies and technical findings
  • Ability to commit approximately 35 hours per week on a consistent basis

Nice to Haves

  • Experience developing or evaluating large language models, foundation models, or generative AI systems
  • Background in reinforcement learning, deep learning, distributed training, or model optimization
  • Familiarity with benchmark design, AI safety evaluations, or research-quality experimentation
  • Experience contributing to research publications, open-source machine learning projects, or advanced AI systems

Benefits

  • Fully remote with flexible working hours
  • Payments are issued weekly based on approved work completed
  • Independent contractor engagement

More jobs like this