Summary
10a Labs provides AI safety and threat-intelligence services for organizations developing and deploying advanced AI systems. The Machine Learning Engineer will design, build, and evaluate machine learning systems for AI safety and model evaluation applications, including experiments, evaluation pipelines, model training, and scalable infrastructure.
Responsibilities
- Design and run ML experiments to evaluate the capabilities, behavior, robustness, and limitations of advanced AI systems
- Develop and evaluate models acrossreinforcement learning, NLP/LLMs, computer vision, and multimodal ML
- Build evaluation pipelines, benchmarks, datasets, and metrics for frontier AI systems
- Train, fine-tune, and evaluate models for safety, security, and other high-impact applications
- Develop reliable tooling and infrastructure to run ML experiments and evaluations at scale
- Analyze results, identify model failure modes, and translate findings into new experiments and technical approaches
Skills
- 3–5+ years of experience in machine learning, research engineering, or a related technical field
- Strong Python skills and experience with ML frameworks such as PyTorch or JAX
- Hands-on experience training, fine-tuning, or evaluating modern ML models
- Strong understanding of experimental design, model evaluation, and quantitative analysis
- Familiarity with agentic AI fundamentals, including common harnesses, Model Context Protocol, agent benchmarks, and security risks to AI agents
- Experience in one or more of the following: reinforcement learning, NLP/LLMs, computer vision, or multimodal ML
- Strong software engineering fundamentals and the ability to work independently on ambiguous technical problems
- Experience with RLHF/RLAIF, reward modeling, policy optimization, or other model post-training techniques
- Experience evaluating frontier language or multimodal models
- Experience with adversarial evaluations, robustness testing, or AI safety
- Experience with distributed training, cloud ML infrastructure, or large-scale ML systems
Qualifications
Must Haves
- 3–5+ years of experience in machine learning, research engineering, or a related technical field
- Strong Python skills and experience with ML frameworks such as PyTorch or JAX
- Hands-on experience training, fine-tuning, or evaluating modern ML models
- Strong understanding of experimental design, model evaluation, and quantitative analysis
- Familiarity with agentic AI fundamentals, including common harnesses, Model Context Protocol, agent benchmarks, and security risks to AI agents
- Experience in one or more of the following: reinforcement learning, NLP/LLMs, computer vision, or multimodal ML
- Strong software engineering fundamentals and the ability to work independently on ambiguous technical problems
Nice to Haves
- Experience with RLHF/RLAIF, reward modeling, policy optimization, or other model post-training techniques
- Experience evaluating frontier language or multimodal models
- Experience with adversarial evaluations, robustness testing, or AI safety
- Experience with distributed training, cloud ML infrastructure, or large-scale ML systems
Benefits
- Performance-based annual bonus
- Support for conferences, continuing education, or leadership training
- Fully remote, U.S.-based
- Comprehensive health, dental, and vision coverage
- Generous PTO and paid holiday schedule