Summary
Weekday AI is a cutting-edge AI research initiative focused on Generative AI and the development of next-generation Large Language Models. They are seeking experienced MLOps Engineers to leverage their expertise in modern machine learning frameworks and large-scale training infrastructure to enhance AI systems and collaborate with research and engineering teams.
Responsibilities
- Partner with research and engineering teams to strengthen AI model capabilities in MLOps, ML infrastructure, and large-scale training systems
- Design challenging, real-world MLOps and machine learning systems tasks that reflect production engineering scenarios
- Develop accurate, well-documented solutions to complex ML infrastructure and training pipeline problems
- Review and evaluate technical tasks and AI-generated solutions, providing clear and actionable written feedback
- Create detailed evaluation rubrics and scoring frameworks for topics including:
- Distributed training architectures
- ML pipeline design
- Infrastructure optimization
- Kernel-level programming
- Performance tuning
- Collaborate with fellow subject matter experts to maintain consistency, quality, and technical accuracy across training datasets
- Contribute domain expertise to improve the reasoning capabilities of advanced AI systems
Skills
- Minimum 2 years of professional experience in MLOps, Machine Learning Infrastructure, or ML Systems Engineering within a recognized technology organization
- Hands-on production experience with JAX and/or PyTorch in large-scale machine learning environments
- Practical experience developing or optimizing custom GPU kernels using Pallas (JAX) or Triton
- Strong understanding of distributed training systems, model optimization, and scalable ML infrastructure
- Demonstrated career growth and increasing technical responsibility
- Availability to work 40 hours per week during standard weekday business hours
- Excellent written communication skills with the ability to clearly explain technical concepts and architectural decisions
- Experience designing and optimizing large-scale ML training pipelines
- Knowledge of distributed computing and GPU performance optimization
- Familiarity with evaluation methodologies for AI models and ML systems
- Experience collaborating with research teams on advanced machine learning projects
- Passion for advancing AI infrastructure and frontier model development
Qualifications
Must Haves
- Minimum 2 years of professional experience in MLOps, Machine Learning Infrastructure, or ML Systems Engineering within a recognized technology organization
- Hands-on production experience with JAX and/or PyTorch in large-scale machine learning environments
- Practical experience developing or optimizing custom GPU kernels using Pallas (JAX) or Triton
- Strong understanding of distributed training systems, model optimization, and scalable ML infrastructure
- Demonstrated career growth and increasing technical responsibility
- Availability to work 40 hours per week during standard weekday business hours
- Excellent written communication skills with the ability to clearly explain technical concepts and architectural decisions
Nice to Haves
- Experience designing and optimizing large-scale ML training pipelines
- Knowledge of distributed computing and GPU performance optimization
- Familiarity with evaluation methodologies for AI models and ML systems
- Experience collaborating with research teams on advanced machine learning projects
- Passion for advancing AI infrastructure and frontier model development
Benefits
- Fully remote engagement
- Independent contractor basis
- 40-hour-per-week remote engagement
- Payments are processed weekly through Stripe or Wise based on approved work completed