Weekday AI (YC W21) logo
Weekday AI (YC W21)
Posted 67 days agoVerified live 2d ago

MLOps Engineer (JAX, PyTorch, Pallas/Triton)

Brief overview

Remote
$70–$110/hrStated range
2+ yrsMinimum
MLOpsMachine Learning InfrastructureML Systems EngineeringJAXPyTorchGPU kernel programmingPallasTritonDistributed training systemsModel optimizationScalable ML infrastructure

About the company

Weekday AI (YC W21) logo
Weekday AI (YC W21)jobs.weekday.works

We are a YC-backed recruitment startup. Find select jobs posted by premium YC as well as VC backed startups here. Hand-curated by Weekday team.

Job description

Summary

Weekday AI is a cutting-edge AI research initiative focused on Generative AI and the development of next-generation Large Language Models. They are seeking experienced MLOps Engineers to leverage their expertise in modern machine learning frameworks and large-scale training infrastructure to enhance AI systems and collaborate with research and engineering teams.

Responsibilities

  • Partner with research and engineering teams to strengthen AI model capabilities in MLOps, ML infrastructure, and large-scale training systems
  • Design challenging, real-world MLOps and machine learning systems tasks that reflect production engineering scenarios
  • Develop accurate, well-documented solutions to complex ML infrastructure and training pipeline problems
  • Review and evaluate technical tasks and AI-generated solutions, providing clear and actionable written feedback
  • Create detailed evaluation rubrics and scoring frameworks for topics including:
  • Distributed training architectures
  • ML pipeline design
  • Infrastructure optimization
  • Kernel-level programming
  • Performance tuning
  • Collaborate with fellow subject matter experts to maintain consistency, quality, and technical accuracy across training datasets
  • Contribute domain expertise to improve the reasoning capabilities of advanced AI systems

Skills

  • Minimum 2 years of professional experience in MLOps, Machine Learning Infrastructure, or ML Systems Engineering within a recognized technology organization
  • Hands-on production experience with JAX and/or PyTorch in large-scale machine learning environments
  • Practical experience developing or optimizing custom GPU kernels using Pallas (JAX) or Triton
  • Strong understanding of distributed training systems, model optimization, and scalable ML infrastructure
  • Demonstrated career growth and increasing technical responsibility
  • Availability to work 40 hours per week during standard weekday business hours
  • Excellent written communication skills with the ability to clearly explain technical concepts and architectural decisions
  • Experience designing and optimizing large-scale ML training pipelines
  • Knowledge of distributed computing and GPU performance optimization
  • Familiarity with evaluation methodologies for AI models and ML systems
  • Experience collaborating with research teams on advanced machine learning projects
  • Passion for advancing AI infrastructure and frontier model development

Qualifications

Must Haves

  • Minimum 2 years of professional experience in MLOps, Machine Learning Infrastructure, or ML Systems Engineering within a recognized technology organization
  • Hands-on production experience with JAX and/or PyTorch in large-scale machine learning environments
  • Practical experience developing or optimizing custom GPU kernels using Pallas (JAX) or Triton
  • Strong understanding of distributed training systems, model optimization, and scalable ML infrastructure
  • Demonstrated career growth and increasing technical responsibility
  • Availability to work 40 hours per week during standard weekday business hours
  • Excellent written communication skills with the ability to clearly explain technical concepts and architectural decisions

Nice to Haves

  • Experience designing and optimizing large-scale ML training pipelines
  • Knowledge of distributed computing and GPU performance optimization
  • Familiarity with evaluation methodologies for AI models and ML systems
  • Experience collaborating with research teams on advanced machine learning projects
  • Passion for advancing AI infrastructure and frontier model development

Benefits

  • Fully remote engagement
  • Independent contractor basis
  • 40-hour-per-week remote engagement
  • Payments are processed weekly through Stripe or Wise based on approved work completed

More jobs like this