Summary
Apple is where individual imaginations gather together, committing to the values that lead to great work. The Special Projects team at Apple is developing novel experiences powered by state-of-the-art agentic vision-language models, and they are seeking a Machine Learning Engineer to help build, fine-tune, and evaluate these systems.
Responsibilities
- Build and evaluate vision-language agents that perceive real-world scenes and incorporate that context into conversational models
- Curate, annotate, and build multimodal datasets to support model training and evaluation
- Develop automated evaluation pipelines including LLM-as-judge frameworks, human evaluation protocols, and domain-specific benchmarks
- Fine-tune Large Language Models (LLMs) and Visual-Language Models (VLMs) to improve performance for specific use cases
- Work closely with other ML Researchers to define evaluation criteria and methodology to systematically evaluate foundation models
- Design controlled experiments to measure model capabilities, identify failure modes, and drive iterative model improvements
- Conduct robust statistical analysis to identify model deficiencies and failure modes and performance gaps
Skills
- BA or Master's degree in Computer Science or Machine Learning
- 2+ years of hands-on experience building and evaluating generative AI or multimodal models
- Experience working with vision-language models or multimodal systems
- Proficiency in Python and ML frameworks (Pytorch or Tensorflow)
- PhD in Computer Science, Machine Learning, Statistics, or other STEM field
- Prior industry internship or research experience applying ML to product use cases
- Experience with video understanding, temporal reasoning, or activity recognition
- Familiarity with agentic system design including tool use, grounding, or perceive-act loops
- Experience building or working with large-scale multimodal data and annotation pipelines
- Proficiency in training, fine-tuning, and evaluation of foundation models and frameworks
- Publications or technical presentations in Machine Learning journals or conferences
- Excellent communication skills and cross functional collaboration
Qualifications
Must Haves
- BA or Master's degree in Computer Science or Machine Learning
- 2+ years of hands-on experience building and evaluating generative AI or multimodal models
- Experience working with vision-language models or multimodal systems
- Proficiency in Python and ML frameworks (Pytorch or Tensorflow)
Nice to Haves
- PhD in Computer Science, Machine Learning, Statistics, or other STEM field
- Prior industry internship or research experience applying ML to product use cases
- Experience with video understanding, temporal reasoning, or activity recognition
- Familiarity with agentic system design including tool use, grounding, or perceive-act loops
- Experience building or working with large-scale multimodal data and annotation pipelines
- Proficiency in training, fine-tuning, and evaluation of foundation models and frameworks
- Publications or technical presentations in Machine Learning journals or conferences
- Excellent communication skills and cross functional collaboration
Benefits
- Comprehensive medical and dental coverage
- Retirement benefits
- A range of discounted products and free services
- Reimbursement for certain educational expenses — including tuition
- Discretionary bonuses or commission payments
- Relocation