Apple logo
Apple
Posted 170 days agoVerified live 2d ago

Machine Learning Engineer - Visual Agents - Special Projects

Brief overview

Cupertino, California, United StatesIn-person
UndergradOr in progress
$127k–$221k/yrStated range
2+ yrsMinimum
6,752 H-1B approvalsDept. of Labor
3,817 green cardsCertified filings
Proficiency in PythonDeveloping automated evaluation pipelinesLLM-as-judge frameworksHuman evaluation protocolsDomain-specific benchmarksStatistical analysis for model evaluationAgentic system designTool use, grounding, and perceive-act loopsVideo understandingTemporal reasoningActivity recognitionLarge-scale multimodal data handlingAnnotation pipelinesTraining and evaluation of foundation modelsCross-functional collaborationCommunication skills

About the company

Apple is a technology company that designs, manufactures, and markets consumer electronics, personal computers, and software.

Visa sponsorship history

4 years sponsoring, last filed FY2026

Data powered by U.S. Department of Labor. This does not guarantee sponsorship for this specific role.
6,752H-1B approved
98%approval rate
2,551new H-1B hires
3,817PERM certified
$178,006median wage / yr
H-1B Petition ApprovalsVisas USCIS actually granted: the strongest sign the company sponsors.
20231,940
20241,811
20252,018
2026983
LCA Certified ApplicationsAn early filing step, not a visa approval: it signals intent, not confirmed sponsorship.
2023884
20241,000
20251,480
20261,175
Green Card (PERM) FilingsCertified green card filings: a long-term commitment to international hires.
20231,199
20241,386
20251,224
20268
Top sponsored roles
Software Engineering ApplicationsSoftware Development EngineeringSoftware Development EngineerSoftware Engineering SystemsSoftware Development Engineer - Applications
Sponsored employees from
IndiaChinaCanadaTaiwanUnited Kingdom

Job description

Summary

Apple is where individual imaginations gather together, committing to the values that lead to great work. The Special Projects team at Apple is developing novel experiences powered by state-of-the-art agentic vision-language models, and they are seeking a Machine Learning Engineer to help build, fine-tune, and evaluate these systems.

Responsibilities

  • Build and evaluate vision-language agents that perceive real-world scenes and incorporate that context into conversational models
  • Curate, annotate, and build multimodal datasets to support model training and evaluation
  • Develop automated evaluation pipelines including LLM-as-judge frameworks, human evaluation protocols, and domain-specific benchmarks
  • Fine-tune Large Language Models (LLMs) and Visual-Language Models (VLMs) to improve performance for specific use cases
  • Work closely with other ML Researchers to define evaluation criteria and methodology to systematically evaluate foundation models
  • Design controlled experiments to measure model capabilities, identify failure modes, and drive iterative model improvements
  • Conduct robust statistical analysis to identify model deficiencies and failure modes and performance gaps

Skills

  • BA or Master's degree in Computer Science or Machine Learning
  • 2+ years of hands-on experience building and evaluating generative AI or multimodal models
  • Experience working with vision-language models or multimodal systems
  • Proficiency in Python and ML frameworks (Pytorch or Tensorflow)
  • PhD in Computer Science, Machine Learning, Statistics, or other STEM field
  • Prior industry internship or research experience applying ML to product use cases
  • Experience with video understanding, temporal reasoning, or activity recognition
  • Familiarity with agentic system design including tool use, grounding, or perceive-act loops
  • Experience building or working with large-scale multimodal data and annotation pipelines
  • Proficiency in training, fine-tuning, and evaluation of foundation models and frameworks
  • Publications or technical presentations in Machine Learning journals or conferences
  • Excellent communication skills and cross functional collaboration

Qualifications

Must Haves

  • BA or Master's degree in Computer Science or Machine Learning
  • 2+ years of hands-on experience building and evaluating generative AI or multimodal models
  • Experience working with vision-language models or multimodal systems
  • Proficiency in Python and ML frameworks (Pytorch or Tensorflow)

Nice to Haves

  • PhD in Computer Science, Machine Learning, Statistics, or other STEM field
  • Prior industry internship or research experience applying ML to product use cases
  • Experience with video understanding, temporal reasoning, or activity recognition
  • Familiarity with agentic system design including tool use, grounding, or perceive-act loops
  • Experience building or working with large-scale multimodal data and annotation pipelines
  • Proficiency in training, fine-tuning, and evaluation of foundation models and frameworks
  • Publications or technical presentations in Machine Learning journals or conferences
  • Excellent communication skills and cross functional collaboration

Benefits

  • Comprehensive medical and dental coverage
  • Retirement benefits
  • A range of discounted products and free services
  • Reimbursement for certain educational expenses — including tuition
  • Discretionary bonuses or commission payments
  • Relocation

More jobs like this