Summary
Pinterest is a platform that helps people discover creative ideas and plan meaningful experiences, with Pinterest Labs focused on applied machine learning research and development. The Machine Learning Engineer II, Visual AI will advance vision-centric language models by prototyping architectures, developing evaluation benchmarks, collecting visual training data, and contributing to research publications for production-scale generative models.
Responsibilities
- Prototype new model architectures for Pinterest VLMs. We’re looking for hands-on experience working with finetuning open-source LLM models and improve their visual perception and tool using capabilities
- Develop new evaluation benchmarks that tailors to vision-centric capabilities such as fashion style recommendations
- Read research papers, participate in group discussions, and help brainstorm our overall visual generative strategy at the company
- Help with collection of relevant visual training data for Pinterest Canvas, particularly to conduct RLHF, targeted fine-tuning, etc
- Publish and publicize your work via conferences, paper submissions, blog posts, etc
Skills
- • Research engineers and scientists who have experience working with generative computer vision models, preferably various forms of visual encoders and LLMs
- • 2+ years of industry computer vision experience
- • M.S. or PhD in Machine Learning, Computer Science, or related areas
- • This role will need to be in the office for in-person collaboration 1-2 times/quarter and therefore can be situated anywhere in the country
- •US based applicants only
- • Publications at top ML conferences
- • Experience using Cursor, Copilot, Codex, or similar AI coding assistants for development, debugging, testing, and refactoring
- • Familiarity with LLM-powered productivity tools for documentation search, experiment analysis, SQL/data exploration, and engineering workflow acceleration
Qualifications
Must Haves
- • Research engineers and scientists who have experience working with generative computer vision models, preferably various forms of visual encoders and LLMs
- • 2+ years of industry computer vision experience
- • M.S. or PhD in Machine Learning, Computer Science, or related areas
- • This role will need to be in the office for in-person collaboration 1-2 times/quarter and therefore can be situated anywhere in the country
- •US based applicants only
Nice to Haves
- • Publications at top ML conferences
- • Experience using Cursor, Copilot, Codex, or similar AI coding assistants for development, debugging, testing, and refactoring
- • Familiarity with LLM-powered productivity tools for documentation search, experiment analysis, SQL/data exploration, and engineering workflow acceleration
Benefits
- The position is also eligible for equity.
- This role will need to be in the office for in-person collaboration 1-2 times/quarter and therefore can be situated anywhere in the country.