Summary
Intuitive is a global leader in robotic-assisted surgery and minimally invasive care. They are seeking a Computer Vision Engineering Intern to join their R&D team, focusing on advancing research and development in cutting-edge computer vision technologies for robotic endoscopic video applications.
Responsibilities
- Explore and experiment with state-of-the-art computer vision models, including foundation models and generative diffusion models, with applications to video understanding, multi-modal data, and visual feature extraction
- Prototype novel algorithms and evaluate performance using public and proprietary datasets
- Conduct literature surveys and summarize key findings in reports and presentations
Skills
- Solid understanding and hands-on experience in computer vision, deep learning, and video analysis
- Knowledge in one or more areas: large vision-language models, generative diffusion models, feature detection, scene understanding, video classification, or multimodal learning
- Proficiency in programming with Python or C++, with experience in relevant frameworks (e.g., PyTorch, OpenCV, DINO/CLIP, HuggingFace Transformers, etc.)
- Strong research and communication skills, with the ability to summarize findings and present them clearly
- Passionate about pushing the boundaries of AI technologies to solve complex, real-world problems
- Passion for developing technologies to improve the lives of patients and physicians
- Self-driven, able to work independently and deliver rapid prototyping and experimentation
- Ability to perform fast prototyping iterations; thinking outside the box to solve practical problems
- Must be currently enrolled in and returning to an accredited degree-seeking academic program in the Spring of 2027
- Must be available to work full-time (approximately 40 hours per week) during a 10-12 week period starting August or September 2026
- Current enrollment in a Computer Science, Robotics, Mechanical Engineering, Electrical Engineering, Biomedical Engineering or related degree-seeking program at the Doctorate level. Master's level students would also be considered based on specific relevant experience
Qualifications
Must Haves
- Solid understanding and hands-on experience in computer vision, deep learning, and video analysis
- Knowledge in one or more areas: large vision-language models, generative diffusion models, feature detection, scene understanding, video classification, or multimodal learning
- Proficiency in programming with Python or C++, with experience in relevant frameworks (e.g., PyTorch, OpenCV, DINO/CLIP, HuggingFace Transformers, etc.)
- Strong research and communication skills, with the ability to summarize findings and present them clearly
- Passionate about pushing the boundaries of AI technologies to solve complex, real-world problems
- Passion for developing technologies to improve the lives of patients and physicians
- Self-driven, able to work independently and deliver rapid prototyping and experimentation
- Ability to perform fast prototyping iterations; thinking outside the box to solve practical problems
- Must be currently enrolled in and returning to an accredited degree-seeking academic program in the Spring of 2027
- Must be available to work full-time (approximately 40 hours per week) during a 10-12 week period starting August or September 2026
- Current enrollment in a Computer Science, Robotics, Mechanical Engineering, Electrical Engineering, Biomedical Engineering or related degree-seeking program at the Doctorate level. Master's level students would also be considered based on specific relevant experience
Benefits
- Benefits
- A housing allowance