Summary
Lightning AI builds an end-to-end platform for developing, training, and deploying AI systems, including tools for experimentation, large-scale compute, and production inference. The Research Engineer will focus on post-training models and the systems supporting them, while developing software, tooling, infrastructure, and workflows that improve AI research, development, evaluation, and deployment.
Responsibilities
- Develop and post-train models, while building and improving the systems and workflows needed to run, evaluate, debug, and scale training workloads
- Build software, tooling, and platform capabilities that improve how researchers, developers, and customers develop, train, and deploy AI systems
- Contribute to Lightning’s open-source projects by building new features, improving existing functionality, and collaborating with the broader developer community
- Work across deep learning systems, developer tooling, backend services, and platform infrastructure to solve a wide variety of engineering challenges
- Collaborate directly with customers to understand real-world AI workloads, investigate technical challenges, and translate those learnings into reusable product and platform improvements
- Prototype new ideas, evaluate approaches, and turn successful experiments into production-quality software
- Partner closely with research, product, and infrastructure engineering teams to improve developer experience, AI workflows, and platform capabilities
- Debug complex technical problems spanning machine learning, distributed systems, backend software, and developer tooling
- Learn new technologies quickly and contribute wherever your skills can have the greatest impact as team priorities evolve
Skills
- Experience building, training, evaluating, or experimenting with deep learning models
- Hands-on experience with deep learning frameworks such as PyTorch
- Strong software engineering fundamentals building software and debugging and problem-solving skills, with the ability to investigate unfamiliar technical challenges
- Curiosity, initiative, and a demonstrated ability to quickly learn new technologies and technical domains
- Excellent communication and collaboration skills, including the ability to work effectively across research, product, infrastructure, and customer-facing engagements
- Comfortable working in fast-moving, ambiguous environments where priorities evolve over time
- Bachelor's degree in Computer Science, Engineering, or a related technical field, or equivalent practical experience
- Experience with model training at scale, including distributed training, performance optimization, training stability, and/or large-scale experimentation
- Experience with transformer-based language models or modern generative AI systems
- Experience with distributed systems, cloud infrastructure, or large-scale machine learning workloads
- Familiarity with technologies such as CUDA, Hugging Face, DeepSpeed, FSDP, Triton, vLLM, SGLang, NVIDIA Molt, or related AI infrastructure tooling
- Experience contributing to open-source software or conducting research through academia, industry, or meaningful independent projects
- Startup experience or experience working on highly cross-functional engineering teams
- Master's degree or higher in Computer Science, Machine Learning, AI, or a related field
Qualifications
Must Haves
- Experience building, training, evaluating, or experimenting with deep learning models
- Hands-on experience with deep learning frameworks such as PyTorch
- Strong software engineering fundamentals building software and debugging and problem-solving skills, with the ability to investigate unfamiliar technical challenges
- Curiosity, initiative, and a demonstrated ability to quickly learn new technologies and technical domains
- Excellent communication and collaboration skills, including the ability to work effectively across research, product, infrastructure, and customer-facing engagements
- Comfortable working in fast-moving, ambiguous environments where priorities evolve over time
- Bachelor's degree in Computer Science, Engineering, or a related technical field, or equivalent practical experience
Nice to Haves
- Experience with model training at scale, including distributed training, performance optimization, training stability, and/or large-scale experimentation
- Experience with transformer-based language models or modern generative AI systems
- Experience with distributed systems, cloud infrastructure, or large-scale machine learning workloads
- Familiarity with technologies such as CUDA, Hugging Face, DeepSpeed, FSDP, Triton, vLLM, SGLang, NVIDIA Molt, or related AI infrastructure tooling
- Experience contributing to open-source software or conducting research through academia, industry, or meaningful independent projects
- Startup experience or experience working on highly cross-functional engineering teams
- Master's degree or higher in Computer Science, Machine Learning, AI, or a related field
Benefits
- Comprehensive Health Coverage: Medical, dental, and vision coverage for employees and eligible dependents.
- Meaningful Equity: RSUs that give employees a stake in the company's long-term success.
- Retirement Savings: 401(k) matching (U.S.) and pension contributions (U.K.).
- Flexible Time Off: Unlimited PTO, company holidays, and floating holidays to support work-life balance.
- Company-Wide Winter Break: Two weeks of company closure each winter to disconnect and recharge.
- Paid Parental & Family Leave: Paid leave to support you and your family through life's important moments.
- Professional Development: Annual learning and development allowance to support your professional growth.
- Wellness Benefits: Wellness and work-from-home stipends to support your physical and mental well-being.
- Sabbatical Program: Four weeks of paid sabbatical leave after four years of service.
- Flexible Work: Flexible schedules and a hybrid work model for our office-based teams.
- In-Office Meals: Complimentary meals at our office hubs.