Summary
SemiAnalysis is an independent research and analysis firm specializing in the semiconductor and AI industries. The Member of Technical Staff will develop training and inference benchmarks, model AI compute systems at large scale, author technical research reports, and contribute to newsletter articles and industry partnerships.
Responsibilities
- Conduct training & inference performance benchmarks across various AI hardware (e.g. NVIDIA H100, AMD Mi300X, Google TPUs, AWS Trainium2) using frameworks such as PyTorch, JAX, vLLM, SGLang, etc
- Author detailed technical research reports analyzing benchmark results, hardware performance, scalability, & efficiency
- Develop comprehensive system modelling using Python & NCCL for existing & future AI compute clusters, scaling from single-GPU setups to O(100k) GPU clusters
- Establish and maintain strategic partnerships & collaborations with over 50 leading neocloud providers & AI chip manufacturers, including AMD, NVIDIA, and other industry stakeholders
- Stay current on emerging trends & technologies by attending major industry & academic conferences such as NeurIPS, MLSys, NVIDIA GTC, AMD’s Advancing AI, etc
Skills
- Proactive, self-motivated, and capable of working independently in a global team
- Demonstrated experience in ML frameworks such as PyTorch or JAX through professional experience, personal projects, or personal Substack blogs
- Solid understanding of at least 1 of the following: transformer architecture, LLM parallelism strategies, and/or CUDA parallel programming
- Strong research skills and the ability to synthesize information from various sources to draw insights
- Undergraduate degree in Computer Science, Engineering or other relevant technical field is not required
Qualifications
Must Haves
- Proactive, self-motivated, and capable of working independently in a global team
- Demonstrated experience in ML frameworks such as PyTorch or JAX through professional experience, personal projects, or personal Substack blogs
- Solid understanding of at least 1 of the following: transformer architecture, LLM parallelism strategies, and/or CUDA parallel programming
- Strong research skills and the ability to synthesize information from various sources to draw insights
- Undergraduate degree in Computer Science, Engineering or other relevant technical field is not required
Benefits
- In-office/Remote work setting
- Paid 2-day coding challenge as part of the interview process
- Opportunity for direct authorship recognition through published work
- Develop deep expertise in AI infrastructure, including training and inference optimization across leading hardware platforms and large-scale compute environments
- Gain hands-on experience modelling and scaling distributed systems, from single-node setups to hyperscale GPU clusters
- Build strong technical writing and research capabilities, with opportunities for direct authorship and industry recognition through published work
- Gain exposure to leading AI hardware vendors, neocloud providers, and industry stakeholders through collaborations and partnerships
- Participation in major global conferences and exposure to cutting-edge research
- Opportunities to lead projects and drive independent initiatives
- Build a strong personal brand within the AI and infrastructure community through thought leadership and published analysis