Zyphra logo
Zyphra
Verified live 12h ago

Research Engineer, Language Model Pre-Training

Brief overview

San FranciscoIn-person
MastersOr in progress
4 H-1B approvalsDept. of Labor
Machine learningModel trainingDistributed computingModel parallelizationData parallelismDistributed optimizersExperimental methodologyLarge-scale data processingPyTorchPythonCodebase navigation

About the company

Zyphra is a full stack artificial intelligence company based in Palo Alto, California.

Visa sponsorship history

2 years sponsoring, last filed FY2026

Data powered by U.S. Department of Labor. This does not guarantee sponsorship for this specific role.
4H-1B approved
100%approval rate
$187,574median wage / yr
H-1B Petition ApprovalsVisas USCIS actually granted: the strongest sign the company sponsors.
20251
20263
LCA Certified ApplicationsAn early filing step, not a visa approval: it signals intent, not confirmed sponsorship.
20263
Top sponsored roles
Data Infrastructure Engineer (Member of Technical Staff)Member of Technical Staff (LLM Performance & Optimization)Software Developer (Member of Technical Staff)

Job description

Zyphra is an artificial intelligence company based in San Francisco, California.

The Role:

As a Research Engineer, Language Model Pre-training, you'll shape our language model roadmap through end-to-end pretraining development. You will work extremely closely with our pretraining team, who will integrate your insights into our next-generation models.

You'll work across:

  • Large-scale training runs and model parallelization

  • Performance optimization of our pretraining stack

  • Dataset collection, processing, and evaluation

  • Architecture and methodology research, including optimizer ablations

Requirements:

  • Strong engineering aptitude for rapidly implementing reliable and robust systems

  • Can rapidly learn new fields and are excited to implement new ideas

  • Excellent communication and collaboration skills, and can work effectively on both research and engineering implementation at scale

Ideal Skillset:

  • Deep expertise and intuition for solving machine learning problems and training models

  • Experience with training on large-scale (multi-node) GPU clusters

  • Deep understanding of model training pipelines – including model/data parallelism, distributed optimizers, etc.

  • Strong grasp of proper experimental methodology for running rigorous ablations and other hypothesis testing

  • Understanding of large-scale, highly parallel data processing pipelines

  • High proficiency with PyTorch and Python.

  • Strong ability to dive into large pre-existing codebases and rapidly get up to speed

  • Published machine learning research in well-respected venues is a plus

  • Postgraduate degree in a scientific subject (Computer Science, EE/EECS, Math, Physics)

Why Work at Zyphra:

  • Our research methodology is to make grounded, methodical steps toward ambitious goals. Both deep research and engineering excellence are equally valued

  • We strongly value new and crazy ideas and are very willing to bet big on new ideas

  • We move as quickly as we can; we aim to minimize the bar to impact as low as possible

  • We all enjoy what we do and love discussing AI

Benefits and Perks:

  • Comprehensive medical, dental, vision, and FSA plans

  • Competitive compensation and 401(k)

  • Relocation and immigration support on a case-by-case basis

  • On-site meals prepared by a dedicated culinary team; Thursday Happy Hours

  • In-person team in San Francisco, CA, with a collaborative, high-energy environment

We’re building something new, and we want someone with strong research taste and intuition. If that is you, Apply Today!

More jobs like this