Summary
Boundless is coordinating GPU compute at scale and building toward becoming a leader in AI. As an Applied AI/ML Engineer, you'll ship AI-powered products end-to-end on top of our growing GPU inference fleet, owning everything from serving low-latency inference to standing up reinforcement-learning post-training pipelines.
Responsibilities
- Own AI features and products from prototype through production — model selection, serving, evaluation, and iteration — shipping working software rather than research artifacts
- Deploy and optimize LLM inference across the fleet using vLLM and SGLang. Tune continuous batching, KV-cache management, quantization, speculative decoding, and multi-model routing to maximize throughput and minimize latency and cost per token
- Build and operate reinforcement-learning and post-training pipelines using slime (Megatron-LM + SGLang) and Prime Intellect (prime-rl + the Environments Hub / verifiers). This includes reward and verifier design, rollout orchestration, weight synchronization, and keeping long-running training stable
- Build eval harnesses and benchmarks that measure quality, throughput, and cost together, and use them to drive fast, data-informed iteration
- Partner with Infrastructure on GPU scheduling and fleet utilization, and with Product on what to build next and why
Skills
- 3+ years shipping ML/AI systems to production
- Hands-on experience serving LLM inference with vLLM, SGLang, or TensorRT-LLM
- Experience with RL / post-training methods (GRPO, PPO, DPO, or SFT), or strong adjacent experience and a clear desire to go deep here
- Strong Python and PyTorch
- Working understanding of GPU execution: batching, memory, and basic CUDA concepts
- Comfort operating in ambiguity with a strong bias for action
- Candidates must include a public GitHub profile in their application
- The GitHub profile should demonstrate a minimum of 1 year of activity/history
- Applications that do not include a GitHub profile, or show insufficient activity, will not be considered
- Direct experience with slime, prime-rl, the verifiers library, or Megatron-LM
- Distributed training experience (FSDP, TP/PP/DP parallelism)
- Quantization (FP8/INT8), P/D disaggregation, or speculative decoding
- Experience with verifiable inference or large-scale distributed systems
- Kubernetes and container-based deployment
- Familiarity with GPU fleet orchestration (Ray, SkyPilot, Slurm)
Qualifications
Must Haves
- 3+ years shipping ML/AI systems to production
- Hands-on experience serving LLM inference with vLLM, SGLang, or TensorRT-LLM
- Experience with RL / post-training methods (GRPO, PPO, DPO, or SFT), or strong adjacent experience and a clear desire to go deep here
- Strong Python and PyTorch
- Working understanding of GPU execution: batching, memory, and basic CUDA concepts
- Comfort operating in ambiguity with a strong bias for action
- Candidates must include a public GitHub profile in their application
- The GitHub profile should demonstrate a minimum of 1 year of activity/history
- Applications that do not include a GitHub profile, or show insufficient activity, will not be considered
Nice to Haves
- Direct experience with slime, prime-rl, the verifiers library, or Megatron-LM
- Distributed training experience (FSDP, TP/PP/DP parallelism)
- Quantization (FP8/INT8), P/D disaggregation, or speculative decoding
- Experience with verifiable inference or large-scale distributed systems
- Kubernetes and container-based deployment
- Familiarity with GPU fleet orchestration (Ray, SkyPilot, Slurm)
Benefits
- Health, dental, vision (for U.S. employees; region-adjusted globally)
- Flexible PTO
- Professional development and conference travel budget
- Remote-first with regular off-sites and a high-trust, high-velocity team environment
- Competitive salary + equity allocation