Boundless logo
Boundless
Posted 62 days agoVerified live 1d ago

Applied AI/ML Engineer

Brief overview

Remote
3+ yrsMinimum
Machine learning systems deploymentLLM inference servingvLLMSGLangTensorRT-LLMReinforcement learning methodsGRPOPPOSFTPythonPyTorchGPU execution conceptsCUDAQuantization (FP8/INT8)Distributed trainingKubernetesContainer deployment

About the company

Boundless is a universal zero-knowledge protocol that lets anyone access abundant verifiable compute, regardless of the blockchain they are using.

Job description

Summary

Boundless is coordinating GPU compute at scale and building toward becoming a leader in AI. As an Applied AI/ML Engineer, you'll ship AI-powered products end-to-end on top of our growing GPU inference fleet, owning everything from serving low-latency inference to standing up reinforcement-learning post-training pipelines.

Responsibilities

  • Own AI features and products from prototype through production — model selection, serving, evaluation, and iteration — shipping working software rather than research artifacts
  • Deploy and optimize LLM inference across the fleet using vLLM and SGLang. Tune continuous batching, KV-cache management, quantization, speculative decoding, and multi-model routing to maximize throughput and minimize latency and cost per token
  • Build and operate reinforcement-learning and post-training pipelines using slime (Megatron-LM + SGLang) and Prime Intellect (prime-rl + the Environments Hub / verifiers). This includes reward and verifier design, rollout orchestration, weight synchronization, and keeping long-running training stable
  • Build eval harnesses and benchmarks that measure quality, throughput, and cost together, and use them to drive fast, data-informed iteration
  • Partner with Infrastructure on GPU scheduling and fleet utilization, and with Product on what to build next and why

Skills

  • 3+ years shipping ML/AI systems to production
  • Hands-on experience serving LLM inference with vLLM, SGLang, or TensorRT-LLM
  • Experience with RL / post-training methods (GRPO, PPO, DPO, or SFT), or strong adjacent experience and a clear desire to go deep here
  • Strong Python and PyTorch
  • Working understanding of GPU execution: batching, memory, and basic CUDA concepts
  • Comfort operating in ambiguity with a strong bias for action
  • Candidates must include a public GitHub profile in their application
  • The GitHub profile should demonstrate a minimum of 1 year of activity/history
  • Applications that do not include a GitHub profile, or show insufficient activity, will not be considered
  • Direct experience with slime, prime-rl, the verifiers library, or Megatron-LM
  • Distributed training experience (FSDP, TP/PP/DP parallelism)
  • Quantization (FP8/INT8), P/D disaggregation, or speculative decoding
  • Experience with verifiable inference or large-scale distributed systems
  • Kubernetes and container-based deployment
  • Familiarity with GPU fleet orchestration (Ray, SkyPilot, Slurm)

Qualifications

Must Haves

  • 3+ years shipping ML/AI systems to production
  • Hands-on experience serving LLM inference with vLLM, SGLang, or TensorRT-LLM
  • Experience with RL / post-training methods (GRPO, PPO, DPO, or SFT), or strong adjacent experience and a clear desire to go deep here
  • Strong Python and PyTorch
  • Working understanding of GPU execution: batching, memory, and basic CUDA concepts
  • Comfort operating in ambiguity with a strong bias for action
  • Candidates must include a public GitHub profile in their application
  • The GitHub profile should demonstrate a minimum of 1 year of activity/history
  • Applications that do not include a GitHub profile, or show insufficient activity, will not be considered

Nice to Haves

  • Direct experience with slime, prime-rl, the verifiers library, or Megatron-LM
  • Distributed training experience (FSDP, TP/PP/DP parallelism)
  • Quantization (FP8/INT8), P/D disaggregation, or speculative decoding
  • Experience with verifiable inference or large-scale distributed systems
  • Kubernetes and container-based deployment
  • Familiarity with GPU fleet orchestration (Ray, SkyPilot, Slurm)

Benefits

  • Health, dental, vision (for U.S. employees; region-adjusted globally)
  • Flexible PTO
  • Professional development and conference travel budget
  • Remote-first with regular off-sites and a high-trust, high-velocity team environment
  • Competitive salary + equity allocation

More jobs like this