24-MAG logo
24-MAG
Posted 11 days agoVerified live 1d ago

Remote | MLOps Engineer, LLM Systems (Serving, GPU Kernels, Profiling) — $90–$120/hour

Brief overview

Remote
$90–$120/hrStated range
ML InfrastructureGPU Kernel ProgrammingPerformance ProfilingDistributed Systems DebuggingHigh-Throughput LLM Inference ServingPyTorchJAXCUDATritonvLLMTensorRT-LLMDistributed Training Frameworks

About the company

At 24-MAG, we support emerging AI and consulting platforms by sourcing and connecting qualified professionals with remote, contract-based opportunities.

Job description

Summary

24-MAG LLC connects experienced professionals with remote consulting opportunities across technical, evaluation, and project-based workstreams. The MLOps Engineer will develop ML-systems tasks and reference solutions, evaluate model-generated outputs, and establish standards across GPU kernels, performance profiling, distributed debugging, and high-throughput LLM serving.

Responsibilities

  • Design technically challenging tasks involving GPU and accelerator workloads
  • Develop solutions covering CUDA, Triton, Pallas, or comparable kernel technologies
  • Evaluate kernel-level optimisation approaches for correctness and efficiency
  • Analyse memory, compute, and hardware-utilisation trade-offs
  • Apply practical accelerator engineering judgement to model-generated solutions
  • Develop tasks involving performance profiling and trace interpretation
  • Analyse outputs from tools such as Kineto, torch.profiler, Nsight, XLA, or JAX profilers
  • Identify bottlenecks across compute, memory, communication, and scheduling
  • Evaluate throughput, latency, and utilisation characteristics
  • Produce clear reference analyses explaining observed performance behaviour
  • Design scenarios involving distributed or accelerator-bound ML workloads
  • Diagnose failures across training and inference infrastructure
  • Evaluate reasoning around FSDP, DDP, DeepSpeed, Megatron, and related systems
  • Review framework-level and distributed-system troubleshooting approaches
  • Identify technically plausible but incorrect explanations or proposed fixes
  • Develop and assess tasks involving high-throughput LLM serving
  • Apply expertise with vLLM, SGLang, TensorRT-LLM, Ray Serve, or comparable platforms
  • Evaluate KV-cache, paged-attention, and continuous-batching strategies
  • Analyse serving architectures for latency, throughput, memory, and scalability trade-offs
  • Review production-oriented approaches to large-scale inference deployment
  • Evaluate MLOps and ML-systems tasks and proposed solutions
  • Provide precise written feedback that can withstand technical review
  • Develop detailed rubrics and evaluation frameworks for systems-level work
  • Help research and engineering teams close technical knowledge gaps
  • Collaborate with subject-matter experts to maintain consistent training-data quality

Skills

  • * 2+ years of hands-on professional experience in ML systems, ML infrastructure, model serving, or accelerator-performance engineering
  • * Strong practical experience in at least one of GPU kernel programming, performance profiling, distributed debugging, or high-throughput inference serving
  • * Production experience with JAX and/or PyTorch
  • * Familiarity with CUDA, Triton, Pallas, or comparable accelerator-programming technologies
  • * Experience with profiling tools such as Kineto, torch.profiler, Nsight, XLA, or JAX profiler
  • * Experience debugging distributed or accelerator-bound workloads
  • * Familiarity with vLLM, SGLang, TensorRT-LLM, Ray Serve, KV cache, paged attention, or continuous batching
  • * Framework-level experience with custom operators, FSDP, DDP, DeepSpeed, Megatron, compiler, or graph-level work is highly valuable
  • * Familiarity with accelerators such as A100, H100, B200, or TPU
  • * Ability to reason precisely about throughput, latency, memory, and compute trade-offs
  • * Demonstrable professional progression in ML infrastructure or systems engineering
  • * Strong written communication and ability to explain complex technical decisions clearly
  • * Full-time **40-hour-per-week** engagement
  • * **Remote — Canada, United Kingdom, and United States**
  • * Reliable weekday availability is required
  • * The engagement requires **no conflicting or concurrent professional engagements**
  • * H1-B and STEM OPT candidates cannot currently be supported

Qualifications

Must Haves

  • * 2+ years of hands-on professional experience in ML systems, ML infrastructure, model serving, or accelerator-performance engineering
  • * Strong practical experience in at least one of GPU kernel programming, performance profiling, distributed debugging, or high-throughput inference serving
  • * Production experience with JAX and/or PyTorch
  • * Familiarity with CUDA, Triton, Pallas, or comparable accelerator-programming technologies
  • * Experience with profiling tools such as Kineto, torch.profiler, Nsight, XLA, or JAX profiler
  • * Experience debugging distributed or accelerator-bound workloads
  • * Familiarity with vLLM, SGLang, TensorRT-LLM, Ray Serve, KV cache, paged attention, or continuous batching
  • * Framework-level experience with custom operators, FSDP, DDP, DeepSpeed, Megatron, compiler, or graph-level work is highly valuable
  • * Familiarity with accelerators such as A100, H100, B200, or TPU
  • * Ability to reason precisely about throughput, latency, memory, and compute trade-offs
  • * Demonstrable professional progression in ML infrastructure or systems engineering
  • * Strong written communication and ability to explain complex technical decisions clearly
  • * Full-time **40-hour-per-week** engagement
  • * **Remote — Canada, United Kingdom, and United States**
  • * Reliable weekday availability is required
  • * The engagement requires **no conflicting or concurrent professional engagements**
  • * H1-B and STEM OPT candidates cannot currently be supported

Benefits

  • Full-time 40-hour-per-week engagement
  • Remote work in Canada, the United Kingdom, and the United States

More jobs like this