Pokee AI logo
Pokee AI
Posted 47 days agoVerified live 1d ago

AI Infrastructure Engineer

Brief overview

Remote
3+ yrsMinimum
3 H-1B approvalsDept. of Labor
PythonSystems Programming RustSystems Programming C++Systems ProgrammingSystems Programming GoML Serving Frameworks vLLMML Serving Frameworks TensorRTML Serving Frameworks TritonML Serving FrameworksML Serving Frameworks ONNX RuntimeKubernetes and DockerCloud Infrastructure AWSCloud Infrastructure GCPGPU ComputingDistributed SystemsPerformance ProfilingML Experiment Tracking and Pipeline Orchestration

About the company

Pokee AI logo
Pokee AIpokee.ai

Pokee AI is an AI company that offers an interactive, personalized, and efficient AI agent that helps in planning, reasoning and tool usage.

Visa sponsorship history

1 year sponsoring, last filed FY2025

Data powered by U.S. Department of Labor. This does not guarantee sponsorship for this specific role.
3H-1B approved
100%approval rate
H-1B Petition ApprovalsVisas USCIS actually granted: the strongest sign the company sponsors.
20253

Job description

Summary

Pokee AI develops RL-trained AI agents and enterprise AI infrastructure. The AI Infrastructure Engineer will build and optimize scalable training and inference systems, model-serving infrastructure, data pipelines, and MLOps tooling across cloud, on-premise, and on-device environments while supporting reliability, security, and compliance requirements.

Responsibilities

  • Design, build, and maintain scalable training and inference infrastructure for RL-based AI agent models
  • Optimize model serving for latency, throughput, and cost across cloud (AWS, GCP) and on-premise/on-device environments
  • Develop and manage CI/CD pipelines, experiment tracking, and model versioning systems
  • Implement efficient data pipelines for training data collection, preprocessing, and reward signal computation
  • Collaborate with research scientists to productionize new algorithms and model architectures
  • Ensure infrastructure meets enterprise requirements for reliability, security, and compliance (SOC 2, data residency)

Skills

  • 3+ years of experience in ML infrastructure, ML platform engineering, or a related systems role
  • Strong proficiency in Python and systems-level languages (Rust, C++, or Go)
  • Hands-on experience with ML serving frameworks (vLLM, TensorRT, Triton, ONNX Runtime, or similar)
  • Experience with container orchestration (Kubernetes, Docker) and cloud infrastructure (AWS or GCP)
  • Solid understanding of GPU computing, distributed systems, and performance profiling
  • Familiarity with ML experiment tracking and pipeline orchestration tools (MLflow, Weights & Biases, Airflow, or similar)
  • Engineering Remote (US/Singapore Preferred) Full-time

Qualifications

Must Haves

  • 3+ years of experience in ML infrastructure, ML platform engineering, or a related systems role
  • Strong proficiency in Python and systems-level languages (Rust, C++, or Go)
  • Hands-on experience with ML serving frameworks (vLLM, TensorRT, Triton, ONNX Runtime, or similar)
  • Experience with container orchestration (Kubernetes, Docker) and cloud infrastructure (AWS or GCP)
  • Solid understanding of GPU computing, distributed systems, and performance profiling
  • Familiarity with ML experiment tracking and pipeline orchestration tools (MLflow, Weights & Biases, Airflow, or similar)

Nice to Haves

  • Engineering Remote (US/Singapore Preferred) Full-time

Benefits

  • Remote work (US/Singapore preferred)
  • Direct impact on the product
  • Access to cutting-edge research
  • Opportunity to shape the future of enterprise AI

More jobs like this