Summary
Pokee AI develops RL-trained AI agents and enterprise AI infrastructure. The AI Infrastructure Engineer will build and optimize scalable training and inference systems, model-serving infrastructure, data pipelines, and MLOps tooling across cloud, on-premise, and on-device environments while supporting reliability, security, and compliance requirements.
Responsibilities
- Design, build, and maintain scalable training and inference infrastructure for RL-based AI agent models
- Optimize model serving for latency, throughput, and cost across cloud (AWS, GCP) and on-premise/on-device environments
- Develop and manage CI/CD pipelines, experiment tracking, and model versioning systems
- Implement efficient data pipelines for training data collection, preprocessing, and reward signal computation
- Collaborate with research scientists to productionize new algorithms and model architectures
- Ensure infrastructure meets enterprise requirements for reliability, security, and compliance (SOC 2, data residency)
Skills
- 3+ years of experience in ML infrastructure, ML platform engineering, or a related systems role
- Strong proficiency in Python and systems-level languages (Rust, C++, or Go)
- Hands-on experience with ML serving frameworks (vLLM, TensorRT, Triton, ONNX Runtime, or similar)
- Experience with container orchestration (Kubernetes, Docker) and cloud infrastructure (AWS or GCP)
- Solid understanding of GPU computing, distributed systems, and performance profiling
- Familiarity with ML experiment tracking and pipeline orchestration tools (MLflow, Weights & Biases, Airflow, or similar)
- Engineering Remote (US/Singapore Preferred) Full-time
Qualifications
Must Haves
- 3+ years of experience in ML infrastructure, ML platform engineering, or a related systems role
- Strong proficiency in Python and systems-level languages (Rust, C++, or Go)
- Hands-on experience with ML serving frameworks (vLLM, TensorRT, Triton, ONNX Runtime, or similar)
- Experience with container orchestration (Kubernetes, Docker) and cloud infrastructure (AWS or GCP)
- Solid understanding of GPU computing, distributed systems, and performance profiling
- Familiarity with ML experiment tracking and pipeline orchestration tools (MLflow, Weights & Biases, Airflow, or similar)
Nice to Haves
- Engineering Remote (US/Singapore Preferred) Full-time
Benefits
- Remote work (US/Singapore preferred)
- Direct impact on the product
- Access to cutting-edge research
- Opportunity to shape the future of enterprise AI