HavocAI logo
HavocAI
Posted 20 days agoVerified live 2d ago

ML Cloud Infrastructure Engineer

Brief overview

Remote
UndergradOr in progress
3+ yrsMinimum
PythonAWSInfrastructure as CodeKubernetesContainerizationMachine Learning Operations (MLOps)Data pipelinesMachine learning workflowsProduction services and APIsTesting and automationSecurity and access controlsGPU and accelerator schedulingDistributed trainingMulti-modal data

About the company

Havoc is the leader in all-domain collaborative autonomy.

Job description

Summary

HavocAI develops collaborative autonomy systems for defense and commercial applications across sea, air, and land. The Machine Learning Cloud Infrastructure Engineer will build and operate infrastructure, pipelines, platforms, and integrations for training, evaluating, deploying, and monitoring machine learning models. The role focuses on scalable cloud infrastructure, data and model workflows, reliability, observability, security, and cross-functional enablement.

Responsibilities

  • Build pipelines that transform raw multi-modal data—including telemetry, imagery, video, sensor, and simulation data—into curated, versioned training datasets
  • Develop reproducible training and evaluation workflows that scale across cloud compute and GPU resources
  • Build and maintain model deployment infrastructure for packaging, serving, inference, versioning, and rollback
  • Implement experiment tracking, dataset lineage, model versioning, and other capabilities required for reproducible ML development
  • Own data schema versioning and migration across pipelines, data lakes, and services as datasets and models evolve
  • Design, build, and operate scalable AWS infrastructure using Infrastructure as Code
  • Build and maintain Kubernetes/EKS workloads and containerized environments for training, batch processing, evaluation, and model serving
  • Develop self-service tooling and paved paths for compute scheduling, storage, data access, training, and deployment
  • Improve utilization, scalability, and cost efficiency across cloud and accelerator infrastructure
  • Build infrastructure that enables engineering teams to launch workloads safely without unnecessary operational overhead
  • Build evaluation frameworks and regression testing for model quality, dataset integrity, and pipeline correctness
  • Establish quality and reliability signals that help determine whether models are ready for production use
  • Develop monitoring, logging, tracing, and observability across training jobs, data pipelines, and deployed models
  • Diagnose and resolve performance, scaling, reliability, and infrastructure bottlenecks
  • Maintain high standards for automation, testing, documentation, and operational readiness
  • Partner with Autonomy, Software, Data, Simulation, and Security teams to build ML infrastructure spanning edge data capture through cloud training and model deployment
  • Contribute to CI/CD and release processes for models, datasets, and ML pipelines
  • Translate engineering requirements into scalable platform capabilities that can support multiple teams and use cases
  • Incorporate feedback from engineers and users to continuously improve ML development workflows
  • Implement secure infrastructure practices, including IAM least privilege, secrets management, access controls, and secure handling of sensitive and defense-related data
  • Build data and ML workflows with reproducibility, traceability, and appropriate controls from the start
  • Partner with security and infrastructure teams to ensure ML systems meet applicable operational and compliance requirements

Skills

  • 3+ years of experience in software engineering, infrastructure engineering, data engineering, ML infrastructure, or a related field
  • Strong programming experience in Python, with experience in Go, C++, or another systems-oriented language preferred
  • Experience building and operating production services, APIs, data pipelines, developer platforms, or infrastructure
  • Hands-on experience with ML workflows such as dataset preparation, model training, evaluation, or deployment
  • Experience with cloud infrastructure, preferably AWS, and Infrastructure as Code
  • Hands-on experience with Kubernetes and containerized environments
  • Strong understanding of production engineering fundamentals, including reliability, observability, testing, automation, and maintainability
  • Ability to work effectively across engineering disciplines and solve ambiguous technical problems with a high degree of ownership
  • U.S. Citizenship and ability to obtain and maintain a U.S. Government security clearance
  • Experience in Go, C++, or another systems-oriented language preferred
  • Experience with MLOps and workflow platforms such as MLflow, Weights & Biases, Kubeflow, Ray, Airflow, or Dagster
  • Experience with GPU/accelerator scheduling, distributed training, or large-scale ML workloads
  • Experience working with multi-modal datasets including imagery, video, telemetry, sensor, or simulation data
  • Experience supporting autonomy, robotics, simulation, or real-time systems
  • Experience deploying ML models to edge or embedded environments
  • Experience with AWS GovCloud, GCP Assured Workloads, or compliance-driven environments such as FedRAMP or IL4/IL5

Qualifications

Must Haves

  • 3+ years of experience in software engineering, infrastructure engineering, data engineering, ML infrastructure, or a related field
  • Strong programming experience in Python, with experience in Go, C++, or another systems-oriented language preferred
  • Experience building and operating production services, APIs, data pipelines, developer platforms, or infrastructure
  • Hands-on experience with ML workflows such as dataset preparation, model training, evaluation, or deployment
  • Experience with cloud infrastructure, preferably AWS, and Infrastructure as Code
  • Hands-on experience with Kubernetes and containerized environments
  • Strong understanding of production engineering fundamentals, including reliability, observability, testing, automation, and maintainability
  • Ability to work effectively across engineering disciplines and solve ambiguous technical problems with a high degree of ownership
  • U.S. Citizenship and ability to obtain and maintain a U.S. Government security clearance

Nice to Haves

  • experience in Go, C++, or another systems-oriented language preferred
  • Experience with MLOps and workflow platforms such as MLflow, Weights & Biases, Kubeflow, Ray, Airflow, or Dagster
  • Experience with GPU/accelerator scheduling, distributed training, or large-scale ML workloads
  • Experience working with multi-modal datasets including imagery, video, telemetry, sensor, or simulation data
  • Experience supporting autonomy, robotics, simulation, or real-time systems
  • Experience deploying ML models to edge or embedded environments
  • Experience with AWS GovCloud, GCP Assured Workloads, or compliance-driven environments such as FedRAMP or IL4/IL5

Benefits

  • 100% Employer paid Health, Dental and Vision Insurance for you and your families
  • Life Insurance (Employer Paid)
  • Ability to participate in the companies 401k program (Matching)
  • Unlimited PTO policy with an enforced 2 week minimum
  • Equity Package
  • Work / Home Office Stipend
  • Global Entry
  • 16 Week Paid Parental Leave
  • Monthly Health and Wellness Stipend

More jobs like this