HavocAI logo
HavocAI
Posted 18 days agoVerified live 1d ago

Data and ML Infrastructure Engineer

Brief overview

Remote
UndergradOr in progress
3+ yrsMinimum
Clearance requiredU.S. government
Data EngineeringProduction Data PipelinesPythonSQLData LakesVideo, Telemetry, and Sensor DataML Infrastructure and MLOpsDataset Versioning and Reproducible Data WorkflowsMetadata Management and Data LineageData Quality and ObservabilitySoftware Testing and ReliabilityDebugging

About the company

Havoc is the leader in all-domain collaborative autonomy.

Job description

Summary

HavocAI develops collaborative autonomous systems for complex sea, air, and land environments. The Data and ML Infrastructure Engineer will build and maintain scalable pipelines, data lake infrastructure, and developer tools that transform multimodal operational data into searchable, reproducible, ML-ready datasets. The role also supports data quality, dataset curation, model development workflows, and collaboration across engineering and field operations teams.

Responsibilities

  • Build and maintain infrastructure for video, imagery, telemetry, sensor data, autonomy logs, mission data, and field-test data
  • Own data ingestion, storage, indexing, metadata, access patterns, and lifecycle management within HavocAI’s data lake
  • Develop scalable pipelines that transform raw operational data into curated datasets for ML training, evaluation, debugging, and analysis
  • Build tools for searching, filtering, tagging, and retrieving data across platforms, missions, operating conditions, and events
  • Design infrastructure capable of handling large volumes of multimodal operational data efficiently and reliably
  • Build workflows to select, clean, label, validate, and version datasets
  • Partner with Autonomy, Perception, Software, and Field Operations teams to identify high-value data for model development and system evaluation
  • Support annotation and labeling workflows for video, imagery, tracks, telemetry, and other ML inputs
  • Develop reproducible dataset-generation workflows for training, validation, regression testing, and benchmarking
  • Integrate datasets and data infrastructure with model training, experiment tracking, evaluation, and deployment workflows
  • Support multimodal dataset construction, including synchronization and alignment across sensors and data streams
  • Develop automated checks for missing streams, corrupted files, synchronization issues, metadata gaps, labeling errors, and pipeline failures
  • Establish standards for dataset quality, lineage, versioning, and reproducibility
  • Build monitoring and observability around critical data pipelines and infrastructure
  • Troubleshoot complex data and infrastructure issues and drive them through resolution
  • Use field data, logs, and test results to help engineering teams understand system performance and identify opportunities for improvement
  • Build self-service tools that make operational data easier for engineers to discover, access, analyze, and use
  • Partner closely with Autonomy, Perception, Software, Simulation, Field Operations, and Program teams
  • Translate engineering and ML requirements into scalable data capabilities
  • Improve workflows for replaying, visualizing, analyzing, and comparing operational data
  • Maintain clear documentation, data standards, and best practices for internal data use, governance, and security

Skills

  • Bachelor's degree in Computer Science, Data Science, Machine Learning, Electrical Engineering, Computer Engineering, Robotics, Applied Mathematics, or a related technical field
  • 3+ years of experience in data engineering, ML infrastructure, data platforms, backend systems, MLOps, or related engineering roles
  • Experience designing and operating production data pipelines for large-scale structured, semi-structured, or unstructured datasets
  • Experience working with video, imagery, time-series telemetry, sensor data, logs, or other high-volume operational data
  • Strong programming skills in Python and SQL
  • Experience with cloud storage, object stores, data lakes, databases, distributed processing, or modern data platforms
  • Familiarity with dataset versioning, metadata management, data lineage, access controls, and reproducible data workflows
  • Strong software engineering fundamentals, including testing, reliability, maintainability, and observability
  • Strong debugging skills and comfort working across complex data pipelines and production infrastructure
  • Ability to operate independently and take ownership in a fast-moving engineering environment
  • U.S. citizenship and ability to obtain and maintain a U.S. Government security clearance
  • Experience with ML infrastructure, MLOps, training pipelines, experiment tracking, model evaluation, or model registries
  • Experience managing video, perception, telemetry, or autonomous-system datasets
  • Experience with technologies such as S3-compatible storage, PostgreSQL, Spark, Ray, Airflow, Dagster, Kubernetes, Docker, or Kafka
  • Experience with data catalogs, dataset versioning platforms, feature stores, or labeling tools
  • Experience building search, replay, visualization, or analysis tools for video, telemetry, logs, or sensor data
  • Experience supporting annotation workflows for computer vision, perception, tracking, or autonomy
  • Familiarity with sensor synchronization, timestamp alignment, calibration metadata, log replay, or multimodal dataset construction
  • Experience with security, access controls, auditability, and data-handling requirements in government or defense environments
  • Experience supporting defense, robotics, autonomy, aerospace, or dual-use technology programs
  • Active or prior security clearance

Qualifications

Must Haves

  • Bachelor's degree in Computer Science, Data Science, Machine Learning, Electrical Engineering, Computer Engineering, Robotics, Applied Mathematics, or a related technical field
  • 3+ years of experience in data engineering, ML infrastructure, data platforms, backend systems, MLOps, or related engineering roles
  • Experience designing and operating production data pipelines for large-scale structured, semi-structured, or unstructured datasets
  • Experience working with video, imagery, time-series telemetry, sensor data, logs, or other high-volume operational data
  • Strong programming skills in Python and SQL
  • Experience with cloud storage, object stores, data lakes, databases, distributed processing, or modern data platforms
  • Familiarity with dataset versioning, metadata management, data lineage, access controls, and reproducible data workflows
  • Strong software engineering fundamentals, including testing, reliability, maintainability, and observability
  • Strong debugging skills and comfort working across complex data pipelines and production infrastructure
  • Ability to operate independently and take ownership in a fast-moving engineering environment
  • U.S. citizenship and ability to obtain and maintain a U.S. Government security clearance

Nice to Haves

  • Experience with ML infrastructure, MLOps, training pipelines, experiment tracking, model evaluation, or model registries
  • Experience managing video, perception, telemetry, or autonomous-system datasets
  • Experience with technologies such as S3-compatible storage, PostgreSQL, Spark, Ray, Airflow, Dagster, Kubernetes, Docker, or Kafka
  • Experience with data catalogs, dataset versioning platforms, feature stores, or labeling tools
  • Experience building search, replay, visualization, or analysis tools for video, telemetry, logs, or sensor data
  • Experience supporting annotation workflows for computer vision, perception, tracking, or autonomy
  • Familiarity with sensor synchronization, timestamp alignment, calibration metadata, log replay, or multimodal dataset construction
  • Experience with security, access controls, auditability, and data-handling requirements in government or defense environments
  • Experience supporting defense, robotics, autonomy, aerospace, or dual-use technology programs
  • Active or prior security clearance

Benefits

  • 100% Employer paid Health, Dental and Vision Insurance for you and your families
  • Life Insurance (Employer Paid)
  • Ability to participate in the companies 401k program (Matching)
  • Unlimited PTO policy with an enforced 2 week minimum
  • Equity Package
  • Work / Home Office Stipend
  • Global Entry
  • 16 Week Paid Parental Leave
  • Monthly Health and Wellness Stipend

More jobs like this