Summary
Marathon TS is seeking a Junior Data Engineer to support critical missions within Peraton's Risk Decision Group. The role develops and maintains data ingestion pipelines, transformations, and governed data products for a federal background investigation data and AI platform, while supporting infrastructure-as-code, CI/CD, data quality, governance, and machine learning teams in a regulated environment.
Responsibilities
- Build, test, and maintain data ingestion, transformation, and pipeline code that turns raw and synthetic data into governed, analysis-ready data products
- Extend the platform following the architecture, patterns, and standards set by the lead platform engineer
- Contribute to infrastructure-as-code (Terraform) and CI/CD pipelines that promote work into the accredited production environment
- Monitor data quality, pipeline reliability, and performance; troubleshoot and resolve issues
- Apply data governance, cataloging, and access-control practices within the established framework
- Support the data science and ML team by delivering the governed datasets and tooling they need
- Operate within FedRAMP Moderate, NIST 800-171, and CUI constraints in day-to-day engineering
- Collaborate closely with the engineering team and maintain clear documentation of pipelines and processes
Skills
- U.S. citizenship required
- Must be able to obtain and maintain a T5/SSBI federally adjudicated clearance
- 5-7 years building and maintaining production data pipelines
- Strong proficiency in Python and SQL
- Hands-on ETL / data pipeline development and sound analytical data modeling
- Working familiarity with cloud data platforms (Azure preferred; AWS or GCP fine) and version control / CI/CD workflows
- Solid analytical, troubleshooting, and communication skills
- Dependable, self-directed execution with minimal oversight
- Strong collaboration across technical and non-technical teams
- Clear documentation and knowledge-sharing
- Eagerness to grow ownership and learn the compliance environment
- Bachelor's degree in computer science, software engineering or relevant field required
- Active clearance preferred
- Databricks experience (or a comparable lakehouse stack: Spark, Delta, dbt, Snowflake)
- Exposure to infrastructure-as-code (Terraform) and CI/CD for data workloads
- Experience in regulated or accredited environments: FedRAMP, NIST 800-171, CMMC, or CUI handling
- Active security clearance (T5/SSBI or higher)
- Government or defense contracting experience
- Familiarity with data governance / cataloging tooling (e.g., Unity
Catalog) and orchestration tools (Airflow, dbt)
Qualifications
Must Haves
- U.S. citizenship required
- Must be able to obtain and maintain a T5/SSBI federally adjudicated clearance
- 5-7 years building and maintaining production data pipelines
- Strong proficiency in Python and SQL
- Hands-on ETL / data pipeline development and sound analytical data modeling
- Working familiarity with cloud data platforms (Azure preferred; AWS or GCP fine) and version control / CI/CD workflows
- Solid analytical, troubleshooting, and communication skills
- Dependable, self-directed execution with minimal oversight
- Strong collaboration across technical and non-technical teams
- Clear documentation and knowledge-sharing
- Eagerness to grow ownership and learn the compliance environment
- Bachelor's degree in computer science, software engineering or relevant field required
Nice to Haves
- active clearance preferred
- Databricks experience (or a comparable lakehouse stack: Spark, Delta, dbt, Snowflake)
- Exposure to infrastructure-as-code (Terraform) and CI/CD for data workloads
- Experience in regulated or accredited environments: FedRAMP, NIST 800-171, CMMC, or CUI handling
- Active security clearance (T5/SSBI or higher)
- Government or defense contracting experience
- Familiarity with data governance / cataloging tooling (e.g., Unity
Catalog) and orchestration tools (Airflow, dbt)
Benefits