Summary
Kentro is a technology and mission-support company serving government customers through secure, innovative solutions. The Mid-Level Data Engineer will build governed cloud data pipelines, automate operational tasks, support Azure and Databricks platforms, and contribute to AI-enabled data workflows. The role also involves troubleshooting, infrastructure automation, documentation, and collaboration with technical and stakeholder teams.
Responsibilities
- Develop, test, deploy, and maintain batch and API-based data ingestion, transformation, and publishing workflows
- Write clear, maintainable Python, PySpark, and SQL for data processing, validation, reconciliation, and automation
- Build and support Databricks workflows using notebooks, jobs, compute, Delta tables, catalogs, schemas, permissions, and governed data products
- Implement layered data designs, including landing/raw, Bronze, Silver, and curated or presentation outputs
- Add data-quality checks, lineage, audit metadata, checksums, retry and recovery behavior, logging, and reproducible tests to pipelines
- Support Azure data-platform services, cloud storage, secrets, workload identities, role-based access, monitoring, networking, and private connectivity
- Develop and review Infrastructure as Code, primarily Terraform, using reusable modules and environment-specific configuration
- Deliver traceable changes through Git branches, pull requests, code reviews, issue tracking, and CI/CD workflows
- Troubleshoot data, code, access, authentication, permissions, deployment, networking, and runtime issues using logs, tests, queries, and documented evidence
- Support AI-enabled data workflows, including model endpoints, response validation, evaluation, regression testing, and responsible-use controls
- Create and maintain architecture diagrams, runbooks, implementation notes, decision records, status updates, and handoff documentation
- Communicate progress, risks, availability, and blockers promptly; ask for help early enough to protect delivery and close the loop on commitments
- Participate in stand-ups, design reviews, demonstrations, and stakeholder discussions, translating technical findings for the intended audience
Skills
- Three to five years of relevant experience in data engineering, software engineering, cloud engineering, analytics engineering, or a closely related discipline
- Hands-on programming experience with Python or a comparable language, including reading unfamiliar code and debugging systematically
- Practical SQL experience with joins, aggregations, transformations, and row-level and aggregate validation
- Experience developing or supporting data pipelines, schemas, APIs, databases, cloud storage, or distributed data-processing workflows
- Working knowledge of Git, including commits, branches, pull requests, code review, and ordinary conflict resolution
- Experience testing work, retaining evidence, and writing documentation or runbooks that another engineer can follow
- Demonstrated ownership, learning agility, persistence, responsiveness, and follow-through when solving unfamiliar or ambiguous problems
- Ability to communicate technical status, risks, assumptions, and blockers clearly to both technical and nontechnical stakeholders
- Commitment to security, least privilege, responsible data handling, and compliance with customer requirements
- * US Citizen or Lawful Permanent Resident (Green Card)
- * Willing and able to obtain and maintain Public Trust Clearance or higher
- Experience with Azure, Azure Government, AWS, or another major cloud platform
- Experience with Databricks, Apache Spark, PySpark, Delta Lake, Lakeflow, or a comparable data-processing platform
- Experience with Terraform or another Infrastructure-as-Code tool
- Familiarity with Azure Data Lake Storage, Data Factory, Key Vault, Entra ID, managed identities, RBAC, private endpoints, DNS, or virtual networks
- Knowledge of data governance, catalogs, lineage, stewardship, audit trails, observability, and reproducible processing
- Experience with REST APIs, JSON/JSONL, CSV, Parquet, large-file processing, third-party ingestion, CI/CD, containers, or structured logging
- Exposure to AI/LLM-enabled applications, model-serving endpoints, evaluation, or automated regression testing
- Experience in a government, healthcare, or other regulated or compliance-sensitive environment
- Bachelor's degree in computer science, data science, information systems, engineering, or a related field, or equivalent relevant experience
Qualifications
Must Haves
- Three to five years of relevant experience in data engineering, software engineering, cloud engineering, analytics engineering, or a closely related discipline
- Hands-on programming experience with Python or a comparable language, including reading unfamiliar code and debugging systematically
- Practical SQL experience with joins, aggregations, transformations, and row-level and aggregate validation
- Experience developing or supporting data pipelines, schemas, APIs, databases, cloud storage, or distributed data-processing workflows
- Working knowledge of Git, including commits, branches, pull requests, code review, and ordinary conflict resolution
- Experience testing work, retaining evidence, and writing documentation or runbooks that another engineer can follow
- Demonstrated ownership, learning agility, persistence, responsiveness, and follow-through when solving unfamiliar or ambiguous problems
- Ability to communicate technical status, risks, assumptions, and blockers clearly to both technical and nontechnical stakeholders
- Commitment to security, least privilege, responsible data handling, and compliance with customer requirements
- * US Citizen or Lawful Permanent Resident (Green Card)
- * Willing and able to obtain and maintain Public Trust Clearance or higher
Nice to Haves
- Experience with Azure, Azure Government, AWS, or another major cloud platform
- Experience with Databricks, Apache Spark, PySpark, Delta Lake, Lakeflow, or a comparable data-processing platform
- Experience with Terraform or another Infrastructure-as-Code tool
- Familiarity with Azure Data Lake Storage, Data Factory, Key Vault, Entra ID, managed identities, RBAC, private endpoints, DNS, or virtual networks
- Knowledge of data governance, catalogs, lineage, stewardship, audit trails, observability, and reproducible processing
- Experience with REST APIs, JSON/JSONL, CSV, Parquet, large-file processing, third-party ingestion, CI/CD, containers, or structured logging
- Exposure to AI/LLM-enabled applications, model-serving endpoints, evaluation, or automated regression testing
- Experience in a government, healthcare, or other regulated or compliance-sensitive environment
- Bachelor's degree in computer science, data science, information systems, engineering, or a related field, or equivalent relevant experience
Benefits
- Paid time off
- Healthcare benefits
- Supplemental benefits
- 401k including an employer match
- Discount perks
- Rewards
- Every employee is eligible for education reimbursement for certifications, degrees, or professional development. Reimbursement amounts may fluctuate due to IRS limitations.
- Flexibility to take a course, complete a certification, or pursue other professional growth and networking
- Funds for virtual and in-person activities, including happy hours, holiday events, fitness and wellness events, and annual celebrations
- Remote work within the United States