Summary
REPAY is a financial technology and payment processing company that provides electronic payment and funding solutions. The Data Engineer will design, build, optimize, and maintain cloud-based data infrastructure, data pipelines, data models, and ETL/ELT processes to support analytics, reporting, business intelligence, and operational decision-making.
Responsibilities
- Design, build, and maintain scalable, reliable cloud-based data pipelines and data infrastructure
- Deliver high-quality data models and curated datasets that support analytics, reporting, and data-driven decision-making
- Optimize Spark, PySpark, and SQL workloads to improve performance, reliability, cost efficiency, and scalability
- Support production data pipelines through monitoring, troubleshooting, incident resolution, and continuous improvement
- Implement data engineering standards, CI/CD practices, automated deployment processes, unit testing, and code quality expectations
- Partner with BI, Product, Engineering, and client-facing teams to translate business and reporting requirements into scalable data solutions
- Document technical solutions, data flows, pipeline logic, and operational processes to support knowledge sharing and long-term maintainability
- Design, build, maintain, and optimize data pipelines using Python, SQL, PySpark, Databricks, and AWS-based data services
- Develop ETL/ELT processes that support data warehousing, analytics, reporting, and business intelligence use cases
- Build and optimize Spark jobs, with a focus on performance, scalability, reliability, and efficient resource utilization
- Design and implement data models for structured, semi-structured, and NoSQL data where applicable
- Implement CI/CD practices, automated deployments, unit tests, and code quality standards for data engineering workflows
- Monitor, troubleshoot, and support production data pipelines, resolving issues and recommending improvements
- Collaborate with BI Analysts, Product, Engineering, Data, and client-facing teams to understand requirements and support reporting needs
- Document technical solutions, data flows, pipeline logic, and operational processes
- Share technical knowledge through documentation, mentorship, and team knowledge-sharing sessions
- Stay current with advancements in data engineering, cloud platforms, Spark, Databricks, data warehousing, and analytics technologies
- Participate in client-facing design sessions, technical presentations, workshops, or training as needed
- Other duties as assigned
Skills
- Undergraduate or Masters' degree in Computer Science, Statistics, or Analytics
- Minimum of 3–5 years of experience in Data Engineering, preferably working with AWS-based cloud data platforms
- Hands-on experience building, maintaining, and supporting cloud-based data pipelines
- Strong knowledge of PySpark, preferably on the Databricks platform
- Hands-on experience with Databricks
- Strong proficiency in SQL, including query optimization
- Strong proficiency in Python
- Strong knowledge of data modeling, data warehousing, ETL/ELT, and analytics concepts
- Experience with CI/CD practices, automated deployment processes, unit testing, and code quality standards
- Experience troubleshooting, monitoring, and supporting production data pipelines
- Experience documenting technical solutions, data flows, and pipeline logic
- Strong analytical and problem-solving skills, with the ability to translate business requirements into scalable data solutions
- Excellent written and verbal communication skills, including the ability to explain technical concepts to technical and non-technical stakeholders
- Ability to collaborate effectively across BI, Product, Engineering, Data, and client-facing teams
- Strong organizational skills and ability to manage multiple priorities in a fast-paced environment
- Proactive, ownership-oriented mindset with the ability to work independently and drive solutions from design through production support
- Professionalism and composure when supporting production issues or participating in client-facing discussions
- We are interested in every qualified candidate who is eligible to work in the United States
- This position is not eligible for hire in California
- Additionally, we are not able to sponsor visas
- Apache Kafka experience
- AWS Lambda experience
- AWS Glue experience
- MongoDB experience
- Experience with streaming technologies such as Kafka or Kinesis
- Terraform experience
- Familiarity with an analytics/visualization platform such as Power BI or Tableau
Qualifications
Must Haves
- Undergraduate or Masters' degree in Computer Science, Statistics, or Analytics
- Minimum of 3–5 years of experience in Data Engineering, preferably working with AWS-based cloud data platforms
- Hands-on experience building, maintaining, and supporting cloud-based data pipelines
- Strong knowledge of PySpark, preferably on the Databricks platform
- Hands-on experience with Databricks
- Strong proficiency in SQL, including query optimization
- Strong proficiency in Python
- Strong knowledge of data modeling, data warehousing, ETL/ELT, and analytics concepts
- Experience with CI/CD practices, automated deployment processes, unit testing, and code quality standards
- Experience troubleshooting, monitoring, and supporting production data pipelines
- Experience documenting technical solutions, data flows, and pipeline logic
- Strong analytical and problem-solving skills, with the ability to translate business requirements into scalable data solutions
- Excellent written and verbal communication skills, including the ability to explain technical concepts to technical and non-technical stakeholders
- Ability to collaborate effectively across BI, Product, Engineering, Data, and client-facing teams
- Strong organizational skills and ability to manage multiple priorities in a fast-paced environment
- Proactive, ownership-oriented mindset with the ability to work independently and drive solutions from design through production support
- Professionalism and composure when supporting production issues or participating in client-facing discussions
- We are interested in every qualified candidate who is eligible to work in the United States
- This position is not eligible for hire in California
- Additionally, we are not able to sponsor visas
Nice to Haves
- Apache Kafka experience
- AWS Lambda experience
- AWS Glue experience
- MongoDB experience
- Experience with streaming technologies such as Kafka or Kinesis
- Terraform experience
- Familiarity with an analytics/visualization platform such as Power BI or Tableau
Benefits
- Business casual dress
- Great snacks & beverages
- Open-air collaborative team settings
- Continuing education, including professional conferences and events
- 100% coverage of employee healthcare premiums
- Free life insurance
- Free disability insurance
- Work-life balance resources
- All benefits go into effect day one
- 401(k)-employer match
- Employee Stock Purchase Plan
- Eligibility to participate in the Annual Bonus Program, with the bonus award reflecting excellent individual performance and goals achieved during the past year