Summary
24-MAG connects experienced professionals with remote consulting opportunities across technical, evaluation, and project-based workstreams. The company is seeking an experienced Data Engineer to build scalable ETL pipelines, transform structured and unstructured data, improve data quality, and support research, analytics, and AI/ML initiatives. The role also involves database engineering, exploratory data analysis, automation, documentation, and collaboration with researchers, data scientists, and engineers.
Responsibilities
- Design, develop, and maintain scalable ETL pipelines
- Build reliable workflows for data extraction, transformation, and loading
- Improve pipeline efficiency, maintainability, and scalability
- Automate recurring data-processing tasks
- Monitor pipelines and troubleshoot operational failures
- Collect data from structured and unstructured sources
- Clean, normalise, transform, and organise raw datasets
- Develop repeatable transformation workflows
- Identify inconsistencies and incomplete records
- Prepare reliable datasets for downstream research and analytics
- Conduct exploratory analysis across complex datasets
- Identify patterns, trends, anomalies, and quality issues
- Investigate unexpected behaviour in source data
- Produce analytical summaries supporting technical decisions
- Communicate findings clearly to interdisciplinary stakeholders
- Write and optimise SQL queries for data extraction and transformation
- Work with relational database systems such as PostgreSQL and MySQL
- Design and maintain database schemas
- Improve query performance and data-access patterns
- Support reliable and maintainable database workflows
- Develop logical and physical data models
- Define database structures appropriate to analytical and operational requirements
- Evaluate schema design and transformation strategies
- Support scalable data-storage solutions
- Maintain consistency across related datasets and systems
- Develop automated data-validation workflows
- Monitor accuracy, integrity, consistency, and completeness
- Identify and resolve data-quality issues
- Establish repeatable quality-control procedures
- Ensure datasets meet downstream analytical and modelling requirements
- Build data-processing workflows in Python
- Use Pandas and NumPy for transformation and analysis
- Develop reusable data-processing components
- Improve performance and reliability of analytical workflows
- Support automation of repetitive data-engineering processes
- Prepare datasets for AI and machine-learning initiatives
- Collaborate with data scientists and researchers on data requirements
- Support model-development workflows through reliable data preparation
- Evaluate data suitability for training and evaluation use cases
- Help structure data pipelines supporting AI/ML experimentation
- Automate recurring reporting and data-processing activities
- Build repeatable validation and monitoring processes
- Reduce manual intervention across routine data workflows
- Improve operational visibility into pipeline health
- Support timely delivery of high-quality datasets
- Document pipelines, schemas, workflows, and technical decisions
- Maintain clear operational and development documentation
- Investigate and resolve pipeline failures
- Diagnose data-related technical issues
- Communicate root causes and remediation steps clearly
Skills
- Strong proficiency in Python
- Strong proficiency in SQL
- Hands-on experience designing and maintaining ETL pipelines
- Experience conducting exploratory data analysis
- Proficiency with Pandas and NumPy
- Experience with PostgreSQL and MySQL
- Strong understanding of data modelling and database schemas
- Experience working with structured and unstructured datasets
- Demonstrated ability to maintain data quality, integrity, and reliability
- Familiarity with Jupyter Notebook, VS Code, PyCharm, or comparable development environments
- Strong analytical and problem-solving skills
- Strong written and verbal communication skills
- Experience collaborating with researchers, data scientists, or engineering teams is valuable
- Exposure to AI and machine-learning workflows is advantageous
- Familiarity with scikit-learn is beneficial
- Experience with Hugging Face Transformers is a plus
- Familiarity with AI APIs or comparable AI platforms is advantageous
- Experience preparing datasets for AI/ML model development is highly valuable
Qualifications
Must Haves
- Strong proficiency in Python
- Strong proficiency in SQL
- Hands-on experience designing and maintaining ETL pipelines
- Experience conducting exploratory data analysis
- Proficiency with Pandas and NumPy
- Experience with PostgreSQL and MySQL
- Strong understanding of data modelling and database schemas
- Experience working with structured and unstructured datasets
- Demonstrated ability to maintain data quality, integrity, and reliability
- Familiarity with Jupyter Notebook, VS Code, PyCharm, or comparable development environments
- Strong analytical and problem-solving skills
- Strong written and verbal communication skills
Nice to Haves
- Experience collaborating with researchers, data scientists, or engineering teams is valuable
- Exposure to AI and machine-learning workflows is advantageous
- Familiarity with scikit-learn is beneficial
- Experience with Hugging Face Transformers is a plus
- Familiarity with AI APIs or comparable AI platforms is advantageous
- Experience preparing datasets for AI/ML model development is highly valuable
Benefits
- Full-time engagement
- Fully remote