Summary
1872 Consulting is a fast-growing company seeking seasoned data engineers to join their small team. The role involves developing and enhancing data systems, building reliable data pipelines, and collaborating with data scientists on data models and algorithms.
Responsibilities
- Build reliable data pipelines to clean, aggregate, and transform large volumes of data from multiple sources
- Develop versatile software components to extract useful information from various unstructured or semi-structured text data
- Implement advanced search functionalities and improve the efficiency of search indexing
- Work closely with data scientists to develop, test and iterate data models and algorithms
- Contribute to company-wide data privacy compliance efforts
Skills
- 3+ years of experience with Data Engineering
- Extensive experience in building large scale data pipelines with mainstream big data stack, ideally some experience building data pipelines with Python
- Strong expertise in extracting useful information from unstructured and semi-structured text data
- Software development skills with Java and Python is a plus
- Professional working experience with Elasticsearch, Apache Beam, Spark, and/or GCP Dataflow a big plus
- Strong expertise in NLP or Text Mining is also a plus
- Bachelor's degree or greater in relevant field of study is a plus
Qualifications
Must Haves
- 3+ years of experience with Data Engineering
- Extensive experience in building large scale data pipelines with mainstream big data stack, ideally some experience building data pipelines with Python
- Strong expertise in extracting useful information from unstructured and semi-structured text data
Nice to Haves
- Software development skills with Java and Python is a plus
- Professional working experience with Elasticsearch, Apache Beam, Spark, and/or GCP Dataflow a big plus
- Strong expertise in NLP or Text Mining is also a plus
- Bachelor's degree or greater in relevant field of study is a plus