Summary
Mindlance is hiring a Data Engineer (I) to support data processing, analysis, visualization, and machine learning initiatives. The role focuses on building agentic data pipelines, preparing spatial and research datasets, supporting annotation workflows, developing data models and pipelines, and creating data solutions with guidance from senior colleagues.
Responsibilities
- Build an agentic pipeline for data processing, analysis, visualization, as well as tools for model analysis and debugging
- Clean, structure, and catalog raw spatial/research datasets to ensure high data quality
- Coordinate with MAXO researchers to provide processed data outputs
- Coordinate with a team in India (potentially work odd hours) to communicate task expectation, enqueue and dequeue annotation jobs, analyse annotation quality and provide feedback to the rating team
- Create and/or consult in creating data visualizations, visualization features using internal BI tools like PLX, Datastudio, and external tools like Tableau and Looker, with some guidance
- Provide ongoing support for data users through maintenance of reports, queries, and dashboards with some guidance
- Perform exploratory data analysis and profiling utilizing relevant tools, leveraging custom data infrastructure, or existing data models with some guidance. Work with clients to understand their needs and clarify details of requirements. Enable data-driven decision-making by collecting, transforming, and publishing data
- Leverage, deploy, and continuously train pre-existing machine learning models. Learn and implement new storage/MPP systems/ML serving systems, with some guidance
- Implement business solutions and infrastructure to build and scale common frameworks for use with some guidance. Seek and follow local technical best practices, including making data discoverable, thinking about the lifecycle of data, and managing master data well. Design, build, operationalize, secure, and monitor data processing systems with a particular emphasis on security and compliance, scalability and efficiency, reliability and fidelity, and flexibility and portability
- Develop and maintain data models, pipelines, and exchange formats to assist in the visualization, analysis, interpretation of data and for use of data in ML training/models, with direct guidance
- Consult with users, partners, or decisions makers to identify data sources, required data elements, or data validation standards with some guidance. Consult with application engineers to understand and influence logging/transactional storage. Consult with Data scientists on ML training, feature engineering for ML models
Skills
- Solid programming skills (with python experience), especially on building tools, web UIs, scripts etc for data collection, annotation, and processing
- Data science and engineering skills, including data quality verification, building agentic data processing pipelines, data visualization, and data analysis
- Experienced in conducting human studies or surveys
- Solid communication skills with the ability to discuss technical designs with researchers/SWEs, as well as discuss data annotation with non-technical raters
- You possess a foundational understanding of core data and role-related knowledge, relevant our technologies, and processes
- Proficiency in: - Code comprehension and programming skills - Information gathering skills - Project management - Data exploration - Big data infrastructure - Machine Learning Knowledge - Stakeholder management - Data pipeline(ETL) design and Data Modeling - Statistics & BI tools
- Knowledge of VLM evaluation, multi-modal GenAI model (e.g., image generation, video generation) evaluation
Qualifications
Must Haves
- Solid programming skills (with python experience), especially on building tools, web UIs, scripts etc for data collection, annotation, and processing
- Data science and engineering skills, including data quality verification, building agentic data processing pipelines, data visualization, and data analysis
- Experienced in conducting human studies or surveys
- Solid communication skills with the ability to discuss technical designs with researchers/SWEs, as well as discuss data annotation with non-technical raters
- You possess a foundational understanding of core data and role-related knowledge, relevant our technologies, and processes
- proficiency in: - Code comprehension and programming skills - Information gathering skills - Project management - Data exploration - Big data infrastructure - Machine Learning Knowledge - Stakeholder management - Data pipeline(ETL) design and Data Modeling - Statistics & BI tools
Nice to Haves
- Knowledge of VLM evaluation, multi-modal GenAI model (e.g., image generation, video generation) evaluation
Benefits
- Hybrid work model
- Remote work from Portland, OR