Voxel51 logo
Voxel51
Posted 84 days agoVerified live 11h ago

Data Engineer Intern

Brief overview

Los Angeles, CAIn-person
UndergradOr in progress

About the company

Voxel51: the physical AI data platform to analyze multimodal data, annotate smarter, debug model failures & ship reliable AI.

Job description

Summary

VoxelCloud is a Los Angeles-based leader in artificial intelligence analysis of medical images. The Data Engineer intern will participate in the acquisition and manipulation of large-scale medical and healthcare data, supporting software developers and machine learning engineers while optimizing data systems.

Responsibilities

  • Create and maintain optimal data pipelines to support machine learning research and development
  • Identify, design, and implement internal process improvements: automating data QA, optimizing data delivery, re-designing infrastructure for greater scalability, etc
  • Build the infrastructure required for optimal extraction, transformation, and loading of data from a wide variety of data sources using SQL and AWS/AliCloud big data technologies
  • Build analytics tools that utilize the data pipeline to provide actionable insights into product utilization and operational efficiency
  • Keep our data separated and secure across national boundaries both locally and on cloud storage

Skills

  • Proficient with at least one object-oriented/object function scripting languages: Python, Java, C++, Scala, etc
  • Working SQL knowledge and experience working with relational databases, query authoring (SQL) as well as working familiarity with a variety of databases (Postgres)
  • Experience building and optimizing ‘big data' data pipelines, architectures and data sets
  • Solid understanding of information retrieval, statistics and machine learning
  • Skillful with automation tasks, but willing to get hands dirty for quality control
  • Detail-oriented, well organized and self-motivated with a continuous drive to learn, explore and challenge; good communication skills and team player
  • Experience supporting and working with cross-functional teams in a dynamic environment
  • MS, BA/BS degree in computer science, statistics or related field
  • Prefer 1+ years in big data and related technology (e.g. DFS); experience with high-performance and scalable distributed system
  • Prefer experience with AWS cloud services: EC2, EMR, RDS, Redshift

Qualifications

Must Haves

  • Proficient with at least one object-oriented/object function scripting languages: Python, Java, C++, Scala, etc
  • Working SQL knowledge and experience working with relational databases, query authoring (SQL) as well as working familiarity with a variety of databases (Postgres)
  • Experience building and optimizing ‘big data' data pipelines, architectures and data sets
  • Solid understanding of information retrieval, statistics and machine learning
  • Skillful with automation tasks, but willing to get hands dirty for quality control
  • Detail-oriented, well organized and self-motivated with a continuous drive to learn, explore and challenge; good communication skills and team player
  • Experience supporting and working with cross-functional teams in a dynamic environment
  • MS, BA/BS degree in computer science, statistics or related field

Nice to Haves

  • Prefer 1+ years in big data and related technology (e.g. DFS); experience with high-performance and scalable distributed system
  • Prefer experience with AWS cloud services: EC2, EMR, RDS, Redshift

Benefits

  • An outstanding start-up culture;
  • Transparent, collaborative work environment;
  • Competitive compensation
  • Excellent Medical, Dental, and Vision coverage
  • 401k, paid Vacation and Holiday

More jobs like this