Medpace logo
Medpace
Posted 130 days agoVerified live 10h ago

Data Engineer (AI)

Brief overview

Cincinnati, OHIn-person
UndergradOr in progress
1+ yrsMinimum
87 H-1B approvalsDept. of Labor
15 green cardsCertified filings
ETL development.Natural Language Processing (NLP).Python programming.SQL Server databases.Snowflake cloud data warehouse.Azure cloud platform.REST API usage.Dimensional data modeling.Star Schema and Snowflake schema design.Software development lifecycle (SDLC) practices.Version control.Data visualization.Software validation and testing.User requirements gathering.Data flow management.Security, confidentiality, and privacy of PHI.C# programming (bonus).

About the company

Medpace, Inc., a clinical research organization, provides clinical development services for pharmaceutical and biotechnology

Visa sponsorship history

4 years sponsoring, last filed FY2026

Data powered by U.S. Department of Labor. This does not guarantee sponsorship for this specific role.
87H-1B approved
95%approval rate
46new H-1B hires
15PERM certified
$120,000median wage / yr
H-1B Petition ApprovalsVisas USCIS actually granted: the strongest sign the company sponsors.
202319
202431
202530
20267
LCA Certified ApplicationsAn early filing step, not a visa approval: it signals intent, not confirmed sponsorship.
20235
20241
20255
20262
Green Card (PERM) FilingsCertified green card filings: a long-term commitment to international hires.
20231
20246
20257
20261
Top sponsored roles
Statistical Analyst IIISoftware Engineer IIIBioinformatics Scientist IIIBiostatistician IISoftware Engineer I
Sponsored employees from
ChinaBangladesh

Job description

Summary

Medpace is a full-service clinical contract research organization (CRO) that accelerates the global development of safe and effective medical therapeutics. They are seeking a full-time Data Engineer to join their AI team, focusing on handling unconventional data and supporting AI tools for data extraction and natural language processing.

Responsibilities

  • Utilize skills in handling of more unconventional data such as unstructured content from web-based sites and varying content (documents, images etc) into different data lakes and with different software solutions (Snowflake, Azure, SQL, Python)
  • Provide the handling of (and where needed training of) Large Language Models (LLMs) in the Extract, Transform, and Load (ETL) of large corpus of data into a data lake
  • Participate in the Natural Language Processing (NLP) extract of unstructured data into structured meta-data through the use of tools such as Semantic understanding and meaning (Python, use of REST API)
  • Support ensuring the data flow of any external content coming in is handled to the latest US and EU AI Acts concerning AI which includes security, confidentiality and privacy of PHI
  • Collect, analyze and document user requirements working with AI engineers to align data sources to downstream integration within systems
  • Create software applications that support the understanding and visualization of data flows from inception to derivation whilst maintaining version control by following software development lifecycle process, which includes requirements gathering, design, development, testing, release, and maintenance
  • Participate in software validation process through development, review, and/or execution of test plan/cases/scripts
  • Communicate with team members regarding projects, development, tools, and procedures; and
  • Provide end-user support including setup, installation, and maintenance for application

Skills

  • Bachelor's Degree in Computer Science, Data Science, or a related field
  • 1-3+ years of experience in Data Engineering
  • Background in working with AI tools that support areas such as data extraction and natural language processing and handling of varied unstructured content into structured meta-data
  • Knowledge of developing dimensional data models from unstructured content and awareness of the advantages and limitations of Star Schema and Snowflake schema designs
  • Solid ETL development, reporting knowledge based off intricate understanding of business process and measures
  • Good knowledge of SQL Server databases and Python programming language
  • Excellent analytical, written and oral communication skills
  • Knowledge of Snowflake cloud data warehouse and Azure cloud
  • Knowledge of REST API
  • Knowledge of C# is a bonus as is working with Azure data fabric

Qualifications

Must Haves

  • Bachelor's Degree in Computer Science, Data Science, or a related field
  • 1-3+ years of experience in Data Engineering
  • Background in working with AI tools that support areas such as data extraction and natural language processing and handling of varied unstructured content into structured meta-data
  • Knowledge of developing dimensional data models from unstructured content and awareness of the advantages and limitations of Star Schema and Snowflake schema designs
  • Solid ETL development, reporting knowledge based off intricate understanding of business process and measures
  • Good knowledge of SQL Server databases and Python programming language
  • Excellent analytical, written and oral communication skills

Nice to Haves

  • Knowledge of Snowflake cloud data warehouse and Azure cloud
  • Knowledge of REST API
  • Knowledge of C# is a bonus as is working with Azure data fabric

Benefits

  • Flexible work environment
  • Competitive compensation and benefits package
  • Competitive PTO packages
  • Structured career paths with opportunities for professional growth
  • Company-sponsored employee appreciation events
  • Employee health and wellness initiatives

More jobs like this