Cotiviti logo
Cotiviti
Posted 10 days agoVerified live 10h ago

Intern AI Engineer, Early-Career – LLM Context & Data Layer (Healthcare)

Brief overview

Remote
PhDOr in progress
$32–$40/hrStated range
274 H-1B approvalsDept. of Labor
55 green cardsCertified filings
SQLLarge Language Model APIsLangChainLlamaIndexGitRetrieval-Augmented Generation (RAG)Embedding ModelsVector DatabasesData PipelinesMachine LearningDeep LearningAWS/Azure Cloud Services

About the company

Cotiviti logo
Cotiviticotiviti.com

Cotiviti enables healthcare organizations to deliver better care at lower cost through advanced technology and data analytics that improve the quality and sustainability of healthcare in the United States.

Visa sponsorship history

4 years sponsoring, last filed FY2026

Data powered by U.S. Department of Labor. This does not guarantee sponsorship for this specific role.
274H-1B approved
99%approval rate
40new H-1B hires
55PERM certified
$123,781median wage / yr
H-1B Petition ApprovalsVisas USCIS actually granted: the strongest sign the company sponsors.
202366
2024100
202586
202622
LCA Certified ApplicationsAn early filing step, not a visa approval: it signals intent, not confirmed sponsorship.
202330
202416
202518
202623
Green Card (PERM) FilingsCertified green card filings: a long-term commitment to international hires.
202318
20246
202519
202612
Top sponsored roles
Senior Software EngineerSoftware EngineerTechnical ArchitectSr. Software EngineerSr. Implementation Consultant
Sponsored employees from
IndiaNepalChinaColombia

Job description

Summary

Cotiviti is seeking an Intern AI Engineer to support the development of AI-powered applications for healthcare treatment, payment, and operations. The role involves building LLM context layers, data retrieval systems, governance guardrails, semantic data pipelines, and reliable AI components in collaboration with technical and business teams.

Responsibilities

  • SQL fundamentals - comfortable writing queries against structured data, not just describing them
  • Exposure to LLM APIs (e.g., OpenAI, Azure OpenAI, Anthropic) and familiarity with one orchestration framework (LangChain, LlamaIndex, or similar)
  • Git basics - comfortable with branching and reviewing code as part of a team workflow
  • Experience building or designing semantic layers, knowledge graphs, or context retrieval systems (RAG pipelines); understands embedding models, chunking strategies, and retrieval optimization
  • Ability to reason about data quality, duplication, and freshness - especially critical for training and maintaining embeddings
  • Fetch and ingest data from live external sources (APIs and web sources) into a structured pipeline - handling formats, rate limits, and failures, not just calling a pre-built connector
  • Design and build the schema and storage layer your pipeline writes to, in collaboration with data engineering, comfortable with vector storage or semantic indexing patterns
  • Build and test AI components - embeddings, retrieval logic, prompts - using established patterns and tools; iterate on context relevance and accuracy
  • Communicates technology-related updates and requirements to other departments and contributes to presentations for senior management
  • Collaborates with research and development teams, product management, and strategic analysts to support ongoing projects
  • Partners with Business Units to provide reliable intelligence, validated technology options, and insights on enterprise and industry trends
  • Supports technology project teams by coordinating specific tasks, assisting with day-to-day operations, and contributing to the successful delivery of solutions
  • Assists in collaborative efforts with academic research teams, practicums, internships, and vendor POCs
  • Supports vendor evaluations and contributes to collaborations with vendors
  • Assists as a liaison between academic institutions, professional organizations, and research groups as needed
  • Support the Academic Corporate Engagement efforts to develop research and educational talent, with a focus on enhancing health tech knowledge and skills
  • Complete all responsibilities and goals outlined in the internship program
  • Complete all special projects and other duties as assigned
  • Must be able to perform duties with or without reasonable accommodation
  • Communicating with others to exchange information
  • Assessing the accuracy, neatness, and thoroughness of the work assigned
  • Remaining in a stationary position, often standing or sitting for prolonged periods
  • Repeating motions that may include the wrists, hands, and/or fingers
  • Must be able to provide a dedicated, secure work area
  • Must be able to provide high-speed internet access/connectivity and office setup and maintenance
  • No adverse environmental conditions are expected

Skills

  • SQL fundamentals - comfortable writing queries against structured data, not just describing them
  • Exposure to LLM APIs (e.g., OpenAI, Azure OpenAI, Anthropic) and familiarity with one orchestration framework (LangChain, LlamaIndex, or similar)
  • Git basics - comfortable with branching and reviewing code as part of a team workflow
  • Experience building or designing semantic layers, knowledge graphs, or context retrieval systems (RAG pipelines); understands embedding models, chunking strategies, and retrieval optimization
  • Ability to reason about data quality, duplication, and freshness - especially critical for training and maintaining embeddings
  • Fetch and ingest data from live external sources (APIs and web sources) into a structured pipeline - handling formats, rate limits, and failures, not just calling a pre-built connector
  • Design and build the schema and storage layer your pipeline writes to, in collaboration with data engineering, comfortable with vector storage or semantic indexing patterns
  • Build and test AI components - embeddings, retrieval logic, prompts - using established patterns and tools; iterate on context relevance and accuracy
  • ***We are currently looking for interns who can start immediately, 12-week internship, and work 40 hours per week.***
  • Currently pursuing or recently completed an advanced degree in healthcare, technology, or a related field (e.g., Biomedical Informatics, Computer Science) with preference of a PhD
  • Demonstrated interest or experience in AI, healthcare technology, or informatics research
  • Strong foundational knowledge in generative AI model development, architectures, and vector databases
  • Ability to work collaboratively and communicate effectively with cross-functional teams
  • Communicating with others to exchange information
  • Assessing the accuracy, neatness, and thoroughness of the work assigned
  • Remaining in a stationary position, often standing or sitting for prolonged periods
  • Repeating motions that may include the wrists, hands, and/or fingers
  • Must be able to provide a dedicated, secure work area
  • Must be able to provide high-speed internet access/connectivity and office setup and maintenance
  • No adverse environmental conditions are expected
  • Hands-on experience with working with Machine Learning and Deep Learning models. Experience with LLM/RAG models and LLM fine-tuning is a plus
  • Hands-on experience working with cloud services (AWS/Azure), large data sets, and building data pipelines for ML solutions. Experience with vector embeddings and databases is a plus

Qualifications

Must Haves

  • SQL fundamentals - comfortable writing queries against structured data, not just describing them
  • Exposure to LLM APIs (e.g., OpenAI, Azure OpenAI, Anthropic) and familiarity with one orchestration framework (LangChain, LlamaIndex, or similar)
  • Git basics - comfortable with branching and reviewing code as part of a team workflow
  • Experience building or designing semantic layers, knowledge graphs, or context retrieval systems (RAG pipelines); understands embedding models, chunking strategies, and retrieval optimization
  • Ability to reason about data quality, duplication, and freshness - especially critical for training and maintaining embeddings
  • Fetch and ingest data from live external sources (APIs and web sources) into a structured pipeline - handling formats, rate limits, and failures, not just calling a pre-built connector
  • Design and build the schema and storage layer your pipeline writes to, in collaboration with data engineering, comfortable with vector storage or semantic indexing patterns
  • Build and test AI components - embeddings, retrieval logic, prompts - using established patterns and tools; iterate on context relevance and accuracy
  • ***We are currently looking for interns who can start immediately, 12-week internship, and work 40 hours per week.***
  • Currently pursuing or recently completed an advanced degree in healthcare, technology, or a related field (e.g., Biomedical Informatics, Computer Science) with preference of a PhD
  • Demonstrated interest or experience in AI, healthcare technology, or informatics research
  • Strong foundational knowledge in generative AI model development, architectures, and vector databases
  • Ability to work collaboratively and communicate effectively with cross-functional teams
  • Communicating with others to exchange information
  • Assessing the accuracy, neatness, and thoroughness of the work assigned
  • Remaining in a stationary position, often standing or sitting for prolonged periods
  • Repeating motions that may include the wrists, hands, and/or fingers
  • Must be able to provide a dedicated, secure work area
  • Must be able to provide high-speed internet access/connectivity and office setup and maintenance
  • No adverse environmental conditions are expected

Nice to Haves

  • Hands-on experience with working with Machine Learning and Deep Learning models. Experience with LLM/RAG models and LLM fine-tuning is a plus
  • Hands-on experience working with cloud services (AWS/Azure), large data sets, and building data pipelines for ML solutions. Experience with vector embeddings and databases is a plus

Benefits

  • US-Remote work arrangement
  • Nonexempt employees are eligible to receive overtime pay for hours worked in excess of 40 hours in a given week, or as otherwise required by applicable state law.

More jobs like this