Protege logo
Protege
Posted 30 days agoVerified live 1d ago

Applied Healthcare Researcher

Brief overview

Remote
Machine LearningLarge Language Model SystemsHealthcare DataPythonSQLEvaluation DesignData Quality AssessmentDataset Representativeness AnalysisTechnical Stakeholder CommunicationResearch Planning

About the company

A data-driven technology company delivering AI training data solutions.

Job description

Summary

Protege is building a secure, efficient, and privacy-centric platform for exchanging AI training data. The Applied Healthcare Researcher will partner with AI researchers to identify and validate healthcare data strategies, develop applied research methods, evaluate dataset feasibility, and create reusable research workflows while collaborating across Solutions, Data Partnerships, Product, and Engineering.

Responsibilities

  • Serve as the primary technical and research point of contact for healthcare customer conversations
  • Translate a lab's model-development goals into concrete, feasible data strategies
  • Help customers scope opportunities and identify the highest-value data available to them
  • Explain data limitations, tradeoffs, and potential biases to technically sophisticated stakeholders while grounding conversations in what real-world data actually looks like
  • After delivery, answer the research questions customers raise about the data we provided. Delivery is not the end of the relationship
  • Develop and evaluate methods — fine-tuning, LLM-based extraction, classification, rules-based approaches, or whatever the problem calls for — to demonstrate that a dataset can support a customer's training or evaluation objective
  • Design and run feasibility research pre-contract: can this data support this model objective, at what quality, with what caveats
  • Build the evidence base that makes a data strategy credible — benchmarks, validation analyses, error characterization, and honest assessments of where the data falls short
  • Partner with the Assessments team on healthcare benchmarks across modalities
  • Evaluate whether requested variables, labels, or cohort definitions are achievable with available healthcare data
  • Identify proxy variables or alternative dataset structures when the ideal variable doesn't exist
  • Analyze partner and source datasets — schema, field availability, quality, completeness, and required transformations
  • Contribute to our point of view on which healthcare data matters most for which modality and which stage of model development
  • Help evaluate new data partners and identify datasets worth acquiring before a customer asks for them
  • Produce reusable research, evidence, and technical collateral rather than starting from scratch for each opportunity
  • Identify where a successful one-off approach should become a repeatable workflow, and work with Product and Engineering to operationalize it
  • Help expand proven healthcare datasets across multiple customers instead of selling them once
  • Work with Solutions and FDEs from the beginning of an opportunity
  • Coordinate with Healthcare Data Partnerships on sourcing and with Product and Engineering on tooling

Skills

  • Advanced degree (PhD or Master's plus 2+ years industry experience) in machine learning, computer science, biomedical informatics, epidemiology, statistics, or a related quantitative field — or equivalent applied experience
  • Hands-on experience building and evaluating ML or LLM-based systems for extraction, classification, or prediction on real-world data
  • Experience working with healthcare data: claims, EMR/EHR, clinical notes, imaging, registries, or similar. You understand why real-world clinical data is messy and what that means for model training
  • Strong Python and SQL, with the ability to work independently against large datasets
  • Experience designing evaluations — measuring data quality and dataset representativeness
  • Demonstrated ability to work directly with technical stakeholders and translate ambiguous goals into concrete, defensible research plans
  • Comfort operating on a customer's timeline without lowering the standard of the research

Qualifications

Must Haves

  • Advanced degree (PhD or Master's plus 2+ years industry experience) in machine learning, computer science, biomedical informatics, epidemiology, statistics, or a related quantitative field — or equivalent applied experience
  • Hands-on experience building and evaluating ML or LLM-based systems for extraction, classification, or prediction on real-world data
  • Experience working with healthcare data: claims, EMR/EHR, clinical notes, imaging, registries, or similar. You understand why real-world clinical data is messy and what that means for model training
  • Strong Python and SQL, with the ability to work independently against large datasets
  • Experience designing evaluations — measuring data quality and dataset representativeness
  • Demonstrated ability to work directly with technical stakeholders and translate ambiguous goals into concrete, defensible research plans
  • Comfort operating on a customer's timeline without lowering the standard of the research

More jobs like this