Iambic logo
Iambic
Posted 31 days agoVerified live 1d ago

Software Engineer — Agentic data pipelines

Brief overview

Remote
MastersOr in progress
$110k–$162k/yrStated range
2+ yrsMinimum
PythonLLM APIsAgentic SystemsBiomedical Data EngineeringETL DesignData ValidationPython Testing FrameworksAWSDockerKubernetesData Curation for ML TrainingAgent Security

About the company

Iambic logo
Iambiciambic.ai

Iambic is disrupting the therapeutics landscape with its unique AI-driven drug-discovery platform.

Job description

Summary

Iambic Therapeutics is a clinical-stage life-science and technology company developing medicines through AI-driven discovery technologies. The Software Engineer will build agentic systems and LLM-based data pipelines that acquire, curate, validate, and document large-scale biomedical datasets for the Enchant multimodal transformer model, while maintaining reliable and secure production operations.

Responsibilities

  • Design, build, and maintain agentic systems that turn a pointer to a biomedical data source (web page, S3 path, GitHub repository, a table in a paper) into a reviewed, versioned dataset. The agent's output is committed pipeline code and its test suite, so re-running it later reproduces the same dataset. Develop LLM-based pipelines for data cleaning, normalization, and formatting across diverse data modalities (e.g., molecular, genomic, clinical, literature)
  • Implement automated quality-control workflows that detect anomalies, flag inconsistencies, and enforce data standards
  • Evaluate and iterate on agent architectures, prompting strategies, tool definitions, validation loops, and evaluation harnesses that make agent-generated code trustworthy, improving reliability and throughput over time
  • Collaborate with ML scientists on the Enchant team to understand data requirements and translate them into scalable acquisition and processing systems
  • Monitor and maintain distributed data pipelines in production, diagnosing failures and improving robustness over time
  • Document data provenance, processing decisions, and quality metrics to support reproducibility and auditing
  • Operate the agents safely with sandboxed execution, least-privilege credentials, restricted network access, audit logs, and raising potential security risks to the team

Skills

  • • Master's degree in a computational STEM field, or a Bachelor's with 2+ years of industry experience
  • • Strong Python engineering skills, including experience building and maintaining production-quality software
  • • Hands-on experience with LLM APIs (e.g., Claude, GPT) and agentic patterns such as tool use, orchestration, and multi-step reasoning
  • • Familiarity with biomedical or chemical data sources and formats (e.g., PDB, UniProt, ChEMBL, SDF/MOL, FASTA, or similar)
  • • Comfort with data engineering fundamentals: ETL design, data validation, and working with structured and unstructured data at scale
  • • Hands-on experience with Python testing frameworks (e.g., pytest fixtures, parametrization)
  • • Experience with agent orchestration frameworks, and with evaluation harnesses for LLM-generated code
  • • Familiarity with cloud infrastructure and workflow orchestration (e.g., AWS, Docker, Kubernetes)
  • • Knowledge of multimodal biomedical data—spanning small molecules, proteins, assays, images, ‘omics, and/or clinical records
  • • Experience with large-scale dataset construction or curation for ML model training
  • • Knowledge of agent security practices: sandboxing, scoped credentials, prompt injection
  • • Interest in a longer-term project: a natural language orchestrator that lets drug prosecution team members request inference, fine-tuning, virtual screens, and dataset analysis without writing code using our internal tools

Qualifications

Must Haves

  • • Master's degree in a computational STEM field, or a Bachelor's with 2+ years of industry experience
  • • Strong Python engineering skills, including experience building and maintaining production-quality software
  • • Hands-on experience with LLM APIs (e.g., Claude, GPT) and agentic patterns such as tool use, orchestration, and multi-step reasoning
  • • Familiarity with biomedical or chemical data sources and formats (e.g., PDB, UniProt, ChEMBL, SDF/MOL, FASTA, or similar)
  • • Comfort with data engineering fundamentals: ETL design, data validation, and working with structured and unstructured data at scale
  • • Hands-on experience with Python testing frameworks (e.g., pytest fixtures, parametrization)

Nice to Haves

  • • Experience with agent orchestration frameworks, and with evaluation harnesses for LLM-generated code
  • • Familiarity with cloud infrastructure and workflow orchestration (e.g., AWS, Docker, Kubernetes)
  • • Knowledge of multimodal biomedical data—spanning small molecules, proteins, assays, images, ‘omics, and/or clinical records
  • • Experience with large-scale dataset construction or curation for ML model training
  • • Knowledge of agent security practices: sandboxing, scoped credentials, prompt injection
  • • Interest in a longer-term project: a natural language orchestrator that lets drug prosecution team members request inference, fine-tuning, virtual screens, and dataset analysis without writing code using our internal tools

Benefits

  • Company-paid healthcare
  • Flexible spending accounts
  • Voluntary life insurance
  • 401K matching
  • Uncapped vacation
  • Onsite gym
  • Dining at the San Diego facility

More jobs like this