Summary
Iambic Therapeutics is a clinical-stage life-science and technology company developing medicines through AI-driven discovery technologies. The Software Engineer will build agentic systems and LLM-based data pipelines that acquire, curate, validate, and document large-scale biomedical datasets for the Enchant multimodal transformer model, while maintaining reliable and secure production operations.
Responsibilities
- Design, build, and maintain agentic systems that turn a pointer to a biomedical data source (web page, S3 path, GitHub repository, a table in a paper) into a reviewed, versioned dataset. The agent's output is committed pipeline code and its test suite, so re-running it later reproduces the same dataset. Develop LLM-based pipelines for data cleaning, normalization, and formatting across diverse data modalities (e.g., molecular, genomic, clinical, literature)
- Implement automated quality-control workflows that detect anomalies, flag inconsistencies, and enforce data standards
- Evaluate and iterate on agent architectures, prompting strategies, tool definitions, validation loops, and evaluation harnesses that make agent-generated code trustworthy, improving reliability and throughput over time
- Collaborate with ML scientists on the Enchant team to understand data requirements and translate them into scalable acquisition and processing systems
- Monitor and maintain distributed data pipelines in production, diagnosing failures and improving robustness over time
- Document data provenance, processing decisions, and quality metrics to support reproducibility and auditing
- Operate the agents safely with sandboxed execution, least-privilege credentials, restricted network access, audit logs, and raising potential security risks to the team
Skills
- • Master's degree in a computational STEM field, or a Bachelor's with 2+ years of industry experience
- • Strong Python engineering skills, including experience building and maintaining production-quality software
- • Hands-on experience with LLM APIs (e.g., Claude, GPT) and agentic patterns such as tool use, orchestration, and multi-step reasoning
- • Familiarity with biomedical or chemical data sources and formats (e.g., PDB, UniProt, ChEMBL, SDF/MOL, FASTA, or similar)
- • Comfort with data engineering fundamentals: ETL design, data validation, and working with structured and unstructured data at scale
- • Hands-on experience with Python testing frameworks (e.g., pytest fixtures, parametrization)
- • Experience with agent orchestration frameworks, and with evaluation harnesses for LLM-generated code
- • Familiarity with cloud infrastructure and workflow orchestration (e.g., AWS, Docker, Kubernetes)
- • Knowledge of multimodal biomedical data—spanning small molecules, proteins, assays, images, ‘omics, and/or clinical records
- • Experience with large-scale dataset construction or curation for ML model training
- • Knowledge of agent security practices: sandboxing, scoped credentials, prompt injection
- • Interest in a longer-term project: a natural language orchestrator that lets drug prosecution team members request inference, fine-tuning, virtual screens, and dataset analysis without writing code using our internal tools
Qualifications
Must Haves
- • Master's degree in a computational STEM field, or a Bachelor's with 2+ years of industry experience
- • Strong Python engineering skills, including experience building and maintaining production-quality software
- • Hands-on experience with LLM APIs (e.g., Claude, GPT) and agentic patterns such as tool use, orchestration, and multi-step reasoning
- • Familiarity with biomedical or chemical data sources and formats (e.g., PDB, UniProt, ChEMBL, SDF/MOL, FASTA, or similar)
- • Comfort with data engineering fundamentals: ETL design, data validation, and working with structured and unstructured data at scale
- • Hands-on experience with Python testing frameworks (e.g., pytest fixtures, parametrization)
Nice to Haves
- • Experience with agent orchestration frameworks, and with evaluation harnesses for LLM-generated code
- • Familiarity with cloud infrastructure and workflow orchestration (e.g., AWS, Docker, Kubernetes)
- • Knowledge of multimodal biomedical data—spanning small molecules, proteins, assays, images, ‘omics, and/or clinical records
- • Experience with large-scale dataset construction or curation for ML model training
- • Knowledge of agent security practices: sandboxing, scoped credentials, prompt injection
- • Interest in a longer-term project: a natural language orchestrator that lets drug prosecution team members request inference, fine-tuning, virtual screens, and dataset analysis without writing code using our internal tools
Benefits
- Company-paid healthcare
- Flexible spending accounts
- Voluntary life insurance
- 401K matching
- Uncapped vacation
- Onsite gym
- Dining at the San Diego facility