Summary
Pathway is developing post-transformer AI models focused on continuous learning, long-context reasoning, and real-time adaptation. The AI Benchmark & Datasets Engineer / Researcher will design and execute rigorous benchmarks, curate datasets, build evaluation infrastructure, and establish standards that guide model development and communicate model performance to customers and the market.
Responsibilities
- Proactively identify, prioritize, and curate relevant public and client-driven benchmarks across our target use cases and markets
- Evaluate candidate benchmarks for clarity, data quality, evaluation methodology, and fit with our model roadmap
- Run benchmarks with baseline models to validate setup, uncover edge cases, and de‑risk R&D runs
- Hand off “benchmark-ready” packages to R&D (specs, data, evaluation scripts, expected metrics, constraints)
- Maintain a shared vocabulary and documentation around benchmarks, datasets, and evaluation formats that GTM and R&D can both use
- Track and organize benchmark results, model leaderboards, and “what good looks like” for different customers and scenarios
- Contribute to demos and public‑facing proof points based on benchmark outcomes
Skills
- Candidates based anywhere in the EU, UK, United States, and Canada will be considered
- 1. You have published at least one paper at NeurIPS, ICLR, or ICML - where you were the lead author or made significant conceptual & code contributions
- 2. You have significantly contributed to an LLM training effort which became newsworthy (topped a Hugging Face benchmark, best in class model, etc.), preferably using multiple GPU's
- 3. You have spent at least 6 months working in a leading Machine Learning research center (e.g. at: Google Brain / Deepmind, Apple, Meta, Anthropic, Nvidia, MILA)
- 4. You were an ICPC World Finalist, or an IOI, IMO, or IPhO medalist in High School
- Have **experience with ML/LLM evaluation**, data science, or technical product roles, ideally around **benchmarks** or experimentation
- Are comfortable **reading papers**, **leaderboards**, and **Github** repos, and turning them into clear, r**epeatable benchmark** specs
- Can talk comfortably with both engineers and customers, and **translate between technical detail and business value**
- Care about **high‑quality data**, **reproducible** experiments, and crisp **documentation**
- Are **respectful** of others
- Are **fluent** in English
- Published or open‑sourced work on LLM evaluation, benchmarking or data quality
- Experience designing custom benchmarks or evaluation protocols for novel model capabilities
Qualifications
Must Haves
- Candidates based anywhere in the EU, UK, United States, and Canada will be considered
- 1. You have published at least one paper at NeurIPS, ICLR, or ICML - where you were the lead author or made significant conceptual & code contributions
- 2. You have significantly contributed to an LLM training effort which became newsworthy (topped a Hugging Face benchmark, best in class model, etc.), preferably using multiple GPU's
- 3. You have spent at least 6 months working in a leading Machine Learning research center (e.g. at: Google Brain / Deepmind, Apple, Meta, Anthropic, Nvidia, MILA)
- 4. You were an ICPC World Finalist, or an IOI, IMO, or IPhO medalist in High School
- Have **experience with ML/LLM evaluation**, data science, or technical product roles, ideally around **benchmarks** or experimentation
- Are comfortable **reading papers**, **leaderboards**, and **Github** repos, and turning them into clear, r**epeatable benchmark** specs
- Can talk comfortably with both engineers and customers, and **translate between technical detail and business value**
- Care about **high‑quality data**, **reproducible** experiments, and crisp **documentation**
- Are **respectful** of others
- Are **fluent** in English
Nice to Haves
- Published or open‑sourced work on LLM evaluation, benchmarking or data quality
- Experience designing custom benchmarks or evaluation protocols for novel model capabilities
Benefits
- Remote work.
- Possibility to work or meet with other team members in one of our offices: Palo Alto, CA; Paris, France or Wroclaw, Poland.
- Join an intellectually stimulating work environment.
- Be a pioneer: you get to work with a new type of "Live AI" challenges around long sequences and changing data.
- Be part of one of an early-stage AI startup that believes in impactful research and foundational changes.