Pathway logo
Pathway
Posted 13 days agoVerified live 1d ago

AI Benchmark & Datasets Engineer / Researcher

Brief overview

Remote
MastersOr in progress
6+ yrsMinimum
Benchmark DesignData ScienceLarge Language Model TrainingEvaluation ProtocolsData Quality

Job description

Summary

Pathway is developing post-transformer AI models focused on continuous learning, long-context reasoning, and real-time adaptation. The AI Benchmark & Datasets Engineer / Researcher will design and execute rigorous benchmarks, curate datasets, build evaluation infrastructure, and establish standards that guide model development and communicate model performance to customers and the market.

Responsibilities

  • Proactively identify, prioritize, and curate relevant public and client-driven benchmarks across our target use cases and markets
  • Evaluate candidate benchmarks for clarity, data quality, evaluation methodology, and fit with our model roadmap
  • Run benchmarks with baseline models to validate setup, uncover edge cases, and de‑risk R&D runs
  • Hand off “benchmark-ready” packages to R&D (specs, data, evaluation scripts, expected metrics, constraints)
  • Maintain a shared vocabulary and documentation around benchmarks, datasets, and evaluation formats that GTM and R&D can both use
  • Track and organize benchmark results, model leaderboards, and “what good looks like” for different customers and scenarios
  • Contribute to demos and public‑facing proof points based on benchmark outcomes

Skills

  • Candidates based anywhere in the EU, UK, United States, and Canada will be considered
  • 1. You have published at least one paper at NeurIPS, ICLR, or ICML - where you were the lead author or made significant conceptual & code contributions
  • 2. You have significantly contributed to an LLM training effort which became newsworthy (topped a Hugging Face benchmark, best in class model, etc.), preferably using multiple GPU's
  • 3. You have spent at least 6 months working in a leading Machine Learning research center (e.g. at: Google Brain / Deepmind, Apple, Meta, Anthropic, Nvidia, MILA)
  • 4. You were an ICPC World Finalist, or an IOI, IMO, or IPhO medalist in High School
  • Have **experience with ML/LLM evaluation**, data science, or technical product roles, ideally around **benchmarks** or experimentation
  • Are comfortable **reading papers**, **leaderboards**, and **Github** repos, and turning them into clear, r**epeatable benchmark** specs
  • Can talk comfortably with both engineers and customers, and **translate between technical detail and business value**
  • Care about **high‑quality data**, **reproducible** experiments, and crisp **documentation**
  • Are **respectful** of others
  • Are **fluent** in English
  • Published or open‑sourced work on LLM evaluation, benchmarking or data quality
  • Experience designing custom benchmarks or evaluation protocols for novel model capabilities

Qualifications

Must Haves

  • Candidates based anywhere in the EU, UK, United States, and Canada will be considered
  • 1. You have published at least one paper at NeurIPS, ICLR, or ICML - where you were the lead author or made significant conceptual & code contributions
  • 2. You have significantly contributed to an LLM training effort which became newsworthy (topped a Hugging Face benchmark, best in class model, etc.), preferably using multiple GPU's
  • 3. You have spent at least 6 months working in a leading Machine Learning research center (e.g. at: Google Brain / Deepmind, Apple, Meta, Anthropic, Nvidia, MILA)
  • 4. You were an ICPC World Finalist, or an IOI, IMO, or IPhO medalist in High School
  • Have **experience with ML/LLM evaluation**, data science, or technical product roles, ideally around **benchmarks** or experimentation
  • Are comfortable **reading papers**, **leaderboards**, and **Github** repos, and turning them into clear, r**epeatable benchmark** specs
  • Can talk comfortably with both engineers and customers, and **translate between technical detail and business value**
  • Care about **high‑quality data**, **reproducible** experiments, and crisp **documentation**
  • Are **respectful** of others
  • Are **fluent** in English

Nice to Haves

  • Published or open‑sourced work on LLM evaluation, benchmarking or data quality
  • Experience designing custom benchmarks or evaluation protocols for novel model capabilities

Benefits

  • Remote work.
  • Possibility to work or meet with other team members in one of our offices: Palo Alto, CA; Paris, France or Wroclaw, Poland.
  • Join an intellectually stimulating work environment.
  • Be a pioneer: you get to work with a new type of "Live AI" challenges around long sequences and changing data.
  • Be part of one of an early-stage AI startup that believes in impactful research and foundational changes.

More jobs like this