Office Hours logo
Office Hours
Posted 48 days agoVerified live 2d ago

Software Engineer- Benchmarking

Brief overview

Remote
$160k–$210k/yrStated range
4+ yrsMinimum
PythonDataset Preparation and MaintenanceAI Model EvaluationDockerEvaluation PipelinesReproducible Execution EnvironmentsReactNext.jsTailwindPyTorchHugging FaceMachine Learning

About the company

Office Hours logo
Office Hoursofficehours.com

Get paid to share what you know.

Job description

Summary

Office Hours is an on-demand expert network connecting organizations with trusted experts across knowledge domains. The company is seeking a Software Engineer to build and operate the platform supporting AI model evaluations, including benchmark datasets, reproducible evaluation pipelines, containerized environments, analysis tools, and published scoreboards and leaderboards.

Responsibilities

  • Prepare and maintain benchmark datasets: Own the data work behind our benchmarks, including cleaning, preparation, conversion into runnable formats, and ongoing maintenance. Validate that tasks are complete, consistent, and executable, and flag ambiguities that would compromise results
  • Build and maintain evaluation pipelines: Build the infrastructure that runs evaluations consistently across model APIs and terminal agents, so results are reproducible and comparable
  • Build evaluation environments: Create lightweight, containerized environments and viewers for tasking and for evaluating model performance on tool use
  • Support model experiments: Help fine-tune small open-source LLMs and compare baseline against post-training performance
  • Develop the scoreboard and leaderboard: Build the published views of our results, including model-level, benchmark-level, task-level, domain-level, and rubric-level performance
  • Build analysis tools: Make it easy to identify recurring failure modes, compare models and agent scaffolds fairly, and track capability improvements and regressions over time
  • Collaborate: Work closely with researchers and engineers to make sure evaluation data and outputs are accurate, consistent, and well integrated into what we publish
  • Build tooling for data creation and review: Support expert annotation and data-generation projects by building lightweight HTML viewers and internal tools for task authoring, review, quality control, and structured data collection

Skills

  • 4+ years of professional experience building and maintaining complex systems, with strong Python
  • You write robust, maintainable code and are comfortable diving deep into existing codebases and infrastructure
  • Experience preparing, cleaning, and maintaining datasets, and the care to make sure two results are genuinely comparable
  • Experience with Docker and building reproducible execution environments
  • You work well alongside researchers and scientists and can translate their methodology into working systems
  • Hands-on experience running AI evaluations, or with frameworks like Harbor, Terminal-Bench, or Inspect, is a strong plus
  • Experience fine-tuning or post-training open-source LLMs, or other hands-on machine learning work
  • Experience with agentic, multi-turn, long-context, or tool-use evaluation
  • Experience validating LLM-as-judge or rubric-based grading setups
  • Background or strong interest in a scientific or technical domain
  • Experience building data-heavy dashboards, leaderboards, or visualizations
  • Open-source contributions or published work related to benchmarks and measurement

Qualifications

Must Haves

  • 4+ years of professional experience building and maintaining complex systems, with strong Python
  • You write robust, maintainable code and are comfortable diving deep into existing codebases and infrastructure
  • Experience preparing, cleaning, and maintaining datasets, and the care to make sure two results are genuinely comparable
  • Experience with Docker and building reproducible execution environments
  • You work well alongside researchers and scientists and can translate their methodology into working systems

Nice to Haves

  • Hands-on experience running AI evaluations, or with frameworks like Harbor, Terminal-Bench, or Inspect, is a strong plus
  • Experience fine-tuning or post-training open-source LLMs, or other hands-on machine learning work
  • Experience with agentic, multi-turn, long-context, or tool-use evaluation
  • Experience validating LLM-as-judge or rubric-based grading setups
  • Background or strong interest in a scientific or technical domain
  • Experience building data-heavy dashboards, leaderboards, or visualizations
  • Open-source contributions or published work related to benchmarks and measurement

Benefits

  • Equity
  • Medical, dental, and vision coverage
  • 401(k)
  • Monthly wellness and fitness stipend
  • Paid time off policy, along with company holidays
  • Annual company off-sites (Tahoe, Mendocino, Mexico City, San Diego, Park City)
  • Parent-friendly policies
  • Remote flexibility
  • Paid family leave

More jobs like this