Weekday AI (YC W21) logo
Weekday AI (YC W21)
Posted 51 days agoVerified live 2d ago

QA/Test Engineer

Brief overview

Remote
MastersOr in progress
$60–$90/hrStated range
1+ yrsMinimum
Quality AssuranceTest EngineeringPythonGitDebuggingTest Case DesignAI EvaluationMachine Learning EvaluationAutomation TestingValidation FrameworksSoftware Quality Engineering

About the company

Weekday AI (YC W21) logo
Weekday AI (YC W21)jobs.weekday.works

We are a YC-backed recruitment startup. Find select jobs posted by premium YC as well as VC backed startups here. Hand-curated by Weekday team.

Job description

Summary

Weekday AI is a pioneering company focused on building the next generation of evaluation benchmarks for frontier AI models. They are seeking experienced QA and Test Engineers to ensure every benchmark is reliable and accurately measures real AI capabilities.

Responsibilities

  • Design comprehensive test cases that validate evaluation tasks, including complex edge cases and unexpected scenarios
  • Review benchmark tasks and reference solutions to identify ambiguity, inconsistencies, missing requirements, and grading gaps
  • Debug task environments and Python-based evaluation scripts to ensure reliable execution and accurate results
  • Develop repeatable quality assurance processes, validation checklists, and testing frameworks for benchmark creation
  • Identify potential shortcuts, exploits, or weaknesses that could compromise evaluation accuracy or benchmark integrity
  • Collaborate with AI researchers, engineers, and task authors to improve task quality, reproducibility, and technical rigor

Skills

  • Master's degree, PhD, or equivalent practical experience in a STEM discipline involving software engineering, research, or advanced technical problem solving
  • Minimum 1 year of professional experience in Quality Assurance, Test Engineering, Software Engineering, Research Engineering, or a related technical field with strong quality ownership
  • Proven experience designing test cases, validating complex software systems, and debugging end-to-end workflows
  • Strong proficiency in Python and Git, with the ability to troubleshoot unfamiliar codebases and technical environments
  • Excellent analytical thinking, problem-solving skills, and exceptional attention to detail
  • Experience documenting bugs, test strategies, and technical findings with clear written communication
  • Experience evaluating AI systems, machine learning models, or AI-generated outputs is preferred
  • Ability to work independently while managing multiple complex tasks with minimal supervision
  • Ability to commit approximately 35 hours per week on a consistent basis
  • Experience with AI evaluation, benchmark development, or quality assurance for machine learning systems
  • Background in automation testing, validation frameworks, or software quality engineering
  • Familiarity with large language models, AI agent workflows, or evaluation pipelines
  • Experience creating repeatable QA processes for research or engineering projects

Qualifications

Must Haves

  • Master's degree, PhD, or equivalent practical experience in a STEM discipline involving software engineering, research, or advanced technical problem solving
  • Minimum 1 year of professional experience in Quality Assurance, Test Engineering, Software Engineering, Research Engineering, or a related technical field with strong quality ownership
  • Proven experience designing test cases, validating complex software systems, and debugging end-to-end workflows
  • Strong proficiency in Python and Git, with the ability to troubleshoot unfamiliar codebases and technical environments
  • Excellent analytical thinking, problem-solving skills, and exceptional attention to detail
  • Experience documenting bugs, test strategies, and technical findings with clear written communication
  • Experience evaluating AI systems, machine learning models, or AI-generated outputs is preferred
  • Ability to work independently while managing multiple complex tasks with minimal supervision
  • Ability to commit approximately 35 hours per week on a consistent basis

Nice to Haves

  • Experience with AI evaluation, benchmark development, or quality assurance for machine learning systems
  • Background in automation testing, validation frameworks, or software quality engineering
  • Familiarity with large language models, AI agent workflows, or evaluation pipelines
  • Experience creating repeatable QA processes for research or engineering projects

Benefits

  • Fully remote engagement
  • Flexible working hours
  • Payments are issued weekly based on approved work completed

More jobs like this