Summary
Weekday AI is a pioneering company focused on building the next generation of evaluation benchmarks for frontier AI models. They are seeking experienced QA and Test Engineers to ensure every benchmark is reliable and accurately measures real AI capabilities.
Responsibilities
- Design comprehensive test cases that validate evaluation tasks, including complex edge cases and unexpected scenarios
- Review benchmark tasks and reference solutions to identify ambiguity, inconsistencies, missing requirements, and grading gaps
- Debug task environments and Python-based evaluation scripts to ensure reliable execution and accurate results
- Develop repeatable quality assurance processes, validation checklists, and testing frameworks for benchmark creation
- Identify potential shortcuts, exploits, or weaknesses that could compromise evaluation accuracy or benchmark integrity
- Collaborate with AI researchers, engineers, and task authors to improve task quality, reproducibility, and technical rigor
Skills
- Master's degree, PhD, or equivalent practical experience in a STEM discipline involving software engineering, research, or advanced technical problem solving
- Minimum 1 year of professional experience in Quality Assurance, Test Engineering, Software Engineering, Research Engineering, or a related technical field with strong quality ownership
- Proven experience designing test cases, validating complex software systems, and debugging end-to-end workflows
- Strong proficiency in Python and Git, with the ability to troubleshoot unfamiliar codebases and technical environments
- Excellent analytical thinking, problem-solving skills, and exceptional attention to detail
- Experience documenting bugs, test strategies, and technical findings with clear written communication
- Experience evaluating AI systems, machine learning models, or AI-generated outputs is preferred
- Ability to work independently while managing multiple complex tasks with minimal supervision
- Ability to commit approximately 35 hours per week on a consistent basis
- Experience with AI evaluation, benchmark development, or quality assurance for machine learning systems
- Background in automation testing, validation frameworks, or software quality engineering
- Familiarity with large language models, AI agent workflows, or evaluation pipelines
- Experience creating repeatable QA processes for research or engineering projects
Qualifications
Must Haves
- Master's degree, PhD, or equivalent practical experience in a STEM discipline involving software engineering, research, or advanced technical problem solving
- Minimum 1 year of professional experience in Quality Assurance, Test Engineering, Software Engineering, Research Engineering, or a related technical field with strong quality ownership
- Proven experience designing test cases, validating complex software systems, and debugging end-to-end workflows
- Strong proficiency in Python and Git, with the ability to troubleshoot unfamiliar codebases and technical environments
- Excellent analytical thinking, problem-solving skills, and exceptional attention to detail
- Experience documenting bugs, test strategies, and technical findings with clear written communication
- Experience evaluating AI systems, machine learning models, or AI-generated outputs is preferred
- Ability to work independently while managing multiple complex tasks with minimal supervision
- Ability to commit approximately 35 hours per week on a consistent basis
Nice to Haves
- Experience with AI evaluation, benchmark development, or quality assurance for machine learning systems
- Background in automation testing, validation frameworks, or software quality engineering
- Familiarity with large language models, AI agent workflows, or evaluation pipelines
- Experience creating repeatable QA processes for research or engineering projects
Benefits
- Fully remote engagement
- Flexible working hours
- Payments are issued weekly based on approved work completed