Bigbear.ai logo
Bigbear.ai
Posted 9 days agoVerified live 2d ago

QA Automation Engineer

Brief overview

Remote
$96k–$143k/yrStated range
4+ yrsMinimum
1 green cardsCertified filings
Clearance requiredU.S. government
TypeScriptJavaScriptAPI TestingCI/CD IntegrationGitAuthentication and Authorization TestingRisk-Based TestingLLM Application TestingPythonSQLWritten and Verbal Communication

About the company

Bigbear.ai logo
Bigbear.aibigbear.ai

Provides AI-powered decision intelligence software for national security.

Visa sponsorship history

1 year sponsoring, last filed FY2023

Data powered by U.S. Department of Labor. This does not guarantee sponsorship for this specific role.
1PERM certified
Green Card (PERM) FilingsCertified green card filings: a long-term commitment to international hires.
20231
Sponsored employees from
Canada

Job description

Summary

BigBear.ai provides AI-powered decision intelligence solutions for national security, supply chain management, and digital identity. The QA Automation Engineer will translate product requirements and customer workflows into risk-based test coverage, build and maintain Playwright automation, and evaluate conversational AI behavior. The role partners with Product and Engineering to assess release readiness across browser, application, backend, security, and compliance surfaces.

Responsibilities

  • Translate requirements into testable outcomes. Understand product goals, customer use cases, and complex feature interactions; identify ambiguities and define acceptance criteria with Product and Engineering
  • Own automated test coverage. Design, implement, and maintain Playwright tests covering critical user journeys, feature functionality, and regressions, supported by API and integration testing
  • Develop a risk-based testing strategy. Prioritize coverage according to customer impact, feature dependencies, and the product roadmap; adapt deliberately as priorities change
  • Test conversational AI workflows. Validate streaming responses, conversation history, file uploads, document retrieval and citations, model selection, and tool execution, including interruptions, timeouts, and partial failures
  • Evaluate AI response quality. Build representative evaluation datasets and scoring criteria for accuracy, grounding, instruction following, and appropriate handling of unsafe requests. Account for natural variation in model responses
  • Make automation reliable and useful. Integrate tests into CI/CD, investigate flaky tests, maintain isolated test data, and provide actionable failure diagnostics
  • Communicate release readiness. Report defects with reproducible evidence, customer impact, and severity; explain coverage gaps and residual risks before UAT and release
  • Preserve decision history. Document expected behavior, approved changes, and the rationale behind testing decisions so the team can distinguish intended changes from regressions
  • Cover platform and AI infrastructure surfaces. Extend automation across backend APIs, authentication flows, billing and token behavior, model routing, AI workflow execution, file parsing, MCP/tool execution, agentic harnesses, and passthrough APIs, including provider-facing API compatibility
  • Control test cost and execution footprint. Design AI workflow coverage that minimizes unnecessary token usage, external provider calls, latency, and execution cost without sacrificing signal
  • Automate security-sensitive validation. Build repeatable coverage for user isolation, permission boundaries, input validation, sanitization, rate limits, replay prevention, safe error handling, and layered control behavior
  • Apply AI-assisted testing tools responsibly. Use AI assistance to accelerate test design, generation, triage, and maintenance while critically reviewing generated tests, assertions, and proposed repairs against reliability, reviewability, and deterministic validation standards
  • Scale the automation footprint. Maintain test infrastructure, fixtures, and test data management that must grow with an expanding product surface area and an active engineering team

Skills

  • All applicants must currently reside in the United States
  • • 4+years of experience QA automation engineering experience
  • • Demonstrated experience building and maintaining automated tests with Playwright, including fixtures, resilient locators, assertions, network handling, and trace-based debugging
  • • Strong coding ability in TypeScript or JavaScript, with experience writing maintainable test code and reviewing changes through Git
  • • Experience with API testing, CI/CD integration, test isolation, and diagnosing failures across browser, application, and backend boundaries
  • • Ability to reason about complex requirements, explore edge cases, and balance testing depth with delivery priorities
  • • Clear written and verbal communication: explaining defects, uncertainty, tradeoffs, and release risk to technical and nontechnical stakeholders. Experience testing authentication, authorization, permissions, and separation of customer data
  • • Familiarity with validating Defense in Depth behavior: confirming that multiple layers of controls work together, without this being a dedicated DevSec role
  • • Strong troubleshooting and analytical skills, with the ability to work independently and as part of a team
  • • Ability to obtain a Department of Defense Secret clearance
  • • Experience testing LLM applications, retrieval-augmented generation, or AI agents
  • • Experience with Python for API testing, test utilities, test data generation, or AI evaluation workflows
  • • Experience with enterprise or government platforms and auditable test evidence
  • • Experience testing model gateways, MCP tools, tool-using systems, or workflow automation and no-code/low-code platforms (e.g., Power Automate, Zapier, Make, n8n)
  • • Familiarity with flagship Generative AI provider APIs (Google Vertex AI, AWS Bedrock, Microsoft Azure OpenAI) and models (OpenAI GPT, Anthropic Claude, Google Gemini)
  • • Experience with CI/CD pipelines (e.g., GitHub Actions), Docker, Kubernetes, and observability/monitoring tooling
  • • Experience with PostgreSQL and SQL for test data setup, teardown, and validation
  • • Knowledge of government compliance frameworks (FedRAMP, NIST AI RMF, CMMC 2.0)
  • • Active DoD security clearance at the Secret level or above

Qualifications

Must Haves

  • All applicants must currently reside in the United States
  • • 4+years of experience QA automation engineering experience
  • • Demonstrated experience building and maintaining automated tests with Playwright, including fixtures, resilient locators, assertions, network handling, and trace-based debugging
  • • Strong coding ability in TypeScript or JavaScript, with experience writing maintainable test code and reviewing changes through Git
  • • Experience with API testing, CI/CD integration, test isolation, and diagnosing failures across browser, application, and backend boundaries
  • • Ability to reason about complex requirements, explore edge cases, and balance testing depth with delivery priorities
  • • Clear written and verbal communication: explaining defects, uncertainty, tradeoffs, and release risk to technical and nontechnical stakeholders. Experience testing authentication, authorization, permissions, and separation of customer data
  • • Familiarity with validating Defense in Depth behavior: confirming that multiple layers of controls work together, without this being a dedicated DevSec role
  • • Strong troubleshooting and analytical skills, with the ability to work independently and as part of a team
  • • Ability to obtain a Department of Defense Secret clearance

Nice to Haves

  • • Experience testing LLM applications, retrieval-augmented generation, or AI agents
  • • Experience with Python for API testing, test utilities, test data generation, or AI evaluation workflows
  • • Experience with enterprise or government platforms and auditable test evidence
  • • Experience testing model gateways, MCP tools, tool-using systems, or workflow automation and no-code/low-code platforms (e.g., Power Automate, Zapier, Make, n8n)
  • • Familiarity with flagship Generative AI provider APIs (Google Vertex AI, AWS Bedrock, Microsoft Azure OpenAI) and models (OpenAI GPT, Anthropic Claude, Google Gemini)
  • • Experience with CI/CD pipelines (e.g., GitHub Actions), Docker, Kubernetes, and observability/monitoring tooling
  • • Experience with PostgreSQL and SQL for test data setup, teardown, and validation
  • • Knowledge of government compliance frameworks (FedRAMP, NIST AI RMF, CMMC 2.0)
  • • Active DoD security clearance at the Secret level or above

More jobs like this