Endava logo
Endava
Posted 48 days agoVerified live 8h ago

AI Ops Engineer

Brief overview

Remote
3+ yrsMinimum
1 H-1B approvalsDept. of Labor
PythonRESTful APIsPrompt EngineeringEmbeddings and Vector RetrievalRetrieval-Augmented Generation (RAG)Agentic AI Frameworks LangChain/LangGraphAgentic AI Frameworks AutoGenAgentic AI Frameworks CrewAIAI Coding Assistants DevinAI Coding Assistants WindsurfServiceNow/CMDBDynatraceZabbixGoogle Cloud Platform (GCP)ITSM and Incident ManagementObservability and MonitoringStructured and Unstructured Data

About the company

Endava is a software development outsourcing company that creates dynamic platforms and intelligent digital experiences for businesses.

Visa sponsorship history

1 year sponsoring, last filed FY2023

Data powered by U.S. Department of Labor. This does not guarantee sponsorship for this specific role.
1H-1B approved
100%approval rate
$146,450median wage / yr
H-1B Petition ApprovalsVisas USCIS actually granted: the strongest sign the company sponsors.
20231
LCA Certified ApplicationsAn early filing step, not a visa approval: it signals intent, not confirmed sponsorship.
20231
Top sponsored roles
Infrastructure Consultant

Job description

Summary

Endava is a technology consultancy that combines engineering expertise, industry knowledge, and a people-centric approach to create digital platforms and experiences. The company is seeking an AI Automation Engineer – Level II to build an enterprise operational AI platform, developing LLM-powered agents, integrations, backend services, RAG pipelines, and progressively autonomous Infrastructure and Operations workflows.

Responsibilities

  • Design, build, and enhance LLM-powered and agentic AI solutions for enterprise Infrastructure & Operations use cases
  • Develop and integrate domain-specific AI agents that collaborate to answer questions, investigate operational issues, and execute defined workflows
  • Build Model Context Protocol (MCP) integrations and tool-calling capabilities that securely connect AI agents with enterprise platforms and APIs
  • Develop backend services and integrations primarily using Python and REST APIs
  • Integrate the AI platform with infrastructure and operational systems such as ServiceNow/CMDB, Dynatrace, Zabbix, and GCP
  • Build pipelines and services that ingest, normalize, enrich, and contextualize structured and unstructured operational data for AI consumption
  • Implement Retrieval-Augmented Generation (RAG) and grounding strategies that provide LLMs with accurate enterprise context
  • Develop initial read-only AI workflows supporting incident triage, incident management, root cause analysis (RCA), infrastructure discovery, and Help Desk automation
  • Progressively extend workflows toward controlled automation, including change-window support, maintenance suppression, proactive outage prevention, and coordinated self-healing
  • Apply appropriate guardrails, validation, access controls, and human-in-the-loop patterns as AI workflows move from recommendations toward autonomous actions
  • Use modern AI-assisted engineering tools, such as Devin, Windsurf, or comparable platforms, to accelerate software development and automation
  • Implement AI observability and evaluation capabilities to measure response quality, confidence, token consumption, reliability, latency, and operational outcomes such as MTTR
  • Collaborate with AI architects, platform engineers, observability teams, IT operations, enterprise search, data, and security teams to deliver production-ready solutions
  • Help improve knowledge quality and cross-validation mechanisms to reduce LLM hallucinations and ensure responses are grounded in authoritative enterprise data
  • Contribute to engineering standards and reusable patterns for deploying secure, scalable, observable, and maintainable enterprise AI systems

Skills

  • 3–5 years of overall relevant engineering experience, combining modern AI engineering with a strong software, integration, data, infrastructure, or AIOps foundation
  • Approximately 1–2 years of hands-on experience with Generative AI/LLMs, including building applications or workflows using modern LLM platforms
  • Approximately 2–3 years of foundational engineering experience in one or more areas such as backend software development, Python engineering, API integration, data engineering, cloud engineering, automation, or AIOps
  • Strong programming skills in Python, including experience developing production-quality backend services and automation
  • Strong experience designing, building, and consuming RESTful APIs and integrating multiple enterprise systems
  • Practical knowledge of LLMs, prompt engineering, context management, embeddings, vector retrieval, and Retrieval-Augmented Generation (RAG)
  • Hands-on experience with agentic or multi-agent AI frameworks, such as LangChain/LangGraph, AutoGen, CrewAI, or comparable technologies
  • Experience with or a strong understanding of Model Context Protocol (MCP), function/tool calling, agent registries, and AI orchestration patterns
  • Experience with modern AI coding assistants or autonomous development tools, such as Devin, Windsurf, or comparable solutions
  • Familiarity with enterprise IT and infrastructure platforms, ideally including one or more of ServiceNow/CMDB, Dynatrace, Zabbix, and GCP
  • Understanding of ITSM, incident management, observability, monitoring, infrastructure telemetry, or AIOps concepts
  • Experience working with both structured and unstructured data and preparing enterprise information for AI consumption
  • Understanding of AI safety, data governance, security, access control, grounding, hallucination mitigation, and responsible AI principles
  • Experience designing or operating production systems where reliability, scalability, observability, and maintainability are important
  • Strong systems-thinking and problem-solving skills, with the ability to understand complex enterprise environments and translate operational requirements into practical technical solutions
  • Ability to collaborate effectively with architects, software engineers, infrastructure teams, IT operations, security, and other technical stakeholders
  • Comfortable working iteratively, delivering measurable value through a phased approach from read-only AI assistance to controlled automation and ultimately agentic execution
  • Participation in both internal meetings and external meetings via video calls, as necessary
  • Ability to go into corporate or client offices to work onsite, as necessary
  • Prolonged periods of remaining stationary at a desk and working on a computer, as necessary
  • Ability to bend, kneel, crouch, and reach overhead, as necessary
  • Hand-eye coordination necessary to operate computers and various pieces of office equipment, as necessary
  • Vision abilities including close vision, toleration of fluorescent lighting, and adjusting focus, as necessary
  • For positions that require business travel and/or event attendance, ability to lift 25 lbs, as necessary
  • For positions that require business travel and/or event attendance, a valid driver's license and acceptable driving record are required, as driving is an essential job function

Qualifications

Must Haves

  • 3–5 years of overall relevant engineering experience, combining modern AI engineering with a strong software, integration, data, infrastructure, or AIOps foundation
  • Approximately 1–2 years of hands-on experience with Generative AI/LLMs, including building applications or workflows using modern LLM platforms
  • Approximately 2–3 years of foundational engineering experience in one or more areas such as backend software development, Python engineering, API integration, data engineering, cloud engineering, automation, or AIOps
  • Strong programming skills in Python, including experience developing production-quality backend services and automation
  • Strong experience designing, building, and consuming RESTful APIs and integrating multiple enterprise systems
  • Practical knowledge of LLMs, prompt engineering, context management, embeddings, vector retrieval, and Retrieval-Augmented Generation (RAG)
  • Hands-on experience with agentic or multi-agent AI frameworks, such as LangChain/LangGraph, AutoGen, CrewAI, or comparable technologies
  • Experience with or a strong understanding of Model Context Protocol (MCP), function/tool calling, agent registries, and AI orchestration patterns
  • Experience with modern AI coding assistants or autonomous development tools, such as Devin, Windsurf, or comparable solutions
  • Familiarity with enterprise IT and infrastructure platforms, ideally including one or more of ServiceNow/CMDB, Dynatrace, Zabbix, and GCP
  • Understanding of ITSM, incident management, observability, monitoring, infrastructure telemetry, or AIOps concepts
  • Experience working with both structured and unstructured data and preparing enterprise information for AI consumption
  • Understanding of AI safety, data governance, security, access control, grounding, hallucination mitigation, and responsible AI principles
  • Experience designing or operating production systems where reliability, scalability, observability, and maintainability are important
  • Strong systems-thinking and problem-solving skills, with the ability to understand complex enterprise environments and translate operational requirements into practical technical solutions
  • Ability to collaborate effectively with architects, software engineers, infrastructure teams, IT operations, security, and other technical stakeholders
  • Comfortable working iteratively, delivering measurable value through a phased approach from read-only AI assistance to controlled automation and ultimately agentic execution
  • Participation in both internal meetings and external meetings via video calls, as necessary
  • Ability to go into corporate or client offices to work onsite, as necessary
  • Prolonged periods of remaining stationary at a desk and working on a computer, as necessary
  • Ability to bend, kneel, crouch, and reach overhead, as necessary
  • Hand-eye coordination necessary to operate computers and various pieces of office equipment, as necessary
  • Vision abilities including close vision, toleration of fluorescent lighting, and adjusting focus, as necessary
  • For positions that require business travel and/or event attendance, ability to lift 25 lbs, as necessary
  • For positions that require business travel and/or event attendance, a valid driver's license and acceptable driving record are required, as driving is an essential job function

Benefits

  • Share plan
  • Company performance bonuses
  • Value-based recognition awards
  • Referral bonus
  • Career coaching
  • Global career opportunities
  • Non-linear career paths
  • Internal development programmes for management and technical leadership
  • Complex projects
  • Rotations
  • Internal tech communities
  • Training
  • Certifications
  • Coaching
  • Online learning platform subscriptions
  • Pass-it-on sessions
  • Workshops
  • Conferences
  • Hybrid work and flexible working hours
  • Employee assistance programme
  • Global internal wellbeing programme
  • Access to wellbeing apps
  • Global internal tech communities
  • Hobby clubs and interest groups
  • Inclusion and diversity programmes
  • Events and celebrations
  • For USA full-time roles only: robust healthcare and benefits including Medical, Dental, vision, Disability coverage, and various other benefit options
  • For USA full-time roles only: Flexible Spending Accounts (Medical, Transit, and Dependent Care)
  • For USA full-time roles only: Employer Paid Life Insurance and AD&D Coverages
  • For USA full-time roles only: Health Savings account paired with our low-cost High Deductible Medical Plan
  • For USA full-time roles only: 401(k) Safe Harbor Retirement plan with employer match with immediately vest

More jobs like this