Summary
FIS provides a unified digital origination and decisioning platform for financial institutions. The Production Support Engineer will manage and resolve high-priority production issues across multiple platforms, troubleshoot technical problems, and coordinate incident response. The role also includes monitoring systems, managing ticket queues, conducting post-release validation, maintaining documentation, and meeting partner SLA requirements during non-day shifts.
Responsibilities
- Technical ability to deep dive into issues by querying tables, analyzing data and problem-solving
- Prioritization and triage of incoming requests/issues
- Drive incident resolution and lead conversations with cross-functional groups. Ask the right questions to help determine impact/priority and the correct route for resolution. Oversee a technical bridge, if required. Emphasize performing root cause analysis (RCA), documentation, and runbook development
- Management of all incidents through the incident management lifecycle
- Documentation of all relevant events, getting status reports while driving decision-making and resolution
- Ensure stakeholders are updated according to predefined service level agreements
- Completion and ownership of the postmortem with appropriate root cause analysis performed
- Improvement suggestions to capture preventative measures that will avoid recurrences of incidents
- Investigate patterns that indicate larger overall issues, even if we don’t have the solution
- Compilation of metrics on a weekly and monthly basis. Maintain dashboards for service incidents and ad hoc reporting as requested
- Management of ticket queues, client interaction, and maintenance of batch jobs
- Play an active role during critical incidents which may occur outside of normal business hours. Nights, weekends, and holidays on an on-call rotation basis is a must
- Creation of runbooks or standard operating procedures (SOP) so we can all learn from each other and add to our knowledge base
Skills
- Technical and/or engineering background with advanced DB SQL skills, including table joins and advanced queries
- Experience with Postman, AI, and agents
- Experience working with development teams in a fast-paced environment
- Basic knowledge or interest of any programming language such as Java, Python or Ruby
- 2 years of experience coordinating and executing major incidents, with demonstrated capacity to lead under pressure
- Experience managing a wide spectrum of internal and external stakeholders, collaborating with cross-functional teams like QA, Customer Success, Developers, DevOps, and SRE
- Worked in an organization with a complex business environment
- Leadership skills with the ability to make quick decisions
- Familiar with ITSM/ITIL concepts
- You thrive being a self-starter, who can lead others during stressful situations
- Familiar with tools such as Confluence, Jira, and on-call management software such as PagerDuty and experience with error monitoring software (Sentry, Kibana)
- This role supports our 24/7 staffing model with working hours of 4:00 PM–11:00 PM CST or 11:00 PM–7:00 AM CST
- Remote role; candidate must be located in the U.S
- Current and future sponsorship not available for this position
- Nights, weekends, and holidays on an on-call rotation basis is a must
- Bonus: Experience with Jenkins, Argo, Vault, and Coralogix
Qualifications
Must Haves
- Technical and/or engineering background with advanced DB SQL skills, including table joins and advanced queries
- Experience with Postman, AI, and agents
- Experience working with development teams in a fast-paced environment
- Basic knowledge or interest of any programming language such as Java, Python or Ruby
- 2 years of experience coordinating and executing major incidents, with demonstrated capacity to lead under pressure
- Experience managing a wide spectrum of internal and external stakeholders, collaborating with cross-functional teams like QA, Customer Success, Developers, DevOps, and SRE
- Worked in an organization with a complex business environment
- Leadership skills with the ability to make quick decisions
- Familiar with ITSM/ITIL concepts
- You thrive being a self-starter, who can lead others during stressful situations
- Familiar with tools such as Confluence, Jira, and on-call management software such as PagerDuty and experience with error monitoring software (Sentry, Kibana)
- This role supports our 24/7 staffing model with working hours of 4:00 PM–11:00 PM CST or 11:00 PM–7:00 AM CST
- Remote role; candidate must be located in the U.S
- Current and future sponsorship not available for this position
- Nights, weekends, and holidays on an on-call rotation basis is a must
Nice to Haves
- Bonus: Experience with Jenkins, Argo, Vault, and Coralogix
Benefits
- Remote role; candidate must be located in the U.S.