Summary
FIS is a financial technology company that provides a unified digital origination and decisioning platform. The Production Support Engineer will manage production issues, troubleshoot technical problems, and ensure efficient resolution while interacting with various teams within the organization.
Responsibilities
- Technical ability to deep dive into issues by querying tables, analyzing data and problem-solving
- Prioritization and triage of incoming requests/issues
- Drive incident resolution and lead conversations with cross-functional groups. Ask the right questions to help determine impact/priority and the correct route for resolution. Oversee a technical bridge, if required. Emphasize performing root cause analysis (RCA), documentation, and runbook development
- Management of all incidents through the incident management lifecycle
- Documentation of all relevant events, getting status reports while driving decision-making and resolution
- Ensure stakeholders are updated according to predefined service level agreements
- Completion and ownership of the postmortem with appropriate root cause analysis performed
- Improvement suggestions to capture preventative measures that will avoid recurrences of incidents
- Investigate patterns that indicate larger overall issues, even if we don’t have the solution
- Compilation of metrics on a weekly and monthly basis. Maintain dashboards for service incidents and ad hoc reporting as requested
- Management of ticket queues, client interaction, and maintenance of batch jobs
- Play an active role during critical incidents which may occur outside of normal business hours. Nights, weekends, and holidays on an on-call rotation basis is a must
- Creation of runbooks or standard operating procedures (SOP) so we can all learn from each other and add to our knowledge base
Skills
- Technical ability to deep dive into issues by querying tables, analyzing data and problem-solving
- Prioritization and triage of incoming requests/issues
- Drive incident resolution and lead conversations with cross-functional groups. Ask the right questions to help determine impact/priority and the correct route for resolution. Oversee a technical bridge, if required. Emphasize performing root cause analysis (RCA), documentation, and runbook development
- Management of all incidents through the incident management lifecycle
- Documentation of all relevant events, getting status reports while driving decision-making and resolution
- Ensure stakeholders are updated according to predefined service level agreements
- Completion and ownership of the postmortem with appropriate root cause analysis performed
- Improvement suggestions to capture preventative measures that will avoid recurrences of incidents
- Investigate patterns that indicate larger overall issues, even if we don't have the solution
- Compilation of metrics on a weekly and monthly basis. Maintain dashboards for service incidents and ad hoc reporting as requested
- Management of ticket queues, client interaction, and maintenance of batch jobs
- Play an active role during critical incidents which may occur outside of normal business hours. Nights, weekends, and holidays on an on-call rotation basis is a must
- Creation of runbooks or standard operating procedures (SOP) so we can all learn from each other and add to our knowledge base
- Technical and/or engineering background with advanced DB SQL skills, including table joins and advanced queries
- 2 years of experience coordinating and executing major incidents, with demonstrated capacity to lead under pressure
- Experience managing a wide spectrum of internal and external stakeholders, collaborating with cross-functional teams like QA, Customer Success, Developers, DevOps, and SRE
- Worked in an organization with a complex business environment
- Leadership skills with the ability to make quick decisions
- You thrive being a self-starter, who can lead others during stressful situations
- Experience with Postman, AI, and agents
- Basic knowledge or interest of any programming language such as Java, Python or Ruby
- Familiar with ITSM/ITIL concepts
- Familiar with tools such as Confluence, Jira, and on-call management software such as PagerDuty and experience with error monitoring software (Sentry, Kibana)
- Bonus: Experience with Jenkins, Argo, Vault, and Coralogix
Qualifications
Must Haves
- Technical ability to deep dive into issues by querying tables, analyzing data and problem-solving
- Prioritization and triage of incoming requests/issues
- Drive incident resolution and lead conversations with cross-functional groups. Ask the right questions to help determine impact/priority and the correct route for resolution. Oversee a technical bridge, if required. Emphasize performing root cause analysis (RCA), documentation, and runbook development
- Management of all incidents through the incident management lifecycle
- Documentation of all relevant events, getting status reports while driving decision-making and resolution
- Ensure stakeholders are updated according to predefined service level agreements
- Completion and ownership of the postmortem with appropriate root cause analysis performed
- Improvement suggestions to capture preventative measures that will avoid recurrences of incidents
- Investigate patterns that indicate larger overall issues, even if we don't have the solution
- Compilation of metrics on a weekly and monthly basis. Maintain dashboards for service incidents and ad hoc reporting as requested
- Management of ticket queues, client interaction, and maintenance of batch jobs
- Play an active role during critical incidents which may occur outside of normal business hours. Nights, weekends, and holidays on an on-call rotation basis is a must
- Creation of runbooks or standard operating procedures (SOP) so we can all learn from each other and add to our knowledge base
- Technical and/or engineering background with advanced DB SQL skills, including table joins and advanced queries
- 2 years of experience coordinating and executing major incidents, with demonstrated capacity to lead under pressure
- Experience managing a wide spectrum of internal and external stakeholders, collaborating with cross-functional teams like QA, Customer Success, Developers, DevOps, and SRE
- Worked in an organization with a complex business environment
- Leadership skills with the ability to make quick decisions
- You thrive being a self-starter, who can lead others during stressful situations
Nice to Haves
- Experience with Postman, AI, and agents
- Basic knowledge or interest of any programming language such as Java, Python or Ruby
- Familiar with ITSM/ITIL concepts
- Familiar with tools such as Confluence, Jira, and on-call management software such as PagerDuty and experience with error monitoring software (Sentry, Kibana)
- Bonus: Experience with Jenkins, Argo, Vault, and Coralogix