Summary
FIS is a technology company that powers the world’s economy, and they are seeking a Production Support Engineer to manage production issues efficiently. This role involves troubleshooting technical issues, managing incident lifecycles, and collaborating with cross-functional teams to ensure service level agreements are met.
Responsibilities
- Manage high-priority issues to resolution following industry best practices
- Troubleshoot, fix, and apply workarounds to resolve technical issues across multiple platforms
- Interact with every aspect of the organization to find the best solution for partners
- Management of ticket queues, monitoring for issues and post-release validation
- Drive incident resolution and lead conversations with cross-functional groups
- Ask the right questions to help determine impact/priority and the correct route for resolution
- Oversee a technical bridge, if required
- Management of all incidents through the incident management lifecycle
- Documentation of all relevant events, getting status reports while driving decision-making and resolution
- Ensure stakeholders are updated according to predefined service level agreements
- Completion and ownership of the postmortem with appropriate root cause analysis performed
- Improvement suggestions to capture preventative measures that will avoid recurrences of incidents
- Investigate patterns that indicate larger overall issues, even if we don’t have the solution
- Compilation of metrics on a weekly and monthly basis
- Maintain dashboards for service incidents and ad hoc reporting as requested
- Play an active role during critical incidents which may occur outside of normal business hours
- Creation of runbooks or standard operating procedures (SOP) so we can all learn from each other and add to our knowledge base
Skills
- Technical ability to deep dive into issues by querying tables, analyzing data and problem-solving
- Prioritization and triage of incoming requests/issues
- Drive incident resolution and lead conversations with cross-functional groups. Ask the right questions to help determine impact/priority and the correct route for resolution. Oversee a technical bridge, if required
- Management of all incidents through the incident management lifecycle
- Documentation of all relevant events, getting status reports while driving decision-making and resolution
- Ensure stakeholders are updated according to predefined service level agreements
- Completion and ownership of the postmortem with appropriate root cause analysis performed
- Improvement suggestions to capture preventative measures that will avoid recurrences of incidents
- Investigate patterns that indicate larger overall issues, even if we don't have the solution
- Compilation of metrics on a weekly and monthly basis. Maintain dashboards for service incidents and ad hoc reporting as requested
- Play an active role during critical incidents which may occur outside of normal business hours. Nights, weekends, and holidays on an on-call rotation basis is a must
- Creation of runbooks or standard operating procedures (SOP) so we can all learn from each other and add to our knowledge base
- Technical and/or engineering background, ideally with experience writing SQL queries
- Experience working with development teams in a fast-paced environment
- Basic knowledge or interest of any programming language such as Java, Python or Ruby
- 2 years of experience coordinating and executing major incidents, with demonstrated capacity to lead under pressure
- Previously collaborated with a wide spectrum of internal and external stakeholders
- Worked in an organization with a complex business environment
- Leadership skills with the ability to make quick decisions
- Familiar with ITSM/ITIL concepts
- You thrive being a self-starter, who can lead others during stressful situations
- Familiar with tools such as Confluence, Jira, and on-call management software such as PagerDuty and experience with error monitoring software (Sentry, Kibana)
Qualifications
Must Haves
- Technical ability to deep dive into issues by querying tables, analyzing data and problem-solving
- Prioritization and triage of incoming requests/issues
- Drive incident resolution and lead conversations with cross-functional groups. Ask the right questions to help determine impact/priority and the correct route for resolution. Oversee a technical bridge, if required
- Management of all incidents through the incident management lifecycle
- Documentation of all relevant events, getting status reports while driving decision-making and resolution
- Ensure stakeholders are updated according to predefined service level agreements
- Completion and ownership of the postmortem with appropriate root cause analysis performed
- Improvement suggestions to capture preventative measures that will avoid recurrences of incidents
- Investigate patterns that indicate larger overall issues, even if we don't have the solution
- Compilation of metrics on a weekly and monthly basis. Maintain dashboards for service incidents and ad hoc reporting as requested
- Play an active role during critical incidents which may occur outside of normal business hours. Nights, weekends, and holidays on an on-call rotation basis is a must
- Creation of runbooks or standard operating procedures (SOP) so we can all learn from each other and add to our knowledge base
- Technical and/or engineering background, ideally with experience writing SQL queries
- Experience working with development teams in a fast-paced environment
- Basic knowledge or interest of any programming language such as Java, Python or Ruby
- 2 years of experience coordinating and executing major incidents, with demonstrated capacity to lead under pressure
- Previously collaborated with a wide spectrum of internal and external stakeholders
- Worked in an organization with a complex business environment
- Leadership skills with the ability to make quick decisions
- Familiar with ITSM/ITIL concepts
- You thrive being a self-starter, who can lead others during stressful situations
- Familiar with tools such as Confluence, Jira, and on-call management software such as PagerDuty and experience with error monitoring software (Sentry, Kibana)
Benefits
- Most positions are hybrid (3 days onsite) in our FIS Office locations in Chicago (Illinois), Milwaukee (Wisconsin), Atlanta (Georgia) and Jacksonville (Florida).