Summary
Open Systems Technologies Corporation (OST) is seeking Software Integration Engineers to support a mission-focused Site Reliability Engineering (SRE) team. This role involves supporting critical mission operations infrastructure through automation, monitoring, and troubleshooting activities in modern Linux environments.
Responsibilities
- Support mission operations infrastructure tooling, including automation, alerting, monitoring, and resiliency initiatives
- Configure, troubleshoot, and maintain customer programming use cases
- Perform Linux system administration activities across multiple distributions
- Develop and maintain automation solutions using Salt and Ansible
- Create and support monitoring solutions utilizing Nagios, Thruk, Prometheus, Splunk, and Grafana
- Troubleshoot system performance and operational issues in mission-critical environments
- Support software integration efforts and infrastructure deployments
- Manage and maintain code repositories and development workflows
- Collaborate with engineering and operations teams to improve system reliability and efficiency
Skills
- Active TS/SCI with Polygraph security clearance
- U.S. Citizenship
- Experience administering Linux operating systems, including: Red Hat Enterprise Linux (RHEL), CentOS, Rocky Linux, SUSE Linux Enterprise Server (SLES), Ubuntu
- Experience with one or more of the following programming/scripting languages: Python, C, Bash
- Experience with Linux system administration, troubleshooting, and operational support
- Ability to work effectively in a fast-paced mission environment
- Experience with any of the following technologies is highly desired: Thruk dashboards, Nagios monitoring and plugin development, Splunk dashboard integration and management, Prometheus, Grafana, Salt, Ansible, GitBucket, Jira, Confluence, Slurm
Qualifications
Must Haves
- Active TS/SCI with Polygraph security clearance
- U.S. Citizenship
- Experience administering Linux operating systems, including: Red Hat Enterprise Linux (RHEL), CentOS, Rocky Linux, SUSE Linux Enterprise Server (SLES), Ubuntu
- Experience with one or more of the following programming/scripting languages: Python, C, Bash
- Experience with Linux system administration, troubleshooting, and operational support
- Ability to work effectively in a fast-paced mission environment
Nice to Haves
- Experience with any of the following technologies is highly desired: Thruk dashboards, Nagios monitoring and plugin development, Splunk dashboard integration and management, Prometheus, Grafana, Salt, Ansible, GitBucket, Jira, Confluence, Slurm
Benefits
- 3 Weeks Paid Time Off
- 11 Federal Holidays
- Medical and Dental Coverage
- Short-Term Disability (STD)
- Long-Term Disability (LTD)
- Life Insurance
- Accidental Death & Dismemberment (AD&D) Coverage
- 401(k) with up to 4% Company Match