Summary
Ontrac Solutions is seeking a Network Automation & Reliability Engineer to operate and automate hyperscale data center, backbone, and out-of-band networks. The role emphasizes Python development for network operations, requiring both deep routing and switching knowledge along with strong software engineering skills.
Responsibilities
- Design, write, and maintain Python tooling that provisions, validates, and audits network devices across the global fleet
- Build model-driven provisioning libraries using Jinja2 templating with NETCONF/YANG or REST APIs to codify golden configurations and eliminate configuration drift
- Develop automated remediation for recurring failure classes — route flaps, hardware faults, optical degradation — triggered by syslog and telemetry
- Reduce operational toil: identify manual runbook steps that recur, and replace them with tested, reviewed code
- Operate and scale data center fabrics, backbone links, and out-of-band management networks across a global footprint of data centers and POP sites
- Design and tune BGP policy — local preference, MED, communities, import/export policy, default-route propagation — to control path selection and eliminate single-carrier points of failure
- Support capacity expansion and site turn-ups: high-level design, port channelization, optics and cabling standards, and hardware qualification
- Execute zero-downtime change management in production, including hitless migrations and staged rollbacks
- Operationalize multi-vendor streaming telemetry (gNMI/gRPC, OpenConfig) and tune subscription jobs for signal quality and management-plane efficiency
- Build and maintain observability that supports hop-by-hop path tracing, multi-layer fault isolation, and fast root cause analysis
- Contribute to a real-time network source of truth aggregating BGP, link-state, and drain-state data
- Participate in a 24x7 on-call rotation; lead incident response, write RCAs, and drive the follow-up automation that prevents recurrence
- Maintain and optimize firewall and ACL policy across multi-vendor platforms, keeping rule sets scoped and performant
- Support SIRT/PSIRT CVE remediation through automated regression testing and config-as-code pipelines
- Author and maintain technical documentation — designs, runbooks, and API contracts — for the tooling and networks you own
Skills
- Python — primary requirement: Demonstrated, sustained Python development in a production network or infrastructure environment. You have written and maintained tooling that other engineers depended on: config generation and validation, API integrations, telemetry collectors, automated remediation, or test harnesses. You are comfortable with modules, packaging, testing, code review, and version control — not just single-file scripts
- Routing & switching depth: Production experience with BGP (policy, path selection, multihoming), plus IS-IS or OSPF, ECMP, and VXLAN/EVPN or MPLS overlays
- Multi-vendor hardware: Hands-on operations across at least two of Arista EOS, Juniper Junos (QFX/SRX/PTX/MX), or Cisco IOS-XR / NX-OS
- Automation frameworks: Ansible and Jinja2, plus NETCONF/YANG, RESTCONF, or vendor REST APIs for model-driven configuration management
- Telemetry & monitoring: gNMI/gRPC streaming telemetry, OpenConfig models, SNMP, flow telemetry, and dashboarding/alerting on top of them
- Linux: Comfortable operating on Linux hosts — networking stack, packet capture, systemd services, and shell scripting
- Version control and CI: Git-based workflows with peer review; experience shipping network changes through a pipeline (Jenkins, GitLab CI, or GitHub Actions)
- Production on-call: Experience holding a 24x7 rotation for a live network, including incident command and root cause analysis
- Experience level: Roughly 2–5 years in a network production, network reliability, or network automation role. Exceptional early-career engineers with a strong Python portfolio and hyperscale or carrier exposure are encouraged to apply
- Location & work authorization: Must be located in the United States and authorized to work in the US
- Out-of-band network experience — console server fleets (ZPE Nodegrid, OpenGear), RS-232 configuration management, or OOB build-out for new data center capacity
- Data center or POP build-out and turn-up: new product introduction (NPI) for switching platforms, port channelization and optics selection, cabling and rack density planning, or hardware qualification and stress testing
- MACsec, IPsec at scale, or zero-trust segmentation with 802.1X and NAC
- Optical or transport exposure — DWDM, Ciena, PON/OLT/ONU, or IXIA/Spirent test automation
- Building or contributing to a network source of truth (NetBox or in-house) aggregating BGP, link-state, and drain-state data
- Applying LLM or agentic tooling to on-call workflows — automated triage, runbook execution, or incident summarization
- A master's degree in Network Engineering, Telecommunications, or Computer Science
Qualifications
Must Haves
- Python — primary requirement: Demonstrated, sustained Python development in a production network or infrastructure environment. You have written and maintained tooling that other engineers depended on: config generation and validation, API integrations, telemetry collectors, automated remediation, or test harnesses. You are comfortable with modules, packaging, testing, code review, and version control — not just single-file scripts
- Routing & switching depth: Production experience with BGP (policy, path selection, multihoming), plus IS-IS or OSPF, ECMP, and VXLAN/EVPN or MPLS overlays
- Multi-vendor hardware: Hands-on operations across at least two of Arista EOS, Juniper Junos (QFX/SRX/PTX/MX), or Cisco IOS-XR / NX-OS
- Automation frameworks: Ansible and Jinja2, plus NETCONF/YANG, RESTCONF, or vendor REST APIs for model-driven configuration management
- Telemetry & monitoring: gNMI/gRPC streaming telemetry, OpenConfig models, SNMP, flow telemetry, and dashboarding/alerting on top of them
- Linux: Comfortable operating on Linux hosts — networking stack, packet capture, systemd services, and shell scripting
- Version control and CI: Git-based workflows with peer review; experience shipping network changes through a pipeline (Jenkins, GitLab CI, or GitHub Actions)
- Production on-call: Experience holding a 24x7 rotation for a live network, including incident command and root cause analysis
- Experience level: Roughly 2–5 years in a network production, network reliability, or network automation role. Exceptional early-career engineers with a strong Python portfolio and hyperscale or carrier exposure are encouraged to apply
- Location & work authorization: Must be located in the United States and authorized to work in the US
Nice to Haves
- Out-of-band network experience — console server fleets (ZPE Nodegrid, OpenGear), RS-232 configuration management, or OOB build-out for new data center capacity
- Data center or POP build-out and turn-up: new product introduction (NPI) for switching platforms, port channelization and optics selection, cabling and rack density planning, or hardware qualification and stress testing
- MACsec, IPsec at scale, or zero-trust segmentation with 802.1X and NAC
- Optical or transport exposure — DWDM, Ciena, PON/OLT/ONU, or IXIA/Spirent test automation
- Building or contributing to a network source of truth (NetBox or in-house) aggregating BGP, link-state, and drain-state data
- Applying LLM or agentic tooling to on-call workflows — automated triage, runbook execution, or incident summarization
- A master's degree in Network Engineering, Telecommunications, or Computer Science