Ontrac Solutions logo
Ontrac Solutions
Posted 59 days agoVerified live 14h ago

Network Automation & Reliability Engineer

Brief overview

Remote
MastersOr in progress
$75–$85/hrStated range
2+ yrsMinimum
Python developmentNetmikoNAPALMNornirpyATS/GenieScraplincclientBGP routingLocal preference policyMED policyBGP communitiesImport/export policyRoute reflectionECMPMultihomingArista EOSJuniper Junos

About the company

Ontrac Solutions logo
Ontrac Solutionsontracsolutions.net

Ontrac Solutions provides AI, cloud, and HubSpot consulting, engineering services, and technical talent for enterprise technology platforms.

Job description

Summary

Ontrac Solutions is seeking a Network Automation & Reliability Engineer to operate and automate hyperscale data center, backbone, and out-of-band networks. The role emphasizes Python development for network operations, requiring both deep routing and switching knowledge along with strong software engineering skills.

Responsibilities

  • Design, write, and maintain Python tooling that provisions, validates, and audits network devices across the global fleet
  • Build model-driven provisioning libraries using Jinja2 templating with NETCONF/YANG or REST APIs to codify golden configurations and eliminate configuration drift
  • Develop automated remediation for recurring failure classes — route flaps, hardware faults, optical degradation — triggered by syslog and telemetry
  • Reduce operational toil: identify manual runbook steps that recur, and replace them with tested, reviewed code
  • Operate and scale data center fabrics, backbone links, and out-of-band management networks across a global footprint of data centers and POP sites
  • Design and tune BGP policy — local preference, MED, communities, import/export policy, default-route propagation — to control path selection and eliminate single-carrier points of failure
  • Support capacity expansion and site turn-ups: high-level design, port channelization, optics and cabling standards, and hardware qualification
  • Execute zero-downtime change management in production, including hitless migrations and staged rollbacks
  • Operationalize multi-vendor streaming telemetry (gNMI/gRPC, OpenConfig) and tune subscription jobs for signal quality and management-plane efficiency
  • Build and maintain observability that supports hop-by-hop path tracing, multi-layer fault isolation, and fast root cause analysis
  • Contribute to a real-time network source of truth aggregating BGP, link-state, and drain-state data
  • Participate in a 24x7 on-call rotation; lead incident response, write RCAs, and drive the follow-up automation that prevents recurrence
  • Maintain and optimize firewall and ACL policy across multi-vendor platforms, keeping rule sets scoped and performant
  • Support SIRT/PSIRT CVE remediation through automated regression testing and config-as-code pipelines
  • Author and maintain technical documentation — designs, runbooks, and API contracts — for the tooling and networks you own

Skills

  • Python — primary requirement: Demonstrated, sustained Python development in a production network or infrastructure environment. You have written and maintained tooling that other engineers depended on: config generation and validation, API integrations, telemetry collectors, automated remediation, or test harnesses. You are comfortable with modules, packaging, testing, code review, and version control — not just single-file scripts
  • Routing & switching depth: Production experience with BGP (policy, path selection, multihoming), plus IS-IS or OSPF, ECMP, and VXLAN/EVPN or MPLS overlays
  • Multi-vendor hardware: Hands-on operations across at least two of Arista EOS, Juniper Junos (QFX/SRX/PTX/MX), or Cisco IOS-XR / NX-OS
  • Automation frameworks: Ansible and Jinja2, plus NETCONF/YANG, RESTCONF, or vendor REST APIs for model-driven configuration management
  • Telemetry & monitoring: gNMI/gRPC streaming telemetry, OpenConfig models, SNMP, flow telemetry, and dashboarding/alerting on top of them
  • Linux: Comfortable operating on Linux hosts — networking stack, packet capture, systemd services, and shell scripting
  • Version control and CI: Git-based workflows with peer review; experience shipping network changes through a pipeline (Jenkins, GitLab CI, or GitHub Actions)
  • Production on-call: Experience holding a 24x7 rotation for a live network, including incident command and root cause analysis
  • Experience level: Roughly 2–5 years in a network production, network reliability, or network automation role. Exceptional early-career engineers with a strong Python portfolio and hyperscale or carrier exposure are encouraged to apply
  • Location & work authorization: Must be located in the United States and authorized to work in the US
  • Out-of-band network experience — console server fleets (ZPE Nodegrid, OpenGear), RS-232 configuration management, or OOB build-out for new data center capacity
  • Data center or POP build-out and turn-up: new product introduction (NPI) for switching platforms, port channelization and optics selection, cabling and rack density planning, or hardware qualification and stress testing
  • MACsec, IPsec at scale, or zero-trust segmentation with 802.1X and NAC
  • Optical or transport exposure — DWDM, Ciena, PON/OLT/ONU, or IXIA/Spirent test automation
  • Building or contributing to a network source of truth (NetBox or in-house) aggregating BGP, link-state, and drain-state data
  • Applying LLM or agentic tooling to on-call workflows — automated triage, runbook execution, or incident summarization
  • A master's degree in Network Engineering, Telecommunications, or Computer Science

Qualifications

Must Haves

  • Python — primary requirement: Demonstrated, sustained Python development in a production network or infrastructure environment. You have written and maintained tooling that other engineers depended on: config generation and validation, API integrations, telemetry collectors, automated remediation, or test harnesses. You are comfortable with modules, packaging, testing, code review, and version control — not just single-file scripts
  • Routing & switching depth: Production experience with BGP (policy, path selection, multihoming), plus IS-IS or OSPF, ECMP, and VXLAN/EVPN or MPLS overlays
  • Multi-vendor hardware: Hands-on operations across at least two of Arista EOS, Juniper Junos (QFX/SRX/PTX/MX), or Cisco IOS-XR / NX-OS
  • Automation frameworks: Ansible and Jinja2, plus NETCONF/YANG, RESTCONF, or vendor REST APIs for model-driven configuration management
  • Telemetry & monitoring: gNMI/gRPC streaming telemetry, OpenConfig models, SNMP, flow telemetry, and dashboarding/alerting on top of them
  • Linux: Comfortable operating on Linux hosts — networking stack, packet capture, systemd services, and shell scripting
  • Version control and CI: Git-based workflows with peer review; experience shipping network changes through a pipeline (Jenkins, GitLab CI, or GitHub Actions)
  • Production on-call: Experience holding a 24x7 rotation for a live network, including incident command and root cause analysis
  • Experience level: Roughly 2–5 years in a network production, network reliability, or network automation role. Exceptional early-career engineers with a strong Python portfolio and hyperscale or carrier exposure are encouraged to apply
  • Location & work authorization: Must be located in the United States and authorized to work in the US

Nice to Haves

  • Out-of-band network experience — console server fleets (ZPE Nodegrid, OpenGear), RS-232 configuration management, or OOB build-out for new data center capacity
  • Data center or POP build-out and turn-up: new product introduction (NPI) for switching platforms, port channelization and optics selection, cabling and rack density planning, or hardware qualification and stress testing
  • MACsec, IPsec at scale, or zero-trust segmentation with 802.1X and NAC
  • Optical or transport exposure — DWDM, Ciena, PON/OLT/ONU, or IXIA/Spirent test automation
  • Building or contributing to a network source of truth (NetBox or in-house) aggregating BGP, link-state, and drain-state data
  • Applying LLM or agentic tooling to on-call workflows — automated triage, runbook execution, or incident summarization
  • A master's degree in Network Engineering, Telecommunications, or Computer Science

More jobs like this