Summary
Harvard Medical School is dedicated to improving health and well-being through excellence in teaching, learning, and research. As a High-Performance Computing Engineer, you will support the implementation and management of HPC environments to facilitate computational research, collaborating closely with researchers and team members.
Responsibilities
- Perform provisioning, configuration, and decommissioning of HPC compute clusters
- Support the administration and tuning of workload schedulers (e.g., Slurm) to ensure efficient job management and cluster utilization
- Help maintain secure, regulated compute environments (e.g., NIST 800-171)
- Contribute to the integration of user accounts and identity management with institutional systems
- Maintain and optimize user-facing software environments, including module systems and containerized applications
- Support development and maintenance of scripts, automation, and tools used in cluster operations
- Monitor system health, respond to alerts, and assist with compliance reporting and documentation
- Collaborate with team members and researchers to troubleshoot and improve the computing environment
- Contribute to operational documentation and support knowledge-sharing across the team
- Participate in off-hours on-call rotation
- Perform other duties as assigned
Skills
- Minimum of two years' post-secondary education or relevant work experience
- Experience managing Linux-based systems in a research or academic environment
- Familiarity with workload schedulers (Slurm preferred), cluster provisioning, or performance tuning
- Experience with infrastructure monitoring, configuration management (e.g., Ansible), and containerization (e.g., Apptainer/Singularity, Docker)
- Understanding of security and compliance frameworks relevant to research computing
- Strong troubleshooting, communication, and collaboration skills
- Ability to work in a team-oriented environment and adapt to evolving priorities
- Demonstrated service orientation and commitment to operational reliability
- Willingness to learn and grow technical depth in HPC tools and methodologies
- Effective time management and documentation habits
- Bachelor's degree preferred
- Familiarity with workload schedulers (Slurm preferred)
- Completion of Harvard IT Academy specified foundational courses (or external equivalent) preferred
Qualifications
Must Haves
- Minimum of two years' post-secondary education or relevant work experience
- Experience managing Linux-based systems in a research or academic environment
- Familiarity with workload schedulers (Slurm preferred), cluster provisioning, or performance tuning
- Experience with infrastructure monitoring, configuration management (e.g., Ansible), and containerization (e.g., Apptainer/Singularity, Docker)
- Understanding of security and compliance frameworks relevant to research computing
- Strong troubleshooting, communication, and collaboration skills
- Ability to work in a team-oriented environment and adapt to evolving priorities
- Demonstrated service orientation and commitment to operational reliability
- Willingness to learn and grow technical depth in HPC tools and methodologies
- Effective time management and documentation habits
Nice to Haves
- Bachelor's degree preferred
- Familiarity with workload schedulers (Slurm preferred)
- Completion of Harvard IT Academy specified foundational courses (or external equivalent) preferred
Benefits
- Generous paid time off including parental leave
- Medical, dental, and vision health insurance coverage starting on day one
- Retirement plans with university contributions
- Wellbeing and mental health resources
- Support for families and caregivers
- Professional development opportunities including tuition assistance and reimbursement
- Commuter benefits, discounts and campus perks