Recruiting from Scratch logo
Recruiting from Scratch
Posted 74 days agoVerified live 2d ago

Site Reliability Engineer

Brief overview

Remote
$150k–$250k/yrStated range
3+ yrsMinimum
AWSGCPAzurePythonFastAPIDockerKubernetesPrometheusGrafanaRedisPostgreSQLC++RustLLM APIs

About the company

Recruiting from Scratch logo
Recruiting from Scratchrecruitingfromscratch.com

A recruiting agency working with technology companies to help them hire software engineers, data roles, product managers, and hardware.

Job description

Summary

Recruiting from Scratch is representing a dynamic startup in the Devtools and AI space that is focused on building innovative solutions for software development. The Site Reliability Engineer will design and implement infrastructure solutions, develop applications, and ensure system performance and reliability.

Responsibilities

  • Design and implement robust infrastructure solutions using AWS, GCP, and Azure
  • Develop and maintain applications using Python and FastAPI to ensure high performance and responsiveness
  • Manage container orchestration with Docker and Kubernetes to facilitate seamless deployment and scaling
  • Monitor system performance and reliability using tools like Prometheus and Grafana, ensuring optimal uptime
  • Collaborate with cross-functional teams to integrate and call various LLM APIs in the OpenAI format
  • Establish CI/CD pipelines to automate deployment processes and enhance development efficiency

Skills

  • 3–8 years of experience in Site Reliability Engineering, DevOps, or related fields
  • Proficient in Python, with experience in C/C++ or Rust as a plus
  • Strong understanding of containerization and orchestration technologies, particularly Kubernetes
  • Experience with database management, specifically Redis and PostgreSQL
  • Familiarity with monitoring and logging tools such as Prometheus and Grafana
  • Experience working with AI technologies and APIs, particularly in the context of large language models
  • Knowledge of software development best practices and agile methodologies
  • Previous involvement in early-stage startups or small teams, demonstrating adaptability and initiative

Qualifications

Must Haves

  • 3–8 years of experience in Site Reliability Engineering, DevOps, or related fields
  • Proficient in Python, with experience in C/C++ or Rust as a plus
  • Strong understanding of containerization and orchestration technologies, particularly Kubernetes
  • Experience with database management, specifically Redis and PostgreSQL
  • Familiarity with monitoring and logging tools such as Prometheus and Grafana

Nice to Haves

  • Experience working with AI technologies and APIs, particularly in the context of large language models
  • Knowledge of software development best practices and agile methodologies
  • Previous involvement in early-stage startups or small teams, demonstrating adaptability and initiative

Benefits

  • Competitive equity options

More jobs like this