Cerebras logo
Cerebras
Posted 69 days agoVerified live 1d ago

AI Inference Core - Infrastructure SW Engineer

Brief overview

Remote
UndergradOr in progress
3+ yrsMinimum
PythonSoftware architectureSoftware design patternsConcurrency conceptsOperating systemsDistributed systemsDebuggingAsyncioMultiprocessingConcurrent futuresPytestKubernetesContainerized environmentsCluster schedulersRemote execution systems

Job description

Summary

Cerebras Systems builds the world's largest AI chip, transforming the user experience of AI applications. The Software Engineer will design and build core software for Cerebras engineering infrastructure, focusing on Python frameworks and orchestration systems.

Responsibilities

  • Design, develop, test, and maintain Python frameworks and services used to orchestrate engineering workflows across machines and clusters
  • Build reusable abstractions for scheduling, distributed execution, resource management, test execution, workflow planning, and failure recovery
  • Define clear APIs, module boundaries, extension points, and data models that allow infrastructure systems to evolve without becoming difficult to maintain
  • Reason about concurrency, asynchronous execution, multiprocessing, state management, retries, idempotency, cancellation, and partial failures
  • Debug complex issues spanning Python applications, operating systems, processes, filesystems, networking, remote machines, and distributed services
  • Write high-quality automated tests and documentation for infrastructure that is expected to be reliable and widely reused
  • Partner with platform, CI, release, quality, ML systems, and product engineering teams to understand requirements and translate them into scalable software designs

Skills

  • 3+ years of professional software-engineering experience
  • Strong proficiency in Python and a solid understanding of the language's strengths, limitations, and runtime behavior
  • Experience designing maintainable software systems, libraries, frameworks, backend services, or developer-facing APIs
  • Good judgment around software architecture, abstraction boundaries, design patterns, extensibility, and long-term maintainability
  • Understanding of concurrency concepts such as processes, threads, asynchronous execution, synchronization, and shared state
  • Foundational understanding of operating systems, including processes, signals, filesystems, resource management, and program execution
  • Foundational understanding of distributed-systems concepts such as retries, timeouts, idempotency, partial failure, coordination, and eventual consistency
  • Strong debugging and problem-solving skills, including the ability to form hypotheses, gather evidence, and work through unfamiliar systems independently
  • Experience with Python concurrency technologies such as `asyncio`, multiprocessing, concurrent futures, or event-driven systems
  • Experience building orchestration engines, workflow systems, schedulers, distributed job runners, or control-plane software
  • Experience developing test infrastructure or extensions for frameworks such as pytest
  • Familiarity with CI systems, build systems, release infrastructure, or developer-productivity tooling
  • Experience with Kubernetes, containerized environments, cluster schedulers, or remote execution systems
  • BS/MS in Computer Science or a related field, or equivalent practical experience

Qualifications

Must Haves

  • 3+ years of professional software-engineering experience
  • Strong proficiency in Python and a solid understanding of the language's strengths, limitations, and runtime behavior
  • Experience designing maintainable software systems, libraries, frameworks, backend services, or developer-facing APIs
  • Good judgment around software architecture, abstraction boundaries, design patterns, extensibility, and long-term maintainability
  • Understanding of concurrency concepts such as processes, threads, asynchronous execution, synchronization, and shared state
  • Foundational understanding of operating systems, including processes, signals, filesystems, resource management, and program execution
  • Foundational understanding of distributed-systems concepts such as retries, timeouts, idempotency, partial failure, coordination, and eventual consistency
  • Strong debugging and problem-solving skills, including the ability to form hypotheses, gather evidence, and work through unfamiliar systems independently

Nice to Haves

  • Experience with Python concurrency technologies such as `asyncio`, multiprocessing, concurrent futures, or event-driven systems
  • Experience building orchestration engines, workflow systems, schedulers, distributed job runners, or control-plane software
  • Experience developing test infrastructure or extensions for frameworks such as pytest
  • Familiarity with CI systems, build systems, release infrastructure, or developer-productivity tooling
  • Experience with Kubernetes, containerized environments, cluster schedulers, or remote execution systems
  • BS/MS in Computer Science or a related field, or equivalent practical experience

Benefits

  • HYBRID / Hybrid work model

More jobs like this