f
fal
Posted 65 days agoVerified live 1d ago

Software Engineer, Distributed Systems

Brief overview

Remote
$180k–$250k/yrStated range
PythonRustDistributed systems fundamentalsSchedulingFault toleranceCapacity planningComputational complexityMemory allocationAI/ML inferenceAI/ML training infrastructureHigh-performance systems programmingAsync runtimesZero-copyMemory-safe concurrencyMulti-tenant compute platformsGPU workload schedulingNetworking fundamentals

About the company

f
fal

A platform company that operates AI inference workloads across multiple data centers and cloud providers.

Job description

Summary

fal is the generative media ecosystem powering the next generation of AI products, providing infrastructure and tools for developers and enterprises. They are seeking an experienced Software Engineer who specializes in building large-scale distributed systems, focusing on core platform development and performance optimization.

Responsibilities

  • Build our core Python/Rust platform: request routing, AI workload orchestration, scheduling, GPU autoscaling, large scale file storage, queueing, etc
  • Produce forward designs for platform evolution as we scale to 100x current traffic and need to provide low latency across the world
  • Leverage AI to an extreme level to automate the mundane parts of building complex but reliable systems
  • Profile and tune low level CPU and memory performance

Skills

  • 3+ years experience building distributed compute and orchestration platforms in Python or Rust
  • Strong understanding of distributed systems fundamentals: consensus, scheduling, fault tolerance, capacity planning
  • Deep understanding of computational complexity and memory allocation
  • Track record of designing systems that scale under real production load
  • Experience building and using observability to drive performance and reliability decisions
  • Excellent communication and ability to drive technical decisions across teams
  • Self-starter who executes quickly, takes ownership, and constantly seeks improvement
  • Experience with AI/ML inference or training infrastructure
  • Experience with high-performance systems programming (async runtimes, zero-copy, memory-safe concurrency)
  • Background in building multi-tenant compute platforms
  • Understanding of networking fundamentals and performance characteristics
  • Familiarity with GPU workload characteristics and scheduling constraints

Qualifications

Must Haves

  • 3+ years experience building distributed compute and orchestration platforms in Python or Rust
  • Strong understanding of distributed systems fundamentals: consensus, scheduling, fault tolerance, capacity planning
  • Deep understanding of computational complexity and memory allocation
  • Track record of designing systems that scale under real production load
  • Experience building and using observability to drive performance and reliability decisions
  • Excellent communication and ability to drive technical decisions across teams
  • Self-starter who executes quickly, takes ownership, and constantly seeks improvement

Nice to Haves

  • Experience with AI/ML inference or training infrastructure
  • Experience with high-performance systems programming (async runtimes, zero-copy, memory-safe concurrency)
  • Background in building multi-tenant compute platforms
  • Understanding of networking fundamentals and performance characteristics
  • Familiarity with GPU workload characteristics and scheduling constraints

Benefits

  • Equity
  • Benefits
  • We offer relocation assistance to San Francisco.
  • Health, dental, and vision insurance (US)
  • Regular team events and offsites

More jobs like this