Etched logo
Etched
Verified live 2d ago

Architecture Intern - Inference

Brief overview

San Jose, CAIn-person
UndergradOr in progress
19 H-1B approvalsDept. of Labor
PythonC++Performance-sensitive software systemsLinux internalsAccelerator architectures (e.g. GPUs, TPUs)CompilersHigh-speed interconnects (e.g. NVLink, InfiniBand)Porting applications to non-standard accelerator hardwareTransformer model architecturesInference serving stacks (vLLM, SGLang, etc.)RustLow-latency, high-performance applicationsKernel-level and user-space networking stacksDistributed systems concepts, algorithms, and challengesConsensus protocolsConsistency modelsCommunication patterns

About the company

A company developing AI inference hardware and transformer accelerators.

Visa sponsorship history

3 years sponsoring, last filed FY2026

Data powered by U.S. Department of Labor. This does not guarantee sponsorship for this specific role.
19H-1B approved
100%approval rate
3new H-1B hires
$216,046median wage / yr
H-1B Petition ApprovalsVisas USCIS actually granted: the strongest sign the company sponsors.
202519
LCA Certified ApplicationsAn early filing step, not a visa approval: it signals intent, not confirmed sponsorship.
20242
20251
20261
Top sponsored roles
Member of Technical Staff, Design VerificationASIC Technical Program ManagerMember of Technical Staff (Design Verification Engineer)Member of Technical Staff (Physical Design Engineer)

Job description

Summary

Etched is building the world’s first AI inference system purpose-built for transformers, delivering exceptional performance and efficiency. The Architecture Intern will contribute to the design of next-generation AI accelerators by developing and optimizing compute architectures for transformer workloads.

Responsibilities

  • Support porting state-of-the-art models to our architecture
  • Help build programming abstractions and testing capabilities to rapidly iterate on model porting
  • Assist in building, enhancing, and scaling Sohu’s runtime, including multi-node inference, intra-node execution, state management, and robust error handling
  • Contribute to optimizing routing and communication layers using Sohu’s collectives
  • Utilize performance profiling and debugging tools to identify bottlenecks and correctness issues
  • Develop and leverage a deep understanding of Sohu to co-design both HW instructions and model architecture operations to maximize model performance
  • Implement high-performance software components for the Model Toolkit

Skills

  • Progress towards a Bachelor's, Master's, or PhD degree in computer science, computer engineering, applied mathematics, or a related field
  • Proficiency in Python, C++
  • Understanding of performance-sensitive or complex distributed software systems, e.g. Linux internals, accelerator architectures (e.g. GPUs, TPUs), Compilers, or high-speed interconnects (e.g. NVLink, InfiniBand)
  • Ported applications to non-standard accelerator hardware or hardware platforms
  • Deep knowledge of transformer model architectures and/or inference serving stacks (vLLM, SGLang, etc.)
  • Proficiency in Rust
  • Low-latency, high-performance applications using both kernel-level and user-space networking stacks
  • Deep understanding of distributed systems concepts, algorithms, and challenges, including consensus protocols, consistency models, and communication patterns
  • Solid grasp of Transformer architectures, particularly Mixture-of-Experts (MoE)
  • Built applications with extensive SIMD (Single Instruction, Multiple Data) optimizations for performance-critical paths
  • Familiarity with PyTorch or JAX
  • Math competitions (AIME, AMC, etc)

Qualifications

Must Haves

  • Progress towards a Bachelor's, Master's, or PhD degree in computer science, computer engineering, applied mathematics, or a related field
  • Proficiency in Python, C++
  • Understanding of performance-sensitive or complex distributed software systems, e.g. Linux internals, accelerator architectures (e.g. GPUs, TPUs), Compilers, or high-speed interconnects (e.g. NVLink, InfiniBand)
  • Ported applications to non-standard accelerator hardware or hardware platforms
  • Deep knowledge of transformer model architectures and/or inference serving stacks (vLLM, SGLang, etc.)

Nice to Haves

  • Proficiency in Rust
  • Low-latency, high-performance applications using both kernel-level and user-space networking stacks
  • Deep understanding of distributed systems concepts, algorithms, and challenges, including consensus protocols, consistency models, and communication patterns
  • Solid grasp of Transformer architectures, particularly Mixture-of-Experts (MoE)
  • Built applications with extensive SIMD (Single Instruction, Multiple Data) optimizations for performance-critical paths
  • Familiarity with PyTorch or JAX
  • Math competitions (AIME, AMC, etc)

Benefits

  • Generous housing support for those relocating
  • Daily lunch and dinner in our office

More jobs like this