EnCharge AI logo
EnCharge AI
Posted 15 days agoVerified live 9h ago

AI Runtime Engineer

Brief overview

Remote
MastersOr in progress
3+ yrsMinimum
C/C++Low-Level Systems ProgrammingAI Runtime SoftwareAI AcceleratorsTask SchedulingConcurrencyMemory ManagementMemory HierarchyPCIeDMAShared MemoryONNX RuntimeTensorRTTVMOpenVINODeep Learning Execution FrameworksHardware-Aware Optimization

About the company

EnCharge AI logo
EnCharge AIenchargeai.com

EnCharge AI designs analog in-memory-computing AI chips and develops AI systems for AI computing.

Job description

Summary

EnCharge AI develops advanced AI hardware and software systems for edge-to-cloud computing. The company is seeking an AI Runtime Engineer to develop and optimize low-latency runtime software for executing deep learning models on specialized AI accelerators, collaborating with hardware, compiler, and AI framework teams to improve inference and training performance.

Responsibilities

  • Develop and optimize the AI runtime software stack for executing deep learning workloads on AI accelerators
  • Implement task scheduling, memory management, and kernel execution strategies for efficient computation
  • Optimize data movement between host and device using PCIe, DMA, shared memory
  • Design and implement high-performance APIs for AI Inference frameworks such as OpenVino, ONNX Runtime, vLLM
  • Work on graph execution optimizations, including kernel fusion, pipelining, tensor tiling, and caching
  • Integrate runtime components with AI compilers (LLVM, MLIR, XLA, TVM) for optimized execution
  • Ensure scalability and reliability of the AI runtime for cloud-based and edge AI deployments

Skills

  • Bachelor's or Master's degree in Computer Science, Electrical Engineering, or a related field
  • 3+ years of experience in developing low-level runtime software for AI accelerators, GPUs, or HPC systems
  • Strong proficiency in C/C++ and low-level systems programming
  • Deep understanding of task scheduling, concurrency, and memory hierarchy
  • Experience with hardware-aware optimizations and dataflow architectures
  • Familiarity with deep learning execution frameworks (ONNX Runtime, TensorRT, TVM, OpenVINO)
  • Experience with low-latency, high-throughput workload execution for AI models
  • Strong debugging and profiling skills for optimizing AI execution performance
  • Exposure to AI model deployment pipelines (Triton, TensorFlow Serving)

Qualifications

Must Haves

  • Bachelor's or Master's degree in Computer Science, Electrical Engineering, or a related field
  • 3+ years of experience in developing low-level runtime software for AI accelerators, GPUs, or HPC systems
  • Strong proficiency in C/C++ and low-level systems programming
  • Deep understanding of task scheduling, concurrency, and memory hierarchy
  • Experience with hardware-aware optimizations and dataflow architectures
  • Familiarity with deep learning execution frameworks (ONNX Runtime, TensorRT, TVM, OpenVINO)
  • Experience with low-latency, high-throughput workload execution for AI models
  • Strong debugging and profiling skills for optimizing AI execution performance
  • Exposure to AI model deployment pipelines (Triton, TensorFlow Serving)

More jobs like this