Summary
EnCharge AI is a leader in advanced AI hardware and software systems for edge-to-cloud computing, seeking an AI Runtime Engineer to develop and optimize the execution stack for their next-generation AI accelerator. The role involves creating low-latency, high-performance runtime software for deep learning models while collaborating with hardware and AI framework teams.
Responsibilities
- Develop and optimize the AI runtime software stack for executing deep learning workloads on AI accelerators
- Implement task scheduling, memory management, and kernel execution strategies for efficient computation
- Optimize data movement between host and device using PCIe, DMA, shared memory
- Design and implement high-performance APIs for AI Inference frameworks such as OpenVino, ONNX Runtime, vLLM
- Work on graph execution optimizations, including kernel fusion, pipelining, tensor tiling, and caching
- Integrate runtime components with AI compilers (LLVM, MLIR, XLA, TVM) for optimized execution
- Ensure scalability and reliability of the AI runtime for cloud-based and edge AI deployments
Skills
- Bachelor's or Master's degree in Computer Science, Electrical Engineering, or a related field
- 3+ years of experience in developing low-level runtime software for AI accelerators, GPUs, or HPC systems
- Strong proficiency in C/C++ and low-level systems programming
- Deep understanding of task scheduling, concurrency, and memory hierarchy
- Experience with hardware-aware optimizations and dataflow architectures
- Familiarity with deep learning execution frameworks (ONNX Runtime, TensorRT, TVM, OpenVINO)
- Experience with low-latency, high-throughput workload execution for AI models
- Strong debugging and profiling skills for optimizing AI execution performance
- Exposure to AI model deployment pipelines (Triton, TensorFlow Serving)
Qualifications
Must Haves
- Bachelor's or Master's degree in Computer Science, Electrical Engineering, or a related field
- 3+ years of experience in developing low-level runtime software for AI accelerators, GPUs, or HPC systems
- Strong proficiency in C/C++ and low-level systems programming
- Deep understanding of task scheduling, concurrency, and memory hierarchy
- Experience with hardware-aware optimizations and dataflow architectures
- Familiarity with deep learning execution frameworks (ONNX Runtime, TensorRT, TVM, OpenVINO)
- Experience with low-latency, high-throughput workload execution for AI models
- Strong debugging and profiling skills for optimizing AI execution performance
- Exposure to AI model deployment pipelines (Triton, TensorFlow Serving)