Summary
EnCharge AI develops advanced AI hardware and software systems for edge-to-cloud computing. The company is seeking an AI Runtime Engineer to develop and optimize low-latency runtime software for executing deep learning models on specialized AI accelerators, collaborating with hardware, compiler, and AI framework teams to improve inference and training performance.
Responsibilities
- Develop and optimize the AI runtime software stack for executing deep learning workloads on AI accelerators
- Implement task scheduling, memory management, and kernel execution strategies for efficient computation
- Optimize data movement between host and device using PCIe, DMA, shared memory
- Design and implement high-performance APIs for AI Inference frameworks such as OpenVino, ONNX Runtime, vLLM
- Work on graph execution optimizations, including kernel fusion, pipelining, tensor tiling, and caching
- Integrate runtime components with AI compilers (LLVM, MLIR, XLA, TVM) for optimized execution
- Ensure scalability and reliability of the AI runtime for cloud-based and edge AI deployments
Skills
- Bachelor's or Master's degree in Computer Science, Electrical Engineering, or a related field
- 3+ years of experience in developing low-level runtime software for AI accelerators, GPUs, or HPC systems
- Strong proficiency in C/C++ and low-level systems programming
- Deep understanding of task scheduling, concurrency, and memory hierarchy
- Experience with hardware-aware optimizations and dataflow architectures
- Familiarity with deep learning execution frameworks (ONNX Runtime, TensorRT, TVM, OpenVINO)
- Experience with low-latency, high-throughput workload execution for AI models
- Strong debugging and profiling skills for optimizing AI execution performance
- Exposure to AI model deployment pipelines (Triton, TensorFlow Serving)
Qualifications
Must Haves
- Bachelor's or Master's degree in Computer Science, Electrical Engineering, or a related field
- 3+ years of experience in developing low-level runtime software for AI accelerators, GPUs, or HPC systems
- Strong proficiency in C/C++ and low-level systems programming
- Deep understanding of task scheduling, concurrency, and memory hierarchy
- Experience with hardware-aware optimizations and dataflow architectures
- Familiarity with deep learning execution frameworks (ONNX Runtime, TensorRT, TVM, OpenVINO)
- Experience with low-latency, high-throughput workload execution for AI models
- Strong debugging and profiling skills for optimizing AI execution performance
- Exposure to AI model deployment pipelines (Triton, TensorFlow Serving)