EnCharge AI logo
EnCharge AI
Posted 75 days agoVerified live 1d ago

AI Runtime Engineer

Brief overview

Remote
UndergradOr in progress
3+ yrsMinimum
AI runtime software developmentC/C++ programmingLow-level systems programmingTask schedulingConcurrencyMemory hierarchyHardware-aware optimizationsDataflow architecturesONNX RuntimeTensorRTTVMOpenVINODebugging and profilingAI model deployment pipelinesTritonTensorFlow Serving

About the company

EnCharge AI logo
EnCharge AIenchargeai.com

EnCharge AI designs analog in-memory-computing AI chips and develops AI systems for AI computing.

Job description

Summary

EnCharge AI is a leader in advanced AI hardware and software systems for edge-to-cloud computing, seeking an AI Runtime Engineer to develop and optimize the execution stack for their next-generation AI accelerator. The role involves creating low-latency, high-performance runtime software for deep learning models while collaborating with hardware and AI framework teams.

Responsibilities

  • Develop and optimize the AI runtime software stack for executing deep learning workloads on AI accelerators
  • Implement task scheduling, memory management, and kernel execution strategies for efficient computation
  • Optimize data movement between host and device using PCIe, DMA, shared memory
  • Design and implement high-performance APIs for AI Inference frameworks such as OpenVino, ONNX Runtime, vLLM
  • Work on graph execution optimizations, including kernel fusion, pipelining, tensor tiling, and caching
  • Integrate runtime components with AI compilers (LLVM, MLIR, XLA, TVM) for optimized execution
  • Ensure scalability and reliability of the AI runtime for cloud-based and edge AI deployments

Skills

  • Bachelor's or Master's degree in Computer Science, Electrical Engineering, or a related field
  • 3+ years of experience in developing low-level runtime software for AI accelerators, GPUs, or HPC systems
  • Strong proficiency in C/C++ and low-level systems programming
  • Deep understanding of task scheduling, concurrency, and memory hierarchy
  • Experience with hardware-aware optimizations and dataflow architectures
  • Familiarity with deep learning execution frameworks (ONNX Runtime, TensorRT, TVM, OpenVINO)
  • Experience with low-latency, high-throughput workload execution for AI models
  • Strong debugging and profiling skills for optimizing AI execution performance
  • Exposure to AI model deployment pipelines (Triton, TensorFlow Serving)

Qualifications

Must Haves

  • Bachelor's or Master's degree in Computer Science, Electrical Engineering, or a related field
  • 3+ years of experience in developing low-level runtime software for AI accelerators, GPUs, or HPC systems
  • Strong proficiency in C/C++ and low-level systems programming
  • Deep understanding of task scheduling, concurrency, and memory hierarchy
  • Experience with hardware-aware optimizations and dataflow architectures
  • Familiarity with deep learning execution frameworks (ONNX Runtime, TensorRT, TVM, OpenVINO)
  • Experience with low-latency, high-throughput workload execution for AI models
  • Strong debugging and profiling skills for optimizing AI execution performance
  • Exposure to AI model deployment pipelines (Triton, TensorFlow Serving)

More jobs like this