Summary
Silicon Data builds data infrastructure, benchmarking, and market indices for the global compute economy. The AI Benchmark Engineer will build and run reproducible inference performance benchmarks, define measurement methodology, investigate anomalies, and translate results into defensible technical and economic metrics.
Responsibilities
- Build and maintain the inference benchmarking harness. Automated, containerized, reproducible runs across inference engines (vLLM, TensorRT-LLM, SGLang, TGI), model families, and hardware targets
- Own the inference metrics that matter. Time-to-first-token, inter-token latency, sustained and peak throughput, latency percentiles under concurrency, goodput under SLO constraints, and their expression per dollar and per watt
- Design workload profiles that reflect real usage. Input and output length distributions, concurrency patterns, streaming versus batch, prefill-heavy versus decode-heavy — rather than synthetic best-case runs
- Capture and validate the configuration fingerprint. Engine version, model revision and quantization, driver and firmware, kernel and runtime versions, tensor and pipeline parallelism, scheduler settings. A result without its full configuration is not a result
- Establish run repeatability and statistical rigor. Define warmup and steady-state criteria, variance thresholds, minimum run counts, and outlier handling. Know when a run is invalid and be willing to throw it out
- Investigate performance anomalies. When a cohort underperforms, determine whether the cause is thermal, network, contention, misconfiguration, or a genuine silicon difference — and produce the evidence
- Partner with our data and GPU benchmarking engineers on cohort definition, cross-suite consistency, and how inference results align with hardware-level benchmarks
- Write the methodology. Public-facing write-ups, versioned methodology documents, and technical documentation that a skeptical customer can audit and reproduce
- Work with our pricing products so measured inference performance can be expressed in economic terms — cost of capability per measured token per second
- Track the field. New engines, new serving techniques, new model architectures, and emerging public benchmarks worth incorporating or explicitly rejecting
Skills
- 4+ years of engineering experience, with substantial hands-on work in LLM inference, model serving, or performance engineering
- Strong Python. You write automation and tooling that other people can run and trust, not one-off scripts
- Direct experience with at least one production inference engine (vLLM, TensorRT-LLM, SGLang, TGI, or equivalent), including configuring and tuning it — not just calling its API
- Working knowledge of what actually drives inference performance: batching and scheduling, KV cache behavior, quantization, tensor and pipeline parallelism, memory bandwidth limits, and prefill versus decode characteristics
- Comfort with Linux, containers, and running workloads on GPU infrastructure
- Genuine measurement discipline. You are skeptical of your own results, you control variables, you distinguish signal from noise, and you can explain the uncertainty in a number
- Clear technical writing. You can document a methodology decision well enough that an outside engineer can reproduce it and a hostile reviewer cannot dismiss it
- Self-direction. The methodology here is not written yet. You will be defining the standard, not executing someone else's spec
- Remote. We are open to W-2 employment, onshore 1099 contract, or offshore corp-to-corp engagement — tell us which structure you are looking for and we will work with it
- Overlap requirement. You must be available during 8:00 AM – 1:00 PM Eastern Time, Monday through Friday. Outside that window, work when you work best
- GPU profiling and telemetry experience — Nsight, DCGM, driver-level APIs, or equivalent
- CUDA or kernel-level familiarity
- Exposure to MLPerf Inference or comparable standardized benchmark suites
- Experience with distributed or multi-node inference
- Familiarity with non-NVIDIA accelerators (AMD, TPU, Cerebras, Groq, or custom silicon)
- Experience building software that runs as an agent inside a customer environment, and the deployment, versioning, and trust problems that come with it
- Background in certification, attestation, provenance, or audit-grade data products
- Understanding of GPU hardware economics, cloud compute pricing, or AI infrastructure procurement
- Early-stage startup experience
Qualifications
Must Haves
- 4+ years of engineering experience, with substantial hands-on work in LLM inference, model serving, or performance engineering
- Strong Python. You write automation and tooling that other people can run and trust, not one-off scripts
- Direct experience with at least one production inference engine (vLLM, TensorRT-LLM, SGLang, TGI, or equivalent), including configuring and tuning it — not just calling its API
- Working knowledge of what actually drives inference performance: batching and scheduling, KV cache behavior, quantization, tensor and pipeline parallelism, memory bandwidth limits, and prefill versus decode characteristics
- Comfort with Linux, containers, and running workloads on GPU infrastructure
- Genuine measurement discipline. You are skeptical of your own results, you control variables, you distinguish signal from noise, and you can explain the uncertainty in a number
- Clear technical writing. You can document a methodology decision well enough that an outside engineer can reproduce it and a hostile reviewer cannot dismiss it
- Self-direction. The methodology here is not written yet. You will be defining the standard, not executing someone else's spec
- Remote. We are open to W-2 employment, onshore 1099 contract, or offshore corp-to-corp engagement — tell us which structure you are looking for and we will work with it
- Overlap requirement. You must be available during 8:00 AM – 1:00 PM Eastern Time, Monday through Friday. Outside that window, work when you work best
Nice to Haves
- GPU profiling and telemetry experience — Nsight, DCGM, driver-level APIs, or equivalent
- CUDA or kernel-level familiarity
- Exposure to MLPerf Inference or comparable standardized benchmark suites
- Experience with distributed or multi-node inference
- Familiarity with non-NVIDIA accelerators (AMD, TPU, Cerebras, Groq, or custom silicon)
- Experience building software that runs as an agent inside a customer environment, and the deployment, versioning, and trust problems that come with it
- Background in certification, attestation, provenance, or audit-grade data products
- Understanding of GPU hardware economics, cloud compute pricing, or AI infrastructure procurement
- Early-stage startup experience
Benefits
- For W-2 hires, meaningful equity.
- Ownership of the measurement standard for a market where no standard yet exists.
- A small, senior team where your work has direct, visible impact and your opinions carry real weight.
- Comprehensive health, dental, and vision benefits for W-2 employees.
- A flexible, remote-friendly environment with a culture that values output over presence.
- A founder-led culture built on merit — your growth here is limited only by the quality of your work.