Sentient Labs logo
Sentient Labs
Posted 11 days agoVerified live 1d ago

Applied ML Engineer

Brief overview

Remote
MastersOr in progress
4+ yrsMinimum
PythonPyTorchHugging Face TransformersMachine Learning EvaluationML Research ImplementationProduction Software EngineeringAPI DevelopmentLLM Inference SystemsReactTypeScript

Job description

Summary

Sentient Labs is seeking an Applied ML Engineer to build systems at the intersection of machine learning research and production software. The role involves evaluating research methods, developing model-verification and evaluation infrastructure, working with model internals and inference systems, and turning research workflows into reliable product experiences. The engineer will also ship production-quality software across backend, frontend, experimentation, and observability systems.

Responsibilities

  • Reproduce and evaluate research methods using open-weight and API-accessible models
  • Design evaluation datasets, probes, scoring methods, baselines, calibration tests, and experiment harnesses
  • Work directly with model weights, logits, hidden states, activations, model APIs, and inference infrastructure when required
  • Build and extend our evaluation infrastructure, including runners, judges, persistence, experiment orchestration, and reporting
  • Turn research workflows into product experiences, including experiment configuration, runs, traces, comparisons, reports, and review workflows
  • Investigate how verification methods behave under model modification, including fine-tuning, merging, quantization, distillation, safety removal, and deliberate evasion
  • Design controlled experiments that separate meaningful signals from artifacts or confounders
  • Write clear technical reports that distinguish measured evidence, interpretation, and hypotheses
  • Ship production-quality systems with APIs, background jobs, observability, testing, and documentation
  • Reproduce at least one published model-provenance or verification method and clearly document its capabilities, assumptions, and limitations
  • Build a repeatable model-verification runner with versioned inputs, artifacts, metrics, and reports
  • Add at least one verification workflow to Construct and make it accessible through the Eldros UI
  • Run controlled experiments across base models, fine-tuned models, merged models, quantized models, and known distilled models
  • Improve our ability to understand when verification methods succeed, when they fail, and why
  • Leave behind production-quality code, tests, tooling, and documentation that another engineer can confidently operate and extend

Skills

  • • Strong Python engineering skills and hands-on experience with PyTorch and Hugging Face Transformers
  • • A strong understanding of ML evaluation, including dataset design, baselines, metrics, calibration, false positives, false negatives, statistical uncertainty, and reproducibility
  • • Ability to read ML research papers and implement methods from first principles rather than relying entirely on existing packages
  • • Experience building production software beyond notebooks, including APIs, asynchronous jobs, databases, logging, testing, and deployment
  • • Comfort working with open-weight models and understanding how modern LLM inference systems operate
  • • Ability to work across backend and frontend boundaries. Our product surface is primarily React/TypeScript, and you should be able to make complex experiments and results understandable to users
  • • Strong technical judgment about what experimental evidence does and does not support. For example, evidence that one model was derived from another is not necessarily evidence that it was directly trained on that model's outputs
  • • High agency and a strong sense of ownership. You are comfortable identifying problems, proposing solutions, and driving work forward without waiting for detailed instructions
  • • Comfortable working in a fast-moving startup environment where priorities can evolve quickly and individuals are expected to operate across functions
  • • Model provenance, fingerprinting, watermarking, distillation detection, red-teaming, safety evaluations, or interpretability
  • • Activation and representation analysis, probing, model hooks, logits, hidden states, or other model-internals work
  • • Evaluation and inference infrastructure such as DSPy, LiteLLM, Temporal, Ray, vLLM, PostgreSQL/pgvector, or similar systems
  • • Next.js, React, TypeScript, data visualization, or experiment dashboards
  • • Running and serving open-weight models on GPUs and reasoning about latency, throughput, memory, precision, and cost tradeoffs
  • • Designing adversarial evaluations or testing systems against deliberate attempts to evade detection

Qualifications

Must Haves

  • • Strong Python engineering skills and hands-on experience with PyTorch and Hugging Face Transformers
  • • A strong understanding of ML evaluation, including dataset design, baselines, metrics, calibration, false positives, false negatives, statistical uncertainty, and reproducibility
  • • Ability to read ML research papers and implement methods from first principles rather than relying entirely on existing packages
  • • Experience building production software beyond notebooks, including APIs, asynchronous jobs, databases, logging, testing, and deployment
  • • Comfort working with open-weight models and understanding how modern LLM inference systems operate
  • • Ability to work across backend and frontend boundaries. Our product surface is primarily React/TypeScript, and you should be able to make complex experiments and results understandable to users
  • • Strong technical judgment about what experimental evidence does and does not support. For example, evidence that one model was derived from another is not necessarily evidence that it was directly trained on that model's outputs
  • • High agency and a strong sense of ownership. You are comfortable identifying problems, proposing solutions, and driving work forward without waiting for detailed instructions
  • • Comfortable working in a fast-moving startup environment where priorities can evolve quickly and individuals are expected to operate across functions

Nice to Haves

  • • Model provenance, fingerprinting, watermarking, distillation detection, red-teaming, safety evaluations, or interpretability
  • • Activation and representation analysis, probing, model hooks, logits, hidden states, or other model-internals work
  • • Evaluation and inference infrastructure such as DSPy, LiteLLM, Temporal, Ray, vLLM, PostgreSQL/pgvector, or similar systems
  • • Next.js, React, TypeScript, data visualization, or experiment dashboards
  • • Running and serving open-weight models on GPUs and reasoning about latency, throughput, memory, precision, and cost tradeoffs
  • • Designing adversarial evaluations or testing systems against deliberate attempts to evade detection

Benefits

  • Remote work

More jobs like this