Datarobot logo
Datarobot
Posted 6 days agoVerified live 2d ago

Deep Learning Research Engineer Intern

Brief overview

Remote
UndergradOr in progress
PyTorchDeep LearningTransformers and Attention MechanismsState-Space Sequence ArchitecturesProbabilistic ModelingProbability and StatisticsGPU-Native ImplementationSoftware TestingReproducible ExperimentationStochastic Processes and Stochastic Differential EquationsSynthetic Data GenerationTabular and Mixed-Modality Deep Learning

Job description

Summary

DataRobot delivers AI that maximizes impact and minimizes business risk through predictive and generative AI applications. The Deep Learning Research Engineer Intern will shape the architecture of a probabilistic foundation model, implement and train production-quality PyTorch models, conduct controlled experiments, and evaluate deep learning and stochastic modeling approaches.

Responsibilities

  • Designing and improving the core model architecture: input encoders for temporal, unordered tabular, and mixed-modality data; attention factorizations across variables, rows, and horizons; the distributional output heads and decoding strategies that produce coherent joint samples
  • Running controlled architecture studies, from ablations and scaling behavior to memory and throughput profiles, and turning the results into design decisions
  • Building scalable PyTorch implementations that support larger input and output spaces, better throughput, and tighter memory budgets
  • Studying how architectural choices, data-generation choices, and inference constraints interact to change benchmark quality and real-world usefulness
  • For candidates with the stochastic-modeling background: extending the synthetic-data engine with richer stochastic dynamics, constraints, dependence structures, heavy tails, and regime behavior
  • Turning research ideas into robust implementations and credible empirical results

Skills

  • Strong PyTorch skills and hands-on experience building and training deep models
  • Solid understanding of Transformers, attention variants, and long-context or state-space sequence architectures, including the tradeoffs between quality, memory, and latency
  • Experience with probabilistic modeling in neural networks: distributional output heads, likelihood-based or proper-scoring-rule losses, mixture or flow models, or related uncertainty-aware learning setups
  • Strong foundation in probability and statistics, enough to reason about densities and masses, dependence between variables, and calibration
  • Ability to connect architectural ideas to working GPU-native implementations, controlled experiments, and diagnostics that show why a change helped
  • Strong engineering habits: readable code, tests, reproducible experiments, and disciplined evaluation of model changes
  • Ability to debug training instability, reason about why results changed, and iterate quickly from hypothesis to evidence
  • Working knowledge of stochastic processes and stochastic differential equations: drift and diffusion, jumps, regime switching, mean reversion, heavy tails, and how such dynamics are simulated numerically
  • Experience designing synthetic data generators or simulation-based training curricula, and an understanding of how the training distribution shapes what a model learns
  • Depth in a domain with rich stochastic structure such as finance, energy, commodities, or a similarly quantitative field
  • Familiarity with tabular or mixed-modality deep learning: categorical, ordinal, count, and bounded targets alongside continuous ones

Qualifications

Must Haves

  • Strong PyTorch skills and hands-on experience building and training deep models
  • Solid understanding of Transformers, attention variants, and long-context or state-space sequence architectures, including the tradeoffs between quality, memory, and latency
  • Experience with probabilistic modeling in neural networks: distributional output heads, likelihood-based or proper-scoring-rule losses, mixture or flow models, or related uncertainty-aware learning setups
  • Strong foundation in probability and statistics, enough to reason about densities and masses, dependence between variables, and calibration
  • Ability to connect architectural ideas to working GPU-native implementations, controlled experiments, and diagnostics that show why a change helped
  • Strong engineering habits: readable code, tests, reproducible experiments, and disciplined evaluation of model changes
  • Ability to debug training instability, reason about why results changed, and iterate quickly from hypothesis to evidence

Nice to Haves

  • Working knowledge of stochastic processes and stochastic differential equations: drift and diffusion, jumps, regime switching, mean reversion, heavy tails, and how such dynamics are simulated numerically
  • Experience designing synthetic data generators or simulation-based training curricula, and an understanding of how the training distribution shapes what a model learns
  • Depth in a domain with rich stochastic structure such as finance, energy, commodities, or a similarly quantitative field
  • Familiarity with tabular or mixed-modality deep learning: categorical, ordinal, count, and bounded targets alongside continuous ones

Benefits

  • Medical, Dental & Vision Insurance
  • Flexible Time Off Program
  • Paid Holidays
  • Paid Parental Leave
  • Global Employee Assistance Program (EAP)

More jobs like this