Summary
DataRobot delivers AI that maximizes impact and minimizes business risk through predictive and generative AI applications. The Deep Learning Research Engineer Intern will shape the architecture of a probabilistic foundation model, implement and train production-quality PyTorch models, conduct controlled experiments, and evaluate deep learning and stochastic modeling approaches.
Responsibilities
- Designing and improving the core model architecture: input encoders for temporal, unordered tabular, and mixed-modality data; attention factorizations across variables, rows, and horizons; the distributional output heads and decoding strategies that produce coherent joint samples
- Running controlled architecture studies, from ablations and scaling behavior to memory and throughput profiles, and turning the results into design decisions
- Building scalable PyTorch implementations that support larger input and output spaces, better throughput, and tighter memory budgets
- Studying how architectural choices, data-generation choices, and inference constraints interact to change benchmark quality and real-world usefulness
- For candidates with the stochastic-modeling background: extending the synthetic-data engine with richer stochastic dynamics, constraints, dependence structures, heavy tails, and regime behavior
- Turning research ideas into robust implementations and credible empirical results
Skills
- Strong PyTorch skills and hands-on experience building and training deep models
- Solid understanding of Transformers, attention variants, and long-context or state-space sequence architectures, including the tradeoffs between quality, memory, and latency
- Experience with probabilistic modeling in neural networks: distributional output heads, likelihood-based or proper-scoring-rule losses, mixture or flow models, or related uncertainty-aware learning setups
- Strong foundation in probability and statistics, enough to reason about densities and masses, dependence between variables, and calibration
- Ability to connect architectural ideas to working GPU-native implementations, controlled experiments, and diagnostics that show why a change helped
- Strong engineering habits: readable code, tests, reproducible experiments, and disciplined evaluation of model changes
- Ability to debug training instability, reason about why results changed, and iterate quickly from hypothesis to evidence
- Working knowledge of stochastic processes and stochastic differential equations: drift and diffusion, jumps, regime switching, mean reversion, heavy tails, and how such dynamics are simulated numerically
- Experience designing synthetic data generators or simulation-based training curricula, and an understanding of how the training distribution shapes what a model learns
- Depth in a domain with rich stochastic structure such as finance, energy, commodities, or a similarly quantitative field
- Familiarity with tabular or mixed-modality deep learning: categorical, ordinal, count, and bounded targets alongside continuous ones
Qualifications
Must Haves
- Strong PyTorch skills and hands-on experience building and training deep models
- Solid understanding of Transformers, attention variants, and long-context or state-space sequence architectures, including the tradeoffs between quality, memory, and latency
- Experience with probabilistic modeling in neural networks: distributional output heads, likelihood-based or proper-scoring-rule losses, mixture or flow models, or related uncertainty-aware learning setups
- Strong foundation in probability and statistics, enough to reason about densities and masses, dependence between variables, and calibration
- Ability to connect architectural ideas to working GPU-native implementations, controlled experiments, and diagnostics that show why a change helped
- Strong engineering habits: readable code, tests, reproducible experiments, and disciplined evaluation of model changes
- Ability to debug training instability, reason about why results changed, and iterate quickly from hypothesis to evidence
Nice to Haves
- Working knowledge of stochastic processes and stochastic differential equations: drift and diffusion, jumps, regime switching, mean reversion, heavy tails, and how such dynamics are simulated numerically
- Experience designing synthetic data generators or simulation-based training curricula, and an understanding of how the training distribution shapes what a model learns
- Depth in a domain with rich stochastic structure such as finance, energy, commodities, or a similarly quantitative field
- Familiarity with tabular or mixed-modality deep learning: categorical, ordinal, count, and bounded targets alongside continuous ones
Benefits
- Medical, Dental & Vision Insurance
- Flexible Time Off Program
- Paid Holidays
- Paid Parental Leave
- Global Employee Assistance Program (EAP)