Nebius logo
Nebius
Posted 66 days agoVerified live 1d ago

Senior Applied AI Solutions Engineer

Brief overview

Remote
Large model fine-tuningDistributed training debuggingProduction RAG pipelinesAgentic pipelinesGPU inference optimizationPyTorchHuggingFaceCUDA fundamentalsKubernetes for MLMLflowVector databasesEnterprise ML collaboration

About the company

The Nebius AI Cloud brings powerful full-stack infrastructure for AI developers and practitioners across startups, enterprises and science institutes to build and deploy generative AI applications and rapidly deliver scientific breakthroughs by training and running ML models within a secure, high-performance, and cost-optimized cloud environment.

Job description

Summary

Nebius is leading a new era in cloud infrastructure for the global AI economy, building a full-stack AI cloud platform for developers and enterprises. The Senior Applied AI Solutions Engineer role focuses on prototyping applied AI use cases, accelerating customer onboarding, and providing insights to the product roadmap.

Responsibilities

  • Build prototypes and demos across the product portfolio — serverless inference, databases, MLflow, MLOps, and vertical use cases in Physical AI and HCLS — that become assets for sales, product, and engineering teams
  • Support new customers hands-on through POC design, technical onboarding, and validation; act as the bridge between their ML team and the platform during the critical first months
  • Go deep on emerging applied AI — new training techniques, inference optimizations, agentic architectures, new frameworks — and turn findings into working prototypes, writeups, and product recommendations
  • Feed the product roadmap with specific, grounded feedback; be the voice of "here's what broke in three customer POCs last month and here's what needs to change"
  • Develop reusable technical assets — notebooks, reference architectures, benchmark results — that reduce onboarding friction at scale

Skills

  • You've fine-tuned large models, debugged distributed training jobs, built production RAG or agentic pipelines, and optimized inference on GPU infrastructure — not just read about it
  • You're fluent in the modern ML stack: PyTorch, HuggingFace, CUDA fundamentals, Kubernetes for ML, MLflow or equivalent, vector databases
  • You've worked with enterprise ML teams — whether as a solutions engineer, customer engineer, or an ML engineer who collaborated closely with customers
  • You read papers and implement them — not for credit, but because it's how you stay sharp
  • You communicate with calibration: you can explain activation checkpointing tradeoffs to an ML engineer in the morning and the cost implication to a CTO in the afternoon
  • Experience in any of our vertical domains: Physical AI / robotics / simulation, HCLS (drug discovery, medical imaging, clinical NLP), or enterprise AI application development
  • Familiarity with MLOps at scale (Kubeflow, Metaflow, Argo, Ray)
  • Prior work at a cloud provider or AI infrastructure company
  • You've shared technical work publicly — notebooks, talks, blog posts that people actually use

Qualifications

Must Haves

  • You've fine-tuned large models, debugged distributed training jobs, built production RAG or agentic pipelines, and optimized inference on GPU infrastructure — not just read about it
  • You're fluent in the modern ML stack: PyTorch, HuggingFace, CUDA fundamentals, Kubernetes for ML, MLflow or equivalent, vector databases
  • You've worked with enterprise ML teams — whether as a solutions engineer, customer engineer, or an ML engineer who collaborated closely with customers
  • You read papers and implement them — not for credit, but because it's how you stay sharp
  • You communicate with calibration: you can explain activation checkpointing tradeoffs to an ML engineer in the morning and the cost implication to a CTO in the afternoon

Nice to Haves

  • Experience in any of our vertical domains: Physical AI / robotics / simulation, HCLS (drug discovery, medical imaging, clinical NLP), or enterprise AI application development
  • Familiarity with MLOps at scale (Kubeflow, Metaflow, Argo, Ray)
  • Prior work at a cloud provider or AI infrastructure company
  • You've shared technical work publicly — notebooks, talks, blog posts that people actually use

Benefits

  • Competitive salary and comprehensive benefits package.
  • Opportunities for professional growth within Nebius.
  • Flexible working arrangements.
  • A dynamic and collaborative work environment that values initiative and innovation.
  • Competitive compensation
  • Career growth and learning opportunities
  • Flexibility and work-life balance
  • Collaborative and innovative culture
  • Opportunity to work on impactful AI projects
  • International environment and talented teams

More jobs like this