Microsoft logo
Microsoft
Posted 40 days agoVerified live 5h ago

Member of Technical Staff - Data Flywheel Infra, Frontier Models

Brief overview

Remote
MastersOr in progress
$120k–$235k/yrStated range
3+ yrsMinimum
7,292 H-1B approvalsDept. of Labor
7,082 green cardsCertified filings
PythonSQLApache SparkApache FlinkRayData GovernanceData EngineeringAI Training-Data GovernanceData Provenance and LineageData Privacy and SecurityAccess ControlData Lifecycle Management

About the company

Microsoft logo
Microsoftmicrosoft.com

Microsoft is a software corporation that develops, manufactures, licenses, supports, and sells a range of software products and services.

Visa sponsorship history

4 years sponsoring, last filed FY2026

Data powered by U.S. Department of Labor. This does not guarantee sponsorship for this specific role.
7,292H-1B approved
97%approval rate
1,172new H-1B hires
7,082PERM certified
$169,178median wage / yr
H-1B Petition ApprovalsVisas USCIS actually granted: the strongest sign the company sponsors.
20232,866
20241,446
20251,801
20261,179
LCA Certified ApplicationsAn early filing step, not a visa approval: it signals intent, not confirmed sponsorship.
20231,912
20242,842
20251,798
20261,338
Green Card (PERM) FilingsCertified green card filings: a long-term commitment to international hires.
20232,396
20241,662
20252,928
202696
Top sponsored roles
Software EngineeringSoftware EngineerTechnical Program ManagementProduct ManagementData Science
Sponsored employees from
IndiaChinaCanadaMexicoBrazil

Job description

Summary

Microsoft is seeking a Data Flywheel Infrastructure Engineer to support its AI Superintelligence Team. The role builds infrastructure that transforms first-party, third-party, model-generated, evaluation, and synthetic data into secure, compliant, traceable, and high-quality training data for frontier LLM and multimodal models. Responsibilities include data governance, acquisition and curation, evaluation feedback loops, synthetic data pipelines, and data quality observability.

Responsibilities

  • Build 1P & 3P Data Flywheel Infrastructure Build scalable systems for ingesting, processing, curating, versioning, and serving first-party and third-party data for pre-training and post-training. Connect model failures, evaluations, and product signals back into targeted data acquisition, generation, and improvement workflows
  • Own Data Governance, Security & Compliance Infrastructure Build governance and policy enforcement directly into the data platform, including: * + Data provenance and lineage + Usage rights, licensing, and consent metadata + PII / sensitive-data detection and protection + Access control and data isolation + Retention and deletion enforcement + Geographic and regulatory restrictions + Dataset approval and audit workflows + Training eligibility and purpose-based usage controls
  • Build Policy-Aware Data Acquisition & Curation Systems Develop automated pipelines for 1P and 3P data ingestion, classification, filtering, deduplication, quality scoring, semantic enrichment, and dataset construction. Make governance policies machine-enforceable so that data can automatically be included, excluded, quarantined, or restricted based on its origin, license, sensitivity, consent, geography, and intended model use
  • Build Evaluation-to-Data Feedback Loops Convert model evaluations and real-world failure signals into actionable data tasks through failure clustering, hard-example mining, long-tail discovery, capability-gap detection, and targeted dataset generation. Enable rapid iteration from: **Model Failure → Data Gap → Data Intervention → Training → Evaluation**
  • Build Synthetic & AI-Native Data Pipelines Use LLMs, VLMs, and Agents to automate data generation, labeling, filtering, quality validation, enrichment, and transformation. Maintain clear provenance between **human-created, first-party, third-party, model-generated, and derived data**, and enforce appropriate policies across each category
  • Build Data Quality, Attribution & Observability Develop metrics and infrastructure to measure dataset quality, coverage, diversity, contamination, duplication, policy compliance, and contribution to model capability improvements. Enable researchers to understand **which data improves which capabilities and under what governance constraints**

Skills

  • • Master's Degree in Computer Science, Math, Software Engineering, Computer Engineering, or related field AND 3+ years experience in business analytics, data science, software development, data modeling, or data engineering OR Bachelor's Degree in Computer Science, Math, Software Engineering, Computer Engineering, or related field AND 4+ years experience in business analytics, data science, software development, data modeling, or data engineering OR equivalent experience
  • • Software Engineering experience using Python,SQL, Spark/Flink/Ray
  • * Experience building **AI training-data governance platforms**, including provenance, licensing/rights metadata, consent management, PII handling, policy enforcement, or auditable lineage
  • * Experience managing **third-party datasets**, data partnerships, licensed content, or externally sourced data with complex contractual and usage restrictions
  • * Experience building privacy- and security-aware systems for **first-party product or user data**, including isolation, access controls, retention/deletion, and purpose limitation
  • * Experience with **data clean rooms, privacy-preserving processing, de-identification, confidential computing, or secure data collaboration**
  • * Experience building **evaluation → failure mining → data generation → training** feedback loops
  • * Experience with synthetic data, model graders, reward signals, hard-example mining, active learning, or data-mixture optimization
  • * Experience with multimodal or agentic datasets including text, image, video, audio, web, GUI, tool-use, or interaction trajectories
  • * Understanding of Modern LLM training workflows including Pre-training, SFT, RL/post-training, evaluation, and synthetic data
  • * Strong understanding of **data governance, security, privacy, provenance, access control, and data lifecycle management**

Qualifications

Must Haves

  • • Master's Degree in Computer Science, Math, Software Engineering, Computer Engineering, or related field AND 3+ years experience in business analytics, data science, software development, data modeling, or data engineering OR Bachelor's Degree in Computer Science, Math, Software Engineering, Computer Engineering, or related field AND 4+ years experience in business analytics, data science, software development, data modeling, or data engineering OR equivalent experience
  • • Software Engineering experience using Python,SQL, Spark/Flink/Ray

Nice to Haves

  • * Experience building **AI training-data governance platforms**, including provenance, licensing/rights metadata, consent management, PII handling, policy enforcement, or auditable lineage
  • * Experience managing **third-party datasets**, data partnerships, licensed content, or externally sourced data with complex contractual and usage restrictions
  • * Experience building privacy- and security-aware systems for **first-party product or user data**, including isolation, access controls, retention/deletion, and purpose limitation
  • * Experience with **data clean rooms, privacy-preserving processing, de-identification, confidential computing, or secure data collaboration**
  • * Experience building **evaluation → failure mining → data generation → training** feedback loops
  • * Experience with synthetic data, model graders, reward signals, hard-example mining, active learning, or data-mixture optimization
  • * Experience with multimodal or agentic datasets including text, image, video, audio, web, GUI, tool-use, or interaction trajectories
  • * Understanding of Modern LLM training workflows including Pre-training, SFT, RL/post-training, evaluation, and synthetic data
  • * Strong understanding of **data governance, security, privacy, provenance, access control, and data lifecycle management**

More jobs like this