Summary
Microsoft is seeking a Data Flywheel Infrastructure Engineer to support its AI Superintelligence Team. The role builds infrastructure that transforms first-party, third-party, model-generated, evaluation, and synthetic data into secure, compliant, traceable, and high-quality training data for frontier LLM and multimodal models. Responsibilities include data governance, acquisition and curation, evaluation feedback loops, synthetic data pipelines, and data quality observability.
Responsibilities
- Build 1P & 3P Data Flywheel Infrastructure
Build scalable systems for ingesting, processing, curating, versioning, and serving first-party and third-party data for pre-training and post-training. Connect model failures, evaluations, and product signals back into targeted data acquisition, generation, and improvement workflows
- Own Data Governance, Security & Compliance Infrastructure
Build governance and policy enforcement directly into the data platform, including:
* + Data provenance and lineage
+ Usage rights, licensing, and consent metadata
+ PII / sensitive-data detection and protection
+ Access control and data isolation
+ Retention and deletion enforcement
+ Geographic and regulatory restrictions
+ Dataset approval and audit workflows
+ Training eligibility and purpose-based usage controls
- Build Policy-Aware Data Acquisition & Curation Systems
Develop automated pipelines for 1P and 3P data ingestion, classification, filtering, deduplication, quality scoring, semantic enrichment, and dataset construction.
Make governance policies machine-enforceable so that data can automatically be included, excluded, quarantined, or restricted based on its origin, license, sensitivity, consent, geography, and intended model use
- Build Evaluation-to-Data Feedback Loops
Convert model evaluations and real-world failure signals into actionable data tasks through failure clustering, hard-example mining, long-tail discovery, capability-gap detection, and targeted dataset generation.
Enable rapid iteration from:
**Model Failure → Data Gap → Data Intervention → Training → Evaluation**
- Build Synthetic & AI-Native Data Pipelines
Use LLMs, VLMs, and Agents to automate data generation, labeling, filtering, quality validation, enrichment, and transformation.
Maintain clear provenance between **human-created, first-party, third-party, model-generated, and derived data**, and enforce appropriate policies across each category
- Build Data Quality, Attribution & Observability
Develop metrics and infrastructure to measure dataset quality, coverage, diversity, contamination, duplication, policy compliance, and contribution to model capability improvements.
Enable researchers to understand **which data improves which capabilities and under what governance constraints**
Skills
- • Master's Degree in Computer Science, Math, Software Engineering, Computer Engineering, or related field AND 3+ years experience in business analytics, data science, software development, data modeling, or data engineering OR Bachelor's Degree in Computer Science, Math, Software Engineering, Computer Engineering, or related field AND 4+ years experience in business analytics, data science, software development, data modeling, or data engineering OR equivalent experience
- • Software Engineering experience using Python,SQL, Spark/Flink/Ray
- * Experience building **AI training-data governance platforms**, including provenance, licensing/rights metadata, consent management, PII handling, policy enforcement, or auditable lineage
- * Experience managing **third-party datasets**, data partnerships, licensed content, or externally sourced data with complex contractual and usage restrictions
- * Experience building privacy- and security-aware systems for **first-party product or user data**, including isolation, access controls, retention/deletion, and purpose limitation
- * Experience with **data clean rooms, privacy-preserving processing, de-identification, confidential computing, or secure data collaboration**
- * Experience building **evaluation → failure mining → data generation → training** feedback loops
- * Experience with synthetic data, model graders, reward signals, hard-example mining, active learning, or data-mixture optimization
- * Experience with multimodal or agentic datasets including text, image, video, audio, web, GUI, tool-use, or interaction trajectories
- * Understanding of Modern LLM training workflows including Pre-training, SFT, RL/post-training, evaluation, and synthetic data
- * Strong understanding of **data governance, security, privacy, provenance, access control, and data lifecycle management**
Qualifications
Must Haves
- • Master's Degree in Computer Science, Math, Software Engineering, Computer Engineering, or related field AND 3+ years experience in business analytics, data science, software development, data modeling, or data engineering OR Bachelor's Degree in Computer Science, Math, Software Engineering, Computer Engineering, or related field AND 4+ years experience in business analytics, data science, software development, data modeling, or data engineering OR equivalent experience
- • Software Engineering experience using Python,SQL, Spark/Flink/Ray
Nice to Haves
- * Experience building **AI training-data governance platforms**, including provenance, licensing/rights metadata, consent management, PII handling, policy enforcement, or auditable lineage
- * Experience managing **third-party datasets**, data partnerships, licensed content, or externally sourced data with complex contractual and usage restrictions
- * Experience building privacy- and security-aware systems for **first-party product or user data**, including isolation, access controls, retention/deletion, and purpose limitation
- * Experience with **data clean rooms, privacy-preserving processing, de-identification, confidential computing, or secure data collaboration**
- * Experience building **evaluation → failure mining → data generation → training** feedback loops
- * Experience with synthetic data, model graders, reward signals, hard-example mining, active learning, or data-mixture optimization
- * Experience with multimodal or agentic datasets including text, image, video, audio, web, GUI, tool-use, or interaction trajectories
- * Understanding of Modern LLM training workflows including Pre-training, SFT, RL/post-training, evaluation, and synthetic data
- * Strong understanding of **data governance, security, privacy, provenance, access control, and data lifecycle management**