Summary
Handshake is building AI products, research platforms, and customer-facing AI systems that support frontier model training and post-training. The Member of Technical Staff will build systems for synthetic data generation, data processing, evaluation, and anonymization, while partnering with researchers, domain experts, legal and compliance stakeholders, and customers to turn ambiguous data needs into scalable products and infrastructure.
Responsibilities
- Design and build systems that improve the quality, scale, and safety of the data Handshake generates and acquires for frontier model training—spanning synthetic data generation and data anonymization/PII removal
- Translate ambiguous research, partner, or compliance needs into clear hypotheses, experiments, evaluation plans, and production-quality implementations
- Build and improve data-processing pipelines, evaluation frameworks, benchmarks, and quality-control systems, whether the goal is generating higher-signal synthetic data or verifying that sensitive data has been properly de-identified
- Run fast, rigorous iteration loops: prototype, evaluate, interpret results, and turn learnings into the next system or product
- Partner directly with researchers, domain experts, and—where relevant—legal and compliance teams to ensure data is both high-utility and responsibly handled
- Identify repeatable patterns across engagements and productize them into reusable software and platforms
- Raise the technical bar through strong design judgment, clear communication, code quality, and mentorship
Skills
- 2–10 years of recent, demonstrated experience in one or more of: synthetic/LLM-generated data, post-training and model-evaluation work, privacy engineering, or data anonymization/de-identification at scale
- A hands-on individual contributor track record—this is not a team-lead or engineering-management role
- Strong Python skills and the ability to write clean, efficient, scalable software for large, messy, real-world datasets
- Sound judgment for reasoning about data quality, risk, and utility—forming hypotheses, choosing meaningful metrics, diagnosing failures, and distinguishing signal from noise
- Experience designing systems—not only implementing specifications—including tradeoffs around quality, scale, reliability, and reuse
- Comfort operating in an ambiguous, fast-moving environment with substantial ownership
- Collaborative, low-ego communication and the ability to work effectively with researchers, engineers, domain experts, and customers
- Building or operating large-scale synthetic or LLM-generated data pipelines for model training
- Building or operating large-scale data de-identification or anonymization systems, ideally involving relational or graph-structured data, with experience preserving referential/relationship integrity after anonymization
- Developing LLM/agent benchmarks, evaluation methodologies, annotation systems, or data-quality frameworks
- Research or applied work on reinforcement learning, alignment, model behavior, synthetic data, or human-in-the-loop systems
- Prior work in a regulated or high-sensitivity data environment (healthcare, finance, HR/people data, government), or experience with re-identification risk assessment and privacy auditing
- Published research, meaningful open-source contributions, or evidence of technical leadership in ML systems, data engineering, or AI research
- Experience productizing research or repeated customer work into robust, reusable platforms
Qualifications
Must Haves
- 2–10 years of recent, demonstrated experience in one or more of: synthetic/LLM-generated data, post-training and model-evaluation work, privacy engineering, or data anonymization/de-identification at scale
- A hands-on individual contributor track record—this is not a team-lead or engineering-management role
- Strong Python skills and the ability to write clean, efficient, scalable software for large, messy, real-world datasets
- Sound judgment for reasoning about data quality, risk, and utility—forming hypotheses, choosing meaningful metrics, diagnosing failures, and distinguishing signal from noise
- Experience designing systems—not only implementing specifications—including tradeoffs around quality, scale, reliability, and reuse
- Comfort operating in an ambiguous, fast-moving environment with substantial ownership
- Collaborative, low-ego communication and the ability to work effectively with researchers, engineers, domain experts, and customers
Nice to Haves
- Building or operating large-scale synthetic or LLM-generated data pipelines for model training
- Building or operating large-scale data de-identification or anonymization systems, ideally involving relational or graph-structured data, with experience preserving referential/relationship integrity after anonymization
- Developing LLM/agent benchmarks, evaluation methodologies, annotation systems, or data-quality frameworks
- Research or applied work on reinforcement learning, alignment, model behavior, synthetic data, or human-in-the-loop systems
- Prior work in a regulated or high-sensitivity data environment (healthcare, finance, HR/people data, government), or experience with re-identification risk assessment and privacy auditing
- Published research, meaningful open-source contributions, or evidence of technical leadership in ML systems, data engineering, or AI research
- Experience productizing research or repeated customer work into robust, reusable platforms
Benefits
- For full-time US employees: Equity in a fast-growing company
- For full-time US employees: 401(k) match
- For full-time US employees: financial coaching
- For full-time US employees: Paid parental leave
- For full-time US employees: fertility benefits
- For full-time US employees: parental coaching
- For full-time US employees: Medical, dental, and vision
- For full-time US employees: mental health support
- For full-time US employees: $500 wellness stipend
- For full-time US employees: $2,000 learning stipend
- For full-time US employees: ongoing development
- For full-time US employees: Commuting support
- For full-time US employees: free lunch
- For full-time US employees: gym in our SF office
- For full-time US employees: Flexible PTO
- For full-time US employees: 15 holidays + 2 flex days
- For full-time US employees: Team outings
- For full-time US employees: referral bonuses