Summary
ISC2 is a nonprofit member organization serving cybersecurity professionals and supporting a safer, more secure cyber world. The Lakehouse Machine Learning Engineer will build and maintain governed data pipelines, develop and validate machine learning models, deploy and monitor them in production, and deliver usable results to business stakeholders. The role also includes applied LLM work, security and governance compliance, and creating reusable engineering patterns.
Responsibilities
- Build and maintain Python/Spark pipelines through bronze, silver, and gold layers, plus the semantic datasets and ML models that consume them
- Dig into the data before modeling it. Work out what it can support, build and test the features that come out of that exploration, then coordinate with stakeholders about which questions are worth answering and which aren’t
- Work across a range of model types: survival and time-to-event, forecasting, classification and propensity, sequence models, recommenders, causal evaluation
- Put models into production and keep them there - Experiment tracking, model registry, scheduled inference, and monitoring for drift and decay once they’re live
- Get results to the people who need them. That means landing output in governed semantic tables feeding dashboards and CDP systems, and being able to walk a business team through what the numbers mean
- Take on applied LLM work as the platform matures, including structured extraction from free text, and retrieval over governed data
- Build inside the platform’s security and governance requirements rather than around them. Access controls, data protection, auditability, and human review where model output drives a decision that affects a member
- Turn what works into reusable patterns, including project templates, shared feature and evaluation code, and implementation standards, so the next model doesn’t start from a blank notebook
- Prove things out before they get built for real - Small proofs of concept that establish whether the data supports the theory, whether the approach holds up, and whether the result can actually be operated once it’s live
- Perform miscellaneous duties, as required
Skills
- Highly organized with strong attention to detail and documentation rigor
- Collaborative, intellectually curious, and proactive in identifying analytical opportunities
- Strong Extract/Transform/Load (ETL) skills, with the ability to assemble a dataset rather than request one
- Fluent Python and SQL skills, comfortable working across enterprise source systems, and fluent with the standard ML stack — scikit-learn at minimum, and at least one deep learning framework such as PyTorch or TensorFlow
- Knowledge range in ML, including knowing when timing matters enough to warrant survival analysis over a churn classifier, and being able to tell a causal question from a predictive one
- Familiarity with hyperparameter tuning, cross-validation, and the techniques used to confirm a model holds up on data it has not seen
- Ability to perform careful model validation, and to provide clarity about uncertainty
- Also, to perform model evaluation and bias mitigation, judgment about where a person needs to stay in the loop, and the ability to explain the result to executives in non-technical terms
- Understanding of data security, privacy, and compliance constraints, as well as how they shape what can be built with member data
- Ability to work within the confines of access controls, data protection, and auditability
- 3+ years of hands-on experience in data engineering and applied machine learning, preferably in enterprise environments
- Experience deploying and monitoring models in production. ML flow *o*r an equivalent tracking, registry, and scheduled-inference stack
- Production experience building in a medallion architecture — bronze, silver, and gold, or an equivalent layered model — in a data-catalog-governed environment. Distributed processing with Spark or a comparable engine, an open table format such as Delta or Iceberg, and catalog-managed schemas, lineage, and access control. Demonstrated experience building and operating in that environment, not just querying it
- Up to 5% travel may be required
- Work normal business hours and extended hours when necessary
- Remain in a stationary position, often standing or sitting, for prolonged periods
- Regular use of office equipment in a remote environment such as a computer/laptop and monitor computer screens
- Dexterity of hands and fingers to operate a computer keyboard, mouse, and other computer components
- This position is not available to residents of California
- Knowledge of Databricks, including Unity Catalog, Workflows, MLflow, a plus
- Ability to perform cohort-based or hierarchical forecasting at scale, a plus
- Working knowledge of Salesforce, a plus
- Relevant certifications: Databricks Data Engineer or Machine Learning Associate/Professional, Azure AI Engineer or Data Scientist Associate, or equivalent, a plus
- Bachelor's or Master's degree in an IT field preferred
- Will consider candidates with a high school diploma or equivalent and 7+ years of hands-on experience in data engineering and applied machine learning, preferably in enterprise environments
- Practical Large Language Model (LLM) experience including embeddings and retrieval, structured extraction, evaluation, a plus
- Experience with subscription or membership-lifecycle data, a plus
- Experience running a build-versus-buy evaluation. Hands-on assessment of tools and vendors, and a recommendation that can be defended, a plus
Qualifications
Must Haves
- Highly organized with strong attention to detail and documentation rigor
- Collaborative, intellectually curious, and proactive in identifying analytical opportunities
- Strong Extract/Transform/Load (ETL) skills, with the ability to assemble a dataset rather than request one
- Fluent Python and SQL skills, comfortable working across enterprise source systems, and fluent with the standard ML stack — scikit-learn at minimum, and at least one deep learning framework such as PyTorch or TensorFlow
- Knowledge range in ML, including knowing when timing matters enough to warrant survival analysis over a churn classifier, and being able to tell a causal question from a predictive one
- Familiarity with hyperparameter tuning, cross-validation, and the techniques used to confirm a model holds up on data it has not seen
- Ability to perform careful model validation, and to provide clarity about uncertainty
- Also, to perform model evaluation and bias mitigation, judgment about where a person needs to stay in the loop, and the ability to explain the result to executives in non-technical terms
- Understanding of data security, privacy, and compliance constraints, as well as how they shape what can be built with member data
- Ability to work within the confines of access controls, data protection, and auditability
- 3+ years of hands-on experience in data engineering and applied machine learning, preferably in enterprise environments
- Experience deploying and monitoring models in production. ML flow *o*r an equivalent tracking, registry, and scheduled-inference stack
- Production experience building in a medallion architecture — bronze, silver, and gold, or an equivalent layered model — in a data-catalog-governed environment. Distributed processing with Spark or a comparable engine, an open table format such as Delta or Iceberg, and catalog-managed schemas, lineage, and access control. Demonstrated experience building and operating in that environment, not just querying it
- Up to 5% travel may be required
- Work normal business hours and extended hours when necessary
- Remain in a stationary position, often standing or sitting, for prolonged periods
- Regular use of office equipment in a remote environment such as a computer/laptop and monitor computer screens
- Dexterity of hands and fingers to operate a computer keyboard, mouse, and other computer components
- This position is not available to residents of California
Nice to Haves
- Knowledge of Databricks, including Unity Catalog, Workflows, MLflow, a plus
- Ability to perform cohort-based or hierarchical forecasting at scale, a plus
- Working knowledge of Salesforce, a plus
- Relevant certifications: Databricks Data Engineer or Machine Learning Associate/Professional, Azure AI Engineer or Data Scientist Associate, or equivalent, a plus
- Bachelor's or Master's degree in an IT field preferred
- Will consider candidates with a high school diploma or equivalent and 7+ years of hands-on experience in data engineering and applied machine learning, preferably in enterprise environments
- Practical Large Language Model (LLM) experience including embeddings and retrieval, structured extraction, evaluation, a plus
- Experience with subscription or membership-lifecycle data, a plus
- Experience running a build-versus-buy evaluation. Hands-on assessment of tools and vendors, and a recommendation that can be defended, a plus