Summary
Sown To Grow (STG) is a K12 education technology platform that empowers schools to improve student social, emotional, and academic health. They are seeking a Data Scientist who will work on building and refining machine learning models to support educators in understanding and assisting students, utilizing both classical and modern approaches.
Responsibilities
- Build and refine NLP/ML models over large volumes of unstructured student and educator text, using both classical approaches (feature engineering, classification, CNNs, ensemble methods) and modern LLM-based techniques where they're the right tool
- Partner with data science, product, and engineering to identify, define, and test opportunities to improve the product through ML/NLP - from turning existing models into user-facing features to prototyping entirely new capabilities
- Extend our models beyond 'proof of concept' scope: new reflection prompts, younger students etc, widening accessibility while protecting accuracy
- Design and apply LLMs responsibly for generative and assistive features - for example, contextualized teacher-response suggestions, and resource recommendations - with careful attention to prompting, retrieval, grounding, evaluation, and guardrails
- Build validation, monitoring, and retraining pipelines in partnership with the ML engineering team - including data-drift detection - so models keep performing as code and data change
- Undertake preprocessing of structured and unstructured data, and build reliable, reproducible feature and evaluation workflows
Skills
- Bachelor's or higher degree in Computer Science, Data Science, Machine Learning, Math, Statistics, or a related field
- 2+ years of experience as a Data Scientist, ML Engineer, or Data Engineer, solving real-world problems with machine learning
- Strong proficiency in Python and the ML stack (pandas, numpy, scikit-learn; PyTorch or TensorFlow; Spark a plus)
- Experience building and deploying ML solutions that involve natural language processing of text data
- Working knowledge of core ML techniques such as classification, clustering, prediction, recommender systems, and anomaly detection
- Working knowledge of the complete machine learning lifecycle - data, training, validation, deployment, monitoring, and retraining
- Solid understanding of how modern LLMs work under the hood - transformer architecture, training and fine-tuning, tokenization, embeddings. We care more about curiosity than credentials here: you enjoy digging into why a model behaves the way it does, not just what it returns. Hands-on with at least one of prompting/evaluation, fine-tuning, retrieval-augmented generation (RAG), or agentic/tool-use patterns, with a thoughtful view of when LLMs are and aren't the right approach
- Experience writing and maintaining high-quality production code, and comfort with Git-based workflows
- Strong interest in working in education technology in an impact-driven, mission-first role
- Experience productionizing ML for real-time, low-latency inference (e.g., AWS SageMaker or comparable), including containerization and CI/CD
- Experience building data-drift detection, model monitoring, and automated retraining systems in partnership with ML engineering teams
- Experience building, training, or fine-tuning language models from the ground up - e.g., implementing transformer components, training or adapting models on domain-specific data, or working with open-weight models beyond off-the-shelf APIs. This is a longer-term direction for us, and we value candidates who can grow into it
- Experience with responsible / trustworthy AI: fairness and bias evaluation, privacy-conscious handling of sensitive data, and building guardrails for user-facing generative features
- Experience designing human-in-the-loop evaluation and running online experiments (A/B testing, feature flagging)
- Familiarity with the practical, ethical, and legal considerations of working with student data
Qualifications
Must Haves
- Bachelor's or higher degree in Computer Science, Data Science, Machine Learning, Math, Statistics, or a related field
- 2+ years of experience as a Data Scientist, ML Engineer, or Data Engineer, solving real-world problems with machine learning
- Strong proficiency in Python and the ML stack (pandas, numpy, scikit-learn; PyTorch or TensorFlow; Spark a plus)
- Experience building and deploying ML solutions that involve natural language processing of text data
- Working knowledge of core ML techniques such as classification, clustering, prediction, recommender systems, and anomaly detection
- Working knowledge of the complete machine learning lifecycle - data, training, validation, deployment, monitoring, and retraining
- Solid understanding of how modern LLMs work under the hood - transformer architecture, training and fine-tuning, tokenization, embeddings. We care more about curiosity than credentials here: you enjoy digging into why a model behaves the way it does, not just what it returns. Hands-on with at least one of prompting/evaluation, fine-tuning, retrieval-augmented generation (RAG), or agentic/tool-use patterns, with a thoughtful view of when LLMs are and aren't the right approach
- Experience writing and maintaining high-quality production code, and comfort with Git-based workflows
Nice to Haves
- Strong interest in working in education technology in an impact-driven, mission-first role
- Experience productionizing ML for real-time, low-latency inference (e.g., AWS SageMaker or comparable), including containerization and CI/CD
- Experience building data-drift detection, model monitoring, and automated retraining systems in partnership with ML engineering teams
- Experience building, training, or fine-tuning language models from the ground up - e.g., implementing transformer components, training or adapting models on domain-specific data, or working with open-weight models beyond off-the-shelf APIs. This is a longer-term direction for us, and we value candidates who can grow into it
- Experience with responsible / trustworthy AI: fairness and bias evaluation, privacy-conscious handling of sensitive data, and building guardrails for user-facing generative features
- Experience designing human-in-the-loop evaluation and running online experiments (A/B testing, feature flagging)
- Familiarity with the practical, ethical, and legal considerations of working with student data
Benefits
- Comprehensive health and wellness benefits for you and your family
- Flexible work arrangements and a genuine commitment to work-life balance
- Real pathways for growth - as the platform and the data team expand, so does the scope of this role
- A collaborative, mission-driven community where every voice is heard
- Competitive compensation with performance-based incentives and meaningful equity