Summary
The National Laboratory of the Rockies (NLR) is focused on energy innovation through research and systems integration. They are seeking a graduate student researcher to investigate uncertainty quantification in LLM-based science assistants, emphasizing the detection of vague or underspecified scientific questions.
Responsibilities
- Research and evaluate uncertainty quantification and hallucination detection methods for multi-turn, agentic scientific workflows
- Develop probing methods that predict, from a model's internal representations, when a scientific task specification is incomplete or inconsistent and a clarifying question is warranted
- Build and instrument evaluation pipelines that capture and analyze model internal states over multi-turn scientific dialogue on HPC systems
- Conduct experiments and analyze model behavior across computational science domains and established benchmarks
- Contribute to technical documentation, research reports, publications, and presentations summarizing project progress and findings
- Develop, test, and maintain high-quality research code and evaluation pipelines
Skills
- Minimum of a 3.0 cumulative grade point average
- Undergraduate: Must be enrolled as a full-time student in a bachelor's degree program from an accredited institution
- Post Undergraduate: Earned a bachelor's degree within the past 12 months. Eligible for an internship period of up to one year
- Graduate: Must be enrolled as a full-time student in a master's degree program from an accredited institution
- Post Graduate: Earned a master's degree within the past 12 months. Eligible for an internship period of up to one year
- Graduate + PhD: Completed master's degree and enrolled as PhD student from an accredited institution
- Familiarity with large language models, including agentic, tool-using, or multi-turn conversational LLM systems
- Experience developing or evaluating machine learning models for classification, uncertainty estimation, or related tasks
- Knowledge of probabilistic machine learning or uncertainty quantification concepts
- Hands-on experience with open-weight LLMs and modern deep learning frameworks
- Experience running Python code on HPC or multi-GPU systems
- Strong software engineering and debugging skills
- Ability to work independently while collaborating effectively in a multidisciplinary research environment
- Research experience related to hallucination detection, uncertainty quantification, interpretability, explainability, or trustworthy AI
- Familiarity with representation probing or mechanistic interpretability methods
- Experience with LLM benchmarking and evaluation, including multi-turn or conversational agent evaluation and LLM-as-a-judge protocols
- Experience with scientific question-answering systems, AI for science applications, or scientific agent frameworks
- Coursework or research background in a computational science domain (e.g., fluid mechanics, solid mechanics, materials science, or numerical methods for PDEs)
Qualifications
Must Haves
- Minimum of a 3.0 cumulative grade point average
- Undergraduate: Must be enrolled as a full-time student in a bachelor's degree program from an accredited institution
- Post Undergraduate: Earned a bachelor's degree within the past 12 months. Eligible for an internship period of up to one year
- Graduate: Must be enrolled as a full-time student in a master's degree program from an accredited institution
- Post Graduate: Earned a master's degree within the past 12 months. Eligible for an internship period of up to one year
- Graduate + PhD: Completed master's degree and enrolled as PhD student from an accredited institution
- Familiarity with large language models, including agentic, tool-using, or multi-turn conversational LLM systems
- Experience developing or evaluating machine learning models for classification, uncertainty estimation, or related tasks
- Knowledge of probabilistic machine learning or uncertainty quantification concepts
- Hands-on experience with open-weight LLMs and modern deep learning frameworks
- Experience running Python code on HPC or multi-GPU systems
- Strong software engineering and debugging skills
- Ability to work independently while collaborating effectively in a multidisciplinary research environment
Nice to Haves
- Research experience related to hallucination detection, uncertainty quantification, interpretability, explainability, or trustworthy AI
- Familiarity with representation probing or mechanistic interpretability methods
- Experience with LLM benchmarking and evaluation, including multi-turn or conversational agent evaluation and LLM-as-a-judge protocols
- Experience with scientific question-answering systems, AI for science applications, or scientific agent frameworks
- Coursework or research background in a computational science domain (e.g., fluid mechanics, solid mechanics, materials science, or numerical methods for PDEs)
Benefits
- Medical, dental, and vision insurance
- 403(b) Employee Savings Plan with employer match*
- Sick leave (where required by law)
- NLR employees may be eligible for, but are not guaranteed, performance-, merit-, and achievement- based awards that include a monetary component
- Some positions may be eligible for relocation expense reimbursement
- Internships projected to be less than 20 hours per week are not eligible for medical, dental, or vision benefits
- Based on eligibility rules