Summary
ScienceLogic is redefining IT operations for the modern enterprise. They are seeking a strong Data Scientist to join their growing Data Science team, responsible for optimizing language models and ensuring high-quality outcomes through evaluation and predictive modeling.
Responsibilities
- Design and own evaluation harnesses for LLM and agentic outputs — golden sets, regression suites, and rubric-based scoring
- Build and calibrate LLM-as-judge pipelines; validate judges against human labels and control for their bias and variance
- Define and track response-quality metrics: faithfulness/groundedness, hallucination rate, answer relevance and completeness, instruction-following, and persona adherence
- Curate, version, and grow evaluation datasets as the product and its surfaces evolve
- Benchmark the models in the suite against each other to decide which model handles which task, and quantify the quality cost of running smaller, local models versus larger alternatives
- Red-team the system: prompt injection, jailbreaks, tool-misuse, and edge-case discovery
- Design chaos and stress tests that probe model and agent reliability under degraded or hostile conditions
- Characterize failure modes and feed them back into guardrails and regression coverage
- Evaluate retrieval quality over the document corpus — recall@k, MRR/nDCG, context precision and recall — and run experiments on chunking, indexing, and hybrid retrieval strategies
- Analyze multi-step agent trajectories: tool-call correctness, trajectory efficiency, replayable-state inspection, and guardrail-breach behavior
- Assess intent classification and routing quality as measurable components, not black boxes
- Build standing evaluation that catches quality and behavioral regressions when a model in the suite is swapped, upgraded, or re-quantized, or when prompts and pipelines change
- Monitor output-distribution and quality drift in production; distinguish genuine regressions from noise on stochastic outputs
- Recommend and validate fixes through the levers available with local models — prompt changes, retrieval and grounding adjustments, routing changes, or model selection
- Build, ship, and own production models that forecast and surface trends from operational telemetry — capacity and resource forecasting, anomaly prediction, and early-warning signals on metrics and logs
- Take these from prototype to production and keep them healthy: deployment, monitoring, recalibration, and retraining as data and behavior shift
- Define accuracy and lead-time metrics that matter operationally — precision/recall on predicted incidents, forecast error, how far ahead a signal fires — not just offline scores
- Wire predictive signals into the LLM and agentic layer so forecasts and trends feed reasoning, advisories, and operator-facing recommendations
- Apply AIOps/NOC analysis where it's the product: log anomaly detection, event correlation, and root-cause and problem analysis
- Quantify the economics of the system — cost and token consumption per interaction, interaction-type taxonomies — and connect them to customer-facing value metrics like MTTR and operator-hours
- Communicate findings to engineering and product stakeholders through clear, in-context analysis
- Use LLM-assisted workflows to scale the work itself — drafting analyses, generating synthetic evaluation cases, and bootstrapping labeled data for human refinement
- Track and adopt state-of-the-art evaluation, retrieval, and agentic-analysis techniques; bring the useful ones into the team's workflow
Skills
- Bachelor's or Master's in Data Science, Computer Science, Statistics, Mathematics, or a related field or equivalent experience
- 3+ years in data science, ML, or applied quantitative analysis
- Strong applied statistics, with the judgment to design sound experiments and significance tests on noisy, non-deterministic outputs (not just clean A/B conversion)
- Experience building, deploying, and monitoring predictive or time-series models in production: forecasting, anomaly detection, or trend analysis, including recalibration as data shifts
- Demonstrated work evaluating, analyzing, or improving LLM or NLP systems: eval design, quality measurement, retrieval evaluation, or agent analysis
- Proficiency in Python
- Strong SQL and comfort querying large analytical datasets
- Fluency with foundation models and hands-on experience with the modern LLM evaluation and tooling layer — eval/harness frameworks, judge pipelines, and the libraries used to serve, prompt, and test models
- Ability to build analysis and visualization in code
- Experience getting strong results out of small or self-hosted/local models under compute, memory, or latency constraints — quantization-aware evaluation, prompt and context optimization, or model routing
- Experience with retrieval-augmented systems and retrieval evaluation at scale
- Experience with agentic frameworks and tool-use/orchestration analysis, including human-in-the-loop and replayable-state patterns
- Familiarity with red-teaming or adversarial robustness for LLMs
- Domain background in IT operations — AIOps, NOC, ITSM, observability, or anomaly detection on logs and telemetry
- Experience with large-scale analytical and big-data stores
- Cloud experience for data science and ML workloads
- Exposure to enterprise security and compliance constraints in a delivery context
Qualifications
Must Haves
- Bachelor's or Master's in Data Science, Computer Science, Statistics, Mathematics, or a related field or equivalent experience
- 3+ years in data science, ML, or applied quantitative analysis
- Strong applied statistics, with the judgment to design sound experiments and significance tests on noisy, non-deterministic outputs (not just clean A/B conversion)
- Experience building, deploying, and monitoring predictive or time-series models in production: forecasting, anomaly detection, or trend analysis, including recalibration as data shifts
- Demonstrated work evaluating, analyzing, or improving LLM or NLP systems: eval design, quality measurement, retrieval evaluation, or agent analysis
- Proficiency in Python
- Strong SQL and comfort querying large analytical datasets
- Fluency with foundation models and hands-on experience with the modern LLM evaluation and tooling layer — eval/harness frameworks, judge pipelines, and the libraries used to serve, prompt, and test models
- Ability to build analysis and visualization in code
Nice to Haves
- Experience getting strong results out of small or self-hosted/local models under compute, memory, or latency constraints — quantization-aware evaluation, prompt and context optimization, or model routing
- Experience with retrieval-augmented systems and retrieval evaluation at scale
- Experience with agentic frameworks and tool-use/orchestration analysis, including human-in-the-loop and replayable-state patterns
- Familiarity with red-teaming or adversarial robustness for LLMs
- Domain background in IT operations — AIOps, NOC, ITSM, observability, or anomaly detection on logs and telemetry
- Experience with large-scale analytical and big-data stores
- Cloud experience for data science and ML workloads
- Exposure to enterprise security and compliance constraints in a delivery context
Benefits
- Comprehensive medical, dental and vision plans.
- 401(k) plan with employer match.
- Flexible Paid Time Off (FTO) so that you can take the time that you need to re-energize.
- Volunteer Time Off (VTO) - take two days off per calendar year to volunteer with your preferred charitable organization.
- 5-year Service Milestone Sabbatical.
- Paid parental leave.
- Generous employee referral bonus program.
- Pet insurance.
- HQ Office centrally located in Reston Town Center featuring a well-stocked kitchen with rotating snacks and beverages, and catered lunch on Thursdays.
- Regular virtual company-wide events, including cooking classes, yoga, meditation and more.
- The opportunity to learn and develop from some of the best and brightest minds in the industry!