Ralco logo
Ralco
Posted 2 days agoVerified live 5h ago

Data & Discovery Engineer

Brief overview

Remote
MastersOr in progress
3+ yrsMinimum
Data Pipeline DesignSQLRelational Database DesignLIMS IntegrationMachine LearningAnomaly DetectionClustering and Similarity AlgorithmsPermissioned Data Access SystemsPatent Database Searches

About the company

Since 1971, Ralco has been helping farmers and ranchers raise healthier crops and livestock.

Job description

Summary

Ralco is an agricultural health and nutrition company developing natural solutions for plant and animal health. The Data & Discovery Engineer will build and maintain data infrastructure, databases, and direct-entry systems while ensuring data quality and security. The role will also develop and validate machine learning tools, review research data for insights, and support R&D through literature and patent searches.

Responsibilities

  • Design, build, and continuously improve the pipeline connecting lab instruments, in-vitro work, and barn/field data sources into the shared database
  • Build, or work with IT to build, a LIMS-style direct-entry system, owning the correct entry, tagging, organizing, and storing of it correctly in the shared database
  • Design and maintain the database schema and tagging conventions so the data can be queried and modeled correctly
  • Work with IT to design and validate permission, gated access controls that determine who can see sensitive data, testing them before any sensitive data goes live
  • Ensure new compound and trial data is entered with a complete, validated profile
  • Resolving data-quality issues directly with the originating scientist before any correction is made, and lead periodic data audits to reconcile duplicate or conflicting records and standardize naming conventions
  • Build and maintain anomaly detection, compound similarity/clustering, and pattern-flagging tools that are useful well before a full predictive model is trustworthy
  • Build, retrain, and validate the models used for querying the database on a documented cadence tied to data volume, tracking hit-rate before and after each cycle and documenting the trend
  • Work with the R&D team to make sure the model used for querying stays the most accurate and correct fit for the data, revisiting its underlying design when it no longer fits rather than assuming it always will
  • Actively review data across active research projects on an ongoing basis, bringing forward insights, trends, or connections for the science team to think through and possibly act on
  • Conduct regular cross-project data reviews to help surface new angles or connections for the R&D team on active research tracks
  • Run patent searches and compile literature searches on request to support hypothesis-building, manuscript preparation, and Regulatory's IP review, maintaining a running log so the same ground is never covered twice
  • Produce a quarterly gap-audit report, and a brief report after every model retraining cycle documenting what changed and what the new hit-rate is
  • Follow, and help enforce, Ralco's IP protection practices for anything touching the shared database, escalating promptly any data handling or access issue that could put proprietary information at risk
  • Regularly meet with the R&D and IT teams to stay closely connected to their work, translating between researcher needs and data architecture as new data types and sources come online

Skills

  • Bachelor's or Master's degree in Data Science, Data Engineering, Bioinformatics, Computer Science, or a closely related field
  • 3–6 years of experience in data science, data engineering, bioinformatics, master data management, or a closely related field
  • Real experience designing data pipelines
  • SQL and relational database concepts
  • Practical, hands-on machine learning experience
  • Experience with permissioned/gated data-access systems
  • Experience building or integrating a LIMS-style direct-entry system
  • Direct, hands-on experience building a predictive model end-to-end: preparing training data, training and validating a model, and deploying and retraining it as new data comes in, not just coursework or theoretical exposure
  • Experience designing a database schema and tagging system from scratch for messy, real-world data, deciding how to structure and label information so it stays queryable and useful as it grows, not just working within a schema someone else already built
  • Demonstrated self-directed learning; concrete examples of teaching oneself a new technical or scientific domain without formal instruction
  • Shows curiosity about the underlying science, not just the data structure; asks why data looks the way it does
  • Some background or coursework in life sciences, chemistry, or animal/agricultural science
  • Experience running patent database searches

Qualifications

Must Haves

  • Bachelor's or Master's degree in Data Science, Data Engineering, Bioinformatics, Computer Science, or a closely related field
  • 3–6 years of experience in data science, data engineering, bioinformatics, master data management, or a closely related field
  • Real experience designing data pipelines
  • SQL and relational database concepts
  • practical, hands-on machine learning experience
  • experience with permissioned/gated data-access systems

Nice to Haves

  • Experience building or integrating a LIMS-style direct-entry system
  • Direct, hands-on experience building a predictive model end-to-end: preparing training data, training and validating a model, and deploying and retraining it as new data comes in, not just coursework or theoretical exposure
  • Experience designing a database schema and tagging system from scratch for messy, real-world data, deciding how to structure and label information so it stays queryable and useful as it grows, not just working within a schema someone else already built
  • Demonstrated self-directed learning; concrete examples of teaching oneself a new technical or scientific domain without formal instruction
  • Shows curiosity about the underlying science, not just the data structure; asks why data looks the way it does
  • Some background or coursework in life sciences, chemistry, or animal/agricultural science
  • Experience running patent database searches

Benefits

  • Remote (US)

More jobs like this