Summary
RCH Solutions is seeking a Data Engineer specialized in knowledge graphs and semantic technologies to join its Data and AI Engineering team. The role focuses on designing and deploying Stardog knowledge graphs, modeling ontologies, integrating heterogeneous life science data, automating graph lifecycles, and supporting data quality, governance, and semantic architecture.
Responsibilities
- Design, build and evolve knowledge graphs in Stardog, from conceptual model through to production deployment
- Model domain ontologies, taxonomies and vocabularies using RDF, RDFS, OWL and SKOS, and enforce them with SHACL constraints
- Write, optimise and troubleshoot SPARQL queries, rules and inference over large graphs
- Integrate heterogeneous sources into the graph using virtual graphs and mappings (R2RML and similar) from relational databases, APIs, files and semi-structured data
- Run discovery sessions with subject matter experts, turning business questions into competency questions and a defensible semantic model
- Align internal models with life science standards and public ontologies, and manage identifier mapping and entity resolution across sources
- Automate graph builds, tests and deployments through CI/CD pipelines and Python tooling
- Embed data quality, validation and reconciliation checks into the graph lifecycle
- Document models and enable others — governance, lineage, and reusable semantic assets that outlive the project
- Work in agile teams, contributing to standups, retrospectives, and continuous improvement
Skills
- Hands-on experience delivering production knowledge graph solutions with Stardog. Experience with other RDF triplestores (**GraphDB, Amazon Neptune, Virtuoso, Anzo**) counts as transferable if you're ready to go deep on Stardog
- Strong command of**semantic web standards**: RDF, RDFS, OWL, SKOS, SHACL and SPARQL
- Practical **ontology and taxonomy modelling** - able to move from stakeholder conversations and messy source data to a model that holds up in production
- Experience **mapping and virtualising** relational and semi-structured sources into a graph
- Solid **Python** and **SQL** for data preparation, transformation, automation and troubleshooting
- Comfortable with **Git-based workflows and CI/CD** (GitHub Actions or Azure DevOps)
- Experience working with **life science or healthcare data**, and comfortable with the quality and regulatory expectations that come with it
- **Autonomy and ownership**: you scope your own work, propose an approach, defend it, and bring the team along — rather than waiting for a fully specified ticket
- Working knowledge of **data quality**, validation frameworks, and test-driven data development
- Team-first mindset and experience in **agile environments**(Scrum or Kanban)
- Role is only open to applicants not needing sponsorship now or in the future
- Familiarity with public life science ontologies and terminologies (e.g. SNOMED CT, MeSH, ChEBI, UMLS, LOINC)
- Exposure to at least one life science domain: clinical and clinical trial data (CDISC, SDTM), R&D and drug discovery, regulatory (RIM, IDMP), or manufacturing, supply chain and quality
- Understanding of GxP or other healthcare data regulations
- Familiarity with FAIR data principles
- Experience combining graphs with AI — GraphRAG, vector search, or LLM-assisted ontology work
- Exposure to property graphs (e.g. Neo4j) and how they compare with RDF
- Knowledge of data lineage, catalog and governance tooling
- Infrastructure automation using Terraform, Bash, or PowerShell, and containers (Docker, Kubernetes)
Qualifications
Must Haves
- Hands-on experience delivering production knowledge graph solutions with Stardog. Experience with other RDF triplestores (**GraphDB, Amazon Neptune, Virtuoso, Anzo**) counts as transferable if you're ready to go deep on Stardog
- Strong command of**semantic web standards**: RDF, RDFS, OWL, SKOS, SHACL and SPARQL
- Practical **ontology and taxonomy modelling** - able to move from stakeholder conversations and messy source data to a model that holds up in production
- Experience **mapping and virtualising** relational and semi-structured sources into a graph
- Solid **Python** and **SQL** for data preparation, transformation, automation and troubleshooting
- Comfortable with **Git-based workflows and CI/CD** (GitHub Actions or Azure DevOps)
- Experience working with **life science or healthcare data**, and comfortable with the quality and regulatory expectations that come with it
- **Autonomy and ownership**: you scope your own work, propose an approach, defend it, and bring the team along — rather than waiting for a fully specified ticket
- Working knowledge of **data quality**, validation frameworks, and test-driven data development
- Team-first mindset and experience in **agile environments**(Scrum or Kanban)
- Role is only open to applicants not needing sponsorship now or in the future
Nice to Haves
- Familiarity with public life science ontologies and terminologies (e.g. SNOMED CT, MeSH, ChEBI, UMLS, LOINC)
- Exposure to at least one life science domain: clinical and clinical trial data (CDISC, SDTM), R&D and drug discovery, regulatory (RIM, IDMP), or manufacturing, supply chain and quality
- Understanding of GxP or other healthcare data regulations
- Familiarity with FAIR data principles
- Experience combining graphs with AI — GraphRAG, vector search, or LLM-assisted ontology work
- Exposure to property graphs (e.g. Neo4j) and how they compare with RDF
- Knowledge of data lineage, catalog and governance tooling
- Infrastructure automation using Terraform, Bash, or PowerShell, and containers (Docker, Kubernetes)
Benefits
- A competitive salary and bonus package based on experience.
- Comprehensive health and wellness benefits, including Medical, Dental, and Vision Insurance.
- Company-provided Life and Long-Term Disability Insurance.
- Company-sponsored 401(k) Plan.
- Team-focused culture and unlimited opportunity for advancement.
- Remote role, with availability required to work on an East Coast (US) time schedule.