Principal Data Product Engineer — Data Ontology & Metadata
Be a key member of the Advanced Robotics & Data Orchestration Team, building the semantic data backbone for Alexion’s Product Development & Clinical Supply (PDCS) organization. Help transform our multimodality product development labs from manual, time intensive workflows, into an automated, scalable, and scientist-friendly ecosystem by defining ontologies, vocabularies, and metadata models that accelerate decisions, strengthen data integrity, and ready the organization for advanced automation and AI/ML. As the team’s primary interface to the Alexion Data Office (ADO) and the lead point of contact to Alexion IT, you will drive bidirectional alignment on data standards and platform implementation while partnering closely with Lab Automation and Orchestration leads. Help us accelerate life-changing therapies to patients with rare and ultra-rare diseases.
This is what you will do:
Design ontologies and metadata models for key domains (chromatography, spectroscopy, assays, automation runs, process parameters, samples/lineage) with governance, versioning, and change control.
Define and publish data contracts and canonical schemas; embed standards into ingestion, validation, and transformation; enforce lineage, provenance, and data quality rules.
Build reference implementations (schema registry, conformance tests, enrichment services) integrated with Databricks on AWS and enterprise catalogs; align with ELN/LIMS/SDMS connectors.
Curate data products exposing consistent entities/relationships; enable self-service discovery, governed APIs/views, and downstream ML feature stores.
Lead cross functional standards forums; deliver documentation, examples, and training; measure adoption, conformance, and time to data improvements.
Serve as primary liaison to ADO and Alexion IT for data standards, catalogs, and governance tooling; maintain a bidirectional communication cadence with Lab Automation and Orchestration leads to ensure implementation fidelity.
You will be responsible for:
Enterprise grade ontology/vocabulary releases with governance artifacts (stewardship roles, RACI, version notes) and deprecation/change communications.
Contracted schemas and metadata policies embedded into pipelines and catalogs, including automated conformance checks, issue triage, and remediation guidance.
Cataloged, discoverable data products with lineage and role-based access controls suitable for audit sensitive environments, with periodic quality reviews and reporting.
A roadmap and backlog for semantic standards aligned with product and platform roadmaps, measurable gains in adoption, conformance, time to data, and reuse across teams.
External/internal stakeholder alignment: single point ownership for ADO/IT coordination on standards, catalogs, and reference implementations supporting PDCS use cases.
You will need to have:
Education/experience: PhD + 4 years, or MS + 8 years, or BS + 12 years in Computer Science, Information Systems, Data Engineering, or related— including 7+ years of hands-on data modeling, ontology, and metadata engineering with standards embedded in production pipelines and experience with lakehouse/catalog platforms (for example, Databricks on AWS).
Demonstrated skills in semantic modeling, schema design, data contracts, lineage/provenance, and data quality frameworks/governance tooling.
Proficiency in Python/SQL for implementation and validation.
Strong cross functional communication and the ability to drive consensus across data, IT, and laboratory stakeholders.
Ability to travel to New Haven, CT ~4x/year, 2–4 weeks per trip, to work face to face with laboratory practitioners and gain hands on familiarity with equipment and robotics.
The duties of this role may require periodic work in a laboratory or manufacturing environment. As is typical of such roles, employees must be able, with or without an accommodation to: lift/carry 15/30 pounds unassisted/assisted; work comfortably in a controlled environment with and around hazardous materials; gown/degown PPE; use a computer; engage in communications via phone, video, and electronic messaging; engage in problem solving and non-linear thought, analysis, and dialogue; collaborate with others; maintain general availability during standard business hours.
We would prefer for you to have:
CMC/lab data familiarity (ELN/LIMS/SDMS, instrument domains) and experience operationalizing metadata services and ML feature workflows; domain experience preferred but not required.
Date Posted
22-Jul-2026Closing Date
05-Aug-2026Our mission is to build an inclusive and equitable environment. We want people to feel they belong at AstraZeneca and Alexion, starting with our recruitment process. We welcome and consider applications from all qualified candidates, regardless of characteristics. We offer reasonable adjustments/accommodations to help all candidates to perform at their best. If you have a need for any adjustments/accommodations, please complete the section in the application form.We’ll keep you up-to-date