Role: Data Modeler Location: Remote USA Employment type: Fulltime
Job Overview:
The Senior Data Modeler will design and govern the data architecture for unstructured knowledge assets across Knowledge Management (KM) ecosystem.
This role bridges data engineering discipline with KM domain expertise, translating raw unstructured content (documents, case files, informal knowledge captures, chat/email extracts, etc.) into well-defined, discoverable, and secure data products within Databricks Unity Catalog.
This is a foundational hire for a newly formed KM Data Platform team supporting broader Knowledge and Research Systems strategy.
Key Responsibilities:
Design logical and physical data models for unstructured and semi-structured content (documents, case artifacts, K-Slices, extracted knowledge fragments, metadata records) originating from KM pipelines such as case mining and informal knowledge capture workflows.
Define domain boundaries and ownership for data products — determining what constitutes a discrete, reusable data product versus a raw or intermediate asset.
Establish metadata standards and tagging taxonomies (content type, practice/domain, provenance, confidentiality, freshness, lineage) to ensure consistent classification across knowledge sources.
Assign and enforce security and sensitivity classifications on data products in line with firm data governance, privacy, and legal/risk requirements.
Register, document, and maintain data products in Databricks Unity Catalog, including schemas, access grants, lineage, and catalog-level metadata.
Partner with data engineers building Databricks pipelines to ensure ingestion, transformation, and storage patterns align to the modeled domain structure.
Collaborate with Knowledge Products, Research Products, and Architecture/Data/Technology stakeholders to align data product design with downstream consumption needs (e.g., surfacing in Sage/Glean, AI agent retrieval).
Support privacy and legal review processes by ensuring data products are classified and documented to enable timely sign-off.
Establish and document repeatable modeling standards/playbooks so future data products can be onboarded consistently as the KM platform scales.