Medasource

Senior Data Engineer

⭐ - Featured Role | Apply direct with Data Freelance Hub
This role is for a Senior Data Engineer on a 6-month contract to hire, focusing on clinical and real-world healthcare data. Key skills include Databricks, OMOP, Python, SQL, and cloud experience, particularly in Azure.
🌎 - Country
United States
πŸ’± - Currency
$ USD
-
πŸ’° - Day rate
800
-
πŸ—“οΈ - Date
August 10, 2026
πŸ•’ - Duration
More than 6 months
-
🏝️ - Location
Unknown
-
πŸ“„ - Contract
Unknown
-
πŸ”’ - Security
Unknown
-
πŸ“ - Location detailed
Ohio, United States
-
🧠 - Skills detailed
#Observability #Datasets #"ETL (Extract #Transform #Load)" #Azure #Spark SQL #Redshift #Databricks #Data Transformations #Data Quality #Data Governance #Normalization #PySpark #SQL (Structured Query Language) #ML (Machine Learning) #Spark (Apache Spark) #Data Mapping #Scala #Data Modeling #Data Engineering #Delta Lake #Cloud #Snowflake #Python #Data Warehouse #BigQuery #Data Pipeline #Airflow #dbt (data build tool) #Data Analysis
Role description
Senior Data Engineer - 6 month contract to hire Overview We are seeking a Senior Data Engineer with deep experience working with clinical and real-world healthcare data. This role will focus on building and scaling data pipelines that support analytics, research, and downstream machine learning use cases. The ideal candidate has hands-on experience with OMOP, Databricks, and modern data stacks, and understands the real-world challenges of clinical data harmonization across disparate healthcare sources. Key Responsibilities β€’ Design, build, and maintain scalable data pipelines for large, complex clinical datasets (EHR, pathology, genomics, etc.) β€’ Implement and manage data transformations and analytics workflows using Databricks (Spark, Delta Lake) β€’ Ingest, standardize, and harmonize healthcare data into OMOP Common Data Model β€’ Partner with clinical, analytics, and ML teams to ensure data is reliable, well-documented, and fit for downstream use β€’ Lead data quality, validation, and observability efforts for clinical data pipelines β€’ Develop data models and schemas that support analytics, research, and ML use cases β€’ Optimize performance, cost, and reliability across the data platform β€’ Contribute to best practices around data governance, versioning, lineage, and reproducibility β€’ Taking data analysis requirements from commercial customers and mapping to clinical variables from the OMOP, Epic, or other data models Required Qualifications β€’ 5+ years of experience as a Data Engineer, β€’ Strong hands-on experience with Databricks (Spark SQL, PySpark, Delta Lake) β€’ Deep understanding of OMOP CDM, including: β€’ Standard vocabularies (SNOMED, LOINC, RxNorm, ICD, CPT) β€’ ETL patterns for clinical data mapping and normalization β€’ Experience with clinical data harmonization, including: β€’ Mapping heterogeneous source systems into a common schema β€’ Managing missing, inconsistent, or conflicting clinical data β€’ Understanding clinical workflows and data provenance β€’ Strong cloud experience, preferably in Azure, relating to items such as Data Factory and other data related tooling β€’ Proficiency in Python and SQL β€’ Experience with modern data stacks, including: β€’ Cloud data warehouses or lakehouses (Databricks, Snowflake, BigQuery, Redshift) β€’ Orchestration tools (Airflow, Dagster, Prefect) β€’ Data transformation frameworks (dbt or equivalent) β€’ Strong data modeling and analytics engineering skills β€’ Clarity and Caboodle Certifications