

Xaxis Solutions
LLM Data Engineer
β - Featured Role | Apply direct with Data Freelance Hub
This role is for an LLM Data Engineer on a contract basis, focusing on healthcare data. Key skills include AWS expertise, data pipeline development, and strong AI & LLM experience. Contract length and pay rate are unspecified.
π - Country
United States
π± - Currency
$ USD
-
π° - Day rate
Unknown
-
ποΈ - Date
August 13, 2026
π - Duration
Unknown
-
ποΈ - Location
Unknown
-
π - Contract
Unknown
-
π - Security
Unknown
-
π - Location detailed
United States
-
π§ - Skills detailed
#Athena #AI (Artificial Intelligence) #Data Quality #Logging #Security #Terraform #Batch #Storage #ECR (Elastic Container Registery) #Datasets #API (Application Programming Interface) #S3 (Amazon Simple Storage Service) #Metadata #Data Lifecycle #VPC (Virtual Private Cloud) #IAM (Identity and Access Management) #Data Engineering #Data Modeling #Normalization #Docker #AWS Glue #Spark (Apache Spark) #AWS (Amazon Web Services) #Data Pipeline #SQL (Structured Query Language) #Lambda (AWS Lambda) #"ETL (Extract #Transform #Load)" #FHIR (Fast Healthcare Interoperability Resources) #Infrastructure as Code (IaC) #Data Lineage
Role description
Must-Have Requirements
β’ Strong hands-on experience with AWS
β’ Experience working with data sets, data sources, and AWS data services
β’ Strong AI & LLM experience
β’ Excellent communication skills
β’ Active and complete LinkedIn profile
β’ Healthcare experience is a plus, but not required
Role Overview
We are looking for a Generalist Data Engineer to support a healthcare-focused AI benchmark and evaluation platform.
The engineer will be responsible for the data lifecycle before a model receives it and after the model generates a response. This includes building ingestion, normalization, packaging, storage, and analysis layers to transform raw healthcare data into standardized benchmark inputs and actionable evaluation results.
Key Responsibilities
β’ Build ingestion and normalization pipelines for:
β’ DICOM radiology studies
β’ Whole-slide pathology images
β’ Tabular EHR data, including labs, vitals, encounters, and medication records
β’ Design a canonical benchmark record format that can support different task configurations and model adapters
β’ Develop strategies for packaging large medical imaging datasets within third-party API limitations
β’ Build tiling, region selection, downsampling, and compression workflows while maintaining diagnostic information
β’ Maintain provenance metadata to ensure results are reproducible and defensible
β’ Develop cohort and label pipelines for clinical prediction tasks such as sepsis onset, survival horizons, and longitudinal lab trends
β’ Build results storage and analysis layers for per-run, per-model, and per-task outputs
β’ Enforce PHI handling requirements, including encryption, least-privilege access, audit logging, and de-identification
β’ Ensure clear controls around data leaving the VPC when interacting with third-party APIs
Required Skills:
Data Engineering
β’ Production-grade data pipeline development
β’ Strong testing discipline
β’ Experience building deterministic, idempotent, and re-runnable jobs
AWS Data Stack
β’ Deep hands-on experience with S3
β’ AWS Glue and/or Spark on EMR
β’ Athena
β’ Step Functions
β’ Lambda
β’ AWS Batch
β’ Understanding of storage layout, lifecycle policies, and storage economics at scale
SQL & Data Modeling
β’ Complex temporal joins
β’ Point-in-time correctness
β’ Strong understanding of preventing label leakage in time-series data
Data Quality & Lineage
β’ Data validation frameworks
β’ Schema enforcement
β’ Versioned datasets
β’ Data lineage and reproducibility
AWS Security & Governance
β’ IAM policy design
β’ KMS
β’ VPC endpoints and PrivateLink
β’ Experience working within HIPAA-eligible AWS architectures
Desirable Skills
β’ Healthcare data standards including DICOM, FHIR, HL7v2, OMOP CDM
β’ Familiarity with clinical coding systems such as ICD, LOINC, RxNorm, and SNOMED
β’ Medical imaging experience with pydicom and OpenSlide
β’ Understanding of WSI pyramid structures and tiling
β’ Terraform / Infrastructure as Code
β’ Docker, ECR, and CI/CD
β’ Familiarity with LLM APIs and multimodal payload construction
Must-Have Requirements
β’ Strong hands-on experience with AWS
β’ Experience working with data sets, data sources, and AWS data services
β’ Strong AI & LLM experience
β’ Excellent communication skills
β’ Active and complete LinkedIn profile
β’ Healthcare experience is a plus, but not required
Role Overview
We are looking for a Generalist Data Engineer to support a healthcare-focused AI benchmark and evaluation platform.
The engineer will be responsible for the data lifecycle before a model receives it and after the model generates a response. This includes building ingestion, normalization, packaging, storage, and analysis layers to transform raw healthcare data into standardized benchmark inputs and actionable evaluation results.
Key Responsibilities
β’ Build ingestion and normalization pipelines for:
β’ DICOM radiology studies
β’ Whole-slide pathology images
β’ Tabular EHR data, including labs, vitals, encounters, and medication records
β’ Design a canonical benchmark record format that can support different task configurations and model adapters
β’ Develop strategies for packaging large medical imaging datasets within third-party API limitations
β’ Build tiling, region selection, downsampling, and compression workflows while maintaining diagnostic information
β’ Maintain provenance metadata to ensure results are reproducible and defensible
β’ Develop cohort and label pipelines for clinical prediction tasks such as sepsis onset, survival horizons, and longitudinal lab trends
β’ Build results storage and analysis layers for per-run, per-model, and per-task outputs
β’ Enforce PHI handling requirements, including encryption, least-privilege access, audit logging, and de-identification
β’ Ensure clear controls around data leaving the VPC when interacting with third-party APIs
Required Skills:
Data Engineering
β’ Production-grade data pipeline development
β’ Strong testing discipline
β’ Experience building deterministic, idempotent, and re-runnable jobs
AWS Data Stack
β’ Deep hands-on experience with S3
β’ AWS Glue and/or Spark on EMR
β’ Athena
β’ Step Functions
β’ Lambda
β’ AWS Batch
β’ Understanding of storage layout, lifecycle policies, and storage economics at scale
SQL & Data Modeling
β’ Complex temporal joins
β’ Point-in-time correctness
β’ Strong understanding of preventing label leakage in time-series data
Data Quality & Lineage
β’ Data validation frameworks
β’ Schema enforcement
β’ Versioned datasets
β’ Data lineage and reproducibility
AWS Security & Governance
β’ IAM policy design
β’ KMS
β’ VPC endpoints and PrivateLink
β’ Experience working within HIPAA-eligible AWS architectures
Desirable Skills
β’ Healthcare data standards including DICOM, FHIR, HL7v2, OMOP CDM
β’ Familiarity with clinical coding systems such as ICD, LOINC, RxNorm, and SNOMED
β’ Medical imaging experience with pydicom and OpenSlide
β’ Understanding of WSI pyramid structures and tiling
β’ Terraform / Infrastructure as Code
β’ Docker, ECR, and CI/CD
β’ Familiarity with LLM APIs and multimodal payload construction






