Covetus

Lead Data Engineer

⭐ - Featured Role | Apply direct with Data Freelance Hub
This role is for a Lead Data Engineer, 100% remote, with a contract length of "unknown" and a pay rate of "unknown." Key skills include Databricks, PySpark, Talend, AWS, and healthcare data experience, particularly with EDI processing.
🌎 - Country
United States
💱 - Currency
$ USD
-
💰 - Day rate
360
-
🗓️ - Date
August 11, 2026
🕒 - Duration
Unknown
-
🏝️ - Location
Remote
-
📄 - Contract
Unknown
-
🔒 - Security
Unknown
-
📍 - Location detailed
United States
-
🧠 - Skills detailed
#PySpark #Jenkins #Deployment #Python #Data Processing #Sqoop (Apache Sqoop) #Scala #HBase #Lambda (AWS Lambda) #"ETL (Extract #Transform #Load)" #Terraform #Data Integration #Talend #Databricks #DataStage #Data Lake #Snowflake #Monitoring #Big Data #Batch #Delta Lake #Kafka (Apache Kafka) #API (Application Programming Interface) #Automation #AWS (Amazon Web Services) #Data Lakehouse #Databases #Spark (Apache Spark) #Data Modeling #Data Governance #Migration #SQL (Structured Query Language) #Spark SQL #Data Engineering #S3 (Amazon Simple Storage Service) #Hadoop #Cloud #Data Pipeline #Data Quality #GIT
Role description
Title - Data Engineering Lead Location - 100% Remote Job Summary Seeking a highly experienced Data Engineering Lead with expertise in Databricks, PySpark, Talend, AWS, and Big Data technologies. The ideal candidate will design and implement scalable cloud data platforms, modern data lakehouse architectures, and enterprise ETL solutions. Strong experience with Databricks, PySpark, Talend, AWS, and Big Data technologies is required. Strong experience in healthcare data, EDI processing, and cloud migration is highly preferred. Key Responsibilities: • Design, develop, and optimize scalable data pipelines using Databricks, PySpark, Talend, and Snowflake. • Build and maintain cloud-based Lakehouse architecture using Databricks, Delta Lake, and AWS. • Develop ETL/ELT solutions to ingest, transform, and load data from databases, APIs, Salesforce, S3, and flat files. • Implement and manage Delta Live Tables, Databricks Workflows, and Unity Catalog for data governance. • Develop and optimize large-scale Spark and PySpark workloads for batch and real-time data processing. • Support healthcare data integration initiatives including EDI X12 transactions (834, 835, 837) and HIPAA-compliant processing. • Perform data quality validation, performance tuning, troubleshooting, and root cause analysis. • Collaborate with business, QA, and development teams to gather requirements and deliver scalable data solutions. • Support CI/CD, deployment automation, monitoring, and production support activities. Required Skills & Qualifications: • Strong hands-on experience with Databricks, PySpark, Spark SQL, and Delta Lake. • Extensive experience with Snowflake, including data loading, data modeling, and performance optimization. • Strong ETL development experience using Talend and/or IBM DataStage. • Advanced SQL expertise with strong query tuning and troubleshooting skills. • Experience with AWS services such as S3, Lambda, EMR, and cloud-based data platforms. • Proficiency in Python for data engineering, automation, and API integrations. • Strong knowledge of Data Warehousing, Data Lake, and Lakehouse architecture. • Excellent analytical, problem-solving, and communication skills. Preferred / Nice-to-Have Skills: • Experience with Unity Catalog, Delta Live Tables, and Databricks governance frameworks. • Healthcare domain experience including Claims, Membership, Eligibility, and EDI processing. • Experience with Hadoop ecosystem technologies such as Hive, Kafka, HBase, Sqoop, and Oozie. • Knowledge of CI/CD tools such as Jenkins, Git, and Terraform.