

Optomi
Data Engineer
⭐ - Featured Role | Apply direct with Data Freelance Hub
This role is for a Senior Data Engineer with 7+ years of experience, focusing on large-scale production data pipelines. Contract length is "unknown," pay rate is "unknown," and it requires strong skills in Python, SQL, Apache Spark, Databricks, and AWS.
🌎 - Country
United States
💱 - Currency
$ USD
-
💰 - Day rate
520
-
🗓️ - Date
August 8, 2026
🕒 - Duration
Unknown
-
🏝️ - Location
Hybrid
-
📄 - Contract
Unknown
-
🔒 - Security
Unknown
-
📍 - Location detailed
San Francisco Bay Area
-
🧠 - Skills detailed
#Cloud #Kubernetes #Spark SQL #SQL (Structured Query Language) #SQL Queries #PySpark #Airflow #Data Lake #Delta Lake #Data Transformations #Documentation #Kafka (Apache Kafka) #Apache Spark #Agile #Spark (Apache Spark) #Data Engineering #Scrum #Python #Data Processing #Code Reviews #Docker #Apache Airflow #Scala #Batch #Databricks #Data Pipeline #"ETL (Extract #Transform #Load)" #Data Modeling #Datasets #AWS (Amazon Web Services) #Data Quality
Role description
Hybrid 3 days a week-- non-negotiable
Job Description / Overview
We are seeking a highly experienced Senior Data Engineer to join a Core Data Platform team responsible for building, maintaining, and expanding enterprise-scale data infrastructure and pipelines. This individual will work hands-on with large distributed data systems and help develop reliable, scalable solutions that enable data discovery, lineage, governance, privacy, analytics, and downstream data consumption. The ideal candidate brings deep experience developing production data pipelines using Python, SQL, Apache Spark, Databricks, and Apache Airflow, along with strong cloud and distributed processing expertise. This person should be comfortable owning production workloads, troubleshooting complex data issues, optimizing performance, and working cross-functionally to ensure datasets meet established standards for reliability, accuracy, quality, and SLAs.
Responsibilities
• Design, develop, maintain, and expand large-scale production data pipelines and distributed data processing solutions.
• Develop data processing applications and pipelines using Python, SQL, Apache Spark, and Databricks.
• Build and maintain workflow orchestration using Apache Airflow.
• Work with Databricks and Delta Lake to support scalable data processing and lakehouse architectures.
• Develop and optimize distributed processing workloads handling large volumes of structured and semi-structured data.
• Support enterprise data discovery, lineage, governance, privacy, and data quality capabilities.
• Monitor production data pipelines and troubleshoot failures, performance bottlenecks, and data quality issues.
• Optimize Spark jobs, SQL queries, data transformations, and distributed workloads for performance and scalability.
• Ensure production datasets and pipelines meet established SLAs for availability, reliability, accuracy, and performance.
• Develop and maintain data models supporting both OLTP and OLAP use cases.
• Work with cloud-based data platforms and services, including AWS.
• Support containerized/distributed workloads using technologies such as Docker and Kubernetes.
• Establish and follow engineering standards, development best practices, testing processes, and technical documentation.
• Collaborate with engineering, analytics, product, and other cross-functional teams to understand data requirements and deliver scalable solutions.
• Participate in Agile/Scrum ceremonies, sprint planning, technical discussions, code reviews, and production support.
Qualifications
• 7+ years of Data Engineering experience, with significant experience building and supporting large-scale production data pipelines.
• Strong hands-on Python development experience.
• Advanced SQL skills, including complex joins, CTEs, window functions, aggregations, transformations, and query optimization.
• Strong production experience with Apache Spark / PySpark and distributed data processing.
• Hands-on experience with Databricks in production environments.
• Experience developing and managing data workflows with Apache Airflow.
• Strong understanding of data modeling, including dimensional modeling and OLTP/OLAP concepts.
• Experience working with Delta Lake and modern lakehouse/data lake architectures.
• Strong cloud experience, preferably with AWS and its data ecosystem.
• Experience building and supporting batch and/or near-real-time data pipelines.
• Experience with distributed data services and large-scale data processing architectures.
• Understanding of data quality, lineage, governance, privacy, and production reliability.
• Experience with Docker/Kubernetes or containerized data workloads is highly desirable.
• Experience with Kafka/streaming architectures is highly desirable.
• Strong troubleshooting and performance optimization skills across Spark, SQL, and production data pipelines.
• Experience working within an Agile/Scrum engineering environment.
• Strong communication skills and ability to collaborate effectively with technical and non-technical stakeholders.
Hybrid 3 days a week-- non-negotiable
Job Description / Overview
We are seeking a highly experienced Senior Data Engineer to join a Core Data Platform team responsible for building, maintaining, and expanding enterprise-scale data infrastructure and pipelines. This individual will work hands-on with large distributed data systems and help develop reliable, scalable solutions that enable data discovery, lineage, governance, privacy, analytics, and downstream data consumption. The ideal candidate brings deep experience developing production data pipelines using Python, SQL, Apache Spark, Databricks, and Apache Airflow, along with strong cloud and distributed processing expertise. This person should be comfortable owning production workloads, troubleshooting complex data issues, optimizing performance, and working cross-functionally to ensure datasets meet established standards for reliability, accuracy, quality, and SLAs.
Responsibilities
• Design, develop, maintain, and expand large-scale production data pipelines and distributed data processing solutions.
• Develop data processing applications and pipelines using Python, SQL, Apache Spark, and Databricks.
• Build and maintain workflow orchestration using Apache Airflow.
• Work with Databricks and Delta Lake to support scalable data processing and lakehouse architectures.
• Develop and optimize distributed processing workloads handling large volumes of structured and semi-structured data.
• Support enterprise data discovery, lineage, governance, privacy, and data quality capabilities.
• Monitor production data pipelines and troubleshoot failures, performance bottlenecks, and data quality issues.
• Optimize Spark jobs, SQL queries, data transformations, and distributed workloads for performance and scalability.
• Ensure production datasets and pipelines meet established SLAs for availability, reliability, accuracy, and performance.
• Develop and maintain data models supporting both OLTP and OLAP use cases.
• Work with cloud-based data platforms and services, including AWS.
• Support containerized/distributed workloads using technologies such as Docker and Kubernetes.
• Establish and follow engineering standards, development best practices, testing processes, and technical documentation.
• Collaborate with engineering, analytics, product, and other cross-functional teams to understand data requirements and deliver scalable solutions.
• Participate in Agile/Scrum ceremonies, sprint planning, technical discussions, code reviews, and production support.
Qualifications
• 7+ years of Data Engineering experience, with significant experience building and supporting large-scale production data pipelines.
• Strong hands-on Python development experience.
• Advanced SQL skills, including complex joins, CTEs, window functions, aggregations, transformations, and query optimization.
• Strong production experience with Apache Spark / PySpark and distributed data processing.
• Hands-on experience with Databricks in production environments.
• Experience developing and managing data workflows with Apache Airflow.
• Strong understanding of data modeling, including dimensional modeling and OLTP/OLAP concepts.
• Experience working with Delta Lake and modern lakehouse/data lake architectures.
• Strong cloud experience, preferably with AWS and its data ecosystem.
• Experience building and supporting batch and/or near-real-time data pipelines.
• Experience with distributed data services and large-scale data processing architectures.
• Understanding of data quality, lineage, governance, privacy, and production reliability.
• Experience with Docker/Kubernetes or containerized data workloads is highly desirable.
• Experience with Kafka/streaming architectures is highly desirable.
• Strong troubleshooting and performance optimization skills across Spark, SQL, and production data pipelines.
• Experience working within an Agile/Scrum engineering environment.
• Strong communication skills and ability to collaborate effectively with technical and non-technical stakeholders.






