

Hireproo LLC
Staff Data Engineer
⭐ - Featured Role | Apply direct with Data Freelance Hub
This role is for a Staff Data Engineer with a contract length of "unknown" and a pay rate of "unknown". It requires 10+ years of data engineering experience, strong Scala and Apache Spark skills, and expertise in data security and governance. Remote work option available.
🌎 - Country
United States
💱 - Currency
$ USD
-
💰 - Day rate
Unknown
-
🗓️ - Date
July 24, 2026
🕒 - Duration
Unknown
-
🏝️ - Location
Unknown
-
📄 - Contract
W2 Contractor
-
🔒 - Security
Unknown
-
📍 - Location detailed
Chicago, IL
-
🧠 - Skills detailed
#Data Science #Grafana #Databricks #Azure #Apache Spark #Batch #Scala #Data Warehouse #Data Architecture #Observability #Python #Spark SQL #Data Quality #Strategy #Datasets #Kubernetes #Delta Lake #AWS (Amazon Web Services) #Compliance #Docker #SQL (Structured Query Language) #Spark (Apache Spark) #Data Pipeline #Airflow #Data Engineering #Data Security #Apache Airflow #GCP (Google Cloud Platform) #Classification #Data Lineage #Automation #Programming #Data Lake #Monitoring #Data Governance #Leadership #Code Reviews #Data Catalog #Cloud #Forecasting #Security #Data Processing #AWS Glue
Role description
It is W2 role, Only GC, USC workable for this role.
We are looking for a highly experienced Staff Data Engineer to join our Attribution & Forecasting Data Platform team. In this role, you will define the technical direction for large-scale data engineering solutions powering attribution, measurement, forecasting, and analytics products.
You will design and build secure, scalable, and high-performance data platforms using modern cloud technologies, distributed processing frameworks, and privacy-preserving data solutions. You will work closely with Product, Data Science, Security, Privacy, and Platform Engineering teams to deliver trusted data products in a highly regulated environment.
This role is ideal for a hands-on technical leader who can architect complex data systems, solve large-scale engineering challenges, and mentor senior engineers across teams.
Key Responsibilities
• Lead the architecture and development of large-scale distributed data processing platforms using Scala, Apache Spark, SQL, and cloud-native technologies.
• Design, build, and optimize batch and streaming data pipelines handling large volumes of advertiser, customer, and measurement datasets.
• Define technical strategy and engineering standards for data platforms, lakehouse architectures, and analytics solutions.
• Build reliable data processing workflows using technologies such as Databricks, Delta Lake, Airflow, and cloud orchestration services.
• Design secure data pipelines within trusted environments, clean rooms, and privacy-preserving analytics platforms.
• Implement data governance practices including:
• Data classification
• Data lineage
• Access controls
• Data quality frameworks
• Secure data sharing
• Ensure sensitive datasets, PII, customer identifiers, and advertiser data remain protected within approved security boundaries.
• Develop privacy-aware data processing solutions using techniques such as:
• Aggregation
• Tokenization
• Pseudonymization
• Privacy-preserving analytics
• Build monitoring, observability, and operational tooling to ensure platform reliability and compliance.
• Troubleshoot complex distributed system, performance, and pipeline issues.
• Partner with Security, Privacy, Compliance, Product, and Data Science teams on platform design and governance.
• Mentor senior and lead engineers while driving engineering excellence across teams.
Required Skills & Experience
Data Engineering & Programming
• 10+ years of experience in Data Engineering or Data Platform Engineering.
• Strong expertise in Scala programming.
• Extensive experience with Apache Spark for large-scale distributed data processing.
• Strong Python development skills for automation, tooling, and pipeline development.
• Advanced SQL skills with experience handling large-scale datasets.
• Experience designing and operating complex batch and streaming data pipelines.
Data Platforms & Cloud
• Strong experience with modern data architectures including:
• Data Lakes
• Lakehouse platforms
• Cloud data warehouses
• Distributed processing systems
• Hands-on experience with:
• Databricks
• Delta Lake
• Apache Spark
• Experience with cloud platforms:
• AWS
• GCP
• Azure
• Experience with workflow orchestration tools such as:
• Apache Airflow
• Databricks Workflows
• AWS Step Functions
Data Security & Governance
• Experience working with sensitive data environments involving:
• PII
• Customer identifiers
• Confidential business data
• Regulated datasets
• Strong understanding of:
• Data governance
• Data lineage
• Access controls
• Data classification
• Secure data sharing
• Experience implementing privacy-preserving data processing patterns.
• Familiarity with tools such as:
• Unity Catalog
• AWS Glue Data Catalog
• Apache Atlas
• OpenLineage
Engineering Leadership
• Proven experience leading architecture decisions across multiple teams.
• Ability to mentor senior engineers and establish engineering best practices.
• Strong understanding of:
• CI/CD
• Testing frameworks
• Code reviews
• Observability
• Production support
• Excellent written and verbal communication skills.
Preferred Qualifications
• Experience building solutions in:
• Advertising technology (AdTech)
• Marketing technology (MarTech)
• Attribution platforms
• Audience analytics
• Customer data platforms (CDP)
• Retail media platforms
• Experience with:
• AWS Clean Rooms or similar privacy-enhancing technologies
• Kubernetes and Docker
• ELK Stack
• Grafana
• OpenTelemetry
• Security architecture and threat modelling
It is W2 role, Only GC, USC workable for this role.
We are looking for a highly experienced Staff Data Engineer to join our Attribution & Forecasting Data Platform team. In this role, you will define the technical direction for large-scale data engineering solutions powering attribution, measurement, forecasting, and analytics products.
You will design and build secure, scalable, and high-performance data platforms using modern cloud technologies, distributed processing frameworks, and privacy-preserving data solutions. You will work closely with Product, Data Science, Security, Privacy, and Platform Engineering teams to deliver trusted data products in a highly regulated environment.
This role is ideal for a hands-on technical leader who can architect complex data systems, solve large-scale engineering challenges, and mentor senior engineers across teams.
Key Responsibilities
• Lead the architecture and development of large-scale distributed data processing platforms using Scala, Apache Spark, SQL, and cloud-native technologies.
• Design, build, and optimize batch and streaming data pipelines handling large volumes of advertiser, customer, and measurement datasets.
• Define technical strategy and engineering standards for data platforms, lakehouse architectures, and analytics solutions.
• Build reliable data processing workflows using technologies such as Databricks, Delta Lake, Airflow, and cloud orchestration services.
• Design secure data pipelines within trusted environments, clean rooms, and privacy-preserving analytics platforms.
• Implement data governance practices including:
• Data classification
• Data lineage
• Access controls
• Data quality frameworks
• Secure data sharing
• Ensure sensitive datasets, PII, customer identifiers, and advertiser data remain protected within approved security boundaries.
• Develop privacy-aware data processing solutions using techniques such as:
• Aggregation
• Tokenization
• Pseudonymization
• Privacy-preserving analytics
• Build monitoring, observability, and operational tooling to ensure platform reliability and compliance.
• Troubleshoot complex distributed system, performance, and pipeline issues.
• Partner with Security, Privacy, Compliance, Product, and Data Science teams on platform design and governance.
• Mentor senior and lead engineers while driving engineering excellence across teams.
Required Skills & Experience
Data Engineering & Programming
• 10+ years of experience in Data Engineering or Data Platform Engineering.
• Strong expertise in Scala programming.
• Extensive experience with Apache Spark for large-scale distributed data processing.
• Strong Python development skills for automation, tooling, and pipeline development.
• Advanced SQL skills with experience handling large-scale datasets.
• Experience designing and operating complex batch and streaming data pipelines.
Data Platforms & Cloud
• Strong experience with modern data architectures including:
• Data Lakes
• Lakehouse platforms
• Cloud data warehouses
• Distributed processing systems
• Hands-on experience with:
• Databricks
• Delta Lake
• Apache Spark
• Experience with cloud platforms:
• AWS
• GCP
• Azure
• Experience with workflow orchestration tools such as:
• Apache Airflow
• Databricks Workflows
• AWS Step Functions
Data Security & Governance
• Experience working with sensitive data environments involving:
• PII
• Customer identifiers
• Confidential business data
• Regulated datasets
• Strong understanding of:
• Data governance
• Data lineage
• Access controls
• Data classification
• Secure data sharing
• Experience implementing privacy-preserving data processing patterns.
• Familiarity with tools such as:
• Unity Catalog
• AWS Glue Data Catalog
• Apache Atlas
• OpenLineage
Engineering Leadership
• Proven experience leading architecture decisions across multiple teams.
• Ability to mentor senior engineers and establish engineering best practices.
• Strong understanding of:
• CI/CD
• Testing frameworks
• Code reviews
• Observability
• Production support
• Excellent written and verbal communication skills.
Preferred Qualifications
• Experience building solutions in:
• Advertising technology (AdTech)
• Marketing technology (MarTech)
• Attribution platforms
• Audience analytics
• Customer data platforms (CDP)
• Retail media platforms
• Experience with:
• AWS Clean Rooms or similar privacy-enhancing technologies
• Kubernetes and Docker
• ELK Stack
• Grafana
• OpenTelemetry
• Security architecture and threat modelling






