

Amtex Systems Inc.
Data Engineer
⭐ - Featured Role | Apply direct with Data Freelance Hub
This role is for a Data Engineer with 10+ years of experience, focusing on data processing pipelines in secure environments. It offers a long-term contract, hybrid work from Chicago, IL, with a pay rate of "unknown." Key skills include Scala, Apache Spark, and Python.
🌎 - Country
United States
💱 - Currency
$ USD
-
💰 - Day rate
Unknown
-
🗓️ - Date
July 25, 2026
🕒 - Duration
More than 6 months
-
🏝️ - Location
Hybrid
-
📄 - Contract
Unknown
-
🔒 - Security
Unknown
-
📍 - Location detailed
Chicago, IL
-
🧠 - Skills detailed
#Data Architecture #Classification #Data Pipeline #GIT #Airflow #Data Governance #Databricks #Databases #Apache Spark #Compliance #Code Reviews #Batch #Data Lineage #GCP (Google Cloud Platform) #Data Warehouse #Programming #Azure #Data Catalog #Dimensional Modelling #Observability #PCI (Payment Card Industry) #Data Quality #Data Processing #Data Engineering #Monitoring #Spark (Apache Spark) #Delta Lake #Automated Testing #Python #AWS Glue #Automation #Datasets #Security #Scala #Cloud #AWS (Amazon Web Services) #SQL (Structured Query Language)
Role description
Data Engineer
Hybrid from Chicago, IL
Long Term
Description:
• Design and implement data processing pipelines within trusted data environments, clean rooms, secure data-sharing platforms, or equivalent privacy-preserving analytics environments across AWS, GCP, Azure, and on-premises infrastructure.
• Implement access controls, data classification policies, lineage tracking, and governance controls to ensure PII, PCI-scoped data, customer identifiers, and advertiser-confidential signals are processed only within approved secure environments.
• Collaborate with Security, Privacy, and Compliance teams to define and maintain data handling standards, ensuring sensitive datasets and raw identifiers remain within approved trust boundaries.
• Design data flows that enforce privacy-preserving principles, ensuring only aggregated, anonymised, tokenised, or otherwise approved outputs may leave trusted processing environments.
• Build observability, monitoring, and alerting capabilities to detect anomalous data movement, policy violations, and potential data leakage events.
Skills Required:
• 10+ years of Data Engineering experience with deep Scala programming expertise and extensive Apache Spark experience for large-scale distributed data processing on AWS and/or GCP.
• Strong Python development skills for data pipelines, platform tooling, automation, and infrastructure modules .
• Advanced SQL skills across relational databases, cloud data warehouses, and lakehouse platforms; experience handling TB-scale datasets .
• Experience designing, building, and maintaining batch and streaming data pipelines.
• Strong understanding of data warehousing, dimensional modelling, data quality, partitioning, and performance optimisation.
• Experience with distributed data processing and modern lakehouse architectures (Databricks, Delta Lake, Apache Spark, or equivalent) .
• Experience building and operating distributed data platforms at scale.
• Experience with workflow orchestration platforms such as Airflow, Databricks Workflows, AWS Step Functions, or equivalent DAG-based systems.
• Git or equivalent source control; unit, integration, and automated testing frameworks
• Cloud-native development experience across AWS and/or GCP.
• Strong software engineering practices including CI/CD, code reviews, observability, and production support .
• Proven ability to set technical direction across multiple teams, mentor senior and lead engineers, drive engineering standards, and consistently deliver complex, cross-cutting initiatives.
• Experience designing and operating data pipelines within trusted data environments, clean rooms, secure data-sharing platforms, or equivalent privacy-preserving analytics environments
• Experience working with sensitive datasets containing PII, customer identifiers, advertiser data, or regulated information
• Experience implementing fine-grained access controls, data governance policies, and policy-based enforcement for sensitive datasets
• Working knowledge of data classification frameworks including PII, PCI, regulated data, and sensitivity-tier models
• Familiarity with privacy-preserving data processing techniques including tokenisation, pseudonymisation, aggregation-before-export, and differential privacy concepts
• Experience building or supporting clean-room, measurement, attribution, audience analytics, partner data-sharing, or privacy-preserving reporting solutions
• Experience with data lineage and governance tooling (Unity Catalog, AWS Glue Data Catalog, Apache Atlas, OpenLineage, or equivalent) for auditability and compliance
• Understanding of trust boundaries, secure data-sharing patterns, and zero-trust data architecture principles
Data Engineer
Hybrid from Chicago, IL
Long Term
Description:
• Design and implement data processing pipelines within trusted data environments, clean rooms, secure data-sharing platforms, or equivalent privacy-preserving analytics environments across AWS, GCP, Azure, and on-premises infrastructure.
• Implement access controls, data classification policies, lineage tracking, and governance controls to ensure PII, PCI-scoped data, customer identifiers, and advertiser-confidential signals are processed only within approved secure environments.
• Collaborate with Security, Privacy, and Compliance teams to define and maintain data handling standards, ensuring sensitive datasets and raw identifiers remain within approved trust boundaries.
• Design data flows that enforce privacy-preserving principles, ensuring only aggregated, anonymised, tokenised, or otherwise approved outputs may leave trusted processing environments.
• Build observability, monitoring, and alerting capabilities to detect anomalous data movement, policy violations, and potential data leakage events.
Skills Required:
• 10+ years of Data Engineering experience with deep Scala programming expertise and extensive Apache Spark experience for large-scale distributed data processing on AWS and/or GCP.
• Strong Python development skills for data pipelines, platform tooling, automation, and infrastructure modules .
• Advanced SQL skills across relational databases, cloud data warehouses, and lakehouse platforms; experience handling TB-scale datasets .
• Experience designing, building, and maintaining batch and streaming data pipelines.
• Strong understanding of data warehousing, dimensional modelling, data quality, partitioning, and performance optimisation.
• Experience with distributed data processing and modern lakehouse architectures (Databricks, Delta Lake, Apache Spark, or equivalent) .
• Experience building and operating distributed data platforms at scale.
• Experience with workflow orchestration platforms such as Airflow, Databricks Workflows, AWS Step Functions, or equivalent DAG-based systems.
• Git or equivalent source control; unit, integration, and automated testing frameworks
• Cloud-native development experience across AWS and/or GCP.
• Strong software engineering practices including CI/CD, code reviews, observability, and production support .
• Proven ability to set technical direction across multiple teams, mentor senior and lead engineers, drive engineering standards, and consistently deliver complex, cross-cutting initiatives.
• Experience designing and operating data pipelines within trusted data environments, clean rooms, secure data-sharing platforms, or equivalent privacy-preserving analytics environments
• Experience working with sensitive datasets containing PII, customer identifiers, advertiser data, or regulated information
• Experience implementing fine-grained access controls, data governance policies, and policy-based enforcement for sensitive datasets
• Working knowledge of data classification frameworks including PII, PCI, regulated data, and sensitivity-tier models
• Familiarity with privacy-preserving data processing techniques including tokenisation, pseudonymisation, aggregation-before-export, and differential privacy concepts
• Experience building or supporting clean-room, measurement, attribution, audience analytics, partner data-sharing, or privacy-preserving reporting solutions
• Experience with data lineage and governance tooling (Unity Catalog, AWS Glue Data Catalog, Apache Atlas, OpenLineage, or equivalent) for auditability and compliance
• Understanding of trust boundaries, secure data-sharing patterns, and zero-trust data architecture principles






