Amtex Systems Inc.

Data Engineer

⭐ - Featured Role | Apply direct with Data Freelance Hub
This role is for a Data Engineer with 10+ years of experience, focusing on data processing pipelines in secure environments. It offers a long-term contract, hybrid work from Chicago, IL, with a pay rate of "unknown." Key skills include Scala, Apache Spark, and Python.
🌎 - Country
United States
💱 - Currency
$ USD
-
💰 - Day rate
Unknown
-
🗓️ - Date
July 25, 2026
🕒 - Duration
More than 6 months
-
🏝️ - Location
Hybrid
-
📄 - Contract
Unknown
-
🔒 - Security
Unknown
-
📍 - Location detailed
Chicago, IL
-
🧠 - Skills detailed
#Data Architecture #Classification #Data Pipeline #GIT #Airflow #Data Governance #Databricks #Databases #Apache Spark #Compliance #Code Reviews #Batch #Data Lineage #GCP (Google Cloud Platform) #Data Warehouse #Programming #Azure #Data Catalog #Dimensional Modelling #Observability #PCI (Payment Card Industry) #Data Quality #Data Processing #Data Engineering #Monitoring #Spark (Apache Spark) #Delta Lake #Automated Testing #Python #AWS Glue #Automation #Datasets #Security #Scala #Cloud #AWS (Amazon Web Services) #SQL (Structured Query Language)
Role description
Data Engineer Hybrid from Chicago, IL Long Term Description: • Design and implement data processing pipelines within trusted data environments, clean rooms, secure data-sharing platforms, or equivalent privacy-preserving analytics environments across AWS, GCP, Azure, and on-premises infrastructure. • Implement access controls, data classification policies, lineage tracking, and governance controls to ensure PII, PCI-scoped data, customer identifiers, and advertiser-confidential signals are processed only within approved secure environments. • Collaborate with Security, Privacy, and Compliance teams to define and maintain data handling standards, ensuring sensitive datasets and raw identifiers remain within approved trust boundaries. • Design data flows that enforce privacy-preserving principles, ensuring only aggregated, anonymised, tokenised, or otherwise approved outputs may leave trusted processing environments. • Build observability, monitoring, and alerting capabilities to detect anomalous data movement, policy violations, and potential data leakage events. Skills Required: • 10+ years of Data Engineering experience with deep Scala programming expertise and extensive Apache Spark experience for large-scale distributed data processing on AWS and/or GCP. • Strong Python development skills for data pipelines, platform tooling, automation, and infrastructure modules . • Advanced SQL skills across relational databases, cloud data warehouses, and lakehouse platforms; experience handling TB-scale datasets . • Experience designing, building, and maintaining batch and streaming data pipelines. • Strong understanding of data warehousing, dimensional modelling, data quality, partitioning, and performance optimisation. • Experience with distributed data processing and modern lakehouse architectures (Databricks, Delta Lake, Apache Spark, or equivalent) . • Experience building and operating distributed data platforms at scale. • Experience with workflow orchestration platforms such as Airflow, Databricks Workflows, AWS Step Functions, or equivalent DAG-based systems. • Git or equivalent source control; unit, integration, and automated testing frameworks • Cloud-native development experience across AWS and/or GCP. • Strong software engineering practices including CI/CD, code reviews, observability, and production support . • Proven ability to set technical direction across multiple teams, mentor senior and lead engineers, drive engineering standards, and consistently deliver complex, cross-cutting initiatives. • Experience designing and operating data pipelines within trusted data environments, clean rooms, secure data-sharing platforms, or equivalent privacy-preserving analytics environments • Experience working with sensitive datasets containing PII, customer identifiers, advertiser data, or regulated information • Experience implementing fine-grained access controls, data governance policies, and policy-based enforcement for sensitive datasets • Working knowledge of data classification frameworks including PII, PCI, regulated data, and sensitivity-tier models • Familiarity with privacy-preserving data processing techniques including tokenisation, pseudonymisation, aggregation-before-export, and differential privacy concepts • Experience building or supporting clean-room, measurement, attribution, audience analytics, partner data-sharing, or privacy-preserving reporting solutions • Experience with data lineage and governance tooling (Unity Catalog, AWS Glue Data Catalog, Apache Atlas, OpenLineage, or equivalent) for auditability and compliance • Understanding of trust boundaries, secure data-sharing patterns, and zero-trust data architecture principles