Aptimized

Data Engineer

⭐ - Featured Role | Apply direct with Data Freelance Hub
This role is for a Data Engineer (Pentaho) on a 6-month contract, 100% remote. Requires expertise in Pentaho and AWS Glue, focusing on migration from Pentaho to AWS Glue. Key skills include PySpark, ETL, and AWS services.
🌎 - Country
United States
πŸ’± - Currency
$ USD
-
πŸ’° - Day rate
16
-
πŸ—“οΈ - Date
July 22, 2026
πŸ•’ - Duration
More than 6 months
-
🏝️ - Location
Remote
-
πŸ“„ - Contract
Unknown
-
πŸ”’ - Security
Unknown
-
πŸ“ - Location detailed
United States
-
🧠 - Skills detailed
#IAM (Identity and Access Management) #GitHub #Deployment #Athena #Scala #Migration #Documentation #Lambda (AWS Lambda) #Spark (Apache Spark) #Data Pipeline #"ETL (Extract #Transform #Load)" #AWS (Amazon Web Services) #Logging #Monitoring #S3 (Amazon Simple Storage Service) #Data Modeling #Data Engineering #PySpark #Cloud #Data Integration #AWS Glue
Role description
Title: Data Engineer (Pentaho) 100% REMOTE Term: 6 months Rate: Open. MUST HAVE PENTAHO and AWS GLUE background. Migration from Pentaho to AWS Glue is Mandatory Immediate Contract Resource – Pentaho β†’ AWS Glue Migration Role Overview Seeking an experienced Data Integration Engineer to provide immediate contract support for a Pentaho ETL to AWS Glue migration. This role is focused on accelerating the transition from legacy Pentaho jobs to a modern, scalable, serverless data pipeline on AWS. Key Responsibilities β€’ Assess existing Pentaho ETL jobs, transformations, and schedules. β€’ Re‑architect and migrate workflows into AWS Glue (PySpark/Scala), ensuring performance, reliability, and maintainability. β€’ Build and optimize Glue Jobs, Crawlers, Workflows, and Catalog integrations. β€’ Collaborate with internal data engineering teams to validate data flows, dependencies, and business logic. β€’ Implement CI/CD best practices for Glue deployments (CodeCommit, CodePipeline, or GitHub). β€’ Ensure proper logging, monitoring, and alerting using CloudWatch and related AWS services. β€’ Provide knowledge transfer and documentation for internal teams. Required Skills & Experience β€’ Strong hands-on experience with Pentaho Data Integration (PDI). β€’ Proven expertise in AWS Glue, PySpark, and AWS data ecosystem (S3, Lambda, Athena, IAM). β€’ Solid understanding of ETL/ELT patterns, data modeling, and pipeline orchestration. β€’ Ability to quickly analyze legacy systems and translate them into cloud-native architectures. β€’ Excellent communication and documentation skills.