Largeton Group

Mid-Level Data Engineer

⭐ - Featured Role | Apply direct with Data Freelance Hub
This role is for a Mid-Level Data Engineer, 12-month contract, focusing on GCP Data Lake solutions. Key skills include PySpark, Hadoop, data modeling, and CI/CD processes. Experience with IAM and big data tools is essential. Pay rate is "unknown."
🌎 - Country
United States
💱 - Currency
$ USD
-
💰 - Day rate
Unknown
-
🗓️ - Date
July 25, 2026
🕒 - Duration
More than 6 months
-
🏝️ - Location
Unknown
-
📄 - Contract
Unknown
-
🔒 - Security
Unknown
-
📍 - Location detailed
Columbus, OH
-
🧠 - Skills detailed
#Data Pipeline #Data Storage #Deployment #Hadoop #Data Lake #Batch #GCP (Google Cloud Platform) #IAM (Identity and Access Management) #Neo4J #Data Processing #Big Data #Data Engineering #Spark (Apache Spark) #Storage #Datasets #Data Ingestion #PySpark #Scala #Cloud #HDFS (Hadoop Distributed File System) #Data Modeling #"ETL (Extract #Transform #Load)"
Role description
Job Summary For Mid-Level Data Engineer • Support the IAM Data Lake Engineering initiative. • Build and maintain Data Lake solutions on Google Cloud Platform (GCP) using big data tools and technologies. • Design, develop, and optimize data pipelines for ingestion, processing, and transformation of large-scale datasets. • Implement data modeling and data processing solutions to support analytical and operational requirements. • Utilize PySpark for distributed data processing and transformation tasks. • Work extensively with the Hadoop ecosystem, including HDFS for data storage and management. • Apply deep knowledge of GCP architecture, including bucket structuring, naming conventions, lifecycle policies, and access controls. • Integrate and manage data using Neo4j for graph database requirements. • Ensure robust CI/CD processes for continuous integration and deployment of data engineering solutions. • Build both batch and streaming data ingestion pipelines leveraging GCP-native services. • Develop and maintain data consumption and exposure layers via views, APIs, and curated analytical datasets. • Collaborate with cross-functional teams to ensure data solutions are scalable, secure, and aligned with organizational standards. • Duration of assignment is 12 months, with a possible extension based on project needs and performance.