Arbor TekSystems

Data Engineer

⭐ - Featured Role | Apply direct with Data Freelance Hub
This role is a remote Data Engineer contract position focused on migrating legacy ETL workloads to Databricks Lakehouse. Key skills include Databricks, PySpark, AWS, and Apache Airflow. Requires strong experience in data migration strategies and optimization techniques. Pay rate is W2.
🌎 - Country
United States
💱 - Currency
$ USD
-
💰 - Day rate
Unknown
-
🗓️ - Date
August 5, 2026
🕒 - Duration
Unknown
-
🏝️ - Location
Remote
-
📄 - Contract
W2 Contractor
-
🔒 - Security
Unknown
-
📍 - Location detailed
Dallas, TX
-
🧠 - Skills detailed
#Delta Lake #Shell Scripting #Databricks #Spark (Apache Spark) #Strategy #SQL (Structured Query Language) #Airflow #DataStage #Scripting #Deployment #Migration #Data Engineering #PySpark #DevOps #Cloud #Automation #Apache Airflow #AWS (Amazon Web Services) #Python #"ETL (Extract #Transform #Load)"
Role description
Job Title: Data Engineer with Databricks Location: Remote Job Type: Contract Only W2 Job Summary: The Databricks Data Engineer owns the end-to-end migration strategy, target architecture design, and technical execution of moving legacy ETL workloads to the Databricks Lakehouse. He will establish migration standards, optimize PySpark pipelines, orchestrate complex data workflows, and deploy proprietary automation tools to ensure a seamless, high-performing transition from legacy systems. Key Responsibilities: Architectural Strategy & Governance Design the target Databricks Lakehouse architecture utilizing Delta Lake, Photon, and Unity Catalog. Establish global code refactoring standards, optimization benchmarks, and PySpark best practices. Resolve highly complex dependency mappings and architect seamless, zero-downtime dual-run strategies. Lead the technical deployment and integration of specialized migration accelerators. Hands-on Engineering & Optimization Review automated output from migration tools and manually refactor complex legacy logic into high-performing PySpark notebooks. Eliminate legacy anti-patterns such as massive row-by-row processing and inefficient lookups. Optimize PySpark code performance using advanced Spark features including Z-Ordering, partitioning, and caching. Build robust Databricks Workflows and orchestrate complex DAGs based on comprehensive source lineage. Technical Skills & Competencies: Core Platforms: Databricks, Delta Lake, Unity Catalog, Photon, DataStage. Languages & Frameworks: PySpark, Python, SQL, Shell Scripting. Cloud & DevOps: AWS alongside CI/CD deployment pipelines. Orchestration: Apache Airflow, Databricks Workflows.