Net2Source Inc.

Databricks Architect

⭐ - Featured Role | Apply direct with Data Freelance Hub
This role is for a "Databricks Architect" with a contract length of "12+ months" and a remote work location. Key skills required include "8+ years in Data Engineering," "4+ years with Databricks," and expertise in "Sales data domain."
🌎 - Country
United States
💱 - Currency
$ USD
-
💰 - Day rate
Unknown
-
🗓️ - Date
August 12, 2026
🕒 - Duration
More than 6 months
-
🏝️ - Location
Remote
-
📄 - Contract
Unknown
-
🔒 - Security
Unknown
-
📍 - Location detailed
United States
-
🧠 - Skills detailed
#Databricks #Datasets #dbt (data build tool) #Logging #Leadership #AWS Glue #Data Governance #PySpark #Data Architecture #Delta Lake #SQL (Structured Query Language) #SAP #Data Processing #Data Quality #Spark (Apache Spark) #Data Ingestion #Monitoring #Metadata #S3 (Amazon Simple Storage Service) #Data Lineage #Data Lake #Data Modeling #Cloud #Scala #Batch #"ACID (Atomicity #Consistency #Isolation #Durability)" #CRM (Customer Relationship Management) #Slowly Changing Dimensions #DevOps #Deployment #Data Engineering #AWS (Amazon Web Services) #IAM (Identity and Access Management) #Spark SQL #Storage #Security #Data Pipeline #Debugging #"ETL (Extract #Transform #Load)"
Role description
Role: Databricks Architect Work location: Remote Role Duration: 12+ Months Job Description: Role Summary • We are seeking a Databricks Architect to lead the design and implementation of a scalable Sales Data Platform as part of the OneData initiative on Databricks running on AWS. • The role is hands-on and architecture-driven, focused on data ingestion, transformation, modeling, and optimization using modern lakehouse patterns. • The architect will work closely with data engineers, source system teams, and downstream consumers to deliver high-quality, governed, and performance-optimized sales datasets. Key Responsibilities Architecture & Design • Define end-to-end lakehouse architecture on Databricks (AWS) for Sales data domains • Design medallion architecture (Bronze / Silver / Gold) aligned with OneData standards • Establish data modeling standards for Sales facts, dimensions, hierarchies, and aggregations • Define scalable ingestion patterns for batch and incremental loads • Drive performance, scalability, and cost optimization best practices Data Engineering & Implementation • Build and guide development of PySpark-based data pipelines in Databricks • Implement Delta Lake features: • ACID transactions • Schema evolution & enforcement • Time travel & versioning • Design and optimize large-scale joins, aggregations, and window functions • Implement CDC and incremental processing using watermarking and change detection • Ensure idempotent, restartable, and fault-tolerant pipelines AWS & Platform Integration Architect solutions using AWS services: • Amazon S3 (data lake storage) • IAM (security & access control) • AWS Glue / Glue Catalog • CloudWatch (monitoring & logging) • Optimize Databricks cluster configurations (job vs all-purpose clusters) • Implement secrets management and secure connectivity patterns Data Quality, Governance & Reliability • Define and implement data quality checks and validations • Ensure data lineage and metadata capture • Implement error handling, auditing, and reconciliation frameworks • Support data governance and access control requirements DevOps & Operational Excellence • Implement CI/CD pipelines for Databricks notebooks and jobs • Enforce code versioning, reviews, and deployment standards • Design monitoring, alerting, and SLA tracking for pipelines • Support production stabilization and performance tuning Collaboration & Leadership • Act as technical lead / mentor for Databricks data engineers • Collaborate with: • Source system teams (Sales, CRM, ERP) • Data consumers (analytics, downstream apps) • Cloud/platform teams • Translate business requirements into robust technical designs Required Skills & Experience Core Technical Skills • 8+ years in Data Engineering / Data Architecture roles • 4+ years hands-on experience with Databricks • Strong expertise in PySpark & Spark SQL • Strong expertise in dbt • Deep experience with Delta Lake • Strong knowledge of AWS cloud services (S3, IAM, Glue, CloudWatch) Data Engineering Expertise • Sales data domain experience (orders, revenue, pricing, customers, products) • Strong understanding of: • Fact & dimension modeling • Slowly Changing Dimensions (SCD Type 1 / 2) • Large-scale data processing patterns • Experience handling high-volume, high-velocity datasets Platform & Operational Skills • Databricks job orchestration and scheduling • Cluster sizing and performance tuning • CI/CD for data platforms • Strong troubleshooting and debugging skills Nice-to-Have • Experience with enterprise OneData / Data Mesh programs • Exposure to real-time or near-real-time ingestion patterns • Experience integrating CRM / Sales systems (e.g., Salesforce, SAP Sales data) • AWS certifications or Databricks certifications