

Net2Source Inc.
Databricks Architect
⭐ - Featured Role | Apply direct with Data Freelance Hub
This role is for a "Databricks Architect" with a contract length of "12+ months" and a remote work location. Key skills required include "8+ years in Data Engineering," "4+ years with Databricks," and expertise in "Sales data domain."
🌎 - Country
United States
💱 - Currency
$ USD
-
💰 - Day rate
Unknown
-
🗓️ - Date
August 12, 2026
🕒 - Duration
More than 6 months
-
🏝️ - Location
Remote
-
📄 - Contract
Unknown
-
🔒 - Security
Unknown
-
📍 - Location detailed
United States
-
🧠 - Skills detailed
#Databricks #Datasets #dbt (data build tool) #Logging #Leadership #AWS Glue #Data Governance #PySpark #Data Architecture #Delta Lake #SQL (Structured Query Language) #SAP #Data Processing #Data Quality #Spark (Apache Spark) #Data Ingestion #Monitoring #Metadata #S3 (Amazon Simple Storage Service) #Data Lineage #Data Lake #Data Modeling #Cloud #Scala #Batch #"ACID (Atomicity #Consistency #Isolation #Durability)" #CRM (Customer Relationship Management) #Slowly Changing Dimensions #DevOps #Deployment #Data Engineering #AWS (Amazon Web Services) #IAM (Identity and Access Management) #Spark SQL #Storage #Security #Data Pipeline #Debugging #"ETL (Extract #Transform #Load)"
Role description
Role: Databricks Architect
Work location: Remote Role
Duration: 12+ Months
Job Description:
Role Summary
• We are seeking a Databricks Architect to lead the design and implementation of a scalable Sales Data Platform as part of the OneData initiative on Databricks running on AWS.
• The role is hands-on and architecture-driven, focused on data ingestion, transformation, modeling, and optimization using modern lakehouse patterns.
• The architect will work closely with data engineers, source system teams, and downstream consumers to deliver high-quality, governed, and performance-optimized sales datasets.
Key Responsibilities
Architecture & Design
• Define end-to-end lakehouse architecture on Databricks (AWS) for Sales data domains
• Design medallion architecture (Bronze / Silver / Gold) aligned with OneData standards
• Establish data modeling standards for Sales facts, dimensions, hierarchies, and aggregations
• Define scalable ingestion patterns for batch and incremental loads
• Drive performance, scalability, and cost optimization best practices
Data Engineering & Implementation
• Build and guide development of PySpark-based data pipelines in Databricks
• Implement Delta Lake features:
• ACID transactions
• Schema evolution & enforcement
• Time travel & versioning
• Design and optimize large-scale joins, aggregations, and window functions
• Implement CDC and incremental processing using watermarking and change detection
• Ensure idempotent, restartable, and fault-tolerant pipelines
AWS & Platform Integration
Architect solutions using AWS services:
• Amazon S3 (data lake storage)
• IAM (security & access control)
• AWS Glue / Glue Catalog
• CloudWatch (monitoring & logging)
• Optimize Databricks cluster configurations (job vs all-purpose clusters)
• Implement secrets management and secure connectivity patterns
Data Quality, Governance & Reliability
• Define and implement data quality checks and validations
• Ensure data lineage and metadata capture
• Implement error handling, auditing, and reconciliation frameworks
• Support data governance and access control requirements
DevOps & Operational Excellence
• Implement CI/CD pipelines for Databricks notebooks and jobs
• Enforce code versioning, reviews, and deployment standards
• Design monitoring, alerting, and SLA tracking for pipelines
• Support production stabilization and performance tuning
Collaboration & Leadership
• Act as technical lead / mentor for Databricks data engineers
• Collaborate with:
• Source system teams (Sales, CRM, ERP)
• Data consumers (analytics, downstream apps)
• Cloud/platform teams
• Translate business requirements into robust technical designs
Required Skills & Experience
Core Technical Skills
• 8+ years in Data Engineering / Data Architecture roles
• 4+ years hands-on experience with Databricks
• Strong expertise in PySpark & Spark SQL
• Strong expertise in dbt
• Deep experience with Delta Lake
• Strong knowledge of AWS cloud services (S3, IAM, Glue, CloudWatch)
Data Engineering Expertise
• Sales data domain experience (orders, revenue, pricing, customers, products)
• Strong understanding of:
• Fact & dimension modeling
• Slowly Changing Dimensions (SCD Type 1 / 2)
• Large-scale data processing patterns
• Experience handling high-volume, high-velocity datasets
Platform & Operational Skills
• Databricks job orchestration and scheduling
• Cluster sizing and performance tuning
• CI/CD for data platforms
• Strong troubleshooting and debugging skills
Nice-to-Have
• Experience with enterprise OneData / Data Mesh programs
• Exposure to real-time or near-real-time ingestion patterns
• Experience integrating CRM / Sales systems (e.g., Salesforce, SAP Sales data)
• AWS certifications or Databricks certifications
Role: Databricks Architect
Work location: Remote Role
Duration: 12+ Months
Job Description:
Role Summary
• We are seeking a Databricks Architect to lead the design and implementation of a scalable Sales Data Platform as part of the OneData initiative on Databricks running on AWS.
• The role is hands-on and architecture-driven, focused on data ingestion, transformation, modeling, and optimization using modern lakehouse patterns.
• The architect will work closely with data engineers, source system teams, and downstream consumers to deliver high-quality, governed, and performance-optimized sales datasets.
Key Responsibilities
Architecture & Design
• Define end-to-end lakehouse architecture on Databricks (AWS) for Sales data domains
• Design medallion architecture (Bronze / Silver / Gold) aligned with OneData standards
• Establish data modeling standards for Sales facts, dimensions, hierarchies, and aggregations
• Define scalable ingestion patterns for batch and incremental loads
• Drive performance, scalability, and cost optimization best practices
Data Engineering & Implementation
• Build and guide development of PySpark-based data pipelines in Databricks
• Implement Delta Lake features:
• ACID transactions
• Schema evolution & enforcement
• Time travel & versioning
• Design and optimize large-scale joins, aggregations, and window functions
• Implement CDC and incremental processing using watermarking and change detection
• Ensure idempotent, restartable, and fault-tolerant pipelines
AWS & Platform Integration
Architect solutions using AWS services:
• Amazon S3 (data lake storage)
• IAM (security & access control)
• AWS Glue / Glue Catalog
• CloudWatch (monitoring & logging)
• Optimize Databricks cluster configurations (job vs all-purpose clusters)
• Implement secrets management and secure connectivity patterns
Data Quality, Governance & Reliability
• Define and implement data quality checks and validations
• Ensure data lineage and metadata capture
• Implement error handling, auditing, and reconciliation frameworks
• Support data governance and access control requirements
DevOps & Operational Excellence
• Implement CI/CD pipelines for Databricks notebooks and jobs
• Enforce code versioning, reviews, and deployment standards
• Design monitoring, alerting, and SLA tracking for pipelines
• Support production stabilization and performance tuning
Collaboration & Leadership
• Act as technical lead / mentor for Databricks data engineers
• Collaborate with:
• Source system teams (Sales, CRM, ERP)
• Data consumers (analytics, downstream apps)
• Cloud/platform teams
• Translate business requirements into robust technical designs
Required Skills & Experience
Core Technical Skills
• 8+ years in Data Engineering / Data Architecture roles
• 4+ years hands-on experience with Databricks
• Strong expertise in PySpark & Spark SQL
• Strong expertise in dbt
• Deep experience with Delta Lake
• Strong knowledge of AWS cloud services (S3, IAM, Glue, CloudWatch)
Data Engineering Expertise
• Sales data domain experience (orders, revenue, pricing, customers, products)
• Strong understanding of:
• Fact & dimension modeling
• Slowly Changing Dimensions (SCD Type 1 / 2)
• Large-scale data processing patterns
• Experience handling high-volume, high-velocity datasets
Platform & Operational Skills
• Databricks job orchestration and scheduling
• Cluster sizing and performance tuning
• CI/CD for data platforms
• Strong troubleshooting and debugging skills
Nice-to-Have
• Experience with enterprise OneData / Data Mesh programs
• Exposure to real-time or near-real-time ingestion patterns
• Experience integrating CRM / Sales systems (e.g., Salesforce, SAP Sales data)
• AWS certifications or Databricks certifications






