Rivago Infotech Inc

Data Architect

⭐ - Featured Role | Apply direct with Data Freelance Hub
This role is for a Data Architect in Tarrytown, New York, on a long-term project. Key skills include Databricks Lakehouse, Delta Lake, AWS, PySpark, and data governance. Experience in large-scale data platform modernization and migration is required. Pay rate is unspecified.
🌎 - Country
United States
💱 - Currency
$ USD
-
💰 - Day rate
Unknown
-
🗓️ - Date
August 13, 2026
🕒 - Duration
Unknown
-
🏝️ - Location
On-site
-
📄 - Contract
Unknown
-
🔒 - Security
Unknown
-
📍 - Location detailed
Tarrytown, NY
-
🧠 - Skills detailed
#Athena #AWS S3 (Amazon Simple Storage Service) #AI (Artificial Intelligence) #Data Quality #Security #Batch #Strategy #Scala #Compliance #Data Management #S3 (Amazon Simple Storage Service) #AWS EMR (Amazon Elastic MapReduce) #Metadata #Spark SQL #Databricks #MLflow #IAM (Identity and Access Management) #Apache NiFi #Jenkins #PySpark #Airflow #Redshift #Spark (Apache Spark) #Data Integration #ML (Machine Learning) #Monitoring #AWS (Amazon Web Services) #Data Science #Data Pipeline #SQL (Structured Query Language) #Lambda (AWS Lambda) #Observability #Delta Lake #"ETL (Extract #Transform #Load)" #Data Governance #Data Strategy #Data Architecture #Data Access #Infrastructure as Code (IaC) #Migration #Automation #DevOps #Python #NiFi (Apache NiFi) #Data Processing #GitHub #Leadership #Apache Airflow #Data Migration #Cloud
Role description
Role : (Data) Databrick Architect Location : Tarrytown New York (Onsite) Duration: Long term Project Key Responsibilities & Skills: • Lead enterprise-scale data platform modernization initiatives, driving migration from AWS EMR, Apache NiFi, and legacy ETL frameworks to the Databricks Lakehouse Platform. • Architect and implement scalable Lakehouse solutions using Databricks, Delta Lake, Unity Catalog, and Databricks Workflows. • Design and govern end-to-end data pipelines for batch, streaming, CDC, and real-time data integration workloads. • Provide architecture leadership for large-scale AWS-based data ecosystems leveraging S3, IAM, Redshift, Glue Catalog, Airflow, and Databricks. • Develop enterprise data architecture standards covering data modelling, metadata management, lineage, governance, security, and compliance. • Drive adoption of Databricks best practices including Delta Live Tables (DLT), Auto Loader, Unity Catalog, Serverless Compute, Lakehouse Federation, and advanced optimization techniques. • Lead replatforming and migration programmes involving PySpark applications, Airflow DAGs, EMR workloads, Redshift integrations, and NiFi pipelines. • Partner closely with business stakeholders, enterprise architects, product owners, and data science teams to define technology roadmaps and target architectures. • Establish governance frameworks using Unity Catalog, data quality controls, observability, monitoring, and operational excellence practices • Provide technical leadership, mentoring, architecture reviews, design governance, and solution sign-offs across multiple delivery teams. • Lead architecture workshops, executive presentations, solution assessments, and technology evaluations. Mandatory Technical Skills: Databricks Lakehouse Platform Delta Lake, Delta Live Tables (DLT) Unity Catalog Databricks Workflows Auto Loader PySpark, Spark SQL, Python AWS (S3, IAM, Glue, Redshift, EMR, Lambda) Apache Airflow CDC & Data Migration Frameworks Data Governance & Security CI/CD, GitHub, Jenkins JD: • Lead enterprise-scale Databricks Lakehouse Architecture design and implementation on AWS. • Drive large-scale data platform modernisation and cloud transformation initiatives. • Architect scalable Medallion Architecture (Bronze, Silver, Gold) data platforms. • Lead migration of legacy EMR, NiFi, Redshift, and ETL workloads to Databricks. • Design high-performance batch and real-time data processing solutions. • Build robust ingestion frameworks using Auto Loader, Delta Lake, and Structured Streaming. • Define enterprise data governance and security standards using Unity Catalog. • Architect metadata-driven and reusable PySpark-based ETL/ELT frameworks. • Establish best practices for performance tuning, scalability, and cost optimisation. • Design and implement data quality, lineage, and observability frameworks. • Drive adoption of CI/CD, DevOps, Infrastructure as Code, and automation practices. • Collaborate with business, analytics, and engineering teams to define target-state architectures. • Conduct architecture reviews and provide technical leadership across multiple projects. • Mentor architects and senior engineers on Databricks and AWS best practices. • Design secure and scalable solutions leveraging AWS services (S3, Glue, Athena, Lambda, IAM, Redshift). • Implement and govern enterprise-wide data access, compliance, and security controls. • Evaluate and adopt latest Databricks capabilities such as DLT, Lakehouse Federation, Serverless, and MLflow. • Enable AI/ML, Generative AI, and advanced analytics use cases on the Lakehouse platform. • Create architecture roadmaps, migration strategies, standards, and governance frameworks. • Act as the primary technical advisor for customer leadership on data strategy and platform evolution.