Iris Software Inc.

Data Engineer

⭐ - Featured Role | Apply direct with Data Freelance Hub
This role is for a Data & AI Platform Engineer focusing on Databricks and AWS, with an 18-month hybrid contract in Roseland, NJ. Key skills include Databricks, PySpark, Python, and LLM integration. Requires 3+ years of relevant experience.
🌎 - Country
United States
💱 - Currency
$ USD
-
💰 - Day rate
Unknown
-
🗓️ - Date
August 12, 2026
🕒 - Duration
More than 6 months
-
🏝️ - Location
Hybrid
-
📄 - Contract
Unknown
-
🔒 - Security
Unknown
-
📍 - Location detailed
Roseland, NJ
-
🧠 - Skills detailed
#Databricks #Datasets #Lambda (AWS Lambda) #Python #MLflow #React #PySpark #Data Architecture #Delta Lake #Compliance #SQL (Structured Query Language) #Spark (Apache Spark) #AI (Artificial Intelligence) #Data Ingestion #VPC (Virtual Private Cloud) #AWS S3 (Amazon Simple Storage Service) #S3 (Amazon Simple Storage Service) #Forecasting #Data Lineage #Anomaly Detection #Data Engineering #AWS (Amazon Web Services) #IAM (Identity and Access Management) #Storage #API (Application Programming Interface) #Security #Semantic Models #Data Pipeline #"ETL (Extract #Transform #Load)"
Role description
Iris's direct client, a leader Payroll, is looking to hire a strong Databricks Platform Engineer for a Long Term Contractual role. Job title: Data & AI Platform Engineer — (Databricks + AWS) Location: Hybrid (3 days onsite, Roseland, NJ) Duration: 18 Months Skills: Databricks, Pyspark, Python, Genie Space Or Genie Agent Job Description About the Role Own the Databricks lakehouse layer and AI/LLM platform for Finance Insights. You'll design the data architecture on Databricks Unity Catalog, build Genie Spaces for natural language analytics, and integrate LLMs to power intelligent financial workflows — all running on AWS. Responsibilities • Design and build Delta Lake tables on Databricks (bronze/silver/gold medallion architecture) for financial datasets (AR, GL, invoices, collections) • Build and manage Databricks Genie Spaces — configure semantic models, write GENIE-compatible SQL, and tune NL-to-SQL behavior for finance use cases • Develop PySpark and Python-based data pipelines (Databricks Jobs + Workflows) for ERP data ingestion, transformation, and aggregation • Integrate Databricks Model Serving with LLMs (Claude, GPT-4, or custom fine-tuned models) for AI-driven insights, anomaly detection, and financial forecasting • Build RAG pipelines using Databricks Vector Search over financial documents, contracts, and GL notes • Expose Databricks SQL Warehouse endpoints to Python API services and the React frontend • Manage Unity Catalog governance — row-level security, column masking, data lineage for finance compliance • Deploy infrastructure on AWS (S3 as Databricks external storage, IAM roles, VPC networking, Secrets Manager) Required Skills • Databricks (PySpark, Delta Lake, Jobs/Workflows, SQL Warehouses) — 3+ years • Databricks Unity Catalog and governance • Databricks Genie Spaces (semantic layer, NL analytics) • Python 3.11+ for data engineering and API integration • AWS (S3, IAM, VPC, Secrets Manager, optionally Lambda/Glue) • LLM integration: Databricks Model Serving, MLflow, or direct API calls to Claude/OpenAI • SQL (advanced — window functions, CTEs, query optimization)