Infojini Inc

Data Analytics Engineer

⭐ - Featured Role | Apply direct with Data Freelance Hub
This role is a Data Analytics Engineer focused on financial services, offering a full-time permanent position in the Bay Area, CA. Requires 5+ years of experience, strong SQL, PySpark, and AWS skills, and familiarity with banking regulations.
🌎 - Country
United States
πŸ’± - Currency
$ USD
-
πŸ’° - Day rate
Unknown
-
πŸ—“οΈ - Date
August 7, 2026
πŸ•’ - Duration
More than 6 months
-
🏝️ - Location
Hybrid
-
πŸ“„ - Contract
Unknown
-
πŸ”’ - Security
Unknown
-
πŸ“ - Location detailed
San Francisco, CA
-
🧠 - Skills detailed
#Data Governance #Data Pipeline #Datasets #Scala #Data Security #ML (Machine Learning) #IAM (Identity and Access Management) #Compliance #Airflow #Automation #Databricks #PySpark #dbt (data build tool) #Lambda (AWS Lambda) #S3 (Amazon Simple Storage Service) #Kafka (Apache Kafka) #Data Engineering #Terraform #Computer Science #Data Ingestion #AWS S3 (Amazon Simple Storage Service) #"ETL (Extract #Transform #Load)" #Data Science #Python #Scripting #Deployment #Data Modeling #Monitoring #Delta Lake #SQL (Structured Query Language) #Version Control #Data Processing #Cloud #Data Quality #Data Transformations #PCI (Payment Card Industry) #Security #Automated Testing #GitHub #Infrastructure as Code (IaC) #Spark (Apache Spark) #Redshift #AWS (Amazon Web Services)
Role description
Data Analytics Engineer (Financial) Needed: Banking/fintech, or financial services Location: Bay Area, San Francisco, CA- Hybrid Job Type: Full-Time Permanent Role (No Contract) Role Summary We are seeking an experienced Data Analytics Engineer to design, build, and optimize scalable data pipelines and analytics infrastructure that power critical financial products and business decisions. You will work at the intersection of software engineering and data analytics, building reliable ETL/ELT pipelines, orchestrating cloud-native workflows, and enabling trusted, high-quality data for reporting, risk, and product teams. This role requires strong engineering discipline (version control, CI/CD, infrastructure-as-code) combined with deep SQL and PySpark expertise, ideally within a regulated banking or financial services environment. Key Responsibilities β€’ Design, build, and maintain scalable, reliable ETL/ELT data pipelines across cloud and on-premises sources, ensuring data quality, lineage, and auditability. β€’ Develop and optimize Python, PySpark, and SQL-based data transformations for large-scale, high-volume financial datasets. β€’ Architect and manage data pipeline orchestration using tools such as Airflow, Databricks Workflows, Step Functions, or equivalent to automate data ingestion, transformation, and delivery. β€’ Build and maintain CI/CD pipelines using GitHub and GitHub Actions to support automated testing, deployment, and version-controlled infrastructure changes. β€’ Develop cloud-based solutions on AWS (S3, Glue, EMR, Redshift, Lambda, IAM) supporting analytics, reporting, and downstream machine learning use cases. β€’ Deploy and manage infrastructure and pipelines as code, following best practices for environment promotion, rollback, and monitoring. β€’ Monitor, troubleshoot, and optimize pipeline performance, query efficiency, and cloud cost across the data stack. β€’ Partner with data scientists, analysts, product, and risk/compliance teams to translate business requirements into scalable and reliable data solutions. β€’ Enforce data governance, security, and regulatory compliance standards appropriate for financial data (PII, SOX, PCI, etc.). β€’ Document pipeline architecture, data models, and engineering processes while contributing to coding standards and code review practices. Required Technical Skills β€’ Advanced proficiency in Python for scripting, automation, and data engineering workflows. β€’ Strong hands-on experience with PySpark for distributed data processing at scale. β€’ Expert-level SQL skills, including advanced SQL concepts such as window functions, complex joins, query optimization, and performance tuning. β€’ Solid experience with AWS cloud services, including S3, Glue, EMR, Redshift, Lambda, IAM, and CloudWatch. β€’ Proven expertise building and orchestrating data pipelines using Airflow, Databricks Workflows, Step Functions, or similar technologies. β€’ Hands-on CI/CD experience using GitHub and GitHub Actions for automated build, testing, and deployment. β€’ Deep understanding of ETL/ELT design patterns, data modeling, and data warehousing concepts. β€’ Experience deploying infrastructure and data pipelines as code through version-controlled deployments. β€’ Demonstrated ability to optimize pipeline performance, query execution, and cloud resource utilization while improving cost efficiency. Preferred Skills β€’ Hands-on experience with Databricks, including Delta Lake, Unity Catalog, notebooks, and cluster optimization. β€’ Familiarity with Terraform or AWS CloudFormation for Infrastructure as Code (IaC). β€’ Experience working with streaming technologies such as Kafka, Kinesis, or Spark Structured Streaming. β€’ Exposure to data quality and testing frameworks such as Great Expectations or dbt Tests. β€’ Knowledge of dbt for transformation and analytics engineering workflows. β€’ Understanding of financial data domains, including payments, lending, risk, fraud, or accounting. β€’ Relevant certifications such as AWS Certified Data Analytics, AWS Certified Solutions Architect, or Databricks Certified Data Engineer. Qualifications β€’ Bachelor's degree in Computer Science, Engineering, Data Science, or a related field (or equivalent practical experience). β€’ 5+ years of experience in Data Engineering, Analytics Engineering, or a related technical role. β€’ Prior experience working within banking, fintech, or financial services, with an understanding of regulatory and data security requirements. β€’ Proven track record of delivering production-grade data pipelines in cloud environments.