

Tech Observer
Data Engineer
⭐ - Featured Role | Apply direct with Data Freelance Hub
This role is for a Data Engineer on a 12-month contract in Cambridge, MA, offering a pay rate of "unknown." Key skills include Databricks, Seqera Platform, cloud data engineering, and scientific workflows. A Bachelor's degree and 5+ years of relevant experience are required.
🌎 - Country
United States
💱 - Currency
$ USD
-
💰 - Day rate
Unknown
-
🗓️ - Date
August 19, 2026
🕒 - Duration
More than 6 months
-
🏝️ - Location
On-site
-
📄 - Contract
Unknown
-
🔒 - Security
Unknown
-
📍 - Location detailed
Cambridge, MA
-
🧠 - Skills detailed
#Infrastructure as Code (IaC) #Data Engineering #API (Application Programming Interface) #Data Modeling #EC2 #Docker #VPC (Virtual Private Cloud) #Data Lakehouse #Storage #GitLab #Datasets #Data Catalog #IAM (Identity and Access Management) #Python #Jenkins #Databricks #AWS (Amazon Web Services) #Agile #Automation #Documentation #"ETL (Extract #Transform #Load)" #Data Ingestion #Data Processing #Data Quality #DevOps #GitHub #SQL (Structured Query Language) #Spark (Apache Spark) #ML (Machine Learning) #Delta Lake #PySpark #Cloud #AI (Artificial Intelligence) #Terraform #Data Lifecycle #Lambda (AWS Lambda) #Logging #Kubernetes #Scala #Apache Spark #Data Science #Computer Science #R #Spark SQL #Data Pipeline #S3 (Amazon Simple Storage Service) #Monitoring #Data Lake #Security
Role description
Title : Data Engineer
Duration : 12 - months contract (high possibility of extension)
Location : Cambridge MA
Hours per week : 40 Hours
Job Requirement
We are seeking a highly skilled Data Engineer with expertise in Databricks, Seqera Platform (Nextflow Tower), cloud data engineering, and scientific data workflows to support our Discovery R&D and data science initiatives. The ideal candidate will design, develop, and maintain scalable data platforms and automated bioinformatics/data science pipelines that enable researchers and scientists to efficiently process, analyze, and access large-scale scientific and business datasets. This role requires strong experience in cloud-native architectures, data engineering best practices, workflow orchestration, and collaboration with cross-functional teams including scientists, bioinformaticians, data scientists, and IT infrastructure teams.
About the Role
The role involves designing, developing, and maintaining scalable data platforms and automated bioinformatics/data science pipelines.
Responsibilities
• Data Engineering & Platform Development
• Design, develop, and maintain scalable data pipelines using Databricks, Apache Spark, and cloud-native technologies.
• Build and optimize ETL/ELT processes for structured, semi-structured, and unstructured data.
• Develop data ingestion frameworks for research, laboratory, clinical, and external scientific datasets.
• Implement data quality, validation, monitoring, and governance processes.
• Support enterprise data lakehouse architecture and data platform modernization initiatives.
• Databricks Administration & Development
• Develop and maintain Databricks notebooks, workflows, Delta Live Tables, and Jobs.
• Create optimized Spark-based transformations and data processing solutions.
• Implement Medallion Architecture (Bronze, Silver, Gold) for data lifecycle management.
• Manage Delta Lake environments and optimize performance, scalability, and cost.
• Integrate Databricks with cloud-native services and enterprise applications.
• Seqera Platform & Scientific Workflow Management
• Deploy, configure, and support Seqera Platform (formerly Nextflow Tower).
• Develop and maintain Nextflow pipelines for bioinformatics, genomics, imaging, AI/ML, and scientific computing workloads.
• Integrate Seqera workflows with AWS cloud infrastructure and compute environments.
• Support containerized workflows using Docker and Kubernetes technologies.
• Enable reproducible, scalable, and compliant scientific data processing workflows.
• Cloud Engineering
• Design and implement cloud-based data solutions in AWS.
• Manage cloud storage solutions including S3 and data lifecycle policies.
• Develop Infrastructure-as-Code solutions using Terraform or CloudFormation.
• Implement security controls and access management following enterprise IT standards.
• Partner with data scientists, researchers, bioinformaticians, and business stakeholders to understand data requirements.
• Provide technical guidance on data engineering best practices and workflow automation.
• Troubleshoot pipeline failures, performance issues, and workflow bottlenecks.
• Contribute to platform roadmaps and continuous improvement initiatives.
• Maintain technical documentation, SOPs, and knowledge articles.
Qualifications
Education
• Bachelor's degree in Computer Science, Information Technology, Data Engineering, Bioinformatics, or a related technical field.
• Master's degree preferred.
Experience
• 5+ years of experience in data engineering, cloud engineering, or analytics platform development.
• 3+ years of hands-on experience with Databricks and Apache Spark.
• 2+ years of experience with Seqera Platform (Nextflow Tower) and Nextflow workflows.
• Experience supporting scientific research, life sciences, pharmaceutical, biotech, or healthcare environments preferred.
Required Skills
• Databricks & Data Engineering, Databricks Lakehouse Platform
• Apache Spark (PySpark, Spark SQL)
• Delta Lake, Delta Live Tables (DLT)
• Databricks Workflows
• Unity Catalog
• SQL and Python
• Seqera & Scientific Computing; Seqera Platform / Nextflow Tower
• Nextflow pipeline development
• Bioinformatics workflow automation
• Docker and container technologies
• Kubernetes orchestration
• High-performance computing environments
• Cloud Technologies, AWS (required)
• S3, IAM, EC2, VPC, Lambda
• Terraform or CloudFormation
• Cloud monitoring and logging tools
• Data Technologies
• Data Lake and Lakehouse architectures
• ETL/ELT frameworks
• Data modeling
• Data cataloging and governance
• API integrations
• Data quality frameworks
• DevOps & Automation
• GitHub/GitLab
• CI/CD pipelines
• Jenkins, GitHub Actions, or similar tools
• Infrastructure as Code
• Agile and DevOps methodologies
Title : Data Engineer
Duration : 12 - months contract (high possibility of extension)
Location : Cambridge MA
Hours per week : 40 Hours
Job Requirement
We are seeking a highly skilled Data Engineer with expertise in Databricks, Seqera Platform (Nextflow Tower), cloud data engineering, and scientific data workflows to support our Discovery R&D and data science initiatives. The ideal candidate will design, develop, and maintain scalable data platforms and automated bioinformatics/data science pipelines that enable researchers and scientists to efficiently process, analyze, and access large-scale scientific and business datasets. This role requires strong experience in cloud-native architectures, data engineering best practices, workflow orchestration, and collaboration with cross-functional teams including scientists, bioinformaticians, data scientists, and IT infrastructure teams.
About the Role
The role involves designing, developing, and maintaining scalable data platforms and automated bioinformatics/data science pipelines.
Responsibilities
• Data Engineering & Platform Development
• Design, develop, and maintain scalable data pipelines using Databricks, Apache Spark, and cloud-native technologies.
• Build and optimize ETL/ELT processes for structured, semi-structured, and unstructured data.
• Develop data ingestion frameworks for research, laboratory, clinical, and external scientific datasets.
• Implement data quality, validation, monitoring, and governance processes.
• Support enterprise data lakehouse architecture and data platform modernization initiatives.
• Databricks Administration & Development
• Develop and maintain Databricks notebooks, workflows, Delta Live Tables, and Jobs.
• Create optimized Spark-based transformations and data processing solutions.
• Implement Medallion Architecture (Bronze, Silver, Gold) for data lifecycle management.
• Manage Delta Lake environments and optimize performance, scalability, and cost.
• Integrate Databricks with cloud-native services and enterprise applications.
• Seqera Platform & Scientific Workflow Management
• Deploy, configure, and support Seqera Platform (formerly Nextflow Tower).
• Develop and maintain Nextflow pipelines for bioinformatics, genomics, imaging, AI/ML, and scientific computing workloads.
• Integrate Seqera workflows with AWS cloud infrastructure and compute environments.
• Support containerized workflows using Docker and Kubernetes technologies.
• Enable reproducible, scalable, and compliant scientific data processing workflows.
• Cloud Engineering
• Design and implement cloud-based data solutions in AWS.
• Manage cloud storage solutions including S3 and data lifecycle policies.
• Develop Infrastructure-as-Code solutions using Terraform or CloudFormation.
• Implement security controls and access management following enterprise IT standards.
• Partner with data scientists, researchers, bioinformaticians, and business stakeholders to understand data requirements.
• Provide technical guidance on data engineering best practices and workflow automation.
• Troubleshoot pipeline failures, performance issues, and workflow bottlenecks.
• Contribute to platform roadmaps and continuous improvement initiatives.
• Maintain technical documentation, SOPs, and knowledge articles.
Qualifications
Education
• Bachelor's degree in Computer Science, Information Technology, Data Engineering, Bioinformatics, or a related technical field.
• Master's degree preferred.
Experience
• 5+ years of experience in data engineering, cloud engineering, or analytics platform development.
• 3+ years of hands-on experience with Databricks and Apache Spark.
• 2+ years of experience with Seqera Platform (Nextflow Tower) and Nextflow workflows.
• Experience supporting scientific research, life sciences, pharmaceutical, biotech, or healthcare environments preferred.
Required Skills
• Databricks & Data Engineering, Databricks Lakehouse Platform
• Apache Spark (PySpark, Spark SQL)
• Delta Lake, Delta Live Tables (DLT)
• Databricks Workflows
• Unity Catalog
• SQL and Python
• Seqera & Scientific Computing; Seqera Platform / Nextflow Tower
• Nextflow pipeline development
• Bioinformatics workflow automation
• Docker and container technologies
• Kubernetes orchestration
• High-performance computing environments
• Cloud Technologies, AWS (required)
• S3, IAM, EC2, VPC, Lambda
• Terraform or CloudFormation
• Cloud monitoring and logging tools
• Data Technologies
• Data Lake and Lakehouse architectures
• ETL/ELT frameworks
• Data modeling
• Data cataloging and governance
• API integrations
• Data quality frameworks
• DevOps & Automation
• GitHub/GitLab
• CI/CD pipelines
• Jenkins, GitHub Actions, or similar tools
• Infrastructure as Code
• Agile and DevOps methodologies






