

Insight Global
Scientific Data Engineer
⭐ - Featured Role | Apply direct with Data Freelance Hub
This role is for a Senior Data Engineer – Scientific Data & Drug Discovery Platforms on a contract basis until 2028, offering $60-70/hr. Key skills include data engineering, cheminformatics, Python, GCP, and experience with scientific datasets in a pharmaceutical context.
🌎 - Country
United States
💱 - Currency
$ USD
-
💰 - Day rate
560
-
🗓️ - Date
July 22, 2026
🕒 - Duration
Unknown
-
🏝️ - Location
Remote
-
📄 - Contract
Unknown
-
🔒 - Security
Unknown
-
📍 - Location detailed
United States
-
🧠 - Skills detailed
#Data Quality #GCP (Google Cloud Platform) #Python #R #Data Design #Deployment #Scala #BigQuery #Computer Science #ML (Machine Learning) #React #Data Ingestion #Data Pipeline #Data Science #"ETL (Extract #Transform #Load)" #Schema Design #Automation #AI (Artificial Intelligence) #Leadership #Data Architecture #Datasets #Metadata #Data Modeling #SQL (Structured Query Language) #Data Engineering #Database Architecture #Cloud #Storage #Database Design #Data Management #Observability
Role description
Senior Data Engineer – Cheminformatics & Scientific Data Platforms
Job Details
Title: Senior Data Engineer – Scientific Data & Drug Discovery Platforms
Location: Remote (US) EST Preferred
Type: Contract 2028
Industry: Global Pharmaceutical & Life Sciences Organization
Compensation: $60-70/hr
Overview / Job Description
A leading global pharmaceutical organization is seeking a Senior Data Engineer to help build the data foundation for a next-generation AI-enabled drug discovery platform. This individual will play a critical role in designing and delivering a cloud-native scientific data product that serves as the central source of truth for molecular design, computational chemistry, and AI-driven research workflows.
The ideal candidate combines strong data engineering expertise with experience working in chemistry, cheminformatics, computational chemistry, or scientific research environments. They will collaborate closely with computational chemists, data scientists, ML engineers, and research teams to design scalable data architectures capable of supporting generative drug design, laboratory automation, scientific modeling, and future agentic AI initiatives.
This role will focus on creating robust data pipelines, scalable schemas, and enterprise-level data products that support complex molecular datasets, scientific metadata, design lineage, and traceability across the drug discovery lifecycle.
Key Responsibilities
• Design, build, and own a new enterprise scientific data product supporting computational chemistry and molecular design initiatives.
• Develop scalable cloud-native data architecture capable of managing large scientific and chemistry-derived datasets.
• Create and optimize data models, schemas, and storage structures that support future AI/ML and analytics use cases.
• Build and maintain data ingestion, transformation, and orchestration pipelines using Python.
• Develop data solutions utilizing Google Cloud Platform (GCP) and BigQuery.
• Integrate data from scientific systems, computational chemistry platforms, laboratory environments, and external data sources.
• Design solutions that preserve scientific lineage, design history, traceability, and molecular decision-making workflows.
• Partner with computational chemists and scientists to model molecular design, reaction, synthesis, and experimental data.
• Ensure datasets are analytics-ready, trusted, scalable, and easily consumable by researchers, applications, and machine learning platforms.
• Build and maintain CI/CD processes, testing frameworks, and deployment automation.
• Implement data quality, observability, metadata management, and governance standards.
• Support future AI, machine learning, and agent-driven drug discovery initiatives through scalable data architecture design.
• Provide technical leadership on database design, schema evolution, performance optimization, and long-term platform scalability.
Required Skills
• Bachelor's or Master's degree in Chemistry, Life Sciences, Bioinformatics, Computer Science, Data Engineering, or a related discipline.
• Experience working with scientific datasets, chemistry data, computational chemistry data, or research data platforms.
• Experience developing enterprise-scale ETL/ELT pipelines.
• Experience managing large, complex datasets in scientific, pharmaceutical, or R&D environments.
• Experience with one or more cheminformatics tools:
• - ChemAxon, RDKit, BIOVIA, Pipeline Pilot
• Experience implementing CI/CD pipelines and modern software development practices.
• Strong understanding of cloud-native architecture and data platform scalability.
• Strong experience in: Data Engineering, Data Product Development, Database Architecture, Schema Design, Data Modeling
• Hands-on experience with: Google Cloud Platform (GCP), BigQuery, Python, SQL
What Success Looks Like
• Deliver a scalable scientific data platform that becomes the foundation for future drug discovery initiatives.
• Create trusted, analytics-ready datasets used by scientists, ML engineers, and AI systems.
• Establish robust data traceability from molecular design through experimentation and analysis.
• Enable future agentic AI and machine learning capabilities through well-structured, high-quality scientific data.
Senior Data Engineer – Cheminformatics & Scientific Data Platforms
Job Details
Title: Senior Data Engineer – Scientific Data & Drug Discovery Platforms
Location: Remote (US) EST Preferred
Type: Contract 2028
Industry: Global Pharmaceutical & Life Sciences Organization
Compensation: $60-70/hr
Overview / Job Description
A leading global pharmaceutical organization is seeking a Senior Data Engineer to help build the data foundation for a next-generation AI-enabled drug discovery platform. This individual will play a critical role in designing and delivering a cloud-native scientific data product that serves as the central source of truth for molecular design, computational chemistry, and AI-driven research workflows.
The ideal candidate combines strong data engineering expertise with experience working in chemistry, cheminformatics, computational chemistry, or scientific research environments. They will collaborate closely with computational chemists, data scientists, ML engineers, and research teams to design scalable data architectures capable of supporting generative drug design, laboratory automation, scientific modeling, and future agentic AI initiatives.
This role will focus on creating robust data pipelines, scalable schemas, and enterprise-level data products that support complex molecular datasets, scientific metadata, design lineage, and traceability across the drug discovery lifecycle.
Key Responsibilities
• Design, build, and own a new enterprise scientific data product supporting computational chemistry and molecular design initiatives.
• Develop scalable cloud-native data architecture capable of managing large scientific and chemistry-derived datasets.
• Create and optimize data models, schemas, and storage structures that support future AI/ML and analytics use cases.
• Build and maintain data ingestion, transformation, and orchestration pipelines using Python.
• Develop data solutions utilizing Google Cloud Platform (GCP) and BigQuery.
• Integrate data from scientific systems, computational chemistry platforms, laboratory environments, and external data sources.
• Design solutions that preserve scientific lineage, design history, traceability, and molecular decision-making workflows.
• Partner with computational chemists and scientists to model molecular design, reaction, synthesis, and experimental data.
• Ensure datasets are analytics-ready, trusted, scalable, and easily consumable by researchers, applications, and machine learning platforms.
• Build and maintain CI/CD processes, testing frameworks, and deployment automation.
• Implement data quality, observability, metadata management, and governance standards.
• Support future AI, machine learning, and agent-driven drug discovery initiatives through scalable data architecture design.
• Provide technical leadership on database design, schema evolution, performance optimization, and long-term platform scalability.
Required Skills
• Bachelor's or Master's degree in Chemistry, Life Sciences, Bioinformatics, Computer Science, Data Engineering, or a related discipline.
• Experience working with scientific datasets, chemistry data, computational chemistry data, or research data platforms.
• Experience developing enterprise-scale ETL/ELT pipelines.
• Experience managing large, complex datasets in scientific, pharmaceutical, or R&D environments.
• Experience with one or more cheminformatics tools:
• - ChemAxon, RDKit, BIOVIA, Pipeline Pilot
• Experience implementing CI/CD pipelines and modern software development practices.
• Strong understanding of cloud-native architecture and data platform scalability.
• Strong experience in: Data Engineering, Data Product Development, Database Architecture, Schema Design, Data Modeling
• Hands-on experience with: Google Cloud Platform (GCP), BigQuery, Python, SQL
What Success Looks Like
• Deliver a scalable scientific data platform that becomes the foundation for future drug discovery initiatives.
• Create trusted, analytics-ready datasets used by scientists, ML engineers, and AI systems.
• Establish robust data traceability from molecular design through experimentation and analysis.
• Enable future agentic AI and machine learning capabilities through well-structured, high-quality scientific data.






