

Yochana
Contract Role: Data Scientist at Minneapolis, MN(Remote) || W2 Only
β - Featured Role | Apply direct with Data Freelance Hub
This role is a Data Scientist position based in Minneapolis, MN (Remote) for 6+ months on a W2 contract. Requires a PhD/Master's in Computer Science or related field, 8+ years of AI/ML experience, and deep expertise in LLMs and graph-based systems.
π - Country
United States
π± - Currency
$ USD
-
π° - Day rate
Unknown
-
ποΈ - Date
July 21, 2026
π - Duration
More than 6 months
-
ποΈ - Location
Remote
-
π - Contract
W2 Contractor
-
π - Security
Unknown
-
π - Location detailed
Minneapolis, MN
-
π§ - Skills detailed
#ML (Machine Learning) #HBase #Databases #Indexing #Reinforcement Learning #Computer Science #Data Science #Compliance #AI (Artificial Intelligence) #Deployment #Knowledge Graph
Role description
Data Scientist
Minneapolis, MN (Remote)
6+ Months
Employment Type β W2 Only
Work Experience:
β’ Lead end-to-end training and fine-tuning of Large Language Models (LLMs), including both open-source (e.g., Qwen, LLaMA, Mistral) and closed-source (e.g., OpenAI, Gemini, Anthropic) ecosystems.
β’ Architect and implement GraphRAG pipelines, including knowledge graph representation and retrieval for enhanced contextual grounding.
β’ Design, train, and optimize semantic and dense vector embeddings for document understanding, search, and retrieval.
β’ Develop semantic retrieval systems with advanced document segmentation and indexing strategies.
β’ Build and scale distributed training environments using NCCL and InfiniBand for multi-GPU and multi-node training.
β’ Apply reinforcement learning techniques (e.g., RLHF, RLAIF) to align model behavior with human preferences and domain-specific goals.
β’ Collaborate with cross-functional teams to translate business needs into AI-driven solutions and deploy them in production environments.
Qualifications
β’ PhD or Masterβs degree in Computer Science, Machine Learning, or related field.
β’ 8+ years of experience in applied AI/ML, with a strong track record of delivering production-grade models.
Deep expertise in:
β’ LLM training and fine-tuning (e.g., GPT, LLaMA, Mistral, Qwen)
β’ Graph-based retrieval systems (GraphRAG, knowledge graphs)
β’ Embedding models (e.g., BGE, E5, SimCSE)
β’ Semantic search and vector databases (e.g., FAISS, Weaviate, Milvus)
β’ Document segmentation and preprocessing (OCR, layout parsing)
β’ Distributed training frameworks (NCCL, Horovod, DeepSpeed)
β’ High-performance networking (InfiniBand, RDMA)
β’ Model fusion and ensemble techniques (stacking, boosting, gating)
β’ Optimization algorithms (Bayesian, Particle Swarm, Genetic Algorithms)
β’ Symbolic AI and rule-based systems
β’ Meta-learning and Mixture of Experts architectures
β’ Reinforcement learning (e.g., RLHF, PPO, DPO)
Bonus Skills
β’ Experience with healthcare data and medical coding systems (e.g., CPT, CM, PCS).
β’ Familiarity with regulatory and compliance frameworks in AI deployment.
β’ Contributions to open-source AI projects or published research. And/Or ability to take research papers to poc β production.
Data Scientist
Minneapolis, MN (Remote)
6+ Months
Employment Type β W2 Only
Work Experience:
β’ Lead end-to-end training and fine-tuning of Large Language Models (LLMs), including both open-source (e.g., Qwen, LLaMA, Mistral) and closed-source (e.g., OpenAI, Gemini, Anthropic) ecosystems.
β’ Architect and implement GraphRAG pipelines, including knowledge graph representation and retrieval for enhanced contextual grounding.
β’ Design, train, and optimize semantic and dense vector embeddings for document understanding, search, and retrieval.
β’ Develop semantic retrieval systems with advanced document segmentation and indexing strategies.
β’ Build and scale distributed training environments using NCCL and InfiniBand for multi-GPU and multi-node training.
β’ Apply reinforcement learning techniques (e.g., RLHF, RLAIF) to align model behavior with human preferences and domain-specific goals.
β’ Collaborate with cross-functional teams to translate business needs into AI-driven solutions and deploy them in production environments.
Qualifications
β’ PhD or Masterβs degree in Computer Science, Machine Learning, or related field.
β’ 8+ years of experience in applied AI/ML, with a strong track record of delivering production-grade models.
Deep expertise in:
β’ LLM training and fine-tuning (e.g., GPT, LLaMA, Mistral, Qwen)
β’ Graph-based retrieval systems (GraphRAG, knowledge graphs)
β’ Embedding models (e.g., BGE, E5, SimCSE)
β’ Semantic search and vector databases (e.g., FAISS, Weaviate, Milvus)
β’ Document segmentation and preprocessing (OCR, layout parsing)
β’ Distributed training frameworks (NCCL, Horovod, DeepSpeed)
β’ High-performance networking (InfiniBand, RDMA)
β’ Model fusion and ensemble techniques (stacking, boosting, gating)
β’ Optimization algorithms (Bayesian, Particle Swarm, Genetic Algorithms)
β’ Symbolic AI and rule-based systems
β’ Meta-learning and Mixture of Experts architectures
β’ Reinforcement learning (e.g., RLHF, PPO, DPO)
Bonus Skills
β’ Experience with healthcare data and medical coding systems (e.g., CPT, CM, PCS).
β’ Familiarity with regulatory and compliance frameworks in AI deployment.
β’ Contributions to open-source AI projects or published research. And/Or ability to take research papers to poc β production.






