Yochana

Contract Role: Data Scientist at Minneapolis, MN(Remote) || W2 Only

⭐ - Featured Role | Apply direct with Data Freelance Hub
This role is a Data Scientist position based in Minneapolis, MN (Remote) for 6+ months on a W2 contract. Requires a PhD/Master's in Computer Science or related field, 8+ years of AI/ML experience, and deep expertise in LLMs and graph-based systems.
🌎 - Country
United States
πŸ’± - Currency
$ USD
-
πŸ’° - Day rate
Unknown
-
πŸ—“οΈ - Date
July 21, 2026
πŸ•’ - Duration
More than 6 months
-
🏝️ - Location
Remote
-
πŸ“„ - Contract
W2 Contractor
-
πŸ”’ - Security
Unknown
-
πŸ“ - Location detailed
Minneapolis, MN
-
🧠 - Skills detailed
#ML (Machine Learning) #HBase #Databases #Indexing #Reinforcement Learning #Computer Science #Data Science #Compliance #AI (Artificial Intelligence) #Deployment #Knowledge Graph
Role description
Data Scientist Minneapolis, MN (Remote) 6+ Months Employment Type – W2 Only Work Experience: β€’ Lead end-to-end training and fine-tuning of Large Language Models (LLMs), including both open-source (e.g., Qwen, LLaMA, Mistral) and closed-source (e.g., OpenAI, Gemini, Anthropic) ecosystems. β€’ Architect and implement GraphRAG pipelines, including knowledge graph representation and retrieval for enhanced contextual grounding. β€’ Design, train, and optimize semantic and dense vector embeddings for document understanding, search, and retrieval. β€’ Develop semantic retrieval systems with advanced document segmentation and indexing strategies. β€’ Build and scale distributed training environments using NCCL and InfiniBand for multi-GPU and multi-node training. β€’ Apply reinforcement learning techniques (e.g., RLHF, RLAIF) to align model behavior with human preferences and domain-specific goals. β€’ Collaborate with cross-functional teams to translate business needs into AI-driven solutions and deploy them in production environments. Qualifications β€’ PhD or Master’s degree in Computer Science, Machine Learning, or related field. β€’ 8+ years of experience in applied AI/ML, with a strong track record of delivering production-grade models. Deep expertise in: β€’ LLM training and fine-tuning (e.g., GPT, LLaMA, Mistral, Qwen) β€’ Graph-based retrieval systems (GraphRAG, knowledge graphs) β€’ Embedding models (e.g., BGE, E5, SimCSE) β€’ Semantic search and vector databases (e.g., FAISS, Weaviate, Milvus) β€’ Document segmentation and preprocessing (OCR, layout parsing) β€’ Distributed training frameworks (NCCL, Horovod, DeepSpeed) β€’ High-performance networking (InfiniBand, RDMA) β€’ Model fusion and ensemble techniques (stacking, boosting, gating) β€’ Optimization algorithms (Bayesian, Particle Swarm, Genetic Algorithms) β€’ Symbolic AI and rule-based systems β€’ Meta-learning and Mixture of Experts architectures β€’ Reinforcement learning (e.g., RLHF, PPO, DPO) Bonus Skills β€’ Experience with healthcare data and medical coding systems (e.g., CPT, CM, PCS). β€’ Familiarity with regulatory and compliance frameworks in AI deployment. β€’ Contributions to open-source AI projects or published research. And/Or ability to take research papers to poc – production.