Intracruit Solutions

Senior Data Scientist

⭐ - Featured Role | Apply direct with Data Freelance Hub
This role is for a Senior Data Scientist with 13+ years of experience, focused on NLP and SQL. It requires on-site work in Woodlawn, MD, offers a competitive pay rate, and demands strong communication skills and a bachelor's degree in a relevant field.
🌎 - Country
United States
πŸ’± - Currency
$ USD
-
πŸ’° - Day rate
Unknown
-
πŸ—“οΈ - Date
August 19, 2026
πŸ•’ - Duration
Unknown
-
🏝️ - Location
On-site
-
πŸ“„ - Contract
Unknown
-
πŸ”’ - Security
Unknown
-
πŸ“ - Location detailed
Woodlawn, MD
-
🧠 - Skills detailed
#Data Integrity #Hadoop #Deployment #Data Privacy #NLP (Natural Language Processing) #Clustering #SQL Server #Version Control #Datasets #Data Access #Libraries #Python #Jenkins #SpaCy #Data Cleansing #Automation #"ETL (Extract #Transform #Load)" #Data Processing #Mathematics #Oracle #Statistics #Data Security #Code Reviews #SQL (Structured Query Language) #ML (Machine Learning) #Scala #Leadership #Data Science #Computer Science #Integration Testing #SQL Queries #Data Pipeline #Monitoring #Indexing #PostgreSQL #Security
Role description
Position Title: Senior Data Scientist 13+ yrs ||F2F interview Woodlawn MD Visa : USC,GC,H1B Vendor Notes Candidates need to have excellent communication skills. This will be 5 days working from SSA HQ, so candidates should be local and ready to interview immediately. Key Required Skills: Strong practical experience with Natural Language Processing (NLP), Text Processing, and Information Extraction concepts including Named Entity Recognition, Blocking and Indexing, String Distance Metrics, TF-IDF/Cosine Similarity, Phonetics Encoding, Address Standardization Solid Python, Regex, and SQL experience Excellent Communication skills Position Description: β€’ Develop Analytics Solutions: Design, implement, and maintain advanced data processing and entity resolution pipelines using Python and SQL on enterprise data platforms. β€’ Data Hygiene & Management: Clean, transform, and manage large-scale datasets from diverse, complex sources, ensuring absolute data integrity, reliability, and security. β€’ Performance Optimization: Optimize complex SQL queries and database operations to ensure efficient data access, processing, and scalability. β€’ Engineering Standards: Actively participate in code reviews, enforce version control, and uphold best practices for code quality, reproducibility, and data privacy. β€’ End-to-End Delivery: Support data validation, testing, deployment, and post-implementation monitoring in a fast-paced environment. Skills Requirements: β€’ Foundation for Success (Basic Qualifications) β€’ Bachelor’s degree in Statistics, Applied Mathematics, Computer Science, or Information Science with experience in NLP, Text Processing, Information Extraction, Python, SQL, Regex and specialized libraries/frameworks. β€’ Overall 10+ years’ experience in IT industry Factors To Help You Shine (Required Skills) β€’ Selected candidate must be able to obtain and maintain a public trust clearance β€’ Selected candidate must be willing to work on-site in Woodlawn, MD 5 days a week β€’ Master's and 10+ years of experience, Bachelor's and 12+ years of experience or 18+ years in lieu of a degree β€’ Strong practical experience with Natural Language Processing (NLP), Text Processing, and Information Extraction including: β€’ Practical knowledge of Named Entity Recognition and Address Standardization to extract and clean unstructured text data. β€’ Deep understanding of data matching strategies, including Blocking and Indexing, String Distance Metrics, and Phonetic Encoding. β€’ Experience applying TD-IDF and Cosine Similarity for text comparisons and information retrieval. β€’ Strong Python development skills for building analytics solutions and manipulating data. β€’ Advanced SQL proficiency for complex data querying, optimization, and database operations. β€’ Practical experience using Regex for advanced text processing, data cleansing, and pattern matching. β€’ Familiarity with specialized libraries and frameworks including: β€’ Linkage libraries such as Splink / FastLink, Dedupe, or recordlinkage β€’ Core Python data science libraries, specifically spaCy for NLP tasks and Scikit-Learn for general machine learning and clustering β€’ Familiarity with code reviews, version control, and maintaining data security and reproducibility standards. β€’ Excellent communication skills. How To Stand Out From The Crowd (Desired Skills) β€’ Prior experience delivering IT or data initiatives within federal, state, or local government environments β€’ Proven ability to operate independently, take ownership of data pipeline architectures, and drive projects from data discovery through to post-implementation. β€’ Experience retrieving, migrating, and manipulating data from legacy and distributed systems, including PostgreSQL, DB2, Oracle, SQL Server, and Hadoop, as well as unstructured flat files. β€’ Experience utilizing Jenkins to automate continuous integration, testing, and deployment (CI/CD) for data validation pipelines. β€’ Experience with pipeline automation tools to schedule and monitor complex data cleansing jobs. β€’ Strong ability to translate complex algorithmic decisions (such as probabilistic match thresholds) into clear business logic for executive leadership and non-technical stakeholders. β€’ Excellent problem-solving skills and proven verbal/written communication skills when collaborating across cross-functional teams.