

Triune Infomatics Inc
AI Inference Engineer
⭐ - Featured Role | Apply direct with Data Freelance Hub
This role is for an AI Inference Engineer with a 12+ month contract in San Jose, CA. Key skills include GPU management, KV cache management, and experience with LLM serving frameworks. Proficiency in CUDA or ROCm, Python, and C++ is required.
🌎 - Country
United States
💱 - Currency
$ USD
-
💰 - Day rate
Unknown
-
🗓️ - Date
August 10, 2026
🕒 - Duration
More than 6 months
-
🏝️ - Location
Hybrid
-
📄 - Contract
Unknown
-
🔒 - Security
Unknown
-
📍 - Location detailed
San Jose, CA
-
🧠 - Skills detailed
#"ETL (Extract #Transform #Load)" #Distributed Computing #Documentation #Scala #Python #Computer Science #Storage #C++ #Deployment #AI (Artificial Intelligence) #Programming
Role description
AI Inference Engineer ( CUDA or ROCm)
Client: Samsung Cognos |Location: San Jose, CA (Hybrid/Onsite)
|Duration: 12+ Months |Engagement: Contract
MUST
GPU Management
KV Cache Management
About the Engagement
Samsung Cognos is building a next-generation LLM inference layer in partnership with SGLang, one of the leading open-source serving frameworks in the space. The project addresses one of the hardest problems in large-scale AI deployment: making KV cache memory management fast, efficient, and cost-effective across tiered storage hierarchies at production scale. This is a high-impact engineering engagement where your work will directly influence how AI inference performs for thousands of concurrent users.
Role Overview
As an AI Inference Engineer, you will own the serving stack that powers Samsung's LLM inference layer. Your focus will span prefill and decode optimization, KV cache offload strategies, quantization, speculative decoding, and tensor and pipeline parallelism. You will work within SGLang and vLLM-style frameworks, tune ROCm kernels for AMD hardware (MI300 and MI250 series), and evaluate RDMA and RoCE protocols for cross-machine cache sharing. This role is algorithm and framework focused: the core question you are answering is whether the model serving logic itself is efficient, scalable, and production-ready.
Work Model: This position follows a Hybrid/Onsite schedule at the Samsung Cognos facility in San Jose, CA. Candidates must be willing and able to work onsite as required by the client.
Key Responsibilities
• Design and optimize KV cache offload strategies across a tiered memory hierarchy: GPU HBM, CPU DRAM, and NVMe SSD.
• Implement and tune PagedAttention, RadixAttention, chunked prefill, and prefix caching within SGLang and vLLM-style serving frameworks.
• Drive prefill and decode stage optimization to maximize throughput and minimize latency for long-context and multi-user workloads.
• Apply quantization techniques and speculative decoding to reduce memory footprint and improve inference speed without degrading output quality.
• Architect and implement tensor parallelism and pipeline parallelism for multi-GPU inference deployments.
• Tune ROCm kernels for AMD MI300 and MI250 GPUs, adapting CUDA-based patterns to the AMD software stack.
• Evaluate RDMA and RoCE solutions for cross-node KV cache sharing and assess feasibility in production environments.
• Collaborate closely with the GPU Software Engineering team on memory-path correctness, data movement latency, and scheduler integration.
• Profile, benchmark, and iterate on inference stack performance across diverse model architectures and workload profiles.
• Contribute to design reviews, technical documentation, and team knowledge-sharing sessions.
Required Qualifications
• Bachelor's or Master's degree in Computer Science, Electrical Engineering, or a closely related field.
• Hands-on experience with LLM serving frameworks such as SGLang, vLLM, or TensorRT-LLM in a production or research setting.
• Strong understanding of transformer architecture internals, including attention mechanisms and the KV cache.
• Demonstrated experience with GPU programming: CUDA, ROCm/HIP, or equivalent.
• Proficiency in Python and C++ for systems-level performance work.
• Solid grounding in distributed computing concepts including tensor parallelism, pipeline parallelism, and data parallelism.
• Experience with quantization methods (INT8, FP8, GPTQ, AWQ) and their performance tradeoffs.
• Ability to profile and optimize inference pipelines using tools such as NSight, ROCProfiler, or equivalent.
• Strong analytical skills with the ability to translate benchmarking data into actionable engineering decisions.
Preferred Qualifications
• Direct experience with AMD MI300 or MI250 GPU hardware and the ROCm software ecosystem.
• Familiarity with speculative decoding frameworks (Medusa, EAGLE, or similar).
• Exposure to RDMA, RoCE, or InfiniBand for low-latency cross-node communication.
• Knowledge of NVMe storage characteristics and GPUDirect Storage (GDS) or equivalent technologies.
• Prior work in a customer-facing or product-integrated inference environment.
• Open-source contributions to SGLang, vLLM, or related LLM infrastructure projects.
AI Inference Engineer ( CUDA or ROCm)
Client: Samsung Cognos |Location: San Jose, CA (Hybrid/Onsite)
|Duration: 12+ Months |Engagement: Contract
MUST
GPU Management
KV Cache Management
About the Engagement
Samsung Cognos is building a next-generation LLM inference layer in partnership with SGLang, one of the leading open-source serving frameworks in the space. The project addresses one of the hardest problems in large-scale AI deployment: making KV cache memory management fast, efficient, and cost-effective across tiered storage hierarchies at production scale. This is a high-impact engineering engagement where your work will directly influence how AI inference performs for thousands of concurrent users.
Role Overview
As an AI Inference Engineer, you will own the serving stack that powers Samsung's LLM inference layer. Your focus will span prefill and decode optimization, KV cache offload strategies, quantization, speculative decoding, and tensor and pipeline parallelism. You will work within SGLang and vLLM-style frameworks, tune ROCm kernels for AMD hardware (MI300 and MI250 series), and evaluate RDMA and RoCE protocols for cross-machine cache sharing. This role is algorithm and framework focused: the core question you are answering is whether the model serving logic itself is efficient, scalable, and production-ready.
Work Model: This position follows a Hybrid/Onsite schedule at the Samsung Cognos facility in San Jose, CA. Candidates must be willing and able to work onsite as required by the client.
Key Responsibilities
• Design and optimize KV cache offload strategies across a tiered memory hierarchy: GPU HBM, CPU DRAM, and NVMe SSD.
• Implement and tune PagedAttention, RadixAttention, chunked prefill, and prefix caching within SGLang and vLLM-style serving frameworks.
• Drive prefill and decode stage optimization to maximize throughput and minimize latency for long-context and multi-user workloads.
• Apply quantization techniques and speculative decoding to reduce memory footprint and improve inference speed without degrading output quality.
• Architect and implement tensor parallelism and pipeline parallelism for multi-GPU inference deployments.
• Tune ROCm kernels for AMD MI300 and MI250 GPUs, adapting CUDA-based patterns to the AMD software stack.
• Evaluate RDMA and RoCE solutions for cross-node KV cache sharing and assess feasibility in production environments.
• Collaborate closely with the GPU Software Engineering team on memory-path correctness, data movement latency, and scheduler integration.
• Profile, benchmark, and iterate on inference stack performance across diverse model architectures and workload profiles.
• Contribute to design reviews, technical documentation, and team knowledge-sharing sessions.
Required Qualifications
• Bachelor's or Master's degree in Computer Science, Electrical Engineering, or a closely related field.
• Hands-on experience with LLM serving frameworks such as SGLang, vLLM, or TensorRT-LLM in a production or research setting.
• Strong understanding of transformer architecture internals, including attention mechanisms and the KV cache.
• Demonstrated experience with GPU programming: CUDA, ROCm/HIP, or equivalent.
• Proficiency in Python and C++ for systems-level performance work.
• Solid grounding in distributed computing concepts including tensor parallelism, pipeline parallelism, and data parallelism.
• Experience with quantization methods (INT8, FP8, GPTQ, AWQ) and their performance tradeoffs.
• Ability to profile and optimize inference pipelines using tools such as NSight, ROCProfiler, or equivalent.
• Strong analytical skills with the ability to translate benchmarking data into actionable engineering decisions.
Preferred Qualifications
• Direct experience with AMD MI300 or MI250 GPU hardware and the ROCm software ecosystem.
• Familiarity with speculative decoding frameworks (Medusa, EAGLE, or similar).
• Exposure to RDMA, RoCE, or InfiniBand for low-latency cross-node communication.
• Knowledge of NVMe storage characteristics and GPUDirect Storage (GDS) or equivalent technologies.
• Prior work in a customer-facing or product-integrated inference environment.
• Open-source contributions to SGLang, vLLM, or related LLM infrastructure projects.






