AI Model Optimization Manager
Hạn nộp hồ sơ: 05/10/2026 (Còn 20 ngày)
Ứng tuyển sớm để được ưu tiên
Kết nối với Nhà tuyển dụng để tìm hiểu thông tin và gia tăng cơ hội trúng tuyển
Nhà tuyển dụng đang online
About VinSmart Future
VinSmart Future (VSF) is the leading technology company within the Vingroup Corporation, formed by the merger of the group's entire technology ecosystem, including VinApp, VinIT, VinBigdata, and other tech units. As a core driver of Vingroup's future growth, VSF is at the forefront of technological development, with artificial intelligence (AI) as its foundation. With a talented team of nearly 4,000 local and international technology experts, VSF focuses on creating high-utility technologies that enhance lives and connect data, models, and infrastructure to unlock new possibilities.
AI Model Optimization Manager
● Work Location: HCMC: Vincom Dong Khoi, District 1.
Key Responsibilities:
Fine-tuning & Model Customization
Perform fine-tuning of open-source foundation models (e.g., Llama, Mistral, Qwen) using techniques such as Supervised Fine-Tuning (SFT), LoRA, and QLoRA to optimize model performance for enterprise-specific use cases.
Customize and adapt Large Language Models (LLMs) to meet business requirements, domain knowledge, and operational constraints.
Model Optimization
Apply model compression and optimization techniques, including Quantization (AWQ, GPTQ, GGUF), Pruning, and Draft Model architectures (Speculative Decoding), to reduce inference latency, improve throughput, and lower operational costs.
Optimize model deployment performance across various hardware environments. Alignment & Reinforcement Learning Pipelines
Design, implement, and manage advanced alignment and reinforcement learning pipelines, including RLHF (Reinforcement Learning from Human Feedback), RLAIF (Reinforcement Learning from AI Feedback), DPO (Direct Preference Optimization), and PPO-based training workflows.
Ensure model behavior aligns with business objectives, safety requirements, and user expectations.
Large-Scale Frameworks & Infrastructure
Work extensively with high-performance training and inference frameworks such as NVIDIA NeMo, vLLM, TensorRT-LLM, DeepSpeed, or equivalent technologies.
Leverage distributed training and inference techniques to improve scalability, efficiency, and resource utilization.
LLM Evaluation & Validation
Design, develop, and maintain comprehensive LLM evaluation pipelines to assess model quality before production deployment.
Utilize industry-standard benchmarks (e.g., MMLU, GSM8K) and advanced evaluation methodologies such as Ragas and LLM-as-a-Judge frameworks.
Establish evaluation metrics and processes to measure model performance, reliability, safety, and business effectiveness in real-world environments.
Programming Languages
Python
Rust
C++
Preferred Qualifications
Hands-on Industry Experience
Minimum 4 years of hands-on experience in NLP and Generative AI, with proven expertise in training, fine-tuning, and aligning Large Language Models (LLMs).
Strong practical experience in developing and deploying LLM-based solutions for real-world applications.
Tools & Framework Expertise
Deep expertise in the Hugging Face ecosystem, NVIDIA NeMo, vLLM/SGLang, PyTorch, and distributed training/resource management tools.
Proven ability to manage large-scale model training and inference workloads efficiently.
Optimization Mindset
Strong understanding of Transformer architectures and modern LLM internals.
Deep knowledge of GPU architecture, CUDA, and memory optimization techniques.
Ability to identify and resolve training and deployment bottlenecks related to computational resources and system performance.
Evaluation & Data Engineering Skills
Solid Data Engineering capabilities for dataset collection, cleaning, preprocessing, and preparation for fine-tuning workflows.
Strong understanding of production-grade LLM evaluation metrics and methodologies, including hallucination detection, accuracy measurement, toxicity assessment, and model safety evaluation.
Experience building scalable evaluation and monitoring systems for production environments.
Nice-to-Have Qualifications
Experience with low-level hardware optimization, including custom CUDA kernel development.
Proven track record of deploying LLM-powered products in production environments serving large-scale user bases.
Experience with large-scale AI infrastructure, model serving platforms, and performance engineering for enterprise AI systems.
Why You'll Love Working Here:
Flexible working hours and attendance policy (Work from Home on working Saturdays).
Attractive compensation and bonus packages, highly competitive in the market.
Exclusive employee benefits across the Group's ecosystem in accordance with company policies.
Opportunity to work on large-scale and strategic technology projects.
Professional technology environment with leading scientists, experts, and engineers from top technology companies in Vietnam and around the world.
Free access to learning platforms such as Udemy, Coursera, and O'Reilly; internal workshops; sponsorship for professional certifications; and exclusive mentoring programs from the Group and Company leadership team.
Full statutory insurance coverage in accordance with Vietnamese Labor Law (Social
Insurance, Health Insurance, Unemployment Insurance), along with private healthcare insurance based on job grade and annual health check-ups at reputable hospitals and healthcare centers nationwide.
Participation in internal activities, team-building programs, and annual company events.
VinSmart Future (VSF) is the leading technology company within the Vingroup Corporation, formed by the merger of the group's entire technology ecosystem, including VinApp, VinIT, VinBigdata, and other tech units. As a core driver of Vingroup's future growth, VSF is at the forefront of technological development, with artificial intelligence (AI) as its foundation. With a talented team of nearly 4,000 local and international technology experts, VSF focuses on creating high-utility technologies that enhance lives and connect data, models, and infrastructure to unlock new possibilities.
AI Model Optimization Manager
● Work Location: HCMC: Vincom Dong Khoi, District 1.
Key Responsibilities:
Fine-tuning & Model Customization
Perform fine-tuning of open-source foundation models (e.g., Llama, Mistral, Qwen) using techniques such as Supervised Fine-Tuning (SFT), LoRA, and QLoRA to optimize model performance for enterprise-specific use cases.
Customize and adapt Large Language Models (LLMs) to meet business requirements, domain knowledge, and operational constraints.
Model Optimization
Apply model compression and optimization techniques, including Quantization (AWQ, GPTQ, GGUF), Pruning, and Draft Model architectures (Speculative Decoding), to reduce inference latency, improve throughput, and lower operational costs.
Optimize model deployment performance across various hardware environments. Alignment & Reinforcement Learning Pipelines
Design, implement, and manage advanced alignment and reinforcement learning pipelines, including RLHF (Reinforcement Learning from Human Feedback), RLAIF (Reinforcement Learning from AI Feedback), DPO (Direct Preference Optimization), and PPO-based training workflows.
Ensure model behavior aligns with business objectives, safety requirements, and user expectations.
Large-Scale Frameworks & Infrastructure
Work extensively with high-performance training and inference frameworks such as NVIDIA NeMo, vLLM, TensorRT-LLM, DeepSpeed, or equivalent technologies.
Leverage distributed training and inference techniques to improve scalability, efficiency, and resource utilization.
LLM Evaluation & Validation
Design, develop, and maintain comprehensive LLM evaluation pipelines to assess model quality before production deployment.
Utilize industry-standard benchmarks (e.g., MMLU, GSM8K) and advanced evaluation methodologies such as Ragas and LLM-as-a-Judge frameworks.
Establish evaluation metrics and processes to measure model performance, reliability, safety, and business effectiveness in real-world environments.
Programming Languages
Python
Rust
C++
Preferred Qualifications
Hands-on Industry Experience
Minimum 4 years of hands-on experience in NLP and Generative AI, with proven expertise in training, fine-tuning, and aligning Large Language Models (LLMs).
Strong practical experience in developing and deploying LLM-based solutions for real-world applications.
Tools & Framework Expertise
Deep expertise in the Hugging Face ecosystem, NVIDIA NeMo, vLLM/SGLang, PyTorch, and distributed training/resource management tools.
Proven ability to manage large-scale model training and inference workloads efficiently.
Optimization Mindset
Strong understanding of Transformer architectures and modern LLM internals.
Deep knowledge of GPU architecture, CUDA, and memory optimization techniques.
Ability to identify and resolve training and deployment bottlenecks related to computational resources and system performance.
Evaluation & Data Engineering Skills
Solid Data Engineering capabilities for dataset collection, cleaning, preprocessing, and preparation for fine-tuning workflows.
Strong understanding of production-grade LLM evaluation metrics and methodologies, including hallucination detection, accuracy measurement, toxicity assessment, and model safety evaluation.
Experience building scalable evaluation and monitoring systems for production environments.
Nice-to-Have Qualifications
Experience with low-level hardware optimization, including custom CUDA kernel development.
Proven track record of deploying LLM-powered products in production environments serving large-scale user bases.
Experience with large-scale AI infrastructure, model serving platforms, and performance engineering for enterprise AI systems.
Why You'll Love Working Here:
Flexible working hours and attendance policy (Work from Home on working Saturdays).
Attractive compensation and bonus packages, highly competitive in the market.
Exclusive employee benefits across the Group's ecosystem in accordance with company policies.
Opportunity to work on large-scale and strategic technology projects.
Professional technology environment with leading scientists, experts, and engineers from top technology companies in Vietnam and around the world.
Free access to learning platforms such as Udemy, Coursera, and O'Reilly; internal workshops; sponsorship for professional certifications; and exclusive mentoring programs from the Group and Company leadership team.
Full statutory insurance coverage in accordance with Vietnamese Labor Law (Social
Insurance, Health Insurance, Unemployment Insurance), along with private healthcare insurance based on job grade and annual health check-ups at reputable hospitals and healthcare centers nationwide.
Participation in internal activities, team-building programs, and annual company events.
Thông tin chung
- Thu nhập: Thỏa thuận
Việc làm tương tự khác
Công ty Cổ phần Vàng bạc Đá quý Phú Nhuận - PNJ
Hồ Chí Minh
Thương lượng
Công Ty Cổ Phần Giải Pháp Chuỗi Cung Ứng Smartlog
Hồ Chí Minh
Thỏa thuận
Vega Corporation
Hà Nội, Hồ Chí Minh, An Giang
Thỏa thuận
BETRIMEX - Công Ty Cổ Phần Xuất Nhập Khẩu Bến Tre
Hồ Chí Minh
Cạnh tranh
Công ty Cổ phần Vàng bạc Đá quý Phú Nhuận - PNJ
Hồ Chí Minh
Thương lượng
Thông tin công việc
Vị trí:
Nhân viên
Hình thức làm việc:
Toàn thời gian
Việc làm tương tự
Tuyển dụng NHÂN VIÊN LẬP TRÌNH / IT tại TPHCM - Hiệp Lực Phát Triển Việt
CÔNG TY TNHH HIỆP LỰC PHÁT TRIỂN VIỆT
Hồ Chí Minh, Bà Rịa - Vũng Tàu
8,000,000 - 20,000,000 VNĐ
Cảnh báo dấu hiệu lừa đảo tuyển dụng
Đội ngũ hỗ trợ của JobOKO sẵn sàng đồng hành, tư vấn và giới thiệu những cơ hội việc làm phù hợp, giúp Ứng viên tự tin phát triển sự nghiệp và chinh phục mục tiêu nghề nghiệp bền vững.
Hotline CSKH
1900.63.63.84
Công ty Cổ phần JobOKO Toàn cầu
Đội ngũ hỗ trợ của JobOKO luôn chủ động tư vấn các giải pháp tuyển dụng tối ưu, cam kết đồng hành và hỗ trợ Quý Nhà tuyển dụng đạt được hiệu quả tuyển dụng bền vững.
Hotline CSKH
0962.107.888
Công ty Cổ phần JobOKO Toàn cầu