Mô tả công việc
We are looking for Senior MLOps & Infrastructure Engineers to build and operate the hybrid AI computing platform behind VinFast's ADAS and autonomous driving programme - on-premise GPU clusters running perception model training and large-scale inference, a multi-petabyte sensor data archive, and the platform services used daily by our engineering and annotation teams.
This is a fullstack infrastructure role. You will work across compute, storage, networking, automation and security. We do not divide the team into narrow specialists: everyone owns a part of the platform end to end - provisioning, deployment, monitoring and incident response - and is expected to grow across the whole stack over time. Work is assigned according to what the programme needs at that moment, not by a fixed job boundary.
Much of the platform runs on-premise, including air-gapped segments. Experience operating without a public-cloud safety net is valued here
Key Responsibilities
1. AI Compute & Serving
- Operate and scale GPU clusters across training, inference and annotation workloads; manage scheduling, utilisation and capacity
- Deploy and scale high-throughput model serving; tune GPU memory and runtime performance
- Manage multi-GPU distributed training and reproducible environment sandboxing
- Scale pipeline orchestration on Kubernetes for large data processing jobs
2. Platform Automation & Delivery
- Drive GitOps-based deployment and maintain infrastructure-as-code across the platform
- Build and maintain CI pipelines, secure container builds and release automation
- Build monitoring, logging and alerting so that failures are detected and actionable, never silent
- Lead incident response and drive follow-up actions to closure
3. Data Infrastructure & Security
- Operate storage at multi-petabyte scale: object storage, local storage, tiering, backup and recovery
- Operate the ingestion path for large sensor data deliveries, with integrity verification and metadata extraction
- Implement access control and single sign-on across platform services, with audit logging
- Apply data protection measures to sensitive content before it reaches external users
Yêu cầu
- 4+ years in MLOps, DevOps, SRE or HPC platform engineering, with production ownership of AI/ML infrastructure
- Kubernetes administration at production scale: Helm, ingress (Traefik or Envoy), CNI, and GitOps (Flux or ArgoCD)
- GPU & HPC: SLURM, NVIDIA Container Toolkit, CUDA runtime tuning, multi-GPU memory debugging
- Strong Linux systems skills and infrastructure-as-code (SaltStack or Ansible)
- Python and Bash for automation, including Airflow DAGs and custom operators
- Object storage, and a monitoring and logging stack (Prometheus, Grafana or equivalent)
- Willingness to work across the full stack - compute, storage, networking, automation and security - rather than within a single specialty
- Good communication in English - technical documentation and working with international partners
Quyền lợi
Thưởng
• Competitive income (including 13th-month salary, performance bonuses, and other rewards as regulated by Vingroup).
Chăm sóc sức khoẻ
• High-quality health insurance
Khác
• Working in a safe, modern, civilized, and professional environment with numerous opportunities for personal development, and even take the lead.
Thông tin khác
NGÀY ĐĂNG
23/09/2026
CẤP BẬC
Nhân viên
NGÀNH NGHỀ
Công Nghệ Thông Tin/Viễn Thông > System/Cloud/DevOps Engineer
KỸ NĂNG
Infrastructure-As-Code, DevOps, Kubernetes, SRE, MLOps
LĨNH VỰC
Ô tô
NGÔN NGỮ TRÌNH BÀY HỒ SƠ
Bất kỳ
SỐ NĂM KINH NGHIỆM TỐI THIỂU
4
QUỐC TỊCH
Không hiển thị
Xem thêm
Thông tin chung
Nơi làm việc