Senior Site Reliability Engineer SRE, GCP, Kubernetes

CÔNG TY CỔ PHẦN DATVIET VAC GROUP HOLDINGS

You'll love it
08/10/2026
Toàn thời gian

Mô tả công việc

Top 3 Reasons To Join Us
Build impactful products from the ground up
Tackle high-scale, real-world tech challenges
Grow with a multidisciplinary tech team
The Job

DatVietVAC Group Holdings is looking for a Senior SRE Engineer to lead the infrastructure architecture, platform reliability, and production operations of our Fan Commerce Platform.

The platform operates across Google Cloud for computing and scaling and CMC Cloud for local data residency. You will be responsible for building a secure, scalable, and highly available infrastructure capable of supporting high-traffic on-sale periods and major entertainment events while maintaining operational efficiency and infrastructure costs within the approved budget.

1. Cloud Infrastructure Architecture

  • Design and implement multi-environment infrastructure, including Production, Staging, and Development, using Google Cloud services such as GKE, Cloud SQL, Memorystore, and Pub/Sub.
  • Design the hybrid-cloud architecture between Google Cloud and CMC Cloud, including private connectivity, VPN configuration, network segmentation, and data-flow separation.
  • Build and manage infrastructure as code using Terraform.
  • Develop and maintain CI/CD pipelines and automation tools that support engineering teams throughout the software development and release lifecycle.
  • Establish a comprehensive observability platform covering metrics, logs, traces, dashboards, and alerting.
  • Ensure that the infrastructure architecture supports scalability, maintainability, security, and local data-residency requirements.

2. Site Reliability and Incident Response

  • Define and manage Service Level Indicators, Service Level Objectives, and error budgets for core platform services.
  • Take ownership of platform availability, reliability, scalability, and operational readiness.
  • Lead capacity planning and load testing for high-traffic product launches, ticket or merchandise on-sales, and live events.
  • Design autoscaling strategies, traffic-management mechanisms, and overload-protection measures.
  • Participate in the 24/7 production on-call rotation and lead the response to high-severity incidents.
  • Lead blameless postmortems, identify root causes, and ensure that corrective and preventive actions are completed.
  • Develop incident-response procedures, operational runbooks, and disaster-recovery plans.
  • Coach and mentor engineers in production operations, incident management, and reliability practices.

3. Security and Compliance

  • Implement infrastructure security controls, including WAF, DDoS protection, and bot-management solutions using tools such as Cloudflare and Google Cloud Armor.
  • Manage secrets, encryption, identity, and access controls across cloud environments.
  • Establish backup, recovery, and disaster-recovery processes and conduct periodic recovery drills.
  • Collaborate with relevant teams on penetration testing, vulnerability remediation, security audits, and compliance requirements.
  • Ensure appropriate monitoring and audit logging for sensitive infrastructure and operational activities.

4. Performance and Cost Optimization

  • Monitor, analyze, and optimize cloud infrastructure costs through rightsizing, committed-use discounts, storage lifecycle management, and egress optimization.
  • Investigate performance bottlenecks and implement solutions to improve platform efficiency, scalability, and resilience.
  • Provide infrastructure and reliability recommendations to the Project Lead and engineering teams.
  • Evaluate and propose technologies that support the platform's long-term architecture and business requirements.

Your Skills and Experience

Education

  • Bachelor's degree or higher in Information Technology, Software Engineering, Computer Science, or a related field.

Experience

  • At least five years of experience as a DevOps Engineer, Site Reliability Engineer, System Engineer, Cloud Engineer, or in a similar infrastructure role.
  • At least one year of experience at an equivalent senior level or in a technical leadership role.
  • Proven experience designing or operating production systems with high traffic or significant peak-load events, such as flash sales, on-sales, live events, entertainment platforms, or e-commerce platforms.
  • Hands-on experience operating Kubernetes and cloud infrastructure in a production environment.
  • Production experience with Google Cloud Platform is mandatory.

Technical Knowledge and Skills

  • Advanced knowledge of networking and security concepts, including TCP/IP, HTTP/1.1, HTTP/2, HTTP/3, DNS, gRPC, VPC peering, Cloud VPN, and on-premises-to-cloud connectivity.
  • Strong understanding of distributed systems, microservices, clustering, replication, failover, load balancing, and autoscaling.
  • Strong hands-on experience with Kubernetes, particularly Google Kubernetes Engine, and container technologies in production environments.
  • Proficiency in infrastructure as code and automation tools, particularly Terraform and Ansible.
  • Strong experience building and maintaining CI/CD pipelines using GitHub Actions, Jenkins, or equivalent tools.
  • Strong Linux administration skills, including experience with Ubuntu or CentOS.
  • Hands-on experience with Google Cloud services; additional AWS or Azure experience is an advantage.
  • Experience with performance testing, load testing, bottleneck analysis, capacity planning, and system optimization.
  • Knowledge of API gateways, reverse proxies, ingress controllers, and service mesh architecture.
  • Experience building observability systems using metrics, logs, traces, dashboards, and alerts.
  • Knowledge of SLI/SLO management, error budgets, incident response, postmortems, backup, and disaster recovery.
  • Experience with infrastructure security solutions such as WAF, DDoS protection, bot management, secrets management, and access control.

Soft Skills

  • Strong technical leadership and cross-functional collaboration skills.
  • Ability to mentor engineers and guide teams through incident response and production operations.
  • Calm, structured, and decisive when handling critical production incidents.
  • Strong analytical, troubleshooting, and risk-management skills.
  • Data-driven approach to reliability, performance, capacity, and cost optimization.
  • Strong attention to detail and a high sense of ownership.
  • Ability to explain technical decisions and trade-offs clearly to both technical and non-technical stakeholders.

Language Skills

  • Basic English communication skills.
  • Good ability to read and understand technical documentation in English.

Other Requirements

  • Willingness to participate in the production on-call rotation and support major on-sale periods or live events when required.
  • Ability to work under pressure during critical releases, peak-traffic events, and production incidents.

Why You'll Love Working Here

• Full statutory insurance, including Social Insurance, Health Insurance and Unemployment Insurance, based on 100% of the official salary and in compliance with Vietnamese labor regulations.
• Working hours: Monday to Friday, from 8:30 AM to 5:30 PM, with a one-hour lunch break.
• 14 days of annual leave.
• Company-provided working equipment, including a laptop or desktop computer.
• Employee parking area.
• Annual PMP performance bonus, subject to individual KPI achievement and the Company's business performance.
• Periodic health check-ups.
• Employee engagement programs and internal activities throughout the year.

Yêu cầu

Kubernetes, Terraform, CI/CD, System Architecture, DevOps, GCP

Quyền lợi

• Full statutory insurance, including Social Insurance, Health Insurance and Unemployment Insurance, based on 100% of the official salary and in compliance with Vietnamese labor regulations.
• Working hours: Monday to Friday, from 8:30 AM to 5:30 PM, with a one-hour lunch break.
• 14 days of annual leave.
• Company-provided working equipment, including a laptop or desktop computer.
• Employee parking area.
• Annual PMP performance bonus, subject to individual KPI achievement and the Company's business performance.
• Periodic health check-ups.
• Employee engagement programs and internal activities throughout the year.

Thông tin chung

  • Thu nhập: You'll love it

Nơi làm việc

  • 222 Pasteur, Phường Xuân Hòa, Ho Chi Minh

Việc làm tương tự

Trưởng phòng Kinh doanh Bất động sản

CÔNG TY CỔ PHẦN VINHOMES - TẬP ĐOÀN VINGROUP

25-100 triệu VND
Hồ Chí Minh
07/10/2026

NHÂN VIÊN THIẾT KẾ NGÀNH CỬA NHÔM KÍNH

CÔNG TY TNHH CỬA PHÚ MỸ HƯNG

15.000.000 - 20.000.000 VND
Hồ Chí Minh
08/10/2026

Kỹ Sư Phụ Trách Hồ Sơ Thanh Quyết Toán Công Trình - Đi Làm Ngay Tại TP. Hồ Chí Minh

Công ty TNHH đầu tư và xây lắp FSC

15 - 25 triệu VND + Phụ Cấp
Hồ Chí Minh
01/10/2026

Pricing Specialist/ Nhân viên Làm giá

CÔNG TY TNHH TIẾP VẬN VẬN TẢI QUỐC TẾ VÕ LƯƠNG

15 - 20 triệu VND
Hồ Chí Minh
08/10/2026

Nhân viên/Chuyên viên Thu hồi nợ tại nhà - Hồ Chí Minh

Ngân hàng TMCP Việt Nam Thịnh Vượng - VPBank

20 - 40 triệu VNĐ
Hồ Chí Minh
19/09/2026

Kỹ Sư Shopdrawing - Đi Làm Ngay Tại TP. Hồ Chí Minh

Công ty TNHH đầu tư và xây lắp FSC

15 - 25 triệu VND + Phụ Cấp
Hồ Chí Minh
01/10/2026

Senior Structural Design Engineer - PT Slab Design | HCMC

CÔNG TY TNHH TƯ VẤN KỸ THUẬT APEX SOUTHERN CROSS

Negotiable up to 40M
Hồ Chí Minh
16/09/2026

Kỹ sư sản xuất MFG/DFM Engineer - Remote - 25M

Cennext Co., Ltd

Từ 15 - 25 triệu VND
Hà Nội, Hồ Chí Minh, Đà Nẵng, Hải Phòng, Bà Rịa - Vũng Tàu
08/10/2026

NHÂN VIÊN CHỨNG TỪ HÀNG NHẬP - HỢP ĐỒNG 8 THÁNG

CÔNG TY TNHH TOSHIBA LOGISTICS VIỆT NAM

10-12 triệu VND
Hồ Chí Minh
01/10/2026
Vị trí Senior Site Reliability Engineer SRE, GCP, Kubernetes do công ty CÔNG TY CỔ PHẦN DATVIET VAC GROUP HOLDINGS tuyển dụng tại Hồ Chí Minh, Joboko tự động tổng hợp mức lương You'll love it, tìm thêm việc làm về Senior Site Reliability Engineer SRE, GCP, Kubernetes hoặc công ty CÔNG TY CỔ PHẦN DATVIET VAC GROUP HOLDINGS ở các link phía trên

Giới thiệu công ty

CÔNG TY CỔ PHẦN DATVIET VAC GROUP HOLDINGS

Địa chỉ: 222 Pasteur - Phường Võ Thị Sáu - Quận 3 - TP. Hồ Chí Minh
Quy mô: Từ 101 - 500 nhân viên

Việc làm HOT

CÔNG TY CỔ PHẦN TẬP ĐOÀN FECON
Thoả thuận
Hồ Chí Minh
Công Ty TNHH Hình Tượng Ô Tô Sài Gòn - Motor Image SG
Upto 50 triệu VND++
Hà Nội, Hồ Chí Minh
Aeon Fantasy Vietnam Co.,ltd.
9 - 10 triệu VND
Hồ Chí Minh, Hải Phòng, Hải Dương, Hưng Yên, Long An, Tây Ninh
Công ty Cổ phần Eurowindow
Thỏa thuận
Hà Nội
CÔNG TY TNHH J-LONG VIỆT NAM
20 - 25 triệu VND
Hồ Chí Minh