Ass. Manager, Site Reliability Engineer 5yoe, English
Hạn nộp hồ sơ: 17/09/2026 (Còn 33 ngày)
Ứng tuyển sớm để được ưu tiên
Kết nối với Nhà tuyển dụng để tìm hiểu thông tin và gia tăng cơ hội trúng tuyển
Nhà tuyển dụng đang online
Mô tả công việc
Tóm tắt công việc
Team Leadership and Development
Lead and develop a team of Site Reliability Engineers (Level 6-7), initially 3 direct reports and growing as the Vietnam site consolidates, owning performance, coaching, retention, and day-to-day execution
Build individual development plans that grow Level 6 engineers toward independent Level 7 scope
Establish a high-ownership culture where engineers are accountable for outcomes, not tasks
Run team rituals: 1:1s, shift retrospectives, and development check-ins
Follow-the-Sun Operations
Own the Vietnam shift within GRE's global follow-the-sun coverage model, including schedule design, coverage planning, and holiday/leave management
Ensure clean, structured handoffs to and from US and India teams, with clear ownership transfer on open incidents and in-flight work
Maintain shift readiness: runbooks current, alerts actionable, escalation paths clear
Serve as escalation point for the Vietnam shift during complex or high-severity incidents
Reliability Practice Execution
Own production reliability outcomes for the markets, platforms, and services within the Vietnam team's scope
Drive SRE operating standards within the team: incident response rigor, SLO ownership, runbook quality, and post-incident follow-through
Ensure monitoring coverage, dashboards, and alerting remain accurate and effective across owned services
Enforce GRE-wide reliability standards, ensuring the Vietnam practice operates in alignment with the global model
Platform Engineering, Automation, and AI
Own the Vietnam team's contribution to GRE's reliability modernization roadmap: auto-healing, auto-remediation, and self-service issue mitigation
Treat automation as a delivery commitment, not a byproduct: plan, prioritize, and track toil-reduction and self-service work alongside operational coverage
Ensure the team builds with platform engineering practices: infrastructure as code, runbooks as code, GitOps workflows, and reusable tooling over one-off fixes
Champion AI-first engineering practices within the team, ensuring engineers develop with and through modern AI tooling
Partner with GRE's Foundations and Intelligence pillars on observability, signal detection, and automated response infrastructure
Stakeholder Engagement
Represent the Vietnam team in East GRE planning and reliability forums
Communicate reliability outcomes, coverage status, and team health clearly to GRE leadership
Partner with the Associate Director on headcount planning, team scope, and Vietnam-specific delivery
Attractive Benefits:
100% salary during probation period
Annual Leave: 18 days/ year
2 days WFH/ week
Five "Recharge Days" - Extra days, in addition to company holidays.
Flexible Friday afternoon
Full salary insurance
13th-month bonus
1 day off for birthday
Advanced health insurance (Generali)
Regular engagement activities: sport clubs, internal event...
Support Macbook and Monitor
Team Leadership and Development
Lead and develop a team of Site Reliability Engineers (Level 6-7), initially 3 direct reports and growing as the Vietnam site consolidates, owning performance, coaching, retention, and day-to-day execution
Build individual development plans that grow Level 6 engineers toward independent Level 7 scope
Establish a high-ownership culture where engineers are accountable for outcomes, not tasks
Run team rituals: 1:1s, shift retrospectives, and development check-ins
Follow-the-Sun Operations
Own the Vietnam shift within GRE's global follow-the-sun coverage model, including schedule design, coverage planning, and holiday/leave management
Ensure clean, structured handoffs to and from US and India teams, with clear ownership transfer on open incidents and in-flight work
Maintain shift readiness: runbooks current, alerts actionable, escalation paths clear
Serve as escalation point for the Vietnam shift during complex or high-severity incidents
Reliability Practice Execution
Own production reliability outcomes for the markets, platforms, and services within the Vietnam team's scope
Drive SRE operating standards within the team: incident response rigor, SLO ownership, runbook quality, and post-incident follow-through
Ensure monitoring coverage, dashboards, and alerting remain accurate and effective across owned services
Enforce GRE-wide reliability standards, ensuring the Vietnam practice operates in alignment with the global model
Platform Engineering, Automation, and AI
Own the Vietnam team's contribution to GRE's reliability modernization roadmap: auto-healing, auto-remediation, and self-service issue mitigation
Treat automation as a delivery commitment, not a byproduct: plan, prioritize, and track toil-reduction and self-service work alongside operational coverage
Ensure the team builds with platform engineering practices: infrastructure as code, runbooks as code, GitOps workflows, and reusable tooling over one-off fixes
Champion AI-first engineering practices within the team, ensuring engineers develop with and through modern AI tooling
Partner with GRE's Foundations and Intelligence pillars on observability, signal detection, and automated response infrastructure
Stakeholder Engagement
Represent the Vietnam team in East GRE planning and reliability forums
Communicate reliability outcomes, coverage status, and team health clearly to GRE leadership
Partner with the Associate Director on headcount planning, team scope, and Vietnam-specific delivery
Attractive Benefits:
100% salary during probation period
Annual Leave: 18 days/ year
2 days WFH/ week
Five "Recharge Days" - Extra days, in addition to company holidays.
Flexible Friday afternoon
Full salary insurance
13th-month bonus
1 day off for birthday
Advanced health insurance (Generali)
Regular engagement activities: sport clubs, internal event...
Support Macbook and Monitor
Yêu cầu
5+ years in site reliability engineering, DevOps, infrastructure, or production operations roles
1+ years of people management experience, or 2+ years as a senior technical lead with demonstrated coaching and delivery ownership
Hands-on credibility across incident response, observability, and automation, with the technical depth to guide Level 6-7 engineers
Experience operating in shift-based, on-call, or follow-the-sun coverage models
Working knowledge of at least one major cloud provider (AWS preferred) and modern observability tooling (e.g., Datadog, Prometheus, Grafana)
Proficiency in at least one scripting or programming language sufficient to review and guide automation work
Understanding of SLI/SLO frameworks and reliability engineering fundamentals
Strong written and verbal English communication skills for cross-region collaboration with US and India teams
Experience building or standing up a new team, site, or shift operation
Experience managing engineers across the early-to-mid career range with a track record of promotions or level progression
Kubernetes, container orchestration, and infrastructure as code experience (e.g., Terraform)
Familiarity with AI-assisted operations tooling and automation-first reliability approaches, including auto-healing and auto-remediation patterns
Exposure to platform engineering and internal developer platform concepts: self-service tooling, developer portals (e.g., Port, Backstage), GitOps
Experience in multi-region or globally distributed team models
Relevant certifications (AWS, CKA, or similar)
1+ years of people management experience, or 2+ years as a senior technical lead with demonstrated coaching and delivery ownership
Hands-on credibility across incident response, observability, and automation, with the technical depth to guide Level 6-7 engineers
Experience operating in shift-based, on-call, or follow-the-sun coverage models
Working knowledge of at least one major cloud provider (AWS preferred) and modern observability tooling (e.g., Datadog, Prometheus, Grafana)
Proficiency in at least one scripting or programming language sufficient to review and guide automation work
Understanding of SLI/SLO frameworks and reliability engineering fundamentals
Strong written and verbal English communication skills for cross-region collaboration with US and India teams
Experience building or standing up a new team, site, or shift operation
Experience managing engineers across the early-to-mid career range with a track record of promotions or level progression
Kubernetes, container orchestration, and infrastructure as code experience (e.g., Terraform)
Familiarity with AI-assisted operations tooling and automation-first reliability approaches, including auto-healing and auto-remediation patterns
Exposure to platform engineering and internal developer platform concepts: self-service tooling, developer portals (e.g., Port, Backstage), GitOps
Experience in multi-region or globally distributed team models
Relevant certifications (AWS, CKA, or similar)
Thông tin khác
DevOps
DataDog
Grafana
Observability
AWS
Kubernetes
Terraform
Prometheus
IaC
GitOps
CKA
DataDog
Grafana
Observability
AWS
Kubernetes
Terraform
Prometheus
IaC
GitOps
CKA
Thông tin chung
- Thu nhập: Thỏa Thuận
Nơi làm việc
- Waseco Building - 10 Pho Quang Street, Ward 02, Tan Son Hoa, Ho Chi Minh
Việc làm tương tự khác
Công ty Cổ phần Đầu tư & Phát triển Năng Lượng Mặt Trời Bách Khoa SolarBK
Hồ Chí Minh
Thương lượng
Deliveree On-demand Logistics (southeast Asia) - CÔNG TY CỔ PHẦN DELIVEREE VIỆT NAM
Hồ Chí Minh, Sóc Trăng
800 - 1000
Pizza Hut Digital & Technology
Xem trang công ty- Địa chỉ công ty: Tan Binh, Ho Chi Minh
- Quy mô: Từ 101 - 500 nhân viên
Thông tin công việc
Vị trí:
Quản lý
Hình thức làm việc:
Toàn thời gian
Việc làm tương tự
Cảnh báo dấu hiệu lừa đảo tuyển dụng
Đội ngũ hỗ trợ của JobOKO sẵn sàng đồng hành, tư vấn và giới thiệu những cơ hội việc làm phù hợp, giúp Ứng viên tự tin phát triển sự nghiệp và chinh phục mục tiêu nghề nghiệp bền vững.
Hotline CSKH
1900.63.63.84
Công ty Cổ phần JobOKO Toàn cầu
Đội ngũ hỗ trợ của JobOKO luôn chủ động tư vấn các giải pháp tuyển dụng tối ưu, cam kết đồng hành và hỗ trợ Quý Nhà tuyển dụng đạt được hiệu quả tuyển dụng bền vững.
Hotline CSKH
0962.107.888
Công ty Cổ phần JobOKO Toàn cầu