Tóm tắt công việc
We're hiring an Agent Ops Engineer to scale AI agent capabilities across HRS Domain and products. This is a high-impact role at the intersection of
AI engineering, platform operations, and knowledge enablement. You'll provide directions and build AI agents reliable in production across teams by owning the lifecycle, quality gates, observability, and operational standards-while embedding with teams to accelerate adoption. The larger goal of this centralized Agent Ops model is to enable Ai enablers and product builders within each product team for agent development and at the same time contribution common best practices, guard rails, to MFBS adoption across other domains like ERP and SMB.
What you will do:
1) Agent Engineering & operation
Design, build, and maintain production-grade AI agent systems, including: context engineering and instruction architecture, prompt hardening and safe execution boundaries, tool integrations and multi-step orchestration, memory strategies and reliability patterns.
Own the full agent lifecycle: prototype → evaluate → deploy → monitor → iterate.
Build and maintain an evaluation pipeline to measure agent quality, catch regressions, and enforce deployment gates (golden datasets, scenario suites, automated checks).
Instrument agents and agent platforms for production observability: structured logging, tracing, and metrics; latency and cost monitoring; tool-call success rates and failure analysis.
Define operational readiness standards including: rollback criteria, incident response playbooks, recovery paths for common failure modes.
2) Team Enablement & Coaching
Embed with product engineering teams to identify high-value use cases ready for agent automation. We will be operating in a Central Agent Ops role enabling Ai product builders through AI enablers.
Translate business workflows into agent-executable tasks with clear: contact boundaries/interfaces, assumptions and inputs/outputs, failure modes and safe fallbacks.
Deliver targeted coaching to engineers on: context engineering best practices, harness design and regression testing patterns, agent skill design and tool-contract discipline.
Reduce onboarding time for teams adopting AI capabilities-from first conversation to a production-ready agent.
Train product engineers to extend and maintain agent skills independently.
3) Standards & Knowledge operations
Author and maintain org-level standards for agents, including: naming conventions, context file structures and ownership rules, skill interface contracts (inputs/outputs, invariants, error handling), evaluation criteria and release quality bars.
Establish and enforce "repo-as-discipline" practices so agent knowledge is: versioned, reviewable, discoverable, reusable; not trapped in prompt snippets or individual heads.
Build and grow a shared agent skills library that teams can reuse and extend.
Track and aggregate AI tooling/framework updates and external best practices, serving as a central intake so product teams don't each have to follow the entire AI landscape.
Run internal knowledge-sharing sessions, showcases, and retrospectives to propagate learnings efficiently.
Caring Mental & Physical Recreation:
Hybrid working
Full salary in probation & 13th month salary
Social insurance on full salary from probation
Premium Health insurance from probation
Flexible start 8AM-9AM from Mon-Fri
16 days off annually + 1 Birthday Leave
Paternity leave extra 5 days
Annual company trip; Quarterly team building activities
Club activities
Annual health check
Caring Career & Development:
Clear Career path
Foreign language & International technology-related certifications sponsoring
Well-equipped facility: Macbook pro, additional monitor,..
Soft skill workshops
Tech seminars
Monthly and biannually Recognition Awards
Performance review twice/year