JOB DESCRIPTION:1. Conversational AI Test Strategy
Develop test strategies for conversation flows, intents, and entities of the chatbot/voicebot.
Design test case sets covering a wide range of conversational scenarios: happy path, edge cases, multi-turn conversations, flow interruptions, and context switching.
Establish evaluation criteria for AI response quality (accuracy, relevance, tone, naturalness).
2. Functional & Response Quality Testing
Test the accuracy of NLU (Natural Language Understanding): intent recognition and entity extraction.
Evaluate the quality of AI responses: correctness, coherence, contextual relevance, and avoidance of repetition or rambling.
Test the system's ability to handle ambiguous cases, out-of-scope questions, and multiple languages (if applicable).
Test the integration between the AI conversation engine and backend systems (API, CRM, database).
3. AI Safety & Ethics Testing
Assess the risk of the AI producing incorrect, biased, or inappropriate (harmful/toxic) content.
Test the AI's ability to appropriately decline sensitive or unsafe requests.
Work with the AI/ML team to report and improve cases of AI hallucination.
4. Automated Testing & Large-Scale Evaluation
Build automated test suites for conversation flows.
Design and operate large-scale response quality evaluation processes (batch evaluation, sampling review).
Collaborate on building a golden dataset/benchmark to measure quality over time.
5. Defect Management & Reporting
Log and classify conversation defects (incorrect intent, incorrect context, inappropriate responses, safety issues).
Collaborate with AI/ML Engineers, Prompt Engineers, and Conversation
Designers to improve quality.
Provide periodic reports on AI system quality (accuracy rate, CSAT, escalation rate to human agents).
6. Process Improvement
Propose conversational AI quality evaluation processes tailored to the product's specific needs.
Train and mentor junior/middle QC staff on AI conversation testing methods.
REQUIREMENTS:
Minimum 5 years of software QC/QA experience, including experience testing chatbots/voicebots/conversational AI systems.
Application-level understanding of how NLU, NLP, and LLMs (Large Language Models) work.
Experience evaluating the output quality of AI models (prompt-response evaluation).
Understanding of concepts such as intent, entity, context, multi-turn conversation, and RAG (Retrieval-Augmented Generation) is a strong plus.
Able to write basic test scripts (Python) to automate API calls and evaluate responses.
Proficient with test case & bug management tools: Jira, TestRail, or equivalent.
Experience using AI evaluation tools (LLM-as-judge, evaluation frameworks) is a plus.
Basic understanding of prompt engineering.
Strong language analysis skills, sensitive to nuance and conversational context.
Ability to ask "adversarial" questions to test the limits of the AI system (basic red-teaming).
Good communication skills, able to collaborate with the AI/ML team, Conversation Designers, and Product.
Careful and highly responsible regarding AI safety and ethics.
Good reading comprehension of technical/AI research documents in English.
Nice to have: Experience with common chatbot platforms (Dialogflow, Rasa, Botpress) or LLM API integration (OpenAI, Anthropic, etc.).
BENEFITS:
Competitive compensation package based on experience and qualifications
Work in a high-speed technology environment backed by Vingroup and VinDynamics leadership.
Competitive compensation package aligned with capability and business impact.
Clear ownership, measurable KPIs, and exposure to global partners, US platform models, and frontier robotics businesses.