Full-time · Hanoi
Job Type: Full-time, Onsite
Location: H10 Building, 475 Nguyen Trai, Thanh Xuan, Hanoi
Working Hours: Monday - Friday (08:00 AM - 05:30 PM)
Website: [protected info]
ABOUT WIGIN AI
Wigin is a product-led AI company powered by a team of leading engineers across the artificial intelligence and technology sectors. The company specializes in building production-grade solutions that transform frontier AI research into reliable, scalable systems engineered to deliver real, measurable value for businesses. Rather than just creating theoretical models, Wigin focuses on engineering practical products that win at scale and optimize how businesses actually run.
The company operates across three core domains:
(1) AI Services
(2) AI Investment Intelligence
(3) AI Products
ABOUT THE ROLE:
We are looking for a Full-time AI Engineer to join our team and work on projects involving Large Language Models (LLMs) and Generative AI.
You will work closely with a technical team to design, fine-tune, evaluate, optimize, and deploy advanced language models for real-world applications. This is a great opportunity for candidates who want to work deeply with frontier LLMs, Retrieval-Augmented Generation (RAG) architecture, AI Agent systems, and large-scale GPU infrastructure
WHAT YOU'LL DO:
Research, develop, and apply state-of-the-art Large Language Models (LLMs) and Generative Text/Multimodal AI models.
Build, optimize, and evaluate advanced LLM applications, including RAG (Retrieval-Augmented Generation), AI Agents, and workflow automation systems.
Design, pre-train, fine-tune (SFT, LoRA/QLoRA), and align (RLHF/DPO) open-source LLMs for domain-specific tasks.
Optimize model latency, throughput, and inference cost using modern frameworks and techniques.
Read research papers, experiment with new methods (Prompt Engineering, Context Extension, Function Calling), and apply them to practical AI products.
Collaborate with backend and product engineering teams to integrate LLM pipelines into scalable production systems.
Monitor model safety, hallucinations, performance, and continuously improve model quality.
WHAT WE'RE LOOKING FOR:
Strong foundation in Machine Learning, Deep Learning, and Natural Language Processing (NLP).
Solid understanding of Transformer architectures, Attention mechanisms, and modern LLM paradigms.
Hands-on experience working with LLMs (commercial APIs like OpenAI/Anthropic or open-source models like Llama/Qwen).
Practical experience in building RAG architectures, Vector Databases (e.g., Milvus, Qdrant, Pinecone), or AI Agent frameworks (e.g., LangChain, LlamaIndex, AutoGen/CrewAI).
Ability to read and implement research papers and technical documentation effectively.
Strong problem-solving skills, clean coding practices, and a proactive engineering mindset.
Ability to work independently and collaborate effectively with a multi-disciplinary team.
Nice to have:
Experience in fine-tuning LLMs (LoRA, QLoRA, DeepSpeed, Unsloth, PEFT).
Experience with LLM inference optimization engines (e.g., vLLM, TensorRT-LLM, Ollama, TGI) and techniques (Quantization, Speculative Decoding).
THE BENEFITS & PERKS
1. Competitive Compensation & Comprehensive Benefits
Competitive salary package, up to 40M VND/month
Attractive project bonus for outstanding contributions and successful project delivery.
13th-month salary and full benefits in accordance with Vietnamese labor law
Annual leave, social insurance, health insurance and other statutory benefits
Regular salary review twice a year.
2. Flexible & Modern Working
Monday-Friday, with flexible check-in time from [protected info] WFH day/week
Focus on results, ownership, and sustainable work-life balance
3. Work on Challenging AI Projects
Work on diverse, challenging, and high-impact AI projects
Directly tackle real-world problems across LLM, Generative AI, AI Agents, RAG and AI infrastructure
Work with large-scale GPU infrastructure and modern AI frameworks
Freedom to experiment, research, and turn ideas into production-ready products
4. High Ownership & Direct Impact
Join a small, highly technical team with a high level of autonomy
Work directly with the CEO and core technical team
Take ownership of meaningful problems and see your work go from idea to production
Fast decision-making, minimal bureaucracy, and plenty of room to make an impact
5. Premium AI Tools & Resources
Unlimited access to Claude Max and Codex
Access to large-scale GPU infrastructure for research, experimentation, and model development
Premium technical resources and the latest AI tools to accelerate your work
To apply, send your resume and a short note about what you've built to [protected info] or Contact [protected info] (Ms.Thu Trang)
Thông tin chung