Senior Software Engineer (Performance)

Thỏa thuận
5 năm kinh nghiệm
Hạn nộp hồ sơ: 20/10/2026 (Còn 54 ngày)
Ứng tuyển sớm để được ưu tiên
Kết nối với Nhà tuyển dụng để tìm hiểu thông tin và gia tăng cơ hội trúng tuyển
OMess
Nhà tuyển dụng đang online
Job Description:
We are looking for a
Senior Inference Enginee
r with a strong foundation in software engineering, distributed systems, and performance optimization
to build and optimize inference engines for large-scale LLM serving systems
. You will work across both research and production environments, ensuring our LLM serving systems are fast, scalable, and efficient. The role spans the entire inference stack - from kernel and runtime to scheduling, memory management, and distributed execution
Key Responsibilities:
Profile, benchmark, and analyze bottlenecks for LLM inference workloads across multiple layers: kernel, memory, networking, and scheduler
Optimize inference engines (vLLM, SGLang, TensorRT-LLM) for throughput, latency, memory efficiency, GPU utilization, and cost
Implement and fine-tune inference optimization techniques including batching, KV-cache management, quantization, speculative decoding, parallelism strategies, and disaggregated serving
Build instrumentation and profiling tools to identify bottlenecks
Ensure the reliability of the inference pipeline through A/B launches, rollback, model versioning, and fault tolerance
Collaborate with the Platform Engineering team to improve serving architecture based on performance findings
Document and share knowledge, contributing to internal best practices and AI open-source projects whenever possible
Requirements
1. Mandatory:
At least 5 years of experience as a Software Engineer, Performance Engineer, or equivalent.
Strong foundation in Software Engineering, Software Architecture, and Distributed Systems.
Proficiency in at least one of the following languages: Python, Go, or C++.
Experience developing or optimizing distributed systems, high-throughput backends, or large-scale serving systems.
Experience with benchmarking, profiling, and performance tuning in production environments.
Ability to analyze CPU, Memory, Network, or Storage bottlenecks.
Strong systems thinking, Root Cause Analysis capabilities, and the ability to solve complex performance problems.
Strong ownership mindset and the ability to work independently.
2. Nice to Have:
Experience with Linux internals, kernel tuning, or custom Linux kernel.
Understanding of GPU Architecture or CUDA Programming.
Experience with AI/ML Serving Systems or LLM Inference.- Have worked with one of the inference engines such as vLLM, SGLang, TensorRT-LLM, or Triton Inference Server.
Understanding of batching, KV Cache, quantization, speculative decoding, tensor/pipeline parallelism, or disaggregated serving.
Experience with the NVIDIA inference stack (TensorRT, Triton, CUTLASS, NCCL, cuBLAS, cuDNN).
Experience with observability stacks such as Prometheus, Grafana, or OpenTelemetry.
Open-source contributions or research related to AI Infrastructure, ML Systems, or Performance Optimization.

Thông tin chung

  • Thu nhập: Thỏa thuận
Việc làm tương tự khác
Thông tin công việc
Vị trí:
Nhân viên
Hình thức làm việc:
Toàn thời gian
Việc làm tương tự
Web System Engineer - Hà Nội - Chấp Nhận Fresher
Công Ty TNHH OTANI U.P.
Hà Nội
Công Ty TNHH OTANI U.P.
11.5 - 18.5 triệu VND
Golang Developer
CÔNG TY CỔ PHẦN CÔNG NGHỆ TINH VÂN
Hồ Chí Minh
CÔNG TY CỔ PHẦN CÔNG NGHỆ TINH VÂN
Cạnh tranh
AI Engineer
Công ty Giải pháp Công nghệ Sài Gòn - Saigon Technology Solutions
Hồ Chí Minh
Công ty Giải pháp Công nghệ Sài Gòn - Saigon Technology Solutions
Thỏa Thuận
Nhân Viên Triển Khai Vận Hành (Devops)
CÔNG TY CỔ PHẦN BKAV
Hà Nội
CÔNG TY CỔ PHẦN BKAV
15 - 20 triệu
Lập trình viên AI
Công ty Cổ Phần Tin học Giải Pháp Tích Hợp Mở
Hồ Chí Minh
Công ty Cổ Phần Tin học Giải Pháp Tích Hợp Mở
Thỏa thuân
Cảnh báo dấu hiệu lừa đảo tuyển dụng
Thu phí và cung cấp thông tin
  • Phí hồ sơ, đồng phục, đặt cọc.
  • Yêu cầu nộp bản gốc giấy tờ.
  • Cung cấp mã OTP.
Hứa hẹn trúng tuyển 100%
  • Không yêu cầu trình độ.
  • Không cần thử việc.
Phỏng vấn bất thường
  • Địa điểm xa văn phòng công ty.
  • Phỏng vấn qua Telegram.
Yêu cầu làm nhiệm vụ
  • Tải app, nạp tiền.
  • Làm nhiệm vụ nhận thưởng.
Tin tuyển dụng sơ sài
  • Mô tả công việc chung chung
  • Nhiệm vụ đơn giản, thu nhập khủng
  • Lỗi chính tả, đánh máy.
Đội ngũ hỗ trợ của JobOKO sẵn sàng đồng hành, tư vấn và giới thiệu những cơ hội việc làm phù hợp, giúp Ứng viên tự tin phát triển sự nghiệp và chinh phục mục tiêu nghề nghiệp bền vững.
Đội ngũ hỗ trợ của JobOKO luôn chủ động tư vấn các giải pháp tuyển dụng tối ưu, cam kết đồng hành và hỗ trợ Quý Nhà tuyển dụng đạt được hiệu quả tuyển dụng bền vững.