Computer Vision Leader
Hạn nộp hồ sơ: 14/09/2026 (Còn 29 ngày)
Ứng tuyển sớm để được ưu tiên
Kết nối với Nhà tuyển dụng để tìm hiểu thông tin và gia tăng cơ hội trúng tuyển
Nhà tuyển dụng đang online
Mô tả công việc
Automated Data QA/QC Pipeline: Design and implement an end-to-end automated pipeline to parse, validate, and process incoming video batches against strict programmatic standards (metadata, FPS, aspect ratio, duration, compression, and batch-list matching).
Egocentric Interaction Analysis: Develop robust CV modules to analyze first-person view (FPV) mechanics:
Ensure hands are continuously visible, working, and not occluded at the frame edges.
Calculate spatial approximations to guarantee manipulated objects remain centered in the frame.
Action & Task Verification: Implement video-language models or action classifiers to automatically verify that the recorded actions semantically match the assigned task descriptions.
Visual & Sensor Quality Control: Build heuristic and AI-based checks to filter out videos with excessive camera shake, poor lighting conditions (over/underexposed), or missing/corrupted head-mounted IMU data.
Automated Anonymization: Architect a zero-leakage PII redaction system to scan all frames for human faces, children, and sensitive text (IDs, licenses, logos), applying automated blurring before data enters the training lake.
Audio & Speech Validation: Integrate lightweight audio processing to verify the presence of audio tracks and extract/validate spoken voice commands from actors.
Performance Optimization: Optimize the entire video scanning pipeline for high-throughput execution on AWS infrastructure, minimizing processing time and compute costs per GB of video.
Egocentric Interaction Analysis: Develop robust CV modules to analyze first-person view (FPV) mechanics:
Ensure hands are continuously visible, working, and not occluded at the frame edges.
Calculate spatial approximations to guarantee manipulated objects remain centered in the frame.
Action & Task Verification: Implement video-language models or action classifiers to automatically verify that the recorded actions semantically match the assigned task descriptions.
Visual & Sensor Quality Control: Build heuristic and AI-based checks to filter out videos with excessive camera shake, poor lighting conditions (over/underexposed), or missing/corrupted head-mounted IMU data.
Automated Anonymization: Architect a zero-leakage PII redaction system to scan all frames for human faces, children, and sensitive text (IDs, licenses, logos), applying automated blurring before data enters the training lake.
Audio & Speech Validation: Integrate lightweight audio processing to verify the presence of audio tracks and extract/validate spoken voice commands from actors.
Performance Optimization: Optimize the entire video scanning pipeline for high-throughput execution on AWS infrastructure, minimizing processing time and compute costs per GB of video.
Yêu cầu
Computer Vision & Deep Learning Expertise:
Egocentric Vision & Hand Tracking: Deep expertise in hand pose estimation, object tracking, and egocentric action recognition (e.g., MediaPipe, HaMeR, Ego4D ecosystem). Must be able to detect hand presence, occlusion, and interaction with objects in the center of the frame.
Action Recognition & Semantic Understanding: Strong experience with video understanding models (e.g., SlowFast, TimeSformer, VideoCLIP) to verify if the actor's actions match the assigned tasks.
Image Quality Assessment (IQA): Proficiency in classical and deep-learning-based IQA methods to automatically detect over/underexposure, motion blur, and excessive camera shake.
Privacy & Anonymization: Proven ability to build robust detection pipelines for faces, demographics (children), and PII (ID cards, passports, logos) using models like RetinaFace, YOLO, and OCR (Tesseract/EasyOCR) for automatic blurring/redaction.
Video Processing & Sensor Fusion:
Advanced Video Processing: Mastery of FFmpeg, OpenCV, and GStreamer for programmatic video manipulation, trimming, metadata extraction (FPS, bitrates, compression), and batch validation.
Sensor Data Handling (IMU): Experience working with time-series data and sensor fusion, specifically synchronizing head-mounted IMU telemetry (accelerometer/gyro) with video frames.
Multimodal Processing (Bonus): Familiarity with audio processing and Speech-to-Text (e.g., Whisper) to validate spoken commands in video audio tracks.
Engineering & Deployment:
Programming: 5+ years of software engineering experience with strong proficiency in Python and C++.
Frameworks: Expert in PyTorch; familiar with model optimization and inference acceleration using TensorRT, ONNX, or OpenVINO.
Cloud & MLOps: Experience deploying high-throughput CV pipelines on AWS (EC2 GPU instances, SageMaker, or EKS) and working with large-scale storage (S3).
Egocentric Vision & Hand Tracking: Deep expertise in hand pose estimation, object tracking, and egocentric action recognition (e.g., MediaPipe, HaMeR, Ego4D ecosystem). Must be able to detect hand presence, occlusion, and interaction with objects in the center of the frame.
Action Recognition & Semantic Understanding: Strong experience with video understanding models (e.g., SlowFast, TimeSformer, VideoCLIP) to verify if the actor's actions match the assigned tasks.
Image Quality Assessment (IQA): Proficiency in classical and deep-learning-based IQA methods to automatically detect over/underexposure, motion blur, and excessive camera shake.
Privacy & Anonymization: Proven ability to build robust detection pipelines for faces, demographics (children), and PII (ID cards, passports, logos) using models like RetinaFace, YOLO, and OCR (Tesseract/EasyOCR) for automatic blurring/redaction.
Video Processing & Sensor Fusion:
Advanced Video Processing: Mastery of FFmpeg, OpenCV, and GStreamer for programmatic video manipulation, trimming, metadata extraction (FPS, bitrates, compression), and batch validation.
Sensor Data Handling (IMU): Experience working with time-series data and sensor fusion, specifically synchronizing head-mounted IMU telemetry (accelerometer/gyro) with video frames.
Multimodal Processing (Bonus): Familiarity with audio processing and Speech-to-Text (e.g., Whisper) to validate spoken commands in video audio tracks.
Engineering & Deployment:
Programming: 5+ years of software engineering experience with strong proficiency in Python and C++.
Frameworks: Expert in PyTorch; familiar with model optimization and inference acceleration using TensorRT, ONNX, or OpenVINO.
Cloud & MLOps: Experience deploying high-throughput CV pipelines on AWS (EC2 GPU instances, SageMaker, or EKS) and working with large-scale storage (S3).
Quyền lợi
Competitive compensation package based on experience and qualifications
Opportunity to build a strategic global data marketplace for robotics and AI training data from zero to one.
Work in a high-speed technology environment backed by Vingroup and VinDynamics leadership.
Competitive compensation package aligned with capability and business impact.
Clear ownership, measurable KPIs, and exposure to global partners, US platform models, and frontier robotics businesses.
Opportunity to build a strategic global data marketplace for robotics and AI training data from zero to one.
Work in a high-speed technology environment backed by Vingroup and VinDynamics leadership.
Competitive compensation package aligned with capability and business impact.
Clear ownership, measurable KPIs, and exposure to global partners, US platform models, and frontier robotics businesses.
Thông tin khác
Thời gian làm việc
Thứ 2 - Thứ 6 (từ 08:30 đến 17:30)
Và 2 ngày thứ 7 (cách tuần)
Thứ 2 - Thứ 6 (từ 08:30 đến 17:30)
Và 2 ngày thứ 7 (cách tuần)
Thông tin chung
- Thu nhập: Thoả thuận
Nơi làm việc
- Hà Nội: TechnoPark Tower, Xã Gia Lâm (huyện Gia Lâm cũ)
Việc làm tương tự khác
CÔNG TY CỔ PHẦN TẬP ĐOÀN XÂY DỰNG HÒA BÌNH
Hà Nội
20 - 30 triệu VNĐ
CÔNG TY CỔ PHẦN XÂY DỰNG VÀ PHÁT TRIỂN THƯƠNG MẠI HOÀNG LÂM
Hà Nội
Từ 18 đến 25 triệu VND
Thời trang thể thao Lining - CÔNG TY TNHH QUỐC TẾ HẢI LONG
Hà Nội
18 Tr - 22 Tr VND
CÔNG TY CỔ PHẦN XÂY DỰNG VÀ PHÁT TRIỂN THƯƠNG MẠI HOÀNG LÂM
Hà Nội
15 - 50 triệu VNĐ
Công ty Cổ phần Nghiên cứu, Phát triển và Ứng dụng Robot Hình người VinDynamics
Xem trang công ty- Địa chỉ công ty: Techno Park Tower, Vinhomes Ocean Park, Xã Gia Lâm, Quận Gia Lâm, Hà Nội
- Quy mô: Từ 101 - 500 nhân viên
Thông tin công việc
Vị trí:
Nhân viên
Hình thức làm việc:
Toàn thời gian
Việc làm tương tự
Nhân Viên Bóc Tách Dự Toán QS Nội Thất | Thu Nhập 10-18 Triệu + Thưởng | Hà Nội
Công ty TNHH TM & Trang Trí Nội Thất Trung Á
Hà Nội
10 - 18 triệu VND + Thưởng
Kỹ Sư Xây Dựng (Giám Sát Thi Công) - Nhận Sinh Viên Mới Ra Trường
Công ty Cổ Phần Đầu Tư và Xây dựng số 18.3 (LICOGI 18.3)
Hà Nội, Bắc Ninh, Hải Dương, Hưng Yên, Vĩnh Phúc
12 - 25 triệu VND + Thưởng
Cảnh báo dấu hiệu lừa đảo tuyển dụng
Đội ngũ hỗ trợ của JobOKO sẵn sàng đồng hành, tư vấn và giới thiệu những cơ hội việc làm phù hợp, giúp Ứng viên tự tin phát triển sự nghiệp và chinh phục mục tiêu nghề nghiệp bền vững.
Hotline CSKH
1900.63.63.84
Công ty Cổ phần JobOKO Toàn cầu
Đội ngũ hỗ trợ của JobOKO luôn chủ động tư vấn các giải pháp tuyển dụng tối ưu, cam kết đồng hành và hỗ trợ Quý Nhà tuyển dụng đạt được hiệu quả tuyển dụng bền vững.
Hotline CSKH
0962.107.888
Công ty Cổ phần JobOKO Toàn cầu