마감 20일 전
[마음AI] WoRV팀 | Inference Optimization Engineer 채용
We enable Physical AI to operate reliably in the real world. Our mission is to maximize inference efficiency in on-device environments while meeting accuracy and latency requirements under real-world constraints. By understanding the computational structure of VLAs, building dedicated runtime engines, and engineering the software stack from scratch, we shape the next-generation of serving systems.
Analyze the architecture and operations of generative models, especially VLAs
Deploy and run models on edge devices, utilizing hardware accelerators such as NPUs
Identify performance bottlenecks, then research and apply techniques to resolve them
Proficiency in C/C++ or Python
Fundamental CS knowledge and English proficiency to understand ML systems research papers
A proactive approach to problem-solving
Experience deploying and optimizing generative models on edge devices
Experience profiling performance at kernel and hardware level
Kernel programming experience on hardware accelerators such as GPUs or NPUs