1. Conduct systematic research on model inference acceleration and deployment optimization technologies, focusing on performance improvement and efficiency optimization for generative models, 3D models, and large-scale world model inference pipelines.
2. Mitigate core performance and deployment bottlenecks in industrial scenarios by designing and implementing high-efficiency inference acceleration solutions, so as to facilitate large-scale, high-performance industrial deployment and engineering implementation of generative AI models.
3. Systematically summarize innovative acceleration technologies and engineering methodologies, support intellectual property output and technical standard iteration, and promote internal and external industry technical exchanges and knowledge sharing.
任职要求
1. Master’s degree or above. Candidates shall possess solid theoretical foundations and engineering experience in model acceleration, high-performance computing, and deep learning deployment. Relevant industrial or research practical experience is preferred.
2. Possess solid mathematical and engineering fundamentals, proficient in Python and C++. Familiar with CUDA programming, model quantization, pruning computation, and computational graph optimization, with the ability to independently design and implement efficient acceleration algorithms.
3. Highly self-driven and result-oriented. Capable of proactively identifying model performance bottlenecks and achieving critical breakthroughs in training and inference speedup through systematic technical optimization.
4. Possess an agile entrepreneurial mindset, capable of steadily advancing engineering iteration and multi-task performance optimization in a fast-paced technical deployment environment.