Ivyea Jobs 5948 roles · 69 companies · 刚刚
北京通研院 · 研究机构

研究员(深度学习/多模态/具身智能方向)

北京 社招 · 全职 通用智能体中心 具身智能 / 机器人

职位描述

岗位职责: 1.  围绕(1) Vision-Language Model (VLM), (2) Vision-Language-Action (VLA), (3) AGI Agents (具体描述见下) 这三个方向开展前沿研究,产出有影响力的原创研究成果。 2.  领导团队成员开展研究项目,指导实习生。 3.  配合工程师团队完成技术转化,推动算法落地。

任职要求

1.  具有博士学位,在人工智能领域的国际顶级会议或期刊(CVPR、ICCV、ECCV、NeurIPS、ICML、ICLR、ACL、EMNLP等)发表过论文。 2.  自驱力强,有志于从事人工智能前沿研究工作,有能力领导一个项目和指导实习生。 3.  具备优秀的英文口头沟通和书面表达能力。 机器学习实验室重点研究方向: 1.  Vision-Language Modeling (VLM): advancing 2D and 3D vision-language tasks via innovating model architectures, training algorithms, and collecting high-quality data, through which unifying various tasks in one model. Our recent focus: 3D-VL; VLM for long-form and streaming video understanding; video and 3D generation. References: ● 3D-VisTA https://3d-vista.github.io/ ● UltraEdit https://ultra-editing.github.io/   2.  Vision-Language-Action (VLA): integrating vision and language with action-oriented tasks such as planning& control in simulated engine and real robotics. Our recent focus: dexterous manipulation; mobile manipulation. References: ● LEO https://embodied-generalist.github.io/ ● JARVIS-1, OmniJARVIS,https://omnijarvis.github.io/   3.  AGI Agents: exploring the development of AI agents capable of general tool-use, problem solving, feedback reflection, and continual learning. Our recent focus: multi-modal agents; agent tuning; reflecting and learning from user feedback. References: ● CLOVA https://clova-tool.github.io/ ● VideoAgent https://videoagent.github.io/ 允许3个月内投递2个职位,请选择适合的职位进行投递 立即投递