Ivyea Jobs 5948 roles · 69 companies · 刚刚
北京通研院 · 研究机构

算法研究工程师

北京 社招 · 全职 通用智能体中心 具身智能 / 机器人

职位描述

岗位职责: 1.  参与(1) Vision-Language Model (VLM), (2) Vision-Language-Action (VLA), (3) AGI Agents (具体描述见下) 这三个方向的前沿研究。 2.  与团队成员合作完成研究项目,阅读论文,实现算法,进行大规模实验。 3.  参与或主导系统和平台的设计、开发,推动深度学习算法的落地

任职要求

1.具有硕士或博士学位,有人工智能领域的项目经历。 2.具备基本的机器学习和深度学习知识,了解常见的机器学习和深度学习模型,如Transformer,CNN,Seq2Seq等,能熟练使用PyTorch、Transformers等深度学习框架 3.自驱力强,乐于探索新技术,积极跟团队成员沟通讨论、合作。 4.加分项: ○ 在人工智能领域的会议或期刊发表过论文, 如CVPR、ICCV、ECCV、NeurIPS、ICML、ICLR、ACL、EMNLP等。 ○ 有大模型训练、微调经验,能熟练使用PEFT ○ 有网站或AI应用设计和开发经验,参与的AI开源项目github stars数100+ ○ 有较强的PPT制作、Demo视频剪辑能力 机器学习实验室重点研究方向: 1.  Vision-Language Modeling (VLM): advancing 2D and 3D vision-language tasks via innovating model architectures, training algorithms, and collecting high-quality data, through which unifying various tasks in one model. Our recent focus: 3D-VL; VLM for long-form and streaming video understanding; video and 3D generation. References: ● 3D-VisTA https://3d-vista.github.io/ ● UltraEdit https://ultra-editing.github.io/   2.  Vision-Language-Action (VLA): integrating vision and language with action-oriented tasks such as planning& control in simulated engine and real robotics. Our recent focus: dexterous manipulation; mobile manipulation. References: ● LEO https://embodied-generalist.github.io/ ● JARVIS-1, OmniJARVIS,https://omnijarvis.github.io/   3.  AGI Agents: exploring the development of AI agents capable of general tool-use, problem solving, feedback reflection, and continual learning. Our recent focus: multi-modal agents; agent tuning; reflecting and learning from user feedback. References: ● CLOVA https://clova-tool.github.io/ ● VideoAgent https://videoagent.github.io/ 允许3个月内投递2个职位,请选择适合的职位进行投递 立即投递