Job Description:
Design the overall architecture of our cloud-inference super-node & cluster, includingmulti-chip interconnection (e.g., NVLink-like protocols), memory-subsystem design(such as VRAM, memory, SSD),and network topology(Ethernet/InfiniBand)Optimize system-level performance, latency, and scalability to support large-scaleLLM deployment(scale-up&scale-out).Collaborate with chip architects andmechanical engineers to ensure thermal, power, and form-factor compliance withstandards. Drive system prototyping and validation with top internetdata-centercustomers
任职要求
Requirements:
8+ years of experience in high-performance computing (HPC)or data-center
system-architecture design.
Deep expertise in multi-GPU/chip interconnect technologies, memory-subsystem
design, and data-center network protocols.Familiarity with cloud-inference cluster deployment and large-scale AI-systemoptimization(especially scale-up).
Ability to balance system-level trade-offs between performance, cost, and reliability.