Ivyea Jobs 10985 roles · 194 companies · 刚刚
Zilliz · AI Infra

Senior Software Engineer, Cloud Reliability

上海 社招 · 正职 Engineering 训练系统 / 集群 工程开发

职位描述

Own the reliability, availability, and production stability of Zilliz Cloud as we scale through the next stage of growth Debug complex production issues across Kubernetes, cloud infrastructure, networking, storage, and distributed database systems Build automation and diagnostic tooling; log analysis, alert correlation, incident investigation, runbook automation, and remediation workflows so problems get solved once, not repeatedly Turn recurring incidents into reusable tools, automation, documentation, and product improvements Improve observability across latency, availability, throughput, and resource efficiency Partner with database and infrastructure engineers to make Zilliz Cloud more reliable, scalable, and automated

任职要求

3+ years building or operating production cloud systems, infrastructure platforms, database systems, or large-scale online services Bachelor's degree in Computer Science, Software Engineering, or a related field, or equivalent practical experience Strong hands-on experience with Kubernetes, Docker, and at least one major cloud platform (AWS, GCP, or Azure) Solid understanding of distributed systems; availability, scalability, performance, failure recovery, and operational tradeoffs Experience with distributed databases, storage systems, search systems, or large-scale online systems is a strong plus Experience operating highly multi-tenant systems or large infrastructure fleets; thousands of nodes, clusters, tenants, or customer deployments is especially valuable Familiarity with modern cloud operations tooling such as Terraform, Helm, Argo CD, Prometheus, Grafana, and CI/CD systems Strong bias for action, and the drive to thrive in a fast-paced, rapidly scaling environment
Zilliz去官方页面投递 →