You will be responsible for
1. Lake storage and pipeline scaling: Evolution of Iceberg Lake storage; Build/maintain multi-source batch flow links to ensure stability, consistency, and timeliness.
2. Modeling and indicator system: Use DBT precipitation hierarchical model and unified caliber indicators (user/transaction/traffic/commercialization) to support self-service analysis and business collaboration.
3. Query and cost optimization: Perform performance and cost optimization on Spark/Glue/ClickHouse, etc., to enhance the Kanban and analysis experience.
4. Engineering and Observability+Intelligent Exploration: Improve CI, documentation, quality, and alarm systems; Incremental promotion of data engineering capabilities with "AI as the core" (quality diagnosis, anomaly localization, asset retrieval/Q&A, etc.).
【你将负责】
1. 湖仓与管道规模化:演进 Iceberg 湖仓;建设/维护多源批流链路,保障稳定性/一致性/时效性。
2. 建模与指标体系:用 dbt 沉淀分层模型与统一口径指标(用户/交易/流量/商业化),支撑自助分析与业务协作。
3. 查询与成本优化:在 Spark / Glue / ClickHouse等上做性能与成本优化,提升看板与分析体验。
4. 工程化与可观测 + 智能探索:完善 CI、文档、质量与告警体系;增量推进“AI 为内核”的数据工程能力(质量诊断、异常定位、资产检索/问答等)。
任职要求
We hope you possess
1 year+experience in data engineering: Proficient in SQL and able to abstract business into data models; Understand hierarchical modeling, caliber consistency, increment/backfill/power, and so on.
2. Solid coding ability: Production level coding ability (Python/Java/Go, etc.), strong engineering thinking (logging, exception handling, basic testing awareness).
3. Lake Warehouse/Cloud Practice: Familiar with various common databases and data lake architectures and features; Experience in S3 storage and computing engines.
4. Ownership: Effectively break down and evaluate tasks, quickly locate and repair problems, and conduct retrospective analysis and sedimentation; Align across teams and drive delivery.
【 Bonus Points 】
-Actual implementation and optimization experience of Iceberg/Spark/Athena/Trino/Glue/DuckDB/Clickhouse
-Deep dbt (tests/docs/CI/CD/macros) or ClickHouse/OLAP optimization experience
-Experience in building buried links and analyzing growth (such as using PostHog); Funnel/Retention/Attribution/Experimental Caliber)
-Familiar with third-party data models such as Stripe/advertising channels, or have worked on prototypes for "Data Copilot/Intelligent Tools"
You will receive
-Practical application of new generation data architecture: Iceberg+glue+dbt+Dagster's modern data stack, creating more stable, faster, and cost-effective links.
-AI driven methodology: Using the most advanced AI tools to make quality, troubleshooting, caliber, and asset accumulation more intelligent and efficient.
-High influence+instant feedback=rapid growth: great authority, fast feedback on achievements, and strong sense of accomplishment; Good technical atmosphere, encouraging the accumulation of SOP and best practices.
【我们希望你具备】
1. 1 年+ 数据工程经验:SQL 扎实,能把业务抽象为数据模型;理解分层建模、口径一致性、增量/回填/幂等等。
2. 代码能力扎实:生产级别的代码能力(Python/Java/Go等),工程思维强(日志、异常处理、基础测试意识)。
3. 湖仓/云上实践:熟悉各种常见数据库与数据湖架构、特性;有 S3 类存储 + 计算引擎经验。
4. Ownership:对任务进行有效拆解与评估,遇到问题能快速定位修复,并复盘沉淀;跨团队对齐口径并推进交付。
【加分项】
- Iceberg / Spark / Athena / Trino / Glue / DuckDB / Clickhouse 的实际落地与优化经验
- 深度 dbt(tests/docs/CI/CD/宏)或 ClickHouse/OLAP 优化经验
- 埋点链路构建经验与增长分析经验(例如使用过PostHog;漏斗/留存/归因/实验口径)
- 熟悉 Stripe/广告渠道等第三方数据模型,或做过“数据 Copilot/智能工具”原型
【你将获得】
- 新一代数据架构实战:Iceberg + glue + dbt + Dagster 的现代数据栈,做出更稳、更快、更省的链路。
- AI 驱动的方法论:用最先进的AI工具,把质量、排障、口径与资产沉淀做得更智能、更高效。
- 高影响力+即时反馈 = 高速成长:权限大、成果反馈快、成就感强;技术氛围好,鼓励沉淀 SOP 与最佳实践。