Responsibilities
- Architect and manage distributed data systems at scale
- Lead ownership of compute and storage platforms including MaxCompute, Hologres, and Spark
- Develop multi-location task orchestration with dynamic engine selection and policy enforcement
- Improve reliability and efficiency in batch and streaming data workflows
- Construct the foundational platform layer for AI-native applications
- Design and implement tool interfaces enabling AI agents to interact with system APIs
- Build intelligent scheduling and cost-optimization agents for automatic resource tuning
- Implement telemetry systems to support AI-based SLA tracking and anomaly detection
- Create context retrieval systems using RAG and vector search over SQL and configuration data
- Enhance developer productivity through improved tooling and workflows
- Maintain an internal data development platform with IDE support, automated code reviews, and deployment tools
- Develop API-first utilities for data backfill and ingestion, designed for future agent integration
- Partner with data warehouse and service teams to define interoperable platform standards
- Promote high standards in system operations and service quality
- Define SLA targets, cost indicators, and latency visualizations as inputs for AI optimization
- Develop automated pipelines for incident detection and root cause analysis
- Enforce infrastructure policies across multiple cloud providers
- Design a scheduling agent that optimizes task dependencies, engine choice, and alert levels
- Create an operations agent to monitor pipeline health, detect performance drops, and identify schema changes
- Build an incident response agent to trace SLA violations, assign responsibility, and generate post-mortems
- Maintain the MCP Tool Layer, providing unified interfaces for all agent interactions
Requirements
- Minimum of five years building large-scale data systems using technologies like Hadoop, Spark, Flink, or similar
- In-depth knowledge of distributed storage and processing frameworks such as MaxCompute, Hologres, ClickHouse, or Hive
- Proficient in Java, Scala, or Python with experience in API-centric software design
- Practical experience with workflow schedulers like Airflow, DolphinScheduler, or custom solutions
- Strong grasp of multi-cloud infrastructure and cost management strategies
Nice to Have
- Experience integrating large language models using patterns like tool calling, RAG, and context handling
- Background with MCP or comparable agent-to-tool communication frameworks
- Drive to build systems that significantly boost engineering team productivity
Benefits
- Attractive compensation package including base and incentives
- Learning and development programs with education subsidies
- Regular team-building activities and company-wide events
- Wellness and meal support allowances
- Extensive healthcare coverage for employees and dependents
- Additional perks shared during the hiring process
Compensation
Competitive total compensation package
Work Arrangement
Remote (Worldwide)