Responsibilities
- Own the architecture, development, and optimization of our big data platform, supporting large-scale collection, processing, and analysis of on-chain data, trading behavior, and user profiles.
- Design and build real-time/batch data pipelines for on-chain data — blockchain transactions, smart contract events, wallet address behavior.
- Build and maintain the data warehouse/lake, define data-layering standards, and ensure data quality, consistency, and timeliness.
- Combine big data with LLM capabilities to build intelligent applications — on-chain anomaly/fraud detection, smart risk models, user behavior prediction, and natural-language data querying (Text2SQL).
- Build data infrastructure for AI use cases — vector databases, feature platforms, and Embedding pipelines supporting RAG retrieval and Agent data supply.
- Partner with Algorithm/AI teams on large-scale data processing and pipelines for model training data and feature engineering; partner with Product, Risk, and Growth teams on data needs for trading analytics, anti-fraud, growth, and operations.
- Optimize performance and resource efficiency of big data jobs, ensuring stability and SLAs on core pipelines.
- Track industry developments in big data, Web3 data infrastructure, and AI, driving technology selection and architecture evolution.
- Communicate and document in English with overseas colleagues and partners (exchanges, public chain teams, etc.).
Requirements
- Bachelor's degree or above in Computer Science, Software Engineering, or related field
- 7+ years of big data development experience
- Proficiency with Hadoop, Spark, and Flink, with experience building batch/real-time data warehouses
- Familiarity with Hive, Kafka, HBase, ClickHouse, and Doris, with large-scale cluster tuning experience
- Strong big data development skills in Java/Scala/Python, with solid SQL tuning ability
- Understanding of LLM fundamentals and application patterns, with hands-on experience in prompt engineering and RAG
- Experience with vector databases (e.g. Milvus, Pinecone, Weaviate, pgvector) or Embedding data processing
- Understanding of blockchain fundamentals (account model/UTXO, consensus mechanisms, gas mechanics)
- Understanding of on-chain data structures for major public chains (Ethereum, BSC, Solana, Bitcoin, etc.) — blocks, transactions, event logs, token transfers (ERC-20/ERC-721)
- Fluent English reading/writing
- Able to independently read English technical documentation (e.g. Ethereum Yellow Paper, RFCs/EIPs, AI papers) and write technical proposals and weekly reports in English
- Able to communicate in English with overseas teams, communities, or partners, with strong cross-cultural communication skills
- Clear logical thinking, strong problem decomposition and troubleshooting ability, and a strong appetite for learning new AI technology
- Strong team player, able to deliver consistently in a fast-paced, high-uncertainty Web3 environment
Nice to Have
- Experience with data governance, data quality monitoring, or metadata management is a plus
- Familiarity with emerging AI application architectures such as AI Agents or MCP (Model Context Protocol) is a plus
- Experience connecting big data platforms with AI/ML training and inference pipelines (e.g. feature platforms, real-time feature serving)
- Experience using LLMs to accelerate data engineering (automated data quality checks, intelligent ETL generation, Text2SQL) is a plus
- Basic ML/deep learning knowledge and ability to collaborate effectively with algorithm teams is a plus
- Experience with on-chain data collection/parsing is a plus — e.g. building or integrating with The Graph, Dune, Etherscan API, or node RPCs
- Understanding of DeFi, NFT, or cross-chain bridge business models and their data characteristics is a plus
- Data engineering experience at an exchange, wallet, DeFi protocol, or on-chain analytics company is a plus