Responsibilities
- Architect and maintain scalable batch and real-time data processing systems using Kafka, Flink, and ETL frameworks
- Design and refine analytical data structures for time-series data, financial indicators, and trading behavior
- Deploy and optimize analytical database solutions such as ClickHouse, MongoDB, BigQuery, Snowflake, or comparable systems with cost efficiency
- Construct idempotent data workflows that support reliable backfilling and data reconciliation
- Implement monitoring systems to ensure data accuracy, timeliness, and pipeline stability
- Develop and manage internal AI platforms used by multiple quantitative trading groups
- Build vector search infrastructures with efficient HNSW indexing and hybrid retrieval methods
- Design evaluation systems to measure retrieval performance metrics like Recall@K, MRR, and nDCG, as well as RAG effectiveness
- Develop reusable AI components, including standardized retrieval-augmented generation pipelines, prompt engineering tools, and self-service interfaces
- Design and maintain autonomous agent systems using modern frameworks such as LangGraph, A2A, or MCP, emphasizing control and traceability