Responsibilities
- Manage the complete data ingestion workflow from network capture to data lake storage, transforming raw exchange protocols and third-party feeds into a standardized format with precise timing.
- Develop and maintain batch and streaming data pipelines using Airflow, Spark, and dbt for processing tick-level and reference datasets.
- Reconstruct Level 2 and Level 3 order books, ensuring accurate handling of gaps and sequence inconsistencies in market data streams.
- Design and maintain software development kits in Python and Rust for internal systems to consume real-time market data feeds.
- Oversee the Iceberg-based data lake on S3, optimizing file layout, partitioning, and sorting for high-performance queries.
- Ensure schema evolution, snapshot management, time travel, compaction, and retention policies are properly implemented and maintained.
- Maintain slowly changing dimensions for reference data, guaranteeing historical accuracy for backtesting and audit purposes.
- Implement cost-efficiency strategies including intelligent compaction, storage tiering, and lifecycle management of data snapshots.
- Build tooling for schema governance, data contracts, validation frameworks, and lineage tracking within the Iceberg catalog system.
- Develop shared data access services using Spark and Polars to enable consistent, normalized data consumption across research, trading, and simulation workflows.
- Implement monitoring, alerting, service-level objectives, and continuous integration pipelines across data capture and processing infrastructure running on Kubernetes.
- Own operational excellence through data quality dashboards, incident response playbooks, and reliability documentation for the data ingestion fleet.
Responsibilities (12)
- Own the full capture path from wire to lake: decode and normalize raw exchange feeds (pcap, multicast UDP / ITCH / FIX) and vendor sources (OneTick, Refinitiv, Bloomberg, ICE) into a unified canonical model with nanosecond timestamps.
- Build batch + stream pipelines (Airflow, Spark, dbt) for tick and reference data.
- Own L2/L3 order-book reconstruction with gap handling.
- Provide Python and Rust producer SDKs for internal feed handlers.
- Own the Iceberg-over-S3 lakehouse: design partitioning, sort orders, and row-group layout for fast scans; manage schema evolution, snapshots, time travel, compaction, and TTL.
- Maintain reference data as slowly-changing tables with point-in-time correctness for backtests.
- Drive storage cost optimisation via compaction, tiering, and snapshot expiry.
- Build libraries for schema management, data contracts, validation, and lineage on top of the Iceberg catalog.
- Develop shared access services (Spark + Polars) so Research, backtesting, and trading share one normalized data layer, including gap detection and pcap-vs-lake reconciliation.
- Embed monitoring, alerting, SLAs/SLOs, and CI/CD across capture and pipeline layers on Kubernetes (EKS).
- Own data-quality dashboards and incident runbooks for the capture fleet.
- Partner with Quant Research, Data Science, Backend, and DevOps to translate requirements into platform capabilities and champion market-data engineering best practices.