Responsibilities
- Architect and operate resilient ML job execution frameworks covering training, inference, and post-processing workflows.
- Develop and maintain API services and developer tooling to orchestrate ML workflows on Kubernetes using Argo Workflows, Helm, Terraform.
- Build scalable, efficient batch pipelines with Apache Spark to support large-scale ML training and evaluation.
- Design and maintain robust data infrastructures using Trino, Databricks and other modern database technologies, monitored with Prometheus and Grafana for high availability and observability.
- Develop tooling that streamlines ML experimentation, accelerates production workflows, and empowers cross-functional teams to innovate rapidly.
- Collaborate deeply with ML scientists to transform research prototypes into reliable, scalable, user-facing AI products.
- Lead cloud infrastructure design and operations on GCP, leveraging managed services such as Google Compute Engine (GCE), Google Kubernetes Engine (GKE), Cloud Storage, Cloud Functions, Cloud Pub/Sub, Cloud SQL, BigQuery, and more.
- Define and implement CI/CD pipelines with tools like Jenkins, Github Action, or ArgoCD to enable seamless, automated deployments.
- Harness distributed computing and parallel programming principles to optimize system resource utilization and performance.
Requirements
- Bachelor’s degree in Computer Science, Engineering, or a related technical field (Master’s degree preferred).
- 5+ years of hands-on experience in ML platform engineering, MLOps, or data infrastructure, deploying enterprise-grade machine learning systems at scale.
- Expert proficiency in Python, Java, or Go, with solid foundations in data structures and algorithm design.
- In-depth experience with cloud environments (AWS or GCP) and cloud-native service management.
- Proven mastery of Docker containers and Kubernetes cluster management, including resource provisioning, autoscaling, and deployment best practices.
- Strong understanding of the ML lifecycle—from training and prediction to evaluation, backtesting, and feedback loops.
- Familiarity with Git workflows and Linux-based development environments.
- Passionate about continual learning and innovation, leveraging AI-powered developer tools like GitHub Copilot and ChatGPT to boost productivity.
Nice to Have
- Experience in the MarTech industry or other customer-centric domains, eager to deliver products that delight users and drive business impact.
- Demonstrated architectural leadership and ownership, skillfully driving complex, cross-team platform initiatives.
- Strong grasp of deep learning fundamentals and end-to-end ML workflow platforms such as Kubeflow, MLflow, or AWS SageMaker.
- Hands-on experience with distributed data processing frameworks like Apache Spark, and pipeline orchestration tools such as Apache Airflow, Argo Workflow, or Luigi.
- Expertise in production-level ML applications, including handling data imbalance, preventing data leakage, and optimizing resource consumption for large-scale training and serving.
- Familiarity with real-time online inference architectures and batch processing trade-offs.
- Enthusiastic adopter of “vibe coding” culture—collaborative, transparent, and always pushing technical excellence together.
- Prior experience building or developing applications related to large language models (LLM), multi-agent LLM systems, or natural language processing (NLP).