Responsibilities
- Drive the development of vision and multimodal models for document, image and media translation, ranging from media ingestion and generation to end-to-end models.
- Drive hands-on research and development on post-training for our vision and/or multimodal models: supervised fine-tuning, knowledge distillation, preference optimization, and reinforcement learning tuned to translation quality.
- Build evaluator models for document and design quality, including rubric- and reference-based grading, and investigate and mitigate reward hacking and quality-estimation failure modes.
- Own the full lifecycle of model delivery: prototyping, ablations, training, evaluation, optimization, and production deployment, working closely with engineering to ship into real-time systems at scale.
- Establish strong practices for evaluation, reproducibility, monitoring, and continuous model improvement in production.
Requirements
- Proven experience with developing multimodal models, VLM, and/or vision models.
- Deep, hands-on expertise in model post-training, knowledge distillation (teacher-student training), and/or reinforcement learning (RLHF/RLAIF, PPO/GSPO, and reward modeling).
- Strong data-centric instincts for building synthetic-data and preference-data pipelines, Model-as-judge generation, data curation and filtering, data augmentations, and/or reasoning about data mixtures and ablations.
- A hands-on builder who enjoys training models, running experiments, debugging pipelines, and integrating ML systems into production while staying grounded in product impact and real-world quality.
- Strong coding and experimentation skills (Python, PyTorch/JAX/TensorFlow), and the ability to communicate clearly and align research with product and engineering priorities.
- Ability to lead complex research efforts, to communicate clearly and collaborate across teams, while staying grounded in product impact, user experience, and real-world performance.
Nice to Have
- Experience with machine translation, multilingual NLP, efficient long-context modeling, language quality estimation, or multimodal machine translation.
- Experience designing evaluation and reward signals using automatic metrics, Model-as-judge evaluation, non-verifiable rewards, and human-in-the-loop evaluation.
- Experience with multi-objective optimization, consistency models, unified multimodal generation.
- Experience with diffusion models.
- Publications at top-tier venues.