Responsibilities
- Evaluate LLM Architecture Logic: review AI-generated explanations of model architectures, loss functions, and backpropagation for technical accuracy.
- Audit Code & Notebooks: validate ML-specific code (e.g., training loops, data preprocessing scripts, or model evaluations) for efficiency and correctness.
- Refine RLHF Frameworks: provide the high-quality human feedback necessary to align models with human intent, safety, and helpfulness.
- Analyze Model Reasoning: critically assess how an AI model navigates complex chain-of-thought (CoT) prompts and identify where the reasoning breaks down.
- Benchmark Performance: conduct comparative testing between different model outputs based on specific technical taxonomies and performance metrics.
Requirements
- a BS, MS, or PhD in Computer Science, Artificial Intelligence, Robotics, or a related quantitative field with a focus on Machine Learning.
- experience building, deploying, or fine-tuning ML models in a production environment.
- professional-level understanding of neural network architectures (Transformers, CNNs, RNNs) and optimization techniques.
- hands-on experience with Prompt Engineering, RLHF (Reinforcement Learning from Human Feedback), or RAG (Retrieval-Augmented Generation) workflows.
- the ability to audit complex model logic, identify training data contamination, and evaluate mathematical proofs behind ML algorithms.
- high attention to detail in spotting "hallucinations," biased outputs, or logical failures in AI-generated technical content.
Work Arrangement
On-site — Sacramento, US
Additional Information
- You must be prepared to complete paid tasks that require one hour of uninterrupted work, though many are shorter.
- Once you pass our assessment, you can join Prolific in just 15 minutes, and start enjoying competitive pay rates, flexible hours, and the ability to work from home.