Responsibilities
- Create and update model policies for safety-relevant areas such as dual-use, agentic, and emerging frontier risks.
- Convert risk and harm models into clear behavioral specifications, evaluation criteria, grading guidelines, and system-level protections.
- Define practical distinctions between beneficial AI uses and assistance that could enable harm, exploitation, misuse, or unsafe outcomes.
- Develop policy documents that support model training, evaluation, and deployment.
- Collaborate with safety researchers, engineers, product teams, and other stakeholders to implement policies into scalable model behavior and measurable safeguards.
- Utilize red-teaming results, deployment data, model failures, over-refusals, under-refusals, and ambiguous edge cases to enhance policy and evaluation quality over time.
- Identify emerging capability areas where frontier AI systems could introduce new safety challenges or lower barriers to harm.
- Analyze real-world deployments to determine where model behavior succeeds, fails, or deviates from the intended safety posture.
- Combine long-term safety research with hands-on launch and deployment work.
- Contribute to system cards, safety reports, policy documentation, launch reviews, and external communications regarding the approach to model safety and risk mitigation.
- Design and execute human data campaigns, including gold set construction, labeling guidance, calibration, adjudication, and evaluation coverage analysis, to ensure policies can be reliably measured and improved.
Requirements
- Strong judgment about how advanced AI systems may affect real-world risk, especially in ambiguous, fast-moving, or high-impact areas.
- Experience building or applying policies, taxonomies, harm models, threat models, or risk frameworks for complex technical, social, or adversarial systems.
- Ability to work across domains without needing to be the deepest subject-matter expert in every area, while knowing when to seek expert input.
- Ability to transform vague questions into structured policy frameworks, evaluation criteria, operational guidance, and enforceable model behavior.
- Comfortable using empirical evidence, including evaluations, red-teaming results, deployment observations, and model failure modes, to inform policy decisions.
- Thinks in systems across policy, data, graders, classifiers, training, deployment safeguards, measurement, monitoring, and escalation workflows.
- Technical judgment about what model behavior can realistically be trained, measured, evaluated, and enforced at scale.
- Works well across research, engineering, product, policy, domain experts, and operational teams.
- Writes clearly about complex tradeoffs where safety, user value, and implementation constraints all matter.
- Takes a pragmatic approach to safety, focused on reducing real-world risk while preserving legitimate, beneficial, and socially valuable uses of AI.
- Enjoys fast-paced, collaborative research environments where priorities shift as models, evidence, and risks change.
- Stays grounded in implementation details, empirical results, and what can actually be trained or measured.
Work Arrangement
Hybrid — San Francisco
Other
- Relocation support to new employees.
- Background checks for applicants will be administered in accordance with applicable law.
- Qualified applicants with arrest or conviction records will be considered for employment consistent with applicable laws, including the San Francisco Fair Chance Ordinance, the Los Angeles County Fair Chance Ordinance for Employers, and the California Fair Chance Act, for US-based candidates.
- For unincorporated Los Angeles County workers: criminal history may have a direct, adverse and negative relationship with job duties: protect computer hardware from theft, loss or damage; return all computer hardware upon termination; maintain confidentiality of proprietary, confidential, and non-public information. Job duties require access to secure and protected information technology systems and related data security obligations.
- Reasonable accommodations for applicants with disabilities available.