Responsibilities
- Design and harden multi-step, tool-using agent loops behind current and new agentic experiences
- Turn research and experiments into reliable, observable, and affordable products at enterprise scale
- Bring judgment about where evaluations earn their keep
- Use targeted smoke tests and metrics to avoid noise dressed up as rigor
- Decide which models to run where
- Drive model upgrades
- Fine-tune models when appropriate
- Push on how to ground models in a customer's code through retrieval, ranking, context windows, and citations
- Treat cost and latency as product features
- Profile, distill, cache, and right-size models to ship ambitious features sustainably
- Talk to customers, frame problems, and own them end-to-end
Requirements
- Senior engineer and technical leader with skills in agent engineering combining software engineering, machine learning, and statistics
- Real agent engineering depth: built, trained, evaluated, and operated models in production
- Ability to reason about datasets, evals, baselines, metrics, and error analysis
- Experience designing multi-step agentic systems and making them reliable, observable, and cost-bounded
- Pragmatic about evaluations and conscious of cost
- Autonomous operation on ambiguous problems
- Ownership of high-technical-risk projects end-to-end
- Contribution beyond immediate domain: go into any codebase, recognize cross-cut issues
- Translate between engineering goals and business objectives
- Mentor others by pairing, code/design reviews, and spreading agent engineering literacy
- Customer and product-driven: comfortable on customer calls, turning feedback into requirements
- Pragmatic, not perfectionist: ship smallest correct thing, prefer robust over complicated, maintain high quality with simplicity
- Strong software engineer capable of shipping production services
- Comfortable with Go (backend), TypeScript (frontend), GraphQL, Postgres, Docker
- Fluent with agentic coding tools and own every line they submit
- Comfortable in async-first, multi-service, fast-paced remote environment
Nice to Have
- Shipped an LLM-powered or agentic developer-facing product and can speak opinionatedly about what worked, what didn’t, and what you’d do differently
- Built evaluation systems for LLM or ML products
- Fine-tuned, distilled, or trained models to meet cost, latency, or quality targets in production
- Experience with retrieval, ranking, embeddings, or search relevance
- Experience working directly with enterprise customers and translating their needs into a product
- Experience mentoring or up-leveling engineers, especially raising a team's agent engineering fluency
Work Arrangement
Remote (Worldwide) — Europe, North America
Additional Information
- Working hours must overlap with EST for at least 20 hours/week
- Interview process includes: Recruiter Screen (30m), Hiring Manager Screen / Resume Deep Dive (45m), two 60-minute Technical Interviews, Cross-functional team collaboration / Values (60m), Leadership interview (30m)
- Reference check and background check are conducted
- Open to candidates almost anywhere in the world
- No relocation assistance mentioned
- No probation period mentioned
- No travel requirement mentioned
- No equipment requirement mentioned
- No language requirement beyond English implied
- No security clearance mentioned