LLM Trainer - Agent Function Call
Design and craft multi-turn conversational datasets to improve AI agent reasoning, function calling accuracy, and real-world interaction capabilities.
$60–$70/hour
We are seeking experienced AI Safety Practitioners to evaluate the safety, quality, and alignment of frontier AI models across complex, policy-sensitive, and ambiguous ("grey-area") topics. You will assess AI-generated responses, apply safety policies, and help improve model behavior through structured evaluations and feedback.
Sourced from AIUC via Mercor · original listing · application link last checked 4 Aug 2026
Design and craft multi-turn conversational datasets to improve AI agent reasoning, function calling accuracy, and real-world interaction capabilities.
Evaluate AI model outputs for accuracy, safety and helpfulness across domains.
This is a remote, project-based role for machine learning researchers with deep expertise in mechanistic interpretability.
Tell us what you know — we'll surface the AI training work that fits.