AI Evaluation Specialist
Evaluate AI model outputs for accuracy, safety and helpfulness across domains.
$60-$68/hr
About Turing:
Turing is one of the world’s fastest-growing AI companies accelerating the advancement and deployment of powerful AI systems.
Turing helps customers in two ways: Working with the world’s leading AI labs to advance frontier model capabilities in thinking, reasoning, coding, agentic behavior, multimodality, multilinguality, STEM and frontier knowledge; and leveraging that work to build real-world AI systems that solve mission-critical priorities for companies.
Role Overview:
This position is within a project with one of the foundational LLM companies. The goal is to assist these foundational LLM companies in enhancing their Large Language Models.
One way we help these companies improve their models is by providing them with high-quality proprietary data. This data serves two main purposes: first, as a basis for fine-tuning their models, and second, as an evaluation set to benchmark the performance of their models or competitor models.
For example, in the case of Agent Completion (AC) data generation, your task will be to simulate high-quality multi-turn conversations between a user and a smart assistant that utilizes function-calling tools to accomplish user goals. You will craft these dialogues by playing both the assistant and the user, while simulating tool use where necessary to guide the assistant through complex decision-making and real-world reasoning scenarios.
What does day-to-day look like:
Requirements:
Perks of Freelancing With Turing:
Offer Details:
Location: India, Pakistan, Nigeria, Kenya, Egypt, Ghana, Bangladesh, Turkey, Brazil, Mexico
a:["$","Sourced from Turing · original listing · application link last checked 30 Jul 2026
Evaluate AI model outputs for accuracy, safety and helpfulness across domains.
This is a remote, project-based role for machine learning researchers with deep expertise in mechanistic interpretability.
Evaluate personalization quality in AI systems using cultural and contextual understanding of Japanese user behavior.
Tell us what you know — we'll surface the AI training work that fits.