Senior Software Engineer – LLM Evaluation
Evaluate AI-generated code and improve performance, reliability, and scalability of LLM outputs.
Review AI-written code, grade agent traces, and build evaluation benchmarks — the deepest project pool on the platform.
Evaluate AI-generated code and improve performance, reliability, and scalability of LLM outputs.
Build enterprise-grade generative AI systems using knowledge graphs, LLMs, and scalable architectures for production environments.
UK-based software engineers needed to evaluate and generate AI coding tasks, solutions, and code reviews.
Provide engineering and platform expertise for AI training datasets.
Design and deploy machine learning systems with Python and ETL pipelines.
Label and annotate cybersecurity data to train AI models on threat detection, vulnerabilities, and security analysis.
Coding AI training jobs are the largest category on PARA AI Labs: reviewing AI-generated code for correctness, grading coding-agent trajectories, writing benchmark problems, and hardening evaluation pipelines across Python, TypeScript, C++, Rust, and more.
They're remote developer jobs paid by the hour — from generalist code review up to CUDA, MLOps, and performance engineering at senior rates. Side-project friendly: no meetings, no standups.
One profile covers every partner platform — most people are on a paid coding project within a week.
Get matchedCoding AI training jobs are the largest category on PARA AI Labs: reviewing AI-generated code for correctness, grading coding-agent trajectories, writing benchmark problems, and hardening evaluation pipelines across Python, TypeScript, C++, Rust, and more.
Live coding projects typically pay $25–120/hr, with every rate posted upfront on the listing. Difficulty, seniority, and specialist credentials push rates toward the top of the band.
Yes — every project is remote and asynchronous. Work from anywhere, pick your own hours, no meetings. Apply links go straight to the hiring platform with no middlemen and no fees.
Most roles want working proficiency in at least one mainstream language plus the judgment to spot subtle bugs. Senior and specialist tracks (CUDA, MLOps, SRE, security) list higher bars — and pay for them.