Mercor
$60–90/hr
103 people hired
You will serve as the quality backbone for agentic AI evaluation benchmarks, ensuring complex multi-step tasks are clear, correctly graded, and resistant to shortcuts. | The role involves running benchmark tasks, probing edge cases, reviewing reference solutions, and debugging technical environments. | You will work closely with AI researchers and task authors to improve benchmark reliability and evaluation quality. | This is a full-time W-2 role, fully remote within the United States.
AI Benchmark QA Engineer