Mercor
$60–90/hr
65 people hired
QA/Test Engineer responsible for ensuring complex AI evaluation benchmark tasks are accurate, unambiguous, correctly graded, and resistant to shortcuts. The role involves deeply reviewing technical tasks and reference solutions, executing workflows, designing comprehensive test cases, investigating edge cases, debugging environments, and improving quality-review processes. You will collaborate with AI researchers and task authors to ensure benchmark results reliably measure frontier AI model performance.
QA & Test Engineering, Test Case Development, Python System Debugging