Mercor
$60–90/hr
61 people hired
Machine Learning Engineer responsible for designing complex, multi-step ML evaluation tasks and conducting end-to-end experiments to benchmark frontier AI models. The role involves translating open-ended ML research ideas into reproducible tasks, implementing model or training changes, running experiments, analyzing results, and identifying where advanced AI models fail. You will work closely with AI researchers to create rigorous ground-truth solutions and agentic evaluation benchmarks across machine learning and reinforcement learning workflows.
Machine Learning Experimentation, AI Model Evaluation, Reinforcement Learning Research