Published project

Frontier Model Reasoning & Eval Specialist

$110–150/hr · Independent contractor

Role overview

About the Role

We are seeking domain specialists to evaluate frontier AI reasoning capabilities through rigorous trajectory inspection, verification of formal mathematical proofs, and authoring adversarial test cases.

Key Responsibilities

  • Evaluate step-by-step reasoning traces in code generation and symbolic logic.
  • Grade fidelity against ground-truth specifications.
  • Design red-teaming benchmarks exposing boundary failure modes.

Requirements

  • Advanced expertise in software engineering or computational mathematics.
  • Strong background in algorithmic problem solving.
  • Excellent technical communication and analytical rigor.