Responsibilities:
- Validate task quality: Check that instructions, source materials, reference solutions, and evaluation criteria are consistent, with no hidden requirements or missing information.
- Review agent performance: Inspect execution traces, tool calls, and generated deliverables to determine whether successes and failures are justified.
- Audit grading logic: Identify brittle checks, incorrect expected answers, unsupported rubric criteria, and cases where valid alternative solutions are unfairly penalized.
- Investigate discrepancies: Distinguish genuine model limitations from task defects, grader errors, and environment or tool failures. Assess automated QC findings independently rather than accepting them at face value.
- Document decisions: Provide concise, evidence-backed findings and actionable feedback, flag uncertainty, and verify that revisions resolve identified issues.
What we're looking for:
- Technical fluency: Comfortable reading Python, SQL, shell scripts, structured data, and execution logs to understand task setup and grading behavior.
- Analytical judgment: Able to independently check calculations, reconcile conflicting evidence, and assess the correctness and completeness of professional deliverables.
- Clear communication: Strong written English, attention to detail, and experience providing specific, reproducible feedback.
- Relevant experience preferred: AI evaluation, technical QA, data analysis, or benchmark development; familiarity with Harbor task setup.
Offer Details:
- Remote
- Full-time dedication - 40 hours per week with 8 hours PST overlap per day
- Location: LATAM
- Engagement Length: 10 weeks
Important note: Due to the high volume of applications, we are unable to provide individual feedback at this stage. Only candidates shortlisted for the next steps will be contacted by our team.