About this role
Role Overview
Apply your professional judgment to a structured, one time evaluation of AI generated work in a real business domain. You will complete a realistic task using provided source files, then assess how five AI model outputs handled that same assignment.
Key Responsibilities
- Sign a non disclosure agreement before receiving study materials.
- Complete one realistic, domain specific task using the provided source files.
- Review five model generated responses to that task, rank them from 1 through 5 without ties, assign each a score from 0 to 100, and provide evidence based rationales that reference specific output details.
- Provide structured feedback on the task and the study design.
Qualifications
- At least 5 years of hands on professional experience in Admin or Business Operations, Marketing, Human Resources, or Accounting.
- Currently practicing or recently active in your field and able to evaluate work from a working professional''s perspective.
- Ability to write clear, specific, evidence grounded explanations for your judgments.
- Reliability in meeting a short turnaround window.
Work Terms
- Remote, hourly engagement.
- One time evaluation pilot, not ongoing production work.
- Commitment of up to 13 hours total, including 3 to 10 hours to complete the task, about 2 hours to rank and score five outputs, and about 1 hour to provide feedback.
- No rubric or model answer will be provided. Your professional judgment is the basis of the evaluation.
- Use of LLMs or other AI assistants is prohibited while completing the task, reviewing outputs, scoring, ranking, or writing rationales.
Compensation
$60 to $80 per hour.
Eligibility
You are not eligible if you have worked on Project Alchemy in any capacity, including as a task author, reviewer, world expert, or anyone with access to its source world data. Previous exposure to those materials would invalidate the study.