About this role
Role Overview
Evaluate how AI systems handle personalized, real-world life tasks and help improve their usefulness for everyday decision-making. Your judgment will inform AI assistance across food, health, productivity, careers, and learning, with a focus on context, preferences, constraints, tradeoffs, and practical outcomes.
Key Responsibilities
- Assess AI responses to personal, high-context tasks for usefulness, personalization, realism, safety, completeness, and likelihood of success.
- Apply experience with multi-step planning, research, decision-making, and personal workflows to evaluate AI performance.
- Provide clear written judgment and detailed feedback that helps make AI assistants more trustworthy and effective.
Qualifications
- Extensive personal use of large language model products.
- Experience using AI for multi-step tasks, planning, research, decision-making, or personal workflows.
- Familiarity with tools such as ChatGPT, Claude, Gemini, Perplexity, Cursor, Windsurf, Codex, or other AI agents.
- Ability to explain why an AI output is strong, weak, incomplete, unsafe, or unrealistic.
- Strong written judgment and attention to detail.
Work Terms
- Remote, hourly engagement.
- Expected commitment of 20 to 40 hours per week.
Compensation
- $50 to $200 per hour.