About this role
This role evaluates how well AI systems and people perform real-world digital product design work by defining excellence rather than producing final designs. You will create task-specific grading rubrics for common product design deliverables, then score AI-generated and human work samples with clear, evidence-based written justifications that are reproducible and defensible.
Key Responsibilities- Design precise, task-specific grading criteria for real-world digital product and UX deliverables, including design systems and component libraries, responsive UI layouts and dashboards, end-to-end product flows such as onboarding, checkout, and booking, logged-in customer portals, UX writing and product content, and design-to-engineering or production handoff.
- Score AI-generated and human work samples against those criteria, providing a detailed written justification for every score.
- Apply consistent, evidence-based judgment so scores are reproducible and defensible.
- Receive and incorporate structured feedback from senior reviewers, iterating quickly on rubric design and scoring.
- Minimum 5 years of professional experience in product design, UI/UX, or design systems for real digital products, not brand or marketing sites.
- Experience at leading digital-product and UX studios, in-house product design teams at major product-led companies, or e-commerce and marketplace businesses with sophisticated customer portals and dashboards.
- Relevant background may include roles such as Product Designer, UI Designer, Design Systems Designer, UX Lead, UX Writer or Content Designer, or Product or Business Analyst focused on digital products.
- Deep fluency in day-to-day craft: interaction and visual design, design systems, responsive layouts, end-to-end product flows, prototyping and production handoff in Figma, and the ability to articulate critique at the level of a design or UX lead.
- Exceptionally strong written communication, able to explain precisely why a piece of work does or does not meet the bar for a real product and its users.
- Detail-oriented, consistent, and comfortable having your judgment reviewed and calibrated against peers.
- Prior experience with AI training, evaluation, or human-data projects is a strong plus.
To apply, submit your resume and portfolio or other evidence of relevant professional experience. Qualified candidates may be asked to complete a brief assessment or provide additional materials to demonstrate evaluation approach and judgment.
Work Terms- Location: Remote.
- Engagement model: Hourly.
- Pay range: 80 to 150 hourly.