About this role
Join a GenAI team building foundational large language models by applying deep insurance domain expertise to create, evaluate, and refine training data. You will translate real-world underwriting, claims, and risk assessment judgment into domain-relevant tasks, solutions, and evaluation criteria that improve model reasoning and reliability.
Key Responsibilities- Work with research and engineering teams to close knowledge gaps in underwriting, claims, and risk-assessment reasoning.
- Design challenging, domain-relevant insurance tasks and produce accurate, well-reasoned solutions grounded in real underwriting and claims practice.
- Evaluate AI model outputs against structured rubrics, providing clear written feedback on correctness, judgment, and quality of reasoning.
- Develop and refine evaluation guidelines and scoring rubrics tailored to insurance tasks.
- Collaborate with other subject-matter experts to ensure consistency and accuracy in training and evaluation data.
- At least 8 years of professional experience in insurance, such as underwriting, claims, actuarial, or risk management, at a recognized organization, for example AIG, Chubb, Allstate, Progressive, MetLife, Marsh McLennan, or equivalent.
- Prior hands-on experience evaluating LLM or AI model outputs against rubrics or structured scoring criteria, mandatory. Please describe this experience in your application.
- Demonstrable career progression within insurance roles, for example Underwriter to Senior Underwriter to VP of Underwriting.
- Strong verbal and written communication skills, problem-solving ability, and effective interpersonal skills for cross-functional collaboration.
- Ability to engage reliably for at least 35 hours per week on weekdays.
- Employment type, W-2 hourly position through a staffing company that acts as the employer of record, placing employees on client teams.
- Location, United States; assignment will be with a leading AI research lab as part of their extended workforce.
- This is a W-2 employment relationship, with payroll, benefits, and compliance handled by the employer of record.
- Pay range, $60 to $80 per hour.
- Must be legally authorized to work in the United States for a W-2 employer of record. Specific sponsorship policies are not stated here, please confirm if you require sponsorship before applying.
- The employer is an equal opportunity employer and does not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, gender identity, age, veteran status, disability, genetic information, political views, or other legally protected characteristics.
When you apply, include a clear description of your hands-on experience evaluating LLM or AI model outputs against rubrics or structured scoring criteria, as this is required. Applications that detail your relevant years of experience, employer examples, and career progression will be prioritized for review.