AI Vulnerability Expert - Fully Remote | Upto $90/hr

mercor

Apply Now
United States
$60 - $90 / year
full-time
mid
Posted August 11, 2026
via himalayas

About This Role

About the job Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors include Benchmark, General Catalyst, Peter Thiel, Adam D'Angelo, Larry Summers, and Jack Dorsey. Position: LLM Red Team Specialist - Failure Modes & Edge Cases Type:Contract Compensation:$60-$90/hour Location:Remote Commitment:35 hours/week Role Responsibilities • Evaluate frontier AI models on coding, ML, and analysis tasks to identify vulnerabilities and failure modes. • Design complex tasks that challenge models and are fair for grading. • Document findings with clear evidence and reproducible steps. • Collaborate with task authors to close loopholes and improve grading. • Share insights with researchers to enhance benchmark quality. • Work independently and asynchronously to meet deadlines and improve AI model performance. Qualifications Must-Have • MSc or PhD in a STEM field or equivalent experience. • 1+ years in research, research-engineering, security, or AI evaluation. • Experience identifying vulnerabilities in LLMs or ML systems. • Proficiency in Python and Git. • Familiarity with LLM capabilities and evaluation techniques. • Ability to work 35 hours/week. Preferred • Experience in AI training, model evaluation, or benchmark/task authoring. Resources & Support • For details about the interview process and platform information, please check: • For any help or support, reach out to: PS: Our team reviews applications daily. Please complete your AI interview and application steps to be considered for this opportunity. Originally posted on Himalayas

Ready to Apply?

Click the button below to visit the company's application page.

Apply for this Position