Role Overview
We are seeking experienced US-based Medical Doctors (MD) or international equivalent (MBBS, MBChB, Medical Diploma, etc.) to help evaluate and enhance large language model (LLM) outputs in the medical domain. You’ll assess AI-generated clinical content for medical accuracy, safety, and relevance, contributing to the development of responsible AI in healthcare.
What You'll Do Day-to-Day
- Review AI-generated medical text for accuracy, harmfulness, and completeness, with focus on the medical field.
- Evaluate diagnoses, treatment plans, and summaries for medical cases.
- Provide expert feedback to improve medical integrity and safety of LLM outputs.
- Collaborate with AI researchers and clinicians to align model performance with evidence-based medical standards.
Requirements:
- Based in the US (mandatory).
- Medical Doctor (MD) or international equivalent (MBBS, MBChB, Medical Diploma, etc.)
- Must be fully licensed and practicing medicine within the USA.
- Minimum 1 year post-residency experience practicing as a Medical Doctor / General Physician (no specific specialization required).
- Strong attention to detail
- Technically savvy—able to navigate web-based annotation platforms, use image zoom/tools, follow detailed guidelines, and accurately complete structured questionnaires.
- Experience in medical content review, clinical research, annotation or AI/ML evaluation preferred.
Offer Details:
- Remote
- Commitments Required: At least 4 hours per day and minimum 20 hours per week with overlap of 4 hours with PST.
- Location: United States
Engagement: 12 weeks
Important note: Due to the high volume of applications, we are unable to provide individual feedback at this stage. Only candidates shortlisted for the next steps will be contacted by our team.