AI training & evaluation
Health/Biology Expert
Turing
We are seeking an AI Training & Evaluation Specialist (Biology/Health) to design, curate, and review advanced biological and health science assessment tasks to train and evaluate state-of-the-art AI models.
- Work arrangement & location
- Remote
- Hours
- 40 hours per week
- Science
- Science Research
- Biology
- biology
Employer description and requirements
About Turing
Turing is one of the world’s fastest-growing AI companies, accelerating the advancement and deployment of powerful AI systems.
Turing helps customers in two ways: Working with the world’s leading AI labs to advance frontier model capabilities in thinking, reasoning, coding, agentic behavior, multimodality, multilinguality, STEM, and frontier knowledge; and leveraging that work to build real-world AI systems that solve mission-critical priorities for companies.
Role Overview
We are seeking an AI Training & Evaluation Specialist (Biology/Health) to design, curate, and review advanced biological and health science assessment tasks to train and evaluate state-of-the-art AI models. This dual-focus role involves two key work streams: authoring complex, text-only biology problems from graduate-level concepts and reviewing domain-specific question-answering tasks for scientific rigor and quality. You will play a vital role in identifying failure modes in frontier AI systems, ensuring generated outputs meet high academic and clinical standards.
Requirements
Education: Master’s degree or PhD in Biology, Health Sciences, Medicine, or a closely related biological field.
Domain Expertise: Strong mastery of upper-undergraduate and graduate-level biological systems, health concepts, and clinical or research-based reasoning.
Writing & Scientific Translation: Fluent written English with the ability to write precise, rigorous scientific explanations and convert visual/diagram-heavy concepts into self-contained, text-only problems.
Evaluation Skills: Ability to critically navigate domain topics, judge the quality and accuracy of complex questions/answers, diagnose AI failure modes, and refine tasks to meet high difficulty targets.
Availability & Commitment: Talent must have weekend on-call availability (part-time engagement is acceptable).
Technical Infrastructure: Personal desktop/laptop equipped with a stable, high-speed internet connection in a remote setup.
Responsibilities
Problem Design & Curation: Author and curate upper-undergraduate and graduate-level Biology problems with clear, fully worked solutions, converting diagram-heavy or multi-part questions into text-only, solvable tasks.
Quality Review & Audit: Review, evaluate, and judge biology and health question-answering tasks created by others or generated by AI models, ensuring scientific correctness, terminology precision, and clarity.
AI Model Evaluation: Diagnose AI-generated errors, expose model failure modes, and refine questions to calibrate target difficulty levels.
Collaborative Quality Control: Collaborate with cross-functional teams to maintain consistent task standards across large datasets and peer-review tasks prior to final delivery.
Education & Experience
Prior experience in teaching, grading, university-level exam creation, or research-based analytical work required.
Familiarity with AI training data best practices is preferred; 1+ month of prior experience on Turing AI evaluation projects is a strong plus (not mandatory).
Offer Details:
Commitments Required: at least 4 hours per day and upto 40 hours per week with 4 hours of overlap with PST.
Engagement type: Contractor
Engagement Length: 4 weeks
Evaluation Process -
Shortlisted candidates will be sent a Job Interest Form.
Final selected candidates will be contacted with the next steps.