← Back to jobs

AI training & evaluation

Web Research Specialist

Turing

We are building an evaluation benchmark for frontier AI browsing agents.

Work arrangement & location
Remote
Hours
40 hours per week
  • Science
  • Science Research
  • Research & analysis

Employer description and requirements

About Turing:

Turing is one of the world’s fastest-growing AI companies, accelerating the advancement and deployment of powerful AI systems.

Turing helps customers in two ways: Working with the world’s leading AI labs to advance frontier model capabilities in thinking, reasoning, coding, agentic behavior, multimodality, multilinguality, STEM and frontier knowledge; and leveraging that work to build real-world AI systems that solve mission-critical priorities for companies.

About the Role

We are building an evaluation benchmark for frontier AI browsing agents. Your job is to design research problems that a state-of-the-art AI cannot solve, even with full web access and multiple attempts. This is not a subject matter expert role, nor is it a content-writing role. It is investigative research.

You will start from a verifiable fact, work backwards to construct a question that makes that fact extremely hard to locate, and then prove your work with a complete, auditable evidence trail.

What You Will Produce

A natural-language research question with a short, stable, objectively verifiable answer

Clues, each independently checkable, spanning multiple fact types - including dates, people, places, organisations, works, events, records, and quantities - with specific constraints

A validation record showing the obvious searches you ran and the results they returned

Minimum Qualifications

Master's OR >3 YOE.

Demonstrated open-web research ability, including locating primary records and navigating government and institutional databases, archives, registries, and PDF documents

Precision with sourcing. You cite exact pages, tables, and sections—not just homepages

Comfort researching unfamiliar subjects from scratch

Native or near-native written English

High tolerance for structured documentation. The evidence trail is the majority of the work

Experience with LLM evaluation, red-teaming, or benchmark construction

Experience in one or more of the following domains:

Reference librarianship, archival research, or special collections

Investigative journalism or professional fact-checking

OSINT, due diligence, KYC, or investigative research

Patent, prior-art, or legal-discovery search

Genealogy and records research

Competitive quizzing or puzzle-hunt construction

Nice to Have

Familiarity with JSON and structured data delivery formats

Project Snapshot

Current task mix provides estimated earning potential of approximately $30 per approved task.

8 week project

Ideally upto 40 hours/week

Fully remote | Paid in USD

Start as soon as you successfully pass the assessment

What’s Next?

Submit your application and complete the required assessment. Candidates who successfully pass the assessment will be eligible to proceed with onboarding.