← Back to AI Training Jobs

AI training & evaluation

GenAI Security Evaluation Engineer (Up to $150/hr)

Turing

Experience with OWASP LLM security risks, MCP, threat modeling, or security evaluation is a plus.

Work arrangement & location
Remote
Compensation
up to $150/hr

Track this job

My Jobs
  • Cloud & Infrastructure
  • AI training & evaluation
Hours
20 hours per week

Employer description and requirements

About Turing

Based in San Francisco, California, Turing is the world’s leading research accelerator for frontier AI labs and a trusted partner for global enterprises deploying advanced AI systems. Turing supports customers in two ways: first, by accelerating frontier research with high-quality data, advanced training pipelines, plus top AI researchers who specialize in coding, reasoning, STEM, multilinguality, multimodality, and agents; and second, by applying that expertise to help enterprises transform AI from proof of concept into proprietary intelligence with systems that perform reliably, deliver measurable impact, and drive lasting results on the P&L.

Offer Details

Required commitment: At least 4 hours per day, minimum 20 hours per week, with 4 hours of overlap with Pacific Time.

Engagement type: Contractor

Engagement length: Up to 2 weeks

Pay up to $150/hr depending on internal evaluation and location.

About the Role

We’re looking for experienced security engineers to evaluate how effectively static analysis tools detect vulnerabilities in GenAI applications.

You’ll build small, runnable agent and RAG codebases containing realistic examples of Sensitive Information Disclosure and Excessive Agency, along with fixed and near-miss versions. You’ll then trace, annotate, test, and explain each finding.

What You’ll Do

Build runnable agent/RAG repositories with code-reachable Sensitive Information Disclosure or Excessive Agency vulnerabilities across tool calling, memory, and MCP.

Create vulnerable, fixed, and hard-negative variants with minimal security-relevant differences.

Trace and annotate assets, data/action paths, controls, root causes, severity, and residual risk.

Define authorization contexts and write deterministic tests validating vulnerable, fixed, and negative behavior.

Recommend security controls and participate in calibration and peer review.

What We’re Looking For

5+ years in application/product security or security-focused software engineering, including secure code review.

Experience with source-to-sink analysis, taint analysis, SAST, CodeQL, or Semgrep.

Hands-on experience building LLM agents or RAG systems using frameworks such as LangChain, LlamaIndex, OpenAI/Anthropic SDKs, or MCP.

Strong authorization knowledge, including actors, trust boundaries, tenants, OAuth, IAM, identity, permitted data/actions, purposes, recipients, and document-level access control.

Production coding experience in Python and/or TypeScript, with strong judgment in distinguishing genuine SID/EA findings from non-findings.

Experience with OWASP LLM security risks, MCP, threat modeling, or security evaluation is a plus.

Evaluation Process

AI interview (~25 minutes)

Resume and overall application review

Offer