← Back to jobs

AI training & evaluation

Senior Software Engineer - AI Evaluation / Coding Agents ($100-150/hr)

Turing

We’re looking for experienced, hands-on software engineers to help evaluate and improve AI coding models.

Work arrangement & location
Remote
Compensation
$100-$150/hour
Hours
20–40 hours per week
  • Software development
  • Software Engineering
  • Senior
  • Technical

Employer description and requirements

Freelance · Remote · North America (only)

About Turing

Turing is one of the world’s leading AGI infrastructure companies, working with frontier AI labs to accelerate model development through high-quality training data, evaluations, and engineering talent.

Engagement Details

Compensation: $100-$150/hour, please provide a specific hourly rate expectation

Availability: 20-40 hours/week, with at least 6 hours of Pacific Time overlap/day

Type: Independent contractor

Duration: Approximately 3 months

Start: As soon as possible

Location: North America (Only)

About the Role

We’re looking for experienced, hands-on software engineers to help evaluate and improve AI coding models.

Rather than primarily building production applications, you’ll work with coding agents across real-world repositories and assess the quality of their work. You’ll review generated code and agent behavior, determine whether solutions are technically correct, identify failure modes, and create the evaluation signals and feedback used to improve model performance.

Think of the coding agent as another engineer whose work you’re reviewing: Can it understand the task? Did it choose the right approach? Is the resulting code correct, robust, and maintainable? Can you explain precisely where it succeeded or failed?

What You’ll Do

Design, refine, and iterate on rubrics and evaluation criteria for preference data

Review and label data with a high quality bar, catching subtle issues others might overlook

Build and maintain pipelines and infrastructure supporting data generation, collection, and evaluation workflows

Synthesize findings from data work into clear write-ups, updates, and recommendations for the team

Collaborate closely with researchers and engineers to translate qualitative judgment into scalable processes

What We’re Looking For

5+ years of hands-on software engineering experience

Strong proficiency in Python, TypeScript/JavaScript, Go, or another major production language

Experience working in substantial real-world codebases

Strong code-review skills and technical judgment

Ability to clearly explain why an implementation is correct, incorrect, or could be improved

Strong written communication

Experience using modern LLMs or AI coding tools

Experience with LLM evaluation, coding agents, RLHF, preference data, rubric design, or post-training is a plus, but not required.

Evaluation Process

AI interview (~25 minutes)

Practical code/AI evaluation exercise (~30 minutes)

Hiring manager interview (~20 minutes)

The practical exercise focuses on your ability to review and evaluate AI-generated code, not competitive programming or algorithm puzzles.