AI training & evaluation
Pod lead - Civil and structural engineering
Turing
Experience developing or evaluating AI coding agents, terminal-based agents, or LLM-generated engineering solutions.
- Work arrangement & location
- Remote
Track this job
- Hours
- 40 hours per week
Employer description and requirements
About Turing:
Turing is one of the world’s fastest-growing AI companies, accelerating the advancement and deployment of powerful AI systems.
Turing helps customers in two ways: Working with the world’s leading AI labs to advance frontier model capabilities in thinking, reasoning, coding, agentic behavior, multimodality, multilinguality, STEM and frontier knowledge; and leveraging that work to build real-world AI systems that solve mission-critical priorities for companies.
Role Overview:
We are seeking an experienced Civil or Structural Engineering expert to lead a pod of trainers building realistic, terminal-based technical tasks for Terminal Bench Science. You will review every task your trainers produce for engineering correctness, reproducibility, and grader reliability, and coach trainers to consistently deliver high-quality tasks.
Reporting Structure:
Reports to the Civil & Structural Engineering Team Lead. Leads a pod of approximately 5–10 trainers.
What you'll do:
Review engineering tasks end to end, including problem statements, structural and geotechnical models, load cases, simulation data, computational environments, reference solutions, and automated tests.
Validate tasks involving structural analysis, finite-element methods (FEA), earthquake and wind engineering, structural dynamics, geotechnical modeling, hydraulics, hydrology, transportation, and infrastructure systems.
Ensure tasks require genuine multi-step engineering reasoning and reflect realistic civil and structural engineering workflows.
Verify engineering accuracy across units, equilibrium, boundary conditions, load combinations, design codes, numerical stability, convergence, and safety factors.
Validate computational models, solvers, dependencies, and simulation results for reproducibility and technical correctness.
Review automated graders to ensure they evaluate meaningful engineering outputs such as forces, deflections, drifts, settlements, factors of safety, flow rates, and structural performance.
Identify technical errors, unrealistic assumptions, incorrect tolerances, and opportunities to bypass engineering analysis.
Provide clear, actionable feedback to task developers and track revisions through completion.
Mentor and support technical trainers, allocate work, monitor quality and productivity, and resolve technical blockers.
Maintain quality standards, technical documentation, and best practices across engineering task development.
What we're looking for:
Ph.D., postdoctoral experience, or equivalent advanced technical experience in Civil Engineering, Structural Engineering, Geotechnical Engineering, or a closely related discipline.
Strong programming experience in Python, C/C++, Julia, Fortran, or MATLAB/Octave, with proficiency in Linux environments.
Hands-on expertise in at least one major area: structural analysis, FEA, earthquake engineering, geotechnical modeling, or hydraulic/hydrological modeling.
Strong understanding of engineering principles, numerical methods, simulation workflows, and model validation.
Working knowledge of engineering design standards such as ASCE 7, ACI, AISC, Eurocodes, or IS codes.
Experience reviewing complex technical work, engineering simulations, research outputs, or computational models.
Strong analytical judgment with the ability to identify technical inaccuracies, edge cases, and flawed engineering assumptions.
Experience mentoring, reviewing, or leading small technical teams.
Excellent written communication skills for providing precise technical feedback and documenting quality standards.
Nice to have:
Experience with engineering simulation and analysis tools such as OpenSees/OpenSeesPy, Code_Aster, CalculiX, FEniCS, Gmsh, PyNite, EPANET, SWMM, HEC-RAS, MODFLOW, SUMO, QGIS, or GDAL.
Familiarity with commercial engineering software such as SAP2000, ETABS, Abaqus, PLAXIS, or STAAD.Pro.
Experience with nonlinear time-history analysis, performance-based design, reliability analysis, or structural health monitoring.
Familiarity with Docker, Git, CI/CD pipelines, and automated testing.
Experience developing or evaluating AI coding agents, terminal-based agents, or LLM-generated engineering solutions.
Professional engineering licensure such as PE, SE, CEng, or equivalent.
Industry experience working on buildings, bridges, transportation networks, water systems, or major infrastructure projects.
Prior experience in AI training, technical quality control, benchmark development, or evaluation programs.
Offer Details:
Commitments Required: At least 4 hours per day and minimum 40 hours per week with overlap of 4 hours with PST.
Employment type : Contractor assignment (no medical/paid leave)
Duration of contract : 4 week [expected start date is next week]