Ground every answer in facts on this page and the original listing. We never invent Glassdoor-style reviews or salaries that are not in our data.
Be ready to walk through how you score agent quality, handle edge cases, and document disagreements with automated metrics. Ask how tasks are assigned, reviewed, and paid for freelance evaluation work.
Best fit if you want freelance agent-eval work with Mindrift and can operate carefully against written criteria. Thin public company data means clarify scope, tools, and pay structure before committing.
As an Agent Evaluation Engineer (Freelance) at Mindrift, a typical day would center on running structured checks on AI agent behavior: scoring outputs against rubrics, logging failure modes, and feeding concise findings back to the engagement lead—not shipping core product code.
Useful skills for this title include clear rubric writing, careful prompt/response review, reproducible issue notes, and comfort judging agent reliability. No certification path is specified for this listing.
The title is Agent Evaluation Engineer (Freelance), so expect contract evaluation work more than permanent product ownership.
The listing shows Canada; confirm remote vs on-site and timezone expectations with the recruiter.
No certification family or resources were provided for this listing.
Website: mindrift.com
Public cache only — not an employee review.
FREELANCE AGENT EVALUATION ENGINEER The Opportunity Mindrift is seeking a Freelance Agent Evaluation Engineer to join our team in Canada. We are focused on testing, evaluating, and improving AI systems by building datasets that challenge advanced AI coding agents. This role involves creating realistic developer environments, designing tasks within these environments, writing tests for agent solutions, and iterating based on feedback. Day-to-Day Responsibilities Build realistic developer environments with a virtual company's codebase, infrastructure, and context (tickets, docs, conversations) Create challenging tasks from intermediate states of these environments, crafting the prompt, defining "solved," and ensuring solvability by an AI agent Write tests that verify agent solutions, accepting all valid approaches while rejecting incorrect ones without being too strict or lenient Iterate on tasks and tests based on quality assurance feedback, reviewing agent solutions, analyzing failures, and refining until the evaluation is fair and robust Requirements 5+ years in software development Core stack: Python (FastAPI), JavaScript/TypeScript (React), Docker, Postgres, Kafka, Redis Experience writing tests (functional, integration) English proficiency - B2+ Working at Mindrift At Mindrift, we connect specialists with project-based AI opportunities for leading tech companies. Participation is not permanent employment but involves building a dataset to evaluate AI coding agents' performance in real-world developer tasks. Apply Today To apply, complete your application directly on this page, or you'll be redirected to the employer's application platform to finish submitting there.
Generated for personal interview prep · 2026-08-13 UTC · getajob.ai