Print / Save PDF Back to job

Agent Evaluation Engineer (Freelance)

Mindrift · Canada

How to use this kit

Ground every answer in facts on this page and the original listing. We never invent Glassdoor-style reviews or salaries that are not in our data.

Interview prep

Be ready to walk through how you score agent quality, handle edge cases, and document disagreements with automated metrics. Ask how tasks are assigned, reviewed, and paid for freelance evaluation work.

Fit summary

Best fit if you want freelance agent-eval work with Mindrift and can operate carefully against written criteria. Thin public company data means clarify scope, tools, and pay structure before committing.

Day in the role

As an Agent Evaluation Engineer (Freelance) at Mindrift, a typical day would center on running structured checks on AI agent behavior: scoring outputs against rubrics, logging failure modes, and feeding concise findings back to the engagement lead—not shipping core product code.

Skills to emphasize

Useful skills for this title include clear rubric writing, careful prompt/response review, reproducible issue notes, and comfort judging agent reliability. No certification path is specified for this listing.

FAQ from this listing

Is this a full-time product engineering role?

The title is Agent Evaluation Engineer (Freelance), so expect contract evaluation work more than permanent product ownership.

Where is the role based?

The listing shows Canada; confirm remote vs on-site and timezone expectations with the recruiter.

Are certifications required?

No certification family or resources were provided for this listing.

Company facts (cached)

Website: mindrift.com

Public cache only — not an employee review.

Role overview (listing rewrite)

FREELANCE AGENT EVALUATION ENGINEER The Opportunity Mindrift is seeking a Freelance Agent Evaluation Engineer to join our team in Canada. We are focused on testing, evaluating, and improving AI systems by building datasets that challenge advanced AI coding agents. This role involves creating realistic developer environments, designing tasks within these environments, writing tests for agent solutions, and iterating based on feedback. Day-to-Day Responsibilities Build realistic developer environments with a virtual company's codebase, infrastructure, and context (tickets, docs, conversations) Create challenging tasks from intermediate states of these environments, crafting the prompt, defining "solved," and ensuring solvability by an AI agent Write tests that verify agent solutions, accepting all valid approaches while rejecting incorrect ones without being too strict or lenient Iterate on tasks and tests based on quality assurance feedback, reviewing agent solutions, analyzing failures, and refining until the evaluation is fair and robust Requirements 5+ years in software development Core stack: Python (FastAPI), JavaScript/TypeScript (React), Docker, Postgres, Kafka, Redis Experience writing tests (functional, integration) English proficiency - B2+ Working at Mindrift At Mindrift, we connect specialists with project-based AI opportunities for leading tech companies. Participation is not permanent employment but involves building a dataset to evaluate AI coding agents' performance in real-world developer tasks. Apply Today To apply, complete your application directly on this page, or you'll be redirected to the employer's application platform to finish submitting there.

Full job on Get A Job.AI

Questions to ask them

Generated for personal interview prep · 2026-08-13 UTC · getajob.ai