Software Engineering Evaluation Specialist

Mindrift

Apply Now
Canada
$35 - $35 / year
contract
mid
Posted August 30, 2026
via himalayas

About This Role

Please submit your CV in English and indicate your level of English proficiency. Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. Participation is project-based, not permanent employment. About the Role You ll design coding tasks that challenge frontier AI coding agents. Each task is a self-contained Docker environment with a broken piece of software; an AI agent attempts the fix; automated tests verify the outcome. Your deliverable is the full task package: broken code, tests, instructions, and a reference solution proving the task is solvable. Responsibilities: • Invent a realistic developer scenario - a real bug, a broken ETL, a missing feature - not a toy problem. • Build a reproducible Docker environment with pinned dependencies. • Write a pytest that verifies outcomes, not specific commands - deterministic, non-flaky, and does not leak the fix. • Write an instruction.md that reads like a Jira ticket a developer would receive. • Write a reference solve.sh proving the task is solvable. • Calibrate difficulty so current state-of-the-art agents solve the task 20-60% of the time. • Iterate based on feedback from expert QA reviewers. • Later: review other authors tasks as a QA reviewer. Not in scope • Data labeling, prompt engineering. • Production code to ship - you design problems and verification for AI agents. • Leetcode puzzles - scenarios must look like real developer work. • Not every candidate task ships - quality over quantity. Requirements • 3+ years of production software development in one backend stack - Python, Go, Node.js, Java, or Rust. Depth in one stack beats breadth. • Python + pytest fluency - required regardless of primary stack. The task harness is pytest-based even when the broken app is in another language. Fixtures, parametrize, monkeypatch, timeouts, conftest.py. • Docker authoring - reproducible Dockerfiles, pinned dependencies, multi-stage builds when needed, non-root user. • Linux & Bash - comfort debugging inside containers (strace, lsof, journalctl); shell beyond set -euo pipefail. • AI coding agent experience - Claude Code, Cursor, Roo Code, or similar, on non-trivial work. You can cite a specific time the AI was confidently wrong and how you caught it. • English - B2+ written. Not a fit • Data Science, ML, or Computer Vision engineers without backend-engineering output. • Manual QA testers without automation or test authoring. • Frontend-only, low-code / no-code, IT Support, or Business Analysts. • Engineers who have never written pytest from scratch. • Junior, intern, or assistant as the most recent role. Preferred qualifications • Domain depth in Security, System Administration (nginx / systemd / cron), Scientific Computing (NumPy / PyTorch / SciPy), DevOps, or Git internals. • Modern Python tooling (uv, poetry, pyproject.toml). • Coverage tooling (pytest-cov, coverage.py, gcov, llvm-cov, kcov). • Fuzzing or property-based testing (Hypothesis). • Prior contribution to agent-evaluation benchmarks or related frameworks. Process Apply Pass qualification (90-minute sample-task screen + short behavioral interview) Join a project Complete tasks Get paid. Time commitment • Onboarding: ~10 hours per first task. • Steady state: ~5 hours per task, 2-4 parallel tasks per author. • Realistic weekly load: 8-20 hours. Higher volume available for top performers. • You choose when and how to contribute; tasks must be submitted by the deadline and meet acceptance criteria. Compensation: • Paid contributions, rates up to $35/hour*. • Task-based compensation equivalent to hourly rate, depending on performance and volume. • Some projects include incentive payments. *Rates vary based on expertise, skills assessment, location, project needs, and other factors. Higher rates may be provided to highly specialized experts. Lower rates may apply during onboarding or non-core project phases. Payment details are shared per project. Apply Submit your CV via the Mindrift platform. Indicate your English level, note this role (Software Engineering Evaluation Specialist - Terminal Bench), and include a GitHub profile link if available. Originally posted on Himalayas

Ready to Apply?

Click the button below to visit the company's application page.

Apply for this Position