
In Progress
Posted
AI Agent Evaluator Remote · Contract · Flexible Hours We're hiring evaluators to help assess and benchmark cutting-edge AI agents. This is hands-on work at the frontier of agentic AI — not prompt writing, not basic QA. You will be designing complex real-world tasks, running them across multiple large language models, and producing the structured evaluations that shape how the next generation of AI agents are trained. What You'll Do You will design multi-stage agent tasks across real-world domains including health, finance, productivity, relationships, and personal exploration. Each task must require the agent to coordinate across multiple tools and systems, manage persistent state, handle realistic friction points, and produce a verifiable final artifact. Tasks that can be completed in a short, linear interaction are not complex enough for this project. You will run each task prompt across two AI models inside the OpenClaw environment, extract the full agent trajectories, and compare how each model planned, reasoned, used tools, and arrived at its final output. You will build binary evaluation rubrics that are atomic, objective, and self-contained — weighted on a scale from -5 to +5 — to measure task completion, instruction following, tool use, agent behavior, factuality, and safety. At least one negative-weight criterion is required per rubric set. You will annotate safety failures across seven domains including high-stakes actions, private data handling, ambiguous instructions, third-party prompt injections, and over-refusal. Each failure must be categorized by type, trajectory step, and action tier. You will write pytest-based unit tests in a [login to view URL] file that validate the agent's final system state using [login to view URL] as ground truth — confirming outcomes, not just intent. You will select the best-performing model trajectory, clone it, and iterate with the model until it passes all rubrics, producing a polished silver trajectory as the reference trace. You Must Be Able To Design complex multi-stage agentic tasks with real friction, constraint logic, and verifiable outcomes Write evaluation rubrics that are atomic, self-contained, and expose meaningful differences between models Annotate safety failures with the correct category from a defined failure taxonomy and assign the correct action tier Read and reason critically about full agent trajectories Write and validate basic Python unit tests using pytest Source every task idea from a real, publicly available online post about OpenClaw — no invented scenarios Strong Background In AI evaluation, rubric-based assessment, data annotation, agentic AI systems, AI safety concepts, prompt engineering, Python and pytest, and structured critical thinking. Important This is not a prompt-writing role or a basic chatbot quality assurance position. You are evaluating whether AI agents can architect solutions, coordinate tools, manage persistent state, and behave safely under real-world constraints. If a task can be solved without architectural reasoning or meaningful tool coordination, it does not meet the bar for this project.
Project ID: 40512159
3 proposals
Remote project
Active 6 days ago
Set your budget and timeframe
Get paid for your work
Outline your proposal
It's free to sign up and bid on jobs

In today's era of AI, the necessity to implement robust and meticulous evaluation systems should not be underestimated. With my strong background in AI evaluation and rubric-based assessments, I am confident in my ability to design complex multi-stage tasks and construct rigorous evaluation rubrics that will expose valuable insights into model performance for the cutting-edge AI agents. Through carefully analyzing and reasoning about agent trajectories, I effectively extract information to make informed evaluations. Moreover, Python proficiency and experience with pytest have enabled me to design and validate useful unit tests, ensuring an accurate assessment of the AI agents' system state. On top of my technical abilities is my commitment to structured critical thinking and attention to detail, which are crucial qualities in this field. I assure you that each task on the project will be thoughtfully sourced from real online posts about OpenClaw as specified.
$5 USD in 40 days
0.0
0.0
3 freelancers are bidding on average $8 USD/hour for this job

I understand that you are seeking top-tier AI data annotation and evaluation specialists to assess complex AI agents in real-world scenarios. With over 12 years of experience in AI evaluation, data annotation, and Python programming, I am well-equipped to tackle the challenges outlined in your project. My expertise includes designing intricate multi-stage tasks that simulate realistic friction points, as well as crafting binary evaluation rubrics that provide clear differentiation between model performances. I have a strong background in using tools like pytest for unit testing and can effectively manage agent trajectories while ensuring safety compliance across multiple domains. Furthermore, I excel at sourcing authentic task ideas from publicly available resources to ensure relevance and complexity. Collaborating closely with your team, I will produce structured evaluations that shape the future of agentic AI. Could you clarify how often you would like to iterate on the model trajectories during the evaluation process?
$8 USD in 7 days
0.0
0.0

atlanta, United States
Payment method verified
Member since Oct 24, 2019
$10-30 USD
$2-8 USD / hour
$2-8 USD / hour
$10-30 USD
$2-8 USD / hour
$40 USD
₹12500-37500 INR
₹600-1500 INR
₹600-1500 INR
₹1500-12500 INR
$2-8 USD / hour
$750-1500 USD
min $50 USD / hour
$25 USD
₹600-1500 INR
$10-30 USD
$250-750 USD
£250-750 GBP
$48 USD
₹750-1250 INR / hour
₹600-601 INR
$10-30 USD
₹12500-37500 INR
$10-30 USD
₹75000-150000 INR