home/productivity/evaluator-prompt-v1

Evaluator Prompt v1

GPTClaudeDeepSeek··258 copies·updated 2026-07-14
evaluator-prompt-v1.prompt
# Evaluator Prompt v1

## Purpose

Score an agent, prompt, skill, or model profile against a fixed rubric and representative case, then record repeatable, auditable results.

## Instructions

1. Fix the evaluation setup (artifact, case, rubric, model profile, prompt version) and do not change it mid-run.
2. Produce or collect the output without modifying the artifact under test.
3. Score each rubric criterion with a justification tied to evidence in the output.
4. Separate observations from recommendations.
5. Record every result, including failures, to `evals/results/` and a run record.
6. Recommend acceptance, revision, or rejection with specific next actions.

## Required Output Sections

- Evaluation Setup
- Scores (per rubric criterion, with evidence)
- Findings (severity, observation, recommendation)
- Verdict
- Run Record Reference

## Known Failure Modes

- Editing the rubric after seeing results to make a failing artifact pass.
- Scoring without citing evidence from the output.
- Hiding or omitting failed results.
- Claiming one model or prompt is better without comparable runs.

when to use it

Community prompt sourced from the open-source GitHub repo 3rdAI-admin/th3rdai-harness (MIT). A "Evaluator Prompt v1" style prompt — adapt the placeholders and specifics to your task. Imported as-is and not independently retested here, so check the output before relying on it.

tags

productivitycommunitydeveloper

source

3rdAI-admin/th3rdai-harness · MIT