Evaluator Prompt v1
# Evaluator Prompt v1
## Purpose
Score an agent, prompt, skill, or model profile against a fixed rubric and representative case, then record repeatable, auditable results.
## Instructions
1. Fix the evaluation setup (artifact, case, rubric, model profile, prompt version) and do not change it mid-run.
2. Produce or collect the output without modifying the artifact under test.
3. Score each rubric criterion with a justification tied to evidence in the output.
4. Separate observations from recommendations.
5. Record every result, including failures, to `evals/results/` and a run record.
6. Recommend acceptance, revision, or rejection with specific next actions.
## Required Output Sections
- Evaluation Setup
- Scores (per rubric criterion, with evidence)
- Findings (severity, observation, recommendation)
- Verdict
- Run Record Reference
## Known Failure Modes
- Editing the rubric after seeing results to make a failing artifact pass.
- Scoring without citing evidence from the output.
- Hiding or omitting failed results.
- Claiming one model or prompt is better without comparable runs.when to use it
Community prompt sourced from the open-source GitHub repo 3rdAI-admin/th3rdai-harness (MIT). A "Evaluator Prompt v1" style prompt — adapt the placeholders and specifics to your task. Imported as-is and not independently retested here, so check the output before relying on it.
tags
productivitycommunitydeveloper
source
3rdAI-admin/th3rdai-harness · MIT
more in Productivity
Productivity✓ tested
Summarize a doc into decisions & actions
chief of staff who extracts what to DO, not just what was said
Productivity✓ tested
Draft a reply to a hard email
calm, direct communicator who de-escalates without caving
Productivity✓ tested
Turn a brain-dump into a weekly plan
planning coach who protects your focus, not just your calendar