Implementer–verifier separation
principleactivereported
slug
implementer-verifier-separation
claim
Implementation and verification are distinct roles. A single pass that both writes and accepts its own output weakens alignment pressure. reported [@vidal2026serious-agentic]
mechanism
reported [@vidal2026serious-agentic] RL-trained models are eager to finish and will hallucinate completion. Self-preference bias: models judge their own writing favorably.
Role split: the implementer proposes changes; a verifier in an independent context window owns the goal and grades against criteria (tests, rubrics, secondary model, human). Cross-provider pairing (example cited: Codex implementer + Claude hawk) breaks same-model self-preference.
Verifier quality bounds system quality. Who grades matters. Async variants (hawk reading full transcripts post-tool) stop compounding side effects earlier than late PR review.
This principle is a harness pattern under harness = agent − model, and a Goal-ownership pattern under Memory–Goal–State.
limits
Operational heuristic from talk. Verifier failure (weak tests, captured judges) still permits reward hacking. Human verifiers remain attention-bounded.
relations
- related harness-equals-agent-minus-model
- related memory-goal-state-loop
- related verifiable-goals-prerequisite
- related hawk-async-verifier
- related reward-hacking
topics
- agentic-engineering Agentic engineering
sources
- vidal2026serious-agentic — role split; self-preference; cross-provider pairing
json
{
"slug": "implementer-verifier-separation",
"href": "/entries/implementer-verifier-separation",
"api": "/api/v1/entries/implementer-verifier-separation",
"title": "Implementer–verifier separation",
"type": "principle",
"status": "active",
"certainty": "reported",
"claim": "Implementation and verification are distinct roles. A single pass that both writes and accepts its own output weakens alignment pressure. reported [@vidal2026serious-agentic]",
"mechanism": "reported [@vidal2026serious-agentic] RL-trained models are eager to finish and will hallucinate completion. Self-preference bias: models judge their own writing favorably.\n\nRole split: the implementer proposes changes; a verifier in an independent context window owns the goal and grades against criteria (tests, rubrics, secondary model, human). Cross-provider pairing (example cited: Codex implementer + Claude hawk) breaks same-model self-preference.\n\nVerifier quality bounds system quality. Who grades matters. Async variants (hawk reading full transcripts post-tool) stop compounding side effects earlier than late PR review.\n\nThis principle is a harness pattern under harness = agent − model, and a Goal-ownership pattern under Memory–Goal–State.",
"quantities": [],
"limits": "Operational heuristic from talk. Verifier failure (weak tests, captured judges) still permits reward hacking. Human verifiers remain attention-bounded.",
"inventor_note": "",
"sources": [
{
"key": "vidal2026serious-agentic",
"note": "role split; self-preference; cross-provider pairing"
}
],
"links": [
"harness-equals-agent-minus-model",
"memory-goal-state-loop",
"verifiable-goals-prerequisite",
"hawk-async-verifier",
"reward-hacking"
],
"relations": [
{
"slug": "harness-equals-agent-minus-model",
"rel": "related"
},
{
"slug": "memory-goal-state-loop",
"rel": "related"
},
{
"slug": "verifiable-goals-prerequisite",
"rel": "related"
},
{
"slug": "hawk-async-verifier",
"rel": "related"
},
{
"slug": "reward-hacking",
"rel": "related"
}
],
"topics": [
{
"id": "agentic-engineering",
"title": "Agentic engineering"
}
],
"created_at": "2026-10-01T18:51:07.119Z",
"updated_at": "2026-10-01T18:51:07.176Z"
}