palantir

Implementer–verifier separation

principleactivereported
slug
implementer-verifier-separation
claim
Implementation and verification are distinct roles. A single pass that both writes and accepts its own output weakens alignment pressure. reported [@vidal2026serious-agentic]
mechanism
reported [@vidal2026serious-agentic] RL-trained models are eager to finish and will hallucinate completion. Self-preference bias: models judge their own writing favorably. Role split: the implementer proposes changes; a verifier in an independent context window owns the goal and grades against criteria (tests, rubrics, secondary model, human). Cross-provider pairing (example cited: Codex implementer + Claude hawk) breaks same-model self-preference. Verifier quality bounds system quality. Who grades matters. Async variants (hawk reading full transcripts post-tool) stop compounding side effects earlier than late PR review. This principle is a harness pattern under harness = agent − model, and a Goal-ownership pattern under Memory–Goal–State.
limits
Operational heuristic from talk. Verifier failure (weak tests, captured judges) still permits reward hacking. Human verifiers remain attention-bounded.

relations

topics

sources

  • vidal2026serious-agentic — role split; self-preference; cross-provider pairing

json

{
  "slug": "implementer-verifier-separation",
  "href": "/entries/implementer-verifier-separation",
  "api": "/api/v1/entries/implementer-verifier-separation",
  "title": "Implementer–verifier separation",
  "type": "principle",
  "status": "active",
  "certainty": "reported",
  "claim": "Implementation and verification are distinct roles. A single pass that both writes and accepts its own output weakens alignment pressure. reported [@vidal2026serious-agentic]",
  "mechanism": "reported [@vidal2026serious-agentic] RL-trained models are eager to finish and will hallucinate completion. Self-preference bias: models judge their own writing favorably.\n\nRole split: the implementer proposes changes; a verifier in an independent context window owns the goal and grades against criteria (tests, rubrics, secondary model, human). Cross-provider pairing (example cited: Codex implementer + Claude hawk) breaks same-model self-preference.\n\nVerifier quality bounds system quality. Who grades matters. Async variants (hawk reading full transcripts post-tool) stop compounding side effects earlier than late PR review.\n\nThis principle is a harness pattern under harness = agent − model, and a Goal-ownership pattern under Memory–Goal–State.",
  "quantities": [],
  "limits": "Operational heuristic from talk. Verifier failure (weak tests, captured judges) still permits reward hacking. Human verifiers remain attention-bounded.",
  "inventor_note": "",
  "sources": [
    {
      "key": "vidal2026serious-agentic",
      "note": "role split; self-preference; cross-provider pairing"
    }
  ],
  "links": [
    "harness-equals-agent-minus-model",
    "memory-goal-state-loop",
    "verifiable-goals-prerequisite",
    "hawk-async-verifier",
    "reward-hacking"
  ],
  "relations": [
    {
      "slug": "harness-equals-agent-minus-model",
      "rel": "related"
    },
    {
      "slug": "memory-goal-state-loop",
      "rel": "related"
    },
    {
      "slug": "verifiable-goals-prerequisite",
      "rel": "related"
    },
    {
      "slug": "hawk-async-verifier",
      "rel": "related"
    },
    {
      "slug": "reward-hacking",
      "rel": "related"
    }
  ],
  "topics": [
    {
      "id": "agentic-engineering",
      "title": "Agentic engineering"
    }
  ],
  "created_at": "2026-10-01T18:51:07.119Z",
  "updated_at": "2026-10-01T18:51:07.176Z"
}