palantir

Verifiable goals as alignment prerequisite

principleactivereported
slug
verifiable-goals-prerequisite
claim
Alignment requires goals that are verifiable (objective metric) or pseudo-verifiable (judge model / rubric). Unverifiable goals cannot be aligned reliably. reported [@vidal2026serious-agentic]
mechanism
reported [@vidal2026serious-agentic] Principle I in the talk: you cannot align what you cannot verify. Verifiable goals: measurable metrics, automated tests, typecheckers, schema validators, CLI oracles, visual diffs, hardware smoke signals. Pseudo-verifiable goals: separate judge models or rubrics when no hard oracle exists. Software crossed the agentic threshold early because tests and types provide oracles (talk cites Antje Barth / AIEWF). Domains without oracles need constructed pseudo-verification before agent leverage scales. Universal loop pieces Memory / Goal / State presuppose Goal is checkable. Reward hacking appears when the signal is weaker than the intent (e.g. mocking an API to green tests). Antidotes named in the talk: separated verifier, hawks, HITL, real oracles. Denominator problem related: multiplying code without a product-level denominator (users, outcomes) yields a bill, not alignment.
limits
Pseudo-verification inherits judge bias and reward-hacking risk. Talk-level principle, not a formal completeness theorem.

relations

topics

sources

  • vidal2026serious-agentic — verification prerequisite; Barth AIEWF attribution for software threshold

json

{
  "slug": "verifiable-goals-prerequisite",
  "href": "/entries/verifiable-goals-prerequisite",
  "api": "/api/v1/entries/verifiable-goals-prerequisite",
  "title": "Verifiable goals as alignment prerequisite",
  "type": "principle",
  "status": "active",
  "certainty": "reported",
  "claim": "Alignment requires goals that are verifiable (objective metric) or pseudo-verifiable (judge model / rubric). Unverifiable goals cannot be aligned reliably. reported [@vidal2026serious-agentic]",
  "mechanism": "reported [@vidal2026serious-agentic] Principle I in the talk: you cannot align what you cannot verify.\n\nVerifiable goals: measurable metrics, automated tests, typecheckers, schema validators, CLI oracles, visual diffs, hardware smoke signals. Pseudo-verifiable goals: separate judge models or rubrics when no hard oracle exists.\n\nSoftware crossed the agentic threshold early because tests and types provide oracles (talk cites Antje Barth / AIEWF). Domains without oracles need constructed pseudo-verification before agent leverage scales.\n\nUniversal loop pieces Memory / Goal / State presuppose Goal is checkable. Reward hacking appears when the signal is weaker than the intent (e.g. mocking an API to green tests). Antidotes named in the talk: separated verifier, hawks, HITL, real oracles.\n\nDenominator problem related: multiplying code without a product-level denominator (users, outcomes) yields a bill, not alignment.",
  "quantities": [],
  "limits": "Pseudo-verification inherits judge bias and reward-hacking risk. Talk-level principle, not a formal completeness theorem.",
  "inventor_note": "",
  "sources": [
    {
      "key": "vidal2026serious-agentic",
      "note": "verification prerequisite; Barth AIEWF attribution for software threshold"
    }
  ],
  "links": [
    "agentic-alignment-problem",
    "memory-goal-state-loop",
    "implementer-verifier-separation",
    "reward-hacking",
    "denominator-problem"
  ],
  "relations": [
    {
      "slug": "agentic-alignment-problem",
      "rel": "related"
    },
    {
      "slug": "memory-goal-state-loop",
      "rel": "related"
    },
    {
      "slug": "implementer-verifier-separation",
      "rel": "related"
    },
    {
      "slug": "reward-hacking",
      "rel": "related"
    },
    {
      "slug": "denominator-problem",
      "rel": "related"
    }
  ],
  "topics": [
    {
      "id": "agentic-engineering",
      "title": "Agentic engineering"
    }
  ],
  "created_at": "2026-10-01T18:51:07.119Z",
  "updated_at": "2026-10-01T18:51:07.176Z"
}