Verifiable goals as alignment prerequisite
principleactivereported
slug
verifiable-goals-prerequisite
claim
Alignment requires goals that are verifiable (objective metric) or pseudo-verifiable (judge model / rubric). Unverifiable goals cannot be aligned reliably. reported [@vidal2026serious-agentic]
mechanism
reported [@vidal2026serious-agentic] Principle I in the talk: you cannot align what you cannot verify.
Verifiable goals: measurable metrics, automated tests, typecheckers, schema validators, CLI oracles, visual diffs, hardware smoke signals. Pseudo-verifiable goals: separate judge models or rubrics when no hard oracle exists.
Software crossed the agentic threshold early because tests and types provide oracles (talk cites Antje Barth / AIEWF). Domains without oracles need constructed pseudo-verification before agent leverage scales.
Universal loop pieces Memory / Goal / State presuppose Goal is checkable. Reward hacking appears when the signal is weaker than the intent (e.g. mocking an API to green tests). Antidotes named in the talk: separated verifier, hawks, HITL, real oracles.
Denominator problem related: multiplying code without a product-level denominator (users, outcomes) yields a bill, not alignment.
limits
Pseudo-verification inherits judge bias and reward-hacking risk. Talk-level principle, not a formal completeness theorem.
relations
- related agentic-alignment-problem
- related memory-goal-state-loop
- related implementer-verifier-separation
- related reward-hacking
- related denominator-problem
topics
- agentic-engineering Agentic engineering
sources
- vidal2026serious-agentic — verification prerequisite; Barth AIEWF attribution for software threshold
json
{
"slug": "verifiable-goals-prerequisite",
"href": "/entries/verifiable-goals-prerequisite",
"api": "/api/v1/entries/verifiable-goals-prerequisite",
"title": "Verifiable goals as alignment prerequisite",
"type": "principle",
"status": "active",
"certainty": "reported",
"claim": "Alignment requires goals that are verifiable (objective metric) or pseudo-verifiable (judge model / rubric). Unverifiable goals cannot be aligned reliably. reported [@vidal2026serious-agentic]",
"mechanism": "reported [@vidal2026serious-agentic] Principle I in the talk: you cannot align what you cannot verify.\n\nVerifiable goals: measurable metrics, automated tests, typecheckers, schema validators, CLI oracles, visual diffs, hardware smoke signals. Pseudo-verifiable goals: separate judge models or rubrics when no hard oracle exists.\n\nSoftware crossed the agentic threshold early because tests and types provide oracles (talk cites Antje Barth / AIEWF). Domains without oracles need constructed pseudo-verification before agent leverage scales.\n\nUniversal loop pieces Memory / Goal / State presuppose Goal is checkable. Reward hacking appears when the signal is weaker than the intent (e.g. mocking an API to green tests). Antidotes named in the talk: separated verifier, hawks, HITL, real oracles.\n\nDenominator problem related: multiplying code without a product-level denominator (users, outcomes) yields a bill, not alignment.",
"quantities": [],
"limits": "Pseudo-verification inherits judge bias and reward-hacking risk. Talk-level principle, not a formal completeness theorem.",
"inventor_note": "",
"sources": [
{
"key": "vidal2026serious-agentic",
"note": "verification prerequisite; Barth AIEWF attribution for software threshold"
}
],
"links": [
"agentic-alignment-problem",
"memory-goal-state-loop",
"implementer-verifier-separation",
"reward-hacking",
"denominator-problem"
],
"relations": [
{
"slug": "agentic-alignment-problem",
"rel": "related"
},
{
"slug": "memory-goal-state-loop",
"rel": "related"
},
{
"slug": "implementer-verifier-separation",
"rel": "related"
},
{
"slug": "reward-hacking",
"rel": "related"
},
{
"slug": "denominator-problem",
"rel": "related"
}
],
"topics": [
{
"id": "agentic-engineering",
"title": "Agentic engineering"
}
],
"created_at": "2026-10-01T18:51:07.119Z",
"updated_at": "2026-10-01T18:51:07.176Z"
}