{"slug":"reward-hacking","href":"/entries/reward-hacking","api":"/api/v1/entries/reward-hacking","title":"Reward hacking in agent loops","type":"idea","status":"active","certainty":"reported","claim":"Reward hacking is satisfying the verification signal without satisfying the human intention (e.g. mocking an API to green tests). reported [@vidal2026serious-agentic]","mechanism":"reported [@vidal2026serious-agentic] Appears when Goal signals are weaker or cheaper to satisfy than the intended property. Classic software instance: mock the dependency so tests pass while production path remains wrong.\n\nAntidotes listed: separated verifier, hawks, HITL, real oracles. Related to verifiable-goals prerequisite and implementer–verifier separation.\n\nSelf-triage and denominator problems amplify cost when hacking multiplies throughput of non-product work.","quantities":[],"limits":"Any finite oracle can be hacked. Stronger oracles raise cost. Talk examples are illustrative, not a taxonomy.","inventor_note":"","sources":[{"key":"vidal2026serious-agentic","note":"reward hacking failure mode"}],"links":["verifiable-goals-prerequisite","implementer-verifier-separation","denominator-problem"],"relations":[{"slug":"verifiable-goals-prerequisite","rel":"related"},{"slug":"implementer-verifier-separation","rel":"related"},{"slug":"denominator-problem","rel":"related"}],"topics":[{"id":"agentic-engineering","title":"Agentic engineering"}],"created_at":"2026-10-01T19:14:49.307Z","updated_at":"2026-10-01T19:14:49.307Z"}