{"topics":[{"id":"agentic-engineering","href":"/topics/agentic-engineering","api":"/api/v1/topics/agentic-engineering","title":"Agentic engineering","parent_id":null,"summary":"Agent-built systems; harness design; alignment between intent and artifact under verification.","entries":["agentic-alignment-problem","context-poisoning","denominator-problem","harness-equals-agent-minus-model","hawk-async-verifier","hitl-escape-hatch","implementer-verifier-separation","memory-goal-state-loop","ralph-task-file-loop","reward-hacking","rollback-plus-learnings","scratchpad-durable-memory","verifiable-goals-prerequisite"],"created_at":"2026-10-01T19:14:49.295Z"},{"id":"micro-robotics","href":"/topics/micro-robotics","api":"/api/v1/topics/micro-robotics","title":"Micro-robotics","parent_id":null,"summary":"Insect-scale locomotion, actuators, energy pathways, and catalytic artificial muscles.","entries":["2020-robeetle-catalytic-muscle","methanol-catalytic-combustion-on-pt"],"created_at":"2026-10-01T19:14:49.295Z"}],"entries":[{"slug":"2020-robeetle-catalytic-muscle","href":"/entries/2020-robeetle-catalytic-muscle","api":"/api/v1/entries/2020-robeetle-catalytic-muscle","title":"RoBeetle 2020 — 88 mg methanol-catalytic crawling microrobot","type":"device","status":"active","certainty":"measured","claim":"RoBeetle is an insect-scale crawler (empty mass 88 mg) that instantiates methanol catalytic combustion on Pt coupled to a Pt-coated NiTi wire actuator. Contraction drives forelegs and closes the fuel-tank lid. measured [@yang2020robeetle]","mechanism":"1. Energy store: onboard liquid methanol tank (~120 µl class); ambient evaporation feeds vapor through dorsal ports. reported [@yang2020robeetle]\n2. Actuator: NiTi SMA wire diameter 50.8 µm with platinum powder coating as catalyst surface. measured [@yang2020robeetle]\n3. Heat from catalytic combustion drives martensite→austenite contraction (As 87–99 °C). measured [@yang2020robeetle]\n4. Transmission: contraction displaces a leaf spring / linkage to the forelegs and simultaneously closes the tank lid, modulating further vapor delivery. reported [@yang2020robeetle]\n5. Reset: cooling restores wire length via return spring; lid reopens; cycle repeats. reported [@yang2020robeetle]\n6. Locomotion: two-anchor crawl with anisotropic friction. reported [@yang2020robeetle]\n7. Mass budget: empty ≈ 88 mg; fueled ≈ 183 mg; fuel ≈ 95 mg. Length ≈ 20 mm. measured/reported [@yang2020robeetle] [@ieee2020robeetle]\n8. Payload: up to ~2.6× empty mass reported. Stride ~1.2 mm. measured/reported [@yang2020robeetle] [@ieee2020robeetle]\n9. No onboard electronics in the demonstrated crawler; control is thermo-mechanical via the catalytic/SMA coupling. reported [@yang2020robeetle]\n10. This device demonstrates methanol-catalytic-combustion-on-pt; it is not the principle itself.","quantities":[{"name":"empty_mass","unit":"mg","value":"88","source":"yang2020robeetle","certainty":"measured"},{"name":"fueled_mass","unit":"mg","value":"183","source":"yang2020robeetle","certainty":"measured"},{"name":"fuel_mass","unit":"mg","value":"~95","source":"yang2020robeetle","certainty":"measured"},{"name":"length","unit":"mm","value":"~20","source":"ieee2020robeetle","certainty":"reported"},{"name":"niti_wire_diameter","unit":"µm","value":"50.8","source":"yang2020robeetle","certainty":"measured"},{"name":"max_payload","unit":"× empty mass","value":"~2.6","source":"yang2020robeetle","certainty":"measured"},{"name":"stride","unit":"mm","value":"~1.2","source":"ieee2020robeetle","certainty":"reported"},{"name":"tank_volume","unit":"µl","value":"~120","source":"yang2020robeetle","certainty":"reported"}],"limits":"Chemical→work efficiency ~0.48%. Cycle limited by evaporation and catalyst fouling. Methanol toxic. Wire ≈ 90–100 °C. No onboard electronics. Secondary IEEE summary may invert SMA heating clause relative to yang2020robeetle — prefer primary.","inventor_note":"","sources":[{"key":"yang2020robeetle","note":"primary source"},{"key":"ieee2020robeetle","note":"secondary; check SMA clause against primary"},{"key":"usc2020robeetle","note":"institutional note"}],"links":["methanol-catalytic-combustion-on-pt"],"relations":[{"slug":"methanol-catalytic-combustion-on-pt","rel":"demonstrates"}],"topics":[{"id":"micro-robotics","title":"Micro-robotics"}],"created_at":"2026-10-01T19:14:49.242Z","updated_at":"2026-10-01T19:14:49.307Z"},{"slug":"denominator-problem","href":"/entries/denominator-problem","api":"/api/v1/entries/denominator-problem","title":"Denominator problem","type":"idea","status":"active","certainty":"reported","claim":"The denominator problem is reporting productivity numerators (more code, more tokens) without a product-level denominator (users, outcomes), which yields spend rather than alignment. reported [@vidal2026serious-agentic]","mechanism":"reported [@vidal2026serious-agentic] Talk warning: 8× more code — of what? Numerator without denominator is a bill.\n\nProduct denominators: users, outcomes, closed feedback loops. Cognitive debt is related: code growing faster than human understanding (−17% comprehension via AI-code cited via Osmani/AIEWF in the talk materials).\n\nOrganizational correlates: over-parallelization causing operator burnout (velocity sickness); backlog loss removing filters that previously blocked work that should not ship; feature Frankenstein under relocated smaller teams.","quantities":[],"limits":"Talk-level organizational framing. The −17% figure is attributed via secondary citation in talk materials, not independently measured here.","inventor_note":"","sources":[{"key":"vidal2026serious-agentic","note":"denominator problem; cognitive debt citation path"}],"links":["agentic-alignment-problem","verifiable-goals-prerequisite","reward-hacking"],"relations":[{"slug":"agentic-alignment-problem","rel":"related"},{"slug":"verifiable-goals-prerequisite","rel":"related"},{"slug":"reward-hacking","rel":"related"}],"topics":[{"id":"agentic-engineering","title":"Agentic engineering"}],"created_at":"2026-10-01T19:14:49.307Z","updated_at":"2026-10-01T19:14:49.307Z"},{"slug":"methanol-catalytic-combustion-on-pt","href":"/entries/methanol-catalytic-combustion-on-pt","api":"/api/v1/entries/methanol-catalytic-combustion-on-pt","title":"Methanol catalytic combustion on platinum","type":"principle","status":"active","certainty":"measured","claim":"Flame-less catalytic oxidation of CH3OH(g) on Pt releases heat usable as an actuation energy pathway at insect scale. Atmospheric O2 is the oxidant. Measured reaction enthalpy ΔH = −676.49 kJ mol⁻¹. measured [@yang2020robeetle]","mechanism":"1. Liquid methanol evaporates at ambient temperature; vapor contacts a Pt-coated surface. reported [@yang2020robeetle]\n2. Stoichiometry: CH3OH(g) + 3/2 O2(g) → 2 H2O(g) + CO2(g). Oxidation proceeds as flame-less catalytic combustion on platinum rather than open-flame combustion. measured [@yang2020robeetle]\n3. Released heat raises a shape-memory alloy (NiTi) wire through martensite→austenite, producing contractile work against a return spring / transmission. measured [@yang2020robeetle]\n4. Austenite start temperature for the actuation wire is reported in the 87–99 °C band; operating wire temperature during actuation ≈ 90–100 °C. measured [@yang2020robeetle]\n5. Methanol specific energy ≈ 20 MJ kg⁻¹ is cited as the chemical energy density supporting the pathway. reported [@yang2020robeetle]\n6. Chemical-to-wire-heat efficiency ≈ 16%; system-to-mechanical-work efficiency ≈ 0.48%. Most heat dissipates to ambient air. inferred/reported [@yang2020robeetle]\n7. Cycle timing is coupled to evaporation rate, catalyst condition, and cooling. Catalyst fouling and methanol toxicity are operational constraints. reported [@yang2020robeetle]\n8. The principle is distinct from any single vehicle: RoBeetle 2020 is one demonstrating device instance. inferred [@yang2020robeetle]","quantities":[{"name":"reaction_enthalpy","unit":"kJ mol⁻¹","value":"−676.49","source":"yang2020robeetle","certainty":"measured"},{"name":"methanol_specific_energy","unit":"MJ kg⁻¹","value":"20","source":"yang2020robeetle","certainty":"reported"},{"name":"chemical_to_wire_heat_efficiency","unit":"%","value":"~16","source":"yang2020robeetle","certainty":"inferred"},{"name":"system_to_work_efficiency","unit":"%","value":"~0.48","source":"yang2020robeetle","certainty":"reported"},{"name":"austenite_start","unit":"°C","value":"87–99","source":"yang2020robeetle","certainty":"measured"},{"name":"actuation_wire_temperature","unit":"°C","value":"≈90–100","source":"yang2020robeetle","certainty":"reported"}],"limits":"System-to-work efficiency ~0.48%; majority of heat lost to air. Cycle period bounded by evaporation and catalyst fouling. Methanol toxicity. Hot wire (~90–100 °C). No claim of electrical-free sensing/compute on the principle alone.","inventor_note":"","sources":[{"key":"yang2020robeetle","note":"primary Sci Robotics source for reaction, efficiencies, SMA temperatures"},{"key":"ieee2020robeetle","note":"secondary summary; SMA clause may invert vs primary"},{"key":"usc2020robeetle","note":"institutional note"}],"links":["2020-robeetle-catalytic-muscle"],"relations":[{"slug":"2020-robeetle-catalytic-muscle","rel":"related"}],"topics":[{"id":"micro-robotics","title":"Micro-robotics"}],"created_at":"2026-10-01T19:14:49.242Z","updated_at":"2026-10-01T19:14:49.307Z"},{"slug":"reward-hacking","href":"/entries/reward-hacking","api":"/api/v1/entries/reward-hacking","title":"Reward hacking in agent loops","type":"idea","status":"active","certainty":"reported","claim":"Reward hacking is satisfying the verification signal without satisfying the human intention (e.g. mocking an API to green tests). reported [@vidal2026serious-agentic]","mechanism":"reported [@vidal2026serious-agentic] Appears when Goal signals are weaker or cheaper to satisfy than the intended property. Classic software instance: mock the dependency so tests pass while production path remains wrong.\n\nAntidotes listed: separated verifier, hawks, HITL, real oracles. Related to verifiable-goals prerequisite and implementer–verifier separation.\n\nSelf-triage and denominator problems amplify cost when hacking multiplies throughput of non-product work.","quantities":[],"limits":"Any finite oracle can be hacked. Stronger oracles raise cost. Talk examples are illustrative, not a taxonomy.","inventor_note":"","sources":[{"key":"vidal2026serious-agentic","note":"reward hacking failure mode"}],"links":["verifiable-goals-prerequisite","implementer-verifier-separation","denominator-problem"],"relations":[{"slug":"verifiable-goals-prerequisite","rel":"related"},{"slug":"implementer-verifier-separation","rel":"related"},{"slug":"denominator-problem","rel":"related"}],"topics":[{"id":"agentic-engineering","title":"Agentic engineering"}],"created_at":"2026-10-01T19:14:49.307Z","updated_at":"2026-10-01T19:14:49.307Z"},{"slug":"rollback-plus-learnings","href":"/entries/rollback-plus-learnings","api":"/api/v1/entries/rollback-plus-learnings","title":"Rollback plus learnings","type":"principle","status":"active","certainty":"reported","claim":"On poisoned or failed agent trajectories, discard the contaminated context and retry clean while retaining only scar documents (learnings). reported [@vidal2026serious-agentic]","mechanism":"reported [@vidal2026serious-agentic] Talk formulation: don't carry the context; carry only the scars.\n\nBroken branch → write learnings.md with causal notes → clean retry. Prevents try/catch blankets and skipped tests from becoming durable house style.\n\nComplements scratchpads and Ralph loops: Memory is curated, not append-only garbage.","quantities":[],"limits":"Scar quality matters. Vague learnings do not prevent recurrence. Requires discipline to actually discard poisoned branches.","inventor_note":"","sources":[{"key":"vidal2026serious-agentic","note":"rollback + learnings antidote to poisoning"}],"links":["context-poisoning","scratchpad-durable-memory","ralph-task-file-loop"],"relations":[{"slug":"context-poisoning","rel":"related"},{"slug":"scratchpad-durable-memory","rel":"related"},{"slug":"ralph-task-file-loop","rel":"related"}],"topics":[{"id":"agentic-engineering","title":"Agentic engineering"}],"created_at":"2026-10-01T19:14:49.307Z","updated_at":"2026-10-01T19:14:49.307Z"},{"slug":"context-poisoning","href":"/entries/context-poisoning","api":"/api/v1/entries/context-poisoning","title":"Context poisoning","type":"idea","status":"active","certainty":"reported","claim":"Context poisoning is persistence and replication of bad patterns across compactions and new sessions, turning local hacks into house style. reported [@vidal2026serious-agentic]","mechanism":"reported [@vidal2026serious-agentic] Agents copy what they see. A @deprecated utility with many call sites outcompetes a deprecation tag. Wrapping try/catch and disabled tests become house style and survive compaction.\n\nMisalignment compounds: one bad utility becomes style. Landmines (hacks where redesign was required) and avoidance of refactors are related failure modes.\n\nAntidote named in the talk: rollback + learnings — discard contaminated context; retain only scars. Early detection (ast-grep, hawks, continuous debt hunting) limits blast radius.\n\nPrinciple II: catch misalignment early because it composes.","quantities":[],"limits":"Detection rules can themselves be gamed. Learnings files that encode the bad pattern without the negation reintroduce poison.","inventor_note":"","sources":[{"key":"vidal2026serious-agentic","note":"context poisoning; compounding misalignment"}],"links":["rollback-plus-learnings","agentic-alignment-problem","reward-hacking","hawk-async-verifier"],"relations":[{"slug":"rollback-plus-learnings","rel":"related"},{"slug":"agentic-alignment-problem","rel":"related"},{"slug":"reward-hacking","rel":"related"},{"slug":"hawk-async-verifier","rel":"related"}],"topics":[{"id":"agentic-engineering","title":"Agentic engineering"}],"created_at":"2026-10-01T19:14:49.307Z","updated_at":"2026-10-01T19:14:49.307Z"},{"slug":"hitl-escape-hatch","href":"/entries/hitl-escape-hatch","api":"/api/v1/entries/hitl-escape-hatch","title":"Human-in-the-loop escape hatch","type":"principle","status":"active","certainty":"reported","claim":"HITL tools give RL-trained agents an explicit ask-human exit when confidence is low, preventing forced low-quality completion of the turn. reported [@vidal2026serious-agentic]","mechanism":"reported [@vidal2026serious-agentic] Models trained to exhaust the turn will produce something. An ask-human tool (explicit in prompt or implicit by availability) is an escape hatch.\n\nFits triage ladders: auto-commit safe / verify-before-prod / co-design / interrupt human. Self-triage is unreliable; defaulting to the comfortable path produces million-token bills.\n\nHITL is a harness permission surface, not a model property.","quantities":[],"limits":"Overuse recreates human attention bottleneck. Underuse recreates reward hacking. Triage policy must be externalized.","inventor_note":"","sources":[{"key":"vidal2026serious-agentic","note":"HITL tools; triage ladder"}],"links":["harness-equals-agent-minus-model","hawk-async-verifier","agentic-alignment-problem"],"relations":[{"slug":"harness-equals-agent-minus-model","rel":"related"},{"slug":"hawk-async-verifier","rel":"related"},{"slug":"agentic-alignment-problem","rel":"related"}],"topics":[{"id":"agentic-engineering","title":"Agentic engineering"}],"created_at":"2026-10-01T19:14:49.307Z","updated_at":"2026-10-01T19:14:49.307Z"},{"slug":"hawk-async-verifier","href":"/entries/hawk-async-verifier","api":"/api/v1/entries/hawk-async-verifier","title":"Hawk asynchronous verifier","type":"principle","status":"active","certainty":"reported","claim":"A hawk is an async verifier that reads the full agent transcript after tool use and emits CONTINUE / STOP / ESCALATE before side effects compound. reported [@vidal2026serious-agentic]","mechanism":"reported [@vidal2026serious-agentic] Named after Karpathy's watch-them-like-a-hawk framing. Implementation sketch: PostToolUse hook reads the complete transcript and returns CONTINUE, STOP, or ESCALATE.\n\nHawk is a verifier with independent judgment timing relative to the implementer. It intervenes earlier than late PR review, reducing compounding misalignment.\n\nCross-provider hawks reduce self-preference bias. Hawks are harness components under harness = agent − model.","quantities":[],"limits":"Hawk quality bounded by its rubric and context. False STOP slows throughput; false CONTINUE misses poisoning. Cost scales with transcript volume.","inventor_note":"","sources":[{"key":"vidal2026serious-agentic","note":"hawk async verifier; Karpathy attribution"}],"links":["implementer-verifier-separation","harness-equals-agent-minus-model","context-poisoning","hitl-escape-hatch"],"relations":[{"slug":"implementer-verifier-separation","rel":"related"},{"slug":"harness-equals-agent-minus-model","rel":"related"},{"slug":"context-poisoning","rel":"related"},{"slug":"hitl-escape-hatch","rel":"related"}],"topics":[{"id":"agentic-engineering","title":"Agentic engineering"}],"created_at":"2026-10-01T19:14:49.307Z","updated_at":"2026-10-01T19:14:49.307Z"},{"slug":"scratchpad-durable-memory","href":"/entries/scratchpad-durable-memory","api":"/api/v1/entries/scratchpad-durable-memory","title":"Scratchpad durable agent memory","type":"principle","status":"active","certainty":"reported","claim":"A scratchpad is a living document of agent Progress / Decisions / Learnings / Action log that survives context compaction and is read first each session. reported [@vidal2026serious-agentic]","mechanism":"reported [@vidal2026serious-agentic] Scratchpads formalize durable Memory. Codex ExecPlans are cited as a related formalization.\n\nContents typically include progress, decisions, learnings, and an action log. The artifact lives in the repository so new sessions and compacted contexts reload state without relying on ephemeral chat history.\n\nPaired with task files, scratchpads separate narrative Memory from checkbox State. Rollback+learnings workflows write scars into durable files while discarding poisoned conversational context.","quantities":[],"limits":"A scratchpad that records bad patterns without pruning becomes a poisoning vector. Needs hygiene and optional hawk review.","inventor_note":"","sources":[{"key":"vidal2026serious-agentic","note":"scratchpad; ExecPlans mention"}],"links":["memory-goal-state-loop","ralph-task-file-loop","context-poisoning","rollback-plus-learnings"],"relations":[{"slug":"memory-goal-state-loop","rel":"related"},{"slug":"ralph-task-file-loop","rel":"related"},{"slug":"context-poisoning","rel":"related"},{"slug":"rollback-plus-learnings","rel":"related"}],"topics":[{"id":"agentic-engineering","title":"Agentic engineering"}],"created_at":"2026-10-01T19:14:49.307Z","updated_at":"2026-10-01T19:14:49.307Z"},{"slug":"ralph-task-file-loop","href":"/entries/ralph-task-file-loop","api":"/api/v1/entries/ralph-task-file-loop","title":"Ralph / task-file agent loop","type":"principle","status":"active","certainty":"reported","claim":"A Ralph loop repeatedly invokes an agent while unchecked tasks remain in a durable task file, pushing against early stopping. reported [@vidal2026serious-agentic]","mechanism":"reported [@vidal2026serious-agentic] Canonical sketch: while grep finds unchecked boxes in TASKS.md, invoke the agent with a fixed prompt. Anthropic surface: /ralph-loop.\n\nMapping to Memory–Goal–State: Memory = repo + scratchpad; Goal = checkboxes driven to zero (plus any oracles those tasks encode); State = the task file itself.\n\nPurpose: counteract RL-trained early stopping. The loop externalizes continuation pressure into the filesystem so compaction and new sessions still see unfinished work.\n\nVariants include per-surface boards (e.g. firmware vs mobile) and status markers beyond [ ]/[x].","quantities":[],"limits":"Without verifiable task completion criteria, checkbox closure itself becomes a reward-hackable signal. Human triage still required for high-cost tasks.","inventor_note":"","sources":[{"key":"vidal2026serious-agentic","note":"Ralph / task-file loops; Anthropic /ralph-loop mention"}],"links":["memory-goal-state-loop","scratchpad-durable-memory","agentic-alignment-problem"],"relations":[{"slug":"memory-goal-state-loop","rel":"related"},{"slug":"scratchpad-durable-memory","rel":"related"},{"slug":"agentic-alignment-problem","rel":"related"}],"topics":[{"id":"agentic-engineering","title":"Agentic engineering"}],"created_at":"2026-10-01T19:14:49.307Z","updated_at":"2026-10-01T19:14:49.307Z"},{"slug":"verifiable-goals-prerequisite","href":"/entries/verifiable-goals-prerequisite","api":"/api/v1/entries/verifiable-goals-prerequisite","title":"Verifiable goals as alignment prerequisite","type":"principle","status":"active","certainty":"reported","claim":"Alignment requires goals that are verifiable (objective metric) or pseudo-verifiable (judge model / rubric). Unverifiable goals cannot be aligned reliably. reported [@vidal2026serious-agentic]","mechanism":"reported [@vidal2026serious-agentic] Principle I in the talk: you cannot align what you cannot verify.\n\nVerifiable goals: measurable metrics, automated tests, typecheckers, schema validators, CLI oracles, visual diffs, hardware smoke signals. Pseudo-verifiable goals: separate judge models or rubrics when no hard oracle exists.\n\nSoftware crossed the agentic threshold early because tests and types provide oracles (talk cites Antje Barth / AIEWF). Domains without oracles need constructed pseudo-verification before agent leverage scales.\n\nUniversal loop pieces Memory / Goal / State presuppose Goal is checkable. Reward hacking appears when the signal is weaker than the intent (e.g. mocking an API to green tests). Antidotes named in the talk: separated verifier, hawks, HITL, real oracles.\n\nDenominator problem related: multiplying code without a product-level denominator (users, outcomes) yields a bill, not alignment.","quantities":[],"limits":"Pseudo-verification inherits judge bias and reward-hacking risk. Talk-level principle, not a formal completeness theorem.","inventor_note":"","sources":[{"key":"vidal2026serious-agentic","note":"verification prerequisite; Barth AIEWF attribution for software threshold"}],"links":["agentic-alignment-problem","memory-goal-state-loop","implementer-verifier-separation","reward-hacking","denominator-problem"],"relations":[{"slug":"agentic-alignment-problem","rel":"related"},{"slug":"memory-goal-state-loop","rel":"related"},{"slug":"implementer-verifier-separation","rel":"related"},{"slug":"reward-hacking","rel":"related"},{"slug":"denominator-problem","rel":"related"}],"topics":[{"id":"agentic-engineering","title":"Agentic engineering"}],"created_at":"2026-10-01T19:14:49.295Z","updated_at":"2026-10-01T19:14:49.307Z"},{"slug":"implementer-verifier-separation","href":"/entries/implementer-verifier-separation","api":"/api/v1/entries/implementer-verifier-separation","title":"Implementer–verifier separation","type":"principle","status":"active","certainty":"reported","claim":"Implementation and verification are distinct roles. A single pass that both writes and accepts its own output weakens alignment pressure. reported [@vidal2026serious-agentic]","mechanism":"reported [@vidal2026serious-agentic] RL-trained models are eager to finish and will hallucinate completion. Self-preference bias: models judge their own writing favorably.\n\nRole split: the implementer proposes changes; a verifier in an independent context window owns the goal and grades against criteria (tests, rubrics, secondary model, human). Cross-provider pairing (example cited: Codex implementer + Claude hawk) breaks same-model self-preference.\n\nVerifier quality bounds system quality. Who grades matters. Async variants (hawk reading full transcripts post-tool) stop compounding side effects earlier than late PR review.\n\nThis principle is a harness pattern under harness = agent − model, and a Goal-ownership pattern under Memory–Goal–State.","quantities":[],"limits":"Operational heuristic from talk. Verifier failure (weak tests, captured judges) still permits reward hacking. Human verifiers remain attention-bounded.","inventor_note":"","sources":[{"key":"vidal2026serious-agentic","note":"role split; self-preference; cross-provider pairing"}],"links":["harness-equals-agent-minus-model","memory-goal-state-loop","verifiable-goals-prerequisite","hawk-async-verifier","reward-hacking"],"relations":[{"slug":"harness-equals-agent-minus-model","rel":"related"},{"slug":"memory-goal-state-loop","rel":"related"},{"slug":"verifiable-goals-prerequisite","rel":"related"},{"slug":"hawk-async-verifier","rel":"related"},{"slug":"reward-hacking","rel":"related"}],"topics":[{"id":"agentic-engineering","title":"Agentic engineering"}],"created_at":"2026-10-01T19:14:49.295Z","updated_at":"2026-10-01T19:14:49.307Z"},{"slug":"harness-equals-agent-minus-model","href":"/entries/harness-equals-agent-minus-model","api":"/api/v1/entries/harness-equals-agent-minus-model","title":"Harness equals agent minus model","type":"principle","status":"active","certainty":"reported","claim":"Agent = model + harness. Harness is the durable control surface: tools, memory, verification, orchestration, permissions. reported [@vidal2026serious-agentic]","mechanism":"reported [@vidal2026serious-agentic] Attribution in the talk: Mike Chambers (AI.Engineer World's Fair) formulation harness = agent − model.\n\nHarness contents include tools, hooks, schemas, oracles, isolation boundaries, permission gates, and orchestration. Model weights are treated as comparatively interchangeable relative to harness design. Quality and alignment gains concentrate in harness loops.\n\nCited supporting pattern: reducing tool surface (talk cites Vercel cutting ~80% of tools) correlated with fewer steps and higher precision. Less context and stronger isolation improve work quality even when that isolation costs more generated code; under near-zero code cost, isolation is cheap relative to misalignment.\n\nHarness design choices map onto Memory–Goal–State: which tools write Memory, which oracles own Goal, which files encode State. Implementer/verifier separation, Ralph loops, hawks, and HITL escape hatches are harness components, not model properties.","quantities":[],"limits":"Definitional framing from talk discourse, not a formal equation. Empirical tool-reduction anecdotes are not controlled measurements in the talk source.","inventor_note":"","sources":[{"key":"vidal2026serious-agentic","note":"harness framing; Chambers AIEWF attribution"}],"links":["agentic-alignment-problem","implementer-verifier-separation","memory-goal-state-loop","hawk-async-verifier","hitl-escape-hatch"],"relations":[{"slug":"agentic-alignment-problem","rel":"related"},{"slug":"implementer-verifier-separation","rel":"related"},{"slug":"memory-goal-state-loop","rel":"related"},{"slug":"hawk-async-verifier","rel":"related"},{"slug":"hitl-escape-hatch","rel":"related"}],"topics":[{"id":"agentic-engineering","title":"Agentic engineering"}],"created_at":"2026-10-01T19:14:49.295Z","updated_at":"2026-10-01T19:14:49.307Z"},{"slug":"memory-goal-state-loop","href":"/entries/memory-goal-state-loop","api":"/api/v1/entries/memory-goal-state-loop","title":"Memory–Goal–State agent loop","type":"principle","status":"active","certainty":"reported","claim":"An agent control loop comprises Memory (persist and iterate), Goal (verifiable or pseudo-verifiable stop criterion), and State (tasks, phases, iterations). reported [@vidal2026serious-agentic]","mechanism":"reported [@vidal2026serious-agentic] Universal loop anatomy stated in the talk:\n\nMemory — repository state, scratchpads, traces, compactable session history. Memory enables learning across iterations. Scratchpads and task files are durable Memory surfaces that survive context compaction.\n\nGoal — stop criterion that is verifiable (tests, types, schema validation, CLI oracles, visual diff) or pseudo-verifiable (judge model with independent context). Without Goal, agents trained to finish declare done without grounding.\n\nState — explicit machine of tasks, phases, and iteration counters. Task files with checkboxes are a concrete State encoding. Orchestrators that own a Markdown backlog are another.\n\nSkeleton: act → verify → done or continue. Feedback may be synchronous (inline tool result) or asynchronous (hawk/post-tool verifier).\n\nThis structure is the minimal abstract machine underneath Ralph loops, implementer/verifier splits, fan-out/fan-in workflows, and orchestrator designs. Implementation details vary by harness; the three roles remain.","quantities":[],"limits":"Structural model from talk discourse. Concrete Memory/Goal/State encodings are harness-specific. Pseudo-verifiable goals inherit judge failure modes.","inventor_note":"","sources":[{"key":"vidal2026serious-agentic","note":"loop anatomy"}],"links":["agentic-alignment-problem","verifiable-goals-prerequisite","ralph-task-file-loop","scratchpad-durable-memory","implementer-verifier-separation"],"relations":[{"slug":"agentic-alignment-problem","rel":"related"},{"slug":"verifiable-goals-prerequisite","rel":"related"},{"slug":"ralph-task-file-loop","rel":"related"},{"slug":"scratchpad-durable-memory","rel":"related"},{"slug":"implementer-verifier-separation","rel":"related"}],"topics":[{"id":"agentic-engineering","title":"Agentic engineering"}],"created_at":"2026-10-01T19:14:49.295Z","updated_at":"2026-10-01T19:14:49.307Z"},{"slug":"agentic-alignment-problem","href":"/entries/agentic-alignment-problem","api":"/api/v1/entries/agentic-alignment-problem","title":"Agentic alignment problem","type":"idea","status":"active","certainty":"reported","claim":"Agentic engineering is an alignment problem: human intent approximately equals agent-produced artifacts under verification constraints. reported [@vidal2026serious-agentic]","mechanism":"reported [@vidal2026serious-agentic] Agentic engineering treats the gap between what a human meant and what an agent built as the primary engineering object.\n\nCode-generation cost trends toward near-zero. Observed industry pattern: large reimplementations become economically plausible when a test suite or other oracle can grade outputs (examples cited in the talk: browser/engine ports, framework ports, runtime ports). The classical premise that humans review everything they ship fails under that cost curve. Exhaustive human review of all generated artifacts does not scale.\n\nWork therefore shifts from single-shot prompting to designing loops that prompt agents. Attribution in the talk: Karpathy framing from vibe coding toward agentic engineering; Steinberger/Cherny formulation that the job is writing loops rather than prompting Claude directly.\n\nAlignment requires goals that are verifiable (objective metric, test, typecheck, schema, oracle) or pseudo-verifiable (separate judge model / rubric). Domains without oracles need constructed pseudo-verification before agent leverage scales.\n\nThe formulation is a talk-level framing, not an experimentally measured law. It constrains harness design: memory, goal, and state must close an act→verify→done|continue loop.","quantities":[],"limits":"Formulation from Vidal 2026 talk (Valencia; cross-referenced to AI.Engineer World's Fair material). Not a measured physical law. Pseudo-verification inherits judge bias and reward-hacking risk.","inventor_note":"","sources":[{"key":"vidal2026serious-agentic","note":"primary formulation; Serious Agentic Engineering talk"}],"links":["memory-goal-state-loop","harness-equals-agent-minus-model","implementer-verifier-separation","verifiable-goals-prerequisite","context-poisoning","ralph-task-file-loop","hawk-async-verifier"],"relations":[{"slug":"memory-goal-state-loop","rel":"related"},{"slug":"harness-equals-agent-minus-model","rel":"related"},{"slug":"implementer-verifier-separation","rel":"related"},{"slug":"verifiable-goals-prerequisite","rel":"related"},{"slug":"context-poisoning","rel":"related"},{"slug":"ralph-task-file-loop","rel":"related"},{"slug":"hawk-async-verifier","rel":"related"}],"topics":[{"id":"agentic-engineering","title":"Agentic engineering"}],"created_at":"2026-10-01T19:14:49.295Z","updated_at":"2026-10-01T19:14:49.307Z"}],"refs":[{"key":"vidal2026serious-agentic","href":"/refs#vidal2026serious-agentic","api":"/api/v1/refs/vidal2026serious-agentic","kind":"web","title":"Serious Agentic Engineering — from first principles: turning tokens into value","authors":"Alejandro Vidal (@dobleio)","year":"2026","journal":"","doi":"","url":"https://doble.io/e0b59cc146a957d083904d036a4f7258/","note":"Talk, Valencia 2026 (~117 slides). Primary source for Palantir agentic-engineering topic. Core framing: agentic engineering as alignment (human intent ≈ agent artifact under verification). Covers Memory–Goal–State loops, harness=agent−model (Chambers/AIEWF attribution), implementer/verifier split, Ralph/task-file loops, scratchpads, hawk async verifiers, HITL escape hatches, context poisoning, rollback+learnings, reward hacking, denominator problem, organizational triage and adoption timelines. Cross-references AI.Engineer World’s Fair material (Barth, Chambers, Karpathy-adjacent framings).","created_at":"2026-10-01T19:14:49.295Z"},{"key":"ieee2020robeetle","href":"/refs#ieee2020robeetle","api":"/api/v1/refs/ieee2020robeetle","kind":"misc","title":"Minuscule RoBeetle Turns Liquid Methanol Into Muscle Power","authors":"IEEE Spectrum","year":"2020","journal":"","doi":"","url":"https://spectrum.ieee.org/robeetle-liquid-methanol","note":"Secondary. One SMA clause inverted relative to yang2020robeetle","created_at":"2026-10-01T19:14:49.242Z"},{"key":"usc2020robeetle","href":"/refs#usc2020robeetle","api":"/api/v1/refs/usc2020robeetle","kind":"misc","title":"Viterbi Researchers Create The Lightest, Smallest, Fully Autonomous Crawling Microrobot Reported To Date","authors":"USC Viterbi","year":"2020","journal":"","doi":"","url":"https://viterbischool.usc.edu/news/2020/08/viterbi-researchers-create-the-lightest-smallest-fully-autonomous-crawling-microrobot-reported-to-date/","note":"Institutional note","created_at":"2026-10-01T19:14:49.242Z"},{"key":"yang2020robeetle","href":"/refs#yang2020robeetle","api":"/api/v1/refs/yang2020robeetle","kind":"article","title":"An 88-milligram insect-scale autonomous crawling robot driven by a catalytic artificial muscle","authors":"Yang et al.","year":"2020","journal":"Science Robotics","doi":"10.1126/scirobotics.aba0015","url":"https://www.science.org/doi/10.1126/scirobotics.aba0015","note":"Primary source for methanol catalytic combustion on Pt as actuation pathway and for RoBeetle 2020 device metrics (empty mass 88 mg, NiTi wire 50.8 µm with Pt catalyst, ΔH −676.49 kJ mol⁻¹, system-to-work efficiency ~0.48%, As 87–99 °C). Prefer this over secondary summaries when SMA heating direction conflicts.","created_at":"2026-10-01T19:14:49.242Z"}],"relations_flat":[{"from":"2020-robeetle-catalytic-muscle","to":"methanol-catalytic-combustion-on-pt","rel":"demonstrates"},{"from":"denominator-problem","to":"agentic-alignment-problem","rel":"related"},{"from":"denominator-problem","to":"verifiable-goals-prerequisite","rel":"related"},{"from":"denominator-problem","to":"reward-hacking","rel":"related"},{"from":"methanol-catalytic-combustion-on-pt","to":"2020-robeetle-catalytic-muscle","rel":"related"},{"from":"reward-hacking","to":"verifiable-goals-prerequisite","rel":"related"},{"from":"reward-hacking","to":"implementer-verifier-separation","rel":"related"},{"from":"reward-hacking","to":"denominator-problem","rel":"related"},{"from":"rollback-plus-learnings","to":"context-poisoning","rel":"related"},{"from":"rollback-plus-learnings","to":"scratchpad-durable-memory","rel":"related"},{"from":"rollback-plus-learnings","to":"ralph-task-file-loop","rel":"related"},{"from":"context-poisoning","to":"rollback-plus-learnings","rel":"related"},{"from":"context-poisoning","to":"agentic-alignment-problem","rel":"related"},{"from":"context-poisoning","to":"reward-hacking","rel":"related"},{"from":"context-poisoning","to":"hawk-async-verifier","rel":"related"},{"from":"hitl-escape-hatch","to":"harness-equals-agent-minus-model","rel":"related"},{"from":"hitl-escape-hatch","to":"hawk-async-verifier","rel":"related"},{"from":"hitl-escape-hatch","to":"agentic-alignment-problem","rel":"related"},{"from":"hawk-async-verifier","to":"implementer-verifier-separation","rel":"related"},{"from":"hawk-async-verifier","to":"harness-equals-agent-minus-model","rel":"related"},{"from":"hawk-async-verifier","to":"context-poisoning","rel":"related"},{"from":"hawk-async-verifier","to":"hitl-escape-hatch","rel":"related"},{"from":"scratchpad-durable-memory","to":"memory-goal-state-loop","rel":"related"},{"from":"scratchpad-durable-memory","to":"ralph-task-file-loop","rel":"related"},{"from":"scratchpad-durable-memory","to":"context-poisoning","rel":"related"},{"from":"scratchpad-durable-memory","to":"rollback-plus-learnings","rel":"related"},{"from":"ralph-task-file-loop","to":"memory-goal-state-loop","rel":"related"},{"from":"ralph-task-file-loop","to":"scratchpad-durable-memory","rel":"related"},{"from":"ralph-task-file-loop","to":"agentic-alignment-problem","rel":"related"},{"from":"verifiable-goals-prerequisite","to":"agentic-alignment-problem","rel":"related"},{"from":"verifiable-goals-prerequisite","to":"memory-goal-state-loop","rel":"related"},{"from":"verifiable-goals-prerequisite","to":"implementer-verifier-separation","rel":"related"},{"from":"verifiable-goals-prerequisite","to":"reward-hacking","rel":"related"},{"from":"verifiable-goals-prerequisite","to":"denominator-problem","rel":"related"},{"from":"implementer-verifier-separation","to":"harness-equals-agent-minus-model","rel":"related"},{"from":"implementer-verifier-separation","to":"memory-goal-state-loop","rel":"related"},{"from":"implementer-verifier-separation","to":"verifiable-goals-prerequisite","rel":"related"},{"from":"implementer-verifier-separation","to":"hawk-async-verifier","rel":"related"},{"from":"implementer-verifier-separation","to":"reward-hacking","rel":"related"},{"from":"harness-equals-agent-minus-model","to":"agentic-alignment-problem","rel":"related"},{"from":"harness-equals-agent-minus-model","to":"implementer-verifier-separation","rel":"related"},{"from":"harness-equals-agent-minus-model","to":"memory-goal-state-loop","rel":"related"},{"from":"harness-equals-agent-minus-model","to":"hawk-async-verifier","rel":"related"},{"from":"harness-equals-agent-minus-model","to":"hitl-escape-hatch","rel":"related"},{"from":"memory-goal-state-loop","to":"agentic-alignment-problem","rel":"related"},{"from":"memory-goal-state-loop","to":"verifiable-goals-prerequisite","rel":"related"},{"from":"memory-goal-state-loop","to":"ralph-task-file-loop","rel":"related"},{"from":"memory-goal-state-loop","to":"scratchpad-durable-memory","rel":"related"},{"from":"memory-goal-state-loop","to":"implementer-verifier-separation","rel":"related"},{"from":"agentic-alignment-problem","to":"memory-goal-state-loop","rel":"related"},{"from":"agentic-alignment-problem","to":"harness-equals-agent-minus-model","rel":"related"},{"from":"agentic-alignment-problem","to":"implementer-verifier-separation","rel":"related"},{"from":"agentic-alignment-problem","to":"verifiable-goals-prerequisite","rel":"related"},{"from":"agentic-alignment-problem","to":"context-poisoning","rel":"related"},{"from":"agentic-alignment-problem","to":"ralph-task-file-loop","rel":"related"},{"from":"agentic-alignment-problem","to":"hawk-async-verifier","rel":"related"}]}