Code as Agent Harness: Toward Executable, Verifiable, and Stateful Agent Systems¶
Szerzők: Xuying Ning, Katherine Tieu, Dongqi Fu et al. (UIUC + industry) Dátum: 2026. május 18. ArXiv: 2605.18747 Típus: Survey paper
Core Concept¶
A kód már nemcsak output, hanem az agent működésének alaprétege (harness). A survey három rétegben tárgyalja ezt: harness interface, harness mechanisms, harnees scaling.
Három Réteg¶
1. Harness Interface — Kód a gondolkodáshoz, cselekvéshez, környezethez¶
- Reasoning: program-delegált gondolkodás, formális verifikáció, iteratív code-grounded reasoning
- Acting: kód mint skill szelekció, programmatic policy generálás, lifelong code-based agents
- Environment: strukturált világreprezentációk, execution-trace modellezés, verifikálható környezetek
2. Harness Mechanisms¶
Memory & Context Engineering (3.2) — a legrelevánsabb szekció: - Working Memory: azonnali feladathoz szükséges kontextus - Semantic Memory: strukturált tudás (entitások, tények) - Experiential Memory: múltbeli tapasztalatok, hibák, sikerek - Long-Term Memory: perzisztens tudás session-ökön át - Multi-Agent Memory: agentek közötti megosztott emlékezet - Context Compaction and State Offloading: a context pruning implementációja agent rendszerekben
Plan-Execute-Verify Loop (3.4): - Planning as contract formation - Sandboxed execution + permissioned state transition - Verification through deterministic sensors
Harness Optimization (3.5): - Deep telemetry → evolution agent → governed harness mutation - A harness önmagát optimalizálja visszacsatolás alapján
3. Scaling — Multi-Agent Orchestration¶
Funkcionális szerep specializáció: - Program synthesis, Program understanding, Verification, Execution, Planning agentek
Shared-Harness Synchronization: - Shared blackboard, parallel branches with merge - Structured context scheduling, hierarchical memory - Agent pool scaling
Kulcs Minták (4.4)¶
"Context management is the tax of implicit shared state" — a shared state fenntartásának ára a context management
"Execution feedback as the bridge between linguistic and formal reasoning" — a végrehajtási visszacsatolás köti össze a nyelvi és formális gondolkodást
"Topology complexity inversely correlates with harness-state formality" — minél komplexebb a topológia, annál informálisabb a harness state
Open Problems (5.2)¶
- Harness-Level Evaluation: nem elég a task sikeresség — a harness minőségét is mérni kell
- Semantic Verification: a végrehajtási feedback-en túli verifikáció
- Self-Evolving Harnesses without Regression: önfejlesztő harness regresszió nélkül
- Transactional Shared Program State: tranzakciós megosztott állapot szemantikai konfliktusfeloldással
- Human-in-the-Loop Safety: emberi felügyelet biztonságkritikus akcióknál
- Multimodal Code-Harness Systems
OpenClaw Relevancia¶
- A pipeline rendszerünk = harness: state.json → pipeline template → steps → visszacsatolás. A Plan-Execute-Verify loop-unk a cron → process_inbox.py → inbox-processor hármas
- Memory taxonomy: Working (daily logs) → Semantic (MEMORY.md) → Long-Term (wiki) → Experiential (memory_search). A Context Compaction pontot köti a Redis pruninghez
- Multi-agent: subagent spawn-olás a harness része — a state.json mint shared blackboard működik
- Hiányosságok:
- Nincs formal verification a pipeline lépések között (csak status check)
- Nincs deep telemetry → a harness nem önoptimalizáló
- Nincs transactional state — ha a pipeline félbeszakad, manuális beavatkozás kell
- A shared harness-state konvergencia hiányzik (ha több subagent párhuzamosan dolgozik, nincs merge)
- Következő lépés: a pipeline legalább egy verification step-et kapjon (summary integrity check, topic file duplikátum detektálás)
Idézetek¶
"Code is no longer only a target output — it increasingly serves as an operational substrate for agent reasoning, acting, and execution-based verification."
"Context management is the tax of implicit shared state."
"Execution feedback as the bridge between linguistic and formal reasoning."
Forrás¶
arXiv:2605.18747 — https://arxiv.org/abs/2605.18747