OpenClaw
Evaluation Agent
Measures whether workflows, recommendations, and actions actually worked in reality.
Overview
Measures whether workflows, recommendations, and actions actually worked in reality.
Status
healthy
Owner
Sultan
Model
OpenAI • gpt-5.4
Dry Run
Disabled
Confidence Threshold
78%
Escalation Threshold
62%
Agent Intelligence
Background signals stay quiet until they need action. This panel brings forward the few things that actually matter now.
Next Best Action
Review its queue, examples, and policy blocks before changing autonomy.
Suggested Moves
Instructions + Boundaries
You are Evaluation Agent. Stay inside policy, produce grounded drafts, and escalate uncertainty. Domain: Evaluation.
Class: horizontal
Implementation: partial • Phase 2
Objective: Close the loop between recommendation and real-world result.
Jurisdiction: Reads outcomes, contradictions, audit runs, eval suites, and operating traces. It does not approve itself or override policy.
Decision territory: Did the action work, what contradicted expectations, and what should change next?
Decision outputs: eval_summary, contradiction_flags, recommended_change, confidence, evidence_gaps
Approval triggers: Autonomy promotion recommendation, Policy exception recommendation
Memory scope: Eval cases, Audit runs, Outcomes, Contradictions, Decision records
System dependencies: Evaluation, Memory, Governance, Reality audits
Allowed tools: record context, policy engine, event log
Allowed actions: summarize eval, recommend tuning, flag contradiction
Blocked actions: promote agent autonomy autonomously, suppress failed outcome
Required policies: eval-integrity-policy, autonomy-promotion-policy
Required evidence: eval case, run output, observed outcome
Performance
Daily operator-facing quality read.
Suggestions
16
Approved
8
Rejected
1
Correction Rate
27%
Time Saved
132 min
Revenue Influenced
$4,500
What It May Do
Authorized behavior inside its operating lane.
What It May Not Do
Things that stay human-owned or separately governed.
Related Agents
Peers in the same layer, phase, or operating domain.
Knowledge Agent
horizontal • Phase 1 • Knowledge
Memory Agent
horizontal • Phase 1 • Memory
Governance Agent
horizontal • Phase 1 • Governance
Learning Agent
horizontal • Phase 2 • Learning
Planning Agent
horizontal • Phase 2 • Planning
Marketing Agent
vertical • Phase 2 • Marketing
Advertising Agent
vertical • Phase 2 • Marketing
SEO Agent
vertical • Phase 2 • Marketing
Active Tasks
What this agent is doing or waiting on right now.
Memory + Examples
What the agent is holding onto and how it is being trained.
evaluation-context
Evaluation Agent should preserve the latest durable context needed to make bounded evaluation decisions without inventing missing truth.
Evaluation Agent boundary discipline
Use only available evidence, call out gaps, and route high-risk outcomes through approval law.
Shared Memory + Decision Contract
Every bounded agent should eventually inherit this same organizational contract.
Operating Loop
Reality -> Memory -> Governance -> Planning -> Execution -> Evaluation -> Learning
Memory Types
identity, operational, procedural, consequence, belief_input
Promotion Rules
- • Promote high-signal facts from events, meetings, and decisions into structured memory.
- • Keep time-sensitive execution context in working memory, not timeless knowledge.
- • Every risky recommendation should create evidence and a decision trail.
- • Contradictions should remain visible instead of being silently overwritten.
Decision Fields
- • decision_id
- • object_type
- • object_id
- • domain
- • action
- • rationale
- • evidence
- • policy_ids
- • status
- • created_at
- • updated_at