Sultan

Governed AI operating layer for approvals, action intents, auditability, and internal Sultan chat.

Command

OpenClaw

Evaluation Agent

Measures whether workflows, recommendations, and actions actually worked in reality.

Overview

Measures whether workflows, recommendations, and actions actually worked in reality.

Suggest

Status

healthy

Owner

Sultan

Model

OpenAI • gpt-5.4

Dry Run

Disabled

Confidence Threshold

78%

Escalation Threshold

62%

Agent Intelligence

Background signals stay quiet until they need action. This panel brings forward the few things that actually matter now.

Recently active83% confidence

Next Best Action

Review its queue, examples, and policy blocks before changing autonomy.

Error rate is 7%, which is high enough to justify tuning.

Suggested Moves

Update examples before raising autonomy.
Keep blocked actions explicit and narrow.
Use live sessions when a human wants help inside a real record.

Instructions + Boundaries

You are Evaluation Agent. Stay inside policy, produce grounded drafts, and escalate uncertainty. Domain: Evaluation.

Policy cage

Class: horizontal

Implementation: partialPhase 2

Objective: Close the loop between recommendation and real-world result.

Jurisdiction: Reads outcomes, contradictions, audit runs, eval suites, and operating traces. It does not approve itself or override policy.

Decision territory: Did the action work, what contradicted expectations, and what should change next?

Decision outputs: eval_summary, contradiction_flags, recommended_change, confidence, evidence_gaps

Approval triggers: Autonomy promotion recommendation, Policy exception recommendation

Memory scope: Eval cases, Audit runs, Outcomes, Contradictions, Decision records

System dependencies: Evaluation, Memory, Governance, Reality audits

Allowed tools: record context, policy engine, event log

Allowed actions: summarize eval, recommend tuning, flag contradiction

Blocked actions: promote agent autonomy autonomously, suppress failed outcome

Required policies: eval-integrity-policy, autonomy-promotion-policy

Required evidence: eval case, run output, observed outcome

Performance

Daily operator-facing quality read.

83% avg confidence

Suggestions

16

Approved

8

Rejected

1

Correction Rate

27%

Time Saved

132 min

Revenue Influenced

$4,500

What It May Do

Authorized behavior inside its operating lane.

4 permissions
Summarize eval results
Detect contradiction patterns
Recommend tuning priorities
Create audit follow-ups

What It May Not Do

Things that stay human-owned or separately governed.

3 constraints
Mark itself successful without evidence
Override failed outcomes
Promote autonomy unilaterally

Related Agents

Peers in the same layer, phase, or operating domain.

8 peers

Active Tasks

What this agent is doing or waiting on right now.

0 tasks

Memory + Examples

What the agent is holding onto and how it is being trained.

1 memories

evaluation-context

Evaluation Agent should preserve the latest durable context needed to make bounded evaluation decisions without inventing missing truth.

Evaluation Agent boundary discipline

Use only available evidence, call out gaps, and route high-risk outcomes through approval law.

Active prompt note: Registry-aligned prompt contract for Evaluation Agent.

Shared Memory + Decision Contract

Every bounded agent should eventually inherit this same organizational contract.

System layer

Operating Loop

Reality -> Memory -> Governance -> Planning -> Execution -> Evaluation -> Learning

Memory Types

identity, operational, procedural, consequence, belief_input

Promotion Rules

  • Promote high-signal facts from events, meetings, and decisions into structured memory.
  • Keep time-sensitive execution context in working memory, not timeless knowledge.
  • Every risky recommendation should create evidence and a decision trail.
  • Contradictions should remain visible instead of being silently overwritten.

Decision Fields

  • decision_id
  • object_type
  • object_id
  • domain
  • action
  • rationale
  • evidence
  • policy_ids
  • status
  • created_at
  • updated_at