Sultan

Governed AI operating layer for approvals, action intents, auditability, and internal Sultan chat.

Command

OpenClaw

Eval + Reality Audit

This is the missing reality layer: eval suites, holdouts, replay runs, and audits that can prove whether the machine stayed grounded.

Eval Suites

Purpose-built packs, not vibes.

1

Completed Runs

Scored work the system can point back to.

1

Pass Rate

Judge verdicts across scored cases.

100%

Replay Reproducibility

How often the system can do the same safe thing twice.

92%

Eval Suites

Every serious behavior should have named tests, not just hope.

1 cases

Outbound Governance Eval Pack

sales_outboundactive

1 cases

Prove the system drafts grounded outreach, respects policy, and does not overstate reality.

Runs recorded: 1

Judges + Holdouts

Reality needs a rubric, and regression needs protected difficulty.

1 holdouts

Reality Judge Alpha

realityactive

Score whether the system stayed inside evidence, policy, and operator trust.

High-risk outbound holdout

sales_outbound

Preserve difficult outbound scenarios for future regression checks.

Protected cases: 1

Recent Runs

A run should tell you what was challenged and how it ended.

1 runs

lead_follow_up_intent

seeded_baselinecompleted

1 scores

Shadow mode outbound recommendation respected policy and requested approval.

Replay Layer

Replay is where the system proves it can stay consistent under pressure.

0 failed runs

lead_follow_up_intent

completed

92%

Replay remained inside approval law with no new hallucinated actions.