Skip to main content

External Framework Benchmark Mapping Rollout

Purpose

This rollout tracks how the Runtime Governance Benchmark Suite is mapped across all external framework pages.

The rollout is not a ranking table. It records which benchmark dimensions are applicable, which are not applicable, which require adapters, and which require more evidence.

Standard Mapping Rule

Every mapping must preserve the framework's own scope and avoid turning a benchmark mismatch into a framework-invalidity claim.

not_applicable != failure
adapter_required != failure
insufficient_evidence != failure
observed_boundary != certification
benchmark_mapping != execution_authority

Mapping States

StateMeaning
not_startedNo benchmark mapping has been prepared.
intakeFramework role identified; mapping not yet anchored to benchmark dimensions.
mapped_partialAt least three benchmark dimensions are classified.
mapped_coreCore dimensions are classified and non-claims are present.
fixture_readyMachine-readable companion or fixture exists.
replay_readyInputs and replay instructions are sufficient to rerun.
interoperability_readyResult can route into a Commitment Candidate or SPE path.

Required Dimensions

Each framework-specific benchmark mapping should classify at least these dimensions:

execution_boundary
preparation_boundary
commitment_boundary
semantic_equivalence_boundary
unknown_trajectory_boundary
authority_boundary
evidence_freshness_boundary
reconstruction_boundary
interoperability_path

Allowed dimension states:

applicable
not_applicable
adapter_required
insufficient_evidence
observed_partial
mapped_partial
replay_required
fixture_ready

mapped_partial at the dimension level means that a crosswalk exists, but the dimension still lacks the source, replay, or implementation evidence required for a stronger state.

Current Rollout Table

FrameworkMapping StateNext Action
Morrison Runtimefixture_readyAttach raw audit payloads and replay table.
GLMfixture_readyAttach source-versioned GLM declaration examples and replay through Commitment Candidate.
EVIDEfixture_readyAttach concrete EVIDE evidence package examples and reconstruction receipt.
DecisionAssurefixture_readyAttach canonical policy/delegation examples and trace hashes.
MindForgefixture_readyAttach historical review sample and current-standing reconstruction example.
ASROfixture_readyAttach concrete ASRO attestation examples and reconstruction receipt.
CARE Runtimefixture_readyAttach official CARE Runtime source before any public behavior claim.
AARfixture_readyAttach concrete AAR operational evidence package and reconstruction receipt.
MITRE ATLASfixture_readyAttach source-versioned MITRE ATLAS technique mappings to benchmark cases.
OWASP Top 10 for LLM Applicationsfixture_readyAttach source-versioned OWASP risk category mappings to benchmark cases.
Agent Governance Playbookfixture_readyAttach concrete continuation-state examples and replay through Commitment Candidate.
Emergency Stop Conventionfixture_readyAttach emergency-stop trigger examples and reconstruction receipt.
NIST AI RMFfixture_readyAttach NIST AI RMF control examples and lifecycle evidence package.
ISO/IEC 42001fixture_readyAttach ISO/IEC 42001 control examples and management-record evidence package.
EU AI Actfixture_readyAttach EU AI Act obligation mapping examples and documentation evidence package.
Policy Cardsfixture_readyAttach source-versioned policy-card examples and replay through unknown/preparation cases.
Runtime Governance for AI Agentsfixture_readyAttach concrete path-policy traces and replay through multi-step cases.

Framework Mapping Template

framework_id:
framework_role:
source_status:
benchmark_dimensions:
execution_boundary:
preparation_boundary:
commitment_boundary:
semantic_equivalence_boundary:
unknown_trajectory_boundary:
authority_boundary:
evidence_freshness_boundary:
reconstruction_boundary:
interoperability_path:
non_claims:
next_required_action:

Non-Claims

This rollout does not certify external frameworks.
This rollout does not rank external frameworks.
This rollout does not treat unmapped dimensions as failures.
This rollout does not grant execution authority.
This rollout does not replace framework-native evaluation.

External-framework benchmarking is evidence-governance work. Publication does not create standing. Standing must be reconstructed from source, evidence, authority, admissibility, and current commit-time conditions.