Audit Packet: When Agents Misread the Enterprise
What the evidence supports, how the semantic-action-integrity thesis was tested, and where the argument remains deliberately bounded.
This audit packet supports Brief №015: When Agents Misread the Enterprise. Read the brief first for the full argument.
Autonoma briefs are designed to be inspectable. This packet shows what the brief claims, how each claim was tested, what it does not claim, and where caveats remain — without exposing raw internal logs, prompts, operator notes, source-routing mechanics, hashes, local paths, secrets, or unpublished candidate claims.
Audit Verdict
The central judgment, what it is supported to claim, and what it is not.
Verdict: SUPPORTED FOR PUBLICATION WITH MATERIAL CAVEATS.
The evidence supports the brief’s central mechanism: an LLM agent can possess valid authority and access accurate data while still violating organizational policy or taking an inappropriate action when decisive entity attributes, contextual state, relationships, definitions, or history are absent from its visible context. 1
The evidence also supports the architectural direction that machine-readable semantics can become executable controls. Structured knowledge can be used to gather sufficient evidence before answering, compile ontological specifications into constrained tool interfaces, validate graph state against formal conditions, and include request context in authorization decisions. 2345
The evidence does not support claims that:
- this failure is already common across enterprise production systems;
- every enterprise agent requires a general knowledge graph;
- knowledge graphs reliably eliminate semantic errors;
- formal semantics automatically remain correct or current;
- any named vendor has solved semantic action integrity;
- a particular architecture guarantees compliance, fairness, or business outcomes;
- semantic-context contracts will become universal or legally mandatory.
Confidence is high in the conceptual distinction between authorization and semantic action integrity, moderate in the general enterprise relevance of the mechanism, moderate in the value of executable semantic controls for bounded consequential actions, low in any estimate of production prevalence, and moderate in the 12–24 month platform forecast.
What the Audit Tested
The eight questions under test, and what was deliberately excluded.
The audit tested eight questions:
- Can an agent be authorized and use accurate data while still taking an organizationally wrong action?
- Can policy compliance depend on entity attributes, contextual state, relationships, or history missing from the agent’s visible context?
- Does an authorization decision establish that the agent interpreted the enterprise meaning correctly?
- Can structured knowledge improve bounded reasoning by requiring sufficient evidence or abstention?
- Can ontological or schema specifications be converted into executable constraints on agent tools?
- Can graph validation and contextual policy engines operate before a consequential write?
- Does the evidence justify a universal knowledge-graph requirement?
- What enterprise test can distinguish useful semantic governance from costly modeling without operational value?
The audit also evaluated the brief’s forecast, stakeholder implications, and dissent. It deliberately excluded production prevalence, universal reliability claims, legal mandates, vendor-specific outcome claims, and quantified business impact because the source set does not establish them.
Claim-by-Claim Evidence Audit
Each public claim, the evidence behind it, the verdict, and the boundary.
| Public claim | Evidence | Verdict | Boundary |
|---|---|---|---|
| Policy compliance can depend on entity attributes, contextual state, or history absent from an agent’s visible context. | Policy-Invisible Violations in LLM-Based Agents. 1 | Supported | Diagnostic benchmark and proof-of-concept enforcement; not a production-prevalence study. |
| An agent can be authorized yet still act incorrectly because permission and semantic interpretation are different controls. | Hidden-state mechanism in 1, contextual authorization model in Cedar documentation 5, and Autonoma synthesis. | Supported as analytic synthesis | No single source states the full enterprise conclusion verbatim. |
| Enterprise agents can misinterpret metrics and relationships when data lacks explicit machine-readable context. | Forrester architecture analysis. 6 | Supported as market prior art and architecture framing | Analyst guidance, not independent reliability or outcome proof. |
| Knowledge-graph reasoning can be coupled to evidence sufficiency and abstention. | R2-KG benchmark framework. 2 | Supported with benchmark caveat | Five KG reasoning benchmarks; not enterprise policy-compliance or business-outcome proof. |
| Ontological specifications can be compiled into executable tool interfaces that constrain agent behavior. | Ontology-to-tools proof-of-principle research. 3 | Supported with external-validity caveat | Scientific extraction case study; not general enterprise-agent effectiveness proof. |
| RDF graph state can be validated against machine-readable conditions before use. | W3C SHACL Recommendation. 4 | Supported as standards capability | Defines validation behavior; does not prove that validation improves agent outcomes by itself. |
| Contextual authorization can incorporate principal, action, resource, and request context. | Cedar documentation. 5 | Supported as first-party documented behavior | Product/language documentation; not independent proof of effectiveness. |
| Consequential agent actions should require the minimum sufficient meaning layer and fail closed when it is incomplete. | Combined evidence from 12345 and Autonoma synthesis. | Supported as bounded analytic conclusion | Applies to consequential or semantically ambiguous actions, not every model response or workflow. |
| Leading platforms will increasingly expose semantic-context contracts within 12–24 months. | Architecture prior art and executable-control building blocks in 3456. | Moderate-confidence forecast | Directional forecast; timing, terminology, and adoption are uncertain. |
Autonoma analytic synthesis
Three conclusions are synthesis rather than direct quotations from a single source:
- Authorization integrity and semantic action integrity are distinct gates.
- The relevant semantic control must operate before a consequential action, not merely explain it afterward.
- The correct architecture is the minimum sufficient meaning layer for the action—not a universal mandate for one enterprise-wide knowledge graph.
The first two conclusions combine the hidden-state failure mechanism with contextual policy and executable-validation evidence. 1345 The third is constrained by the dissent and prevents the brief from becoming generic knowledge-graph advocacy.
Moderate-confidence forecast
Within 12–24 months, leading enterprise agent platforms will increasingly expose explicit semantic-context contracts: authoritative entity resolution, governed metric definitions, relationship provenance, policy attributes, temporal state, and pre-action validation.
This is a forecast, not a documented current market condition. It is directionally supported by architecture guidance and existing standards and implementation patterns. 3456
Source Quality and Role
What each source is competent to prove — and its limitation.
| Source | Source class | Evidentiary role | Limitation |
|---|---|---|---|
| Policy-Invisible Violations 1 | Recent arXiv preprint and diagnostic benchmark | Central hidden-state and policy-invisibility mechanism | Not peer-reviewed in the packet; does not establish enterprise prevalence |
| R2-KG 2 | Peer-reviewed Findings paper | Independent evidence on KG-grounded reasoning, evidence sufficiency, and reliability | Benchmark scope; not enterprise policy-compliance or business-outcome evidence |
| Ontology-to-tools 3 | Recent arXiv proof of principle | Demonstrates executable semantic constraints through generated tool interfaces | Scientific extraction case study; generalizability remains uncertain |
| W3C SHACL 4 | International technical standard | Formal graph-validation and constraint capability | Capability standard, not an agent-reliability study |
| Cedar authorization documentation 5 | First-party policy-language documentation | Contextual authorization structure and documented system behavior | Vendor documentation proves behavior, not independent outcomes |
| Forrester architecture analysis 6 | Current analyst commentary | Enterprise architecture context and material external prior art | Market framing, not independent technical validation |
Source-role conclusion
The core failure mechanism is supported by independent research. The reasoning and executable-semantics evidence comes from one peer-reviewed paper, two recent preprints, and a formal W3C standard. Cedar is used only to describe contextual authorization behavior. Forrester is used only for architecture context and prior-art framing.
No vendor or analyst source is used to claim improved compliance, reduced error rates, fairness, or realized business outcomes.
Counterarguments and Falsification Tests
The strongest objections, their weight, and how each could be falsified.
Counterargument 1: Most bounded workflows do not need a knowledge graph.
Weight: Strong. Accepted.
A single-system agent with a stable schema, unambiguous entities, and tightly constrained actions may be governed effectively through typed inputs, conventional validation, and contextual policy. A broad ontology or general graph may add cost without improving the decision. 45
Falsification test: Compare a typed-schema-and-policy implementation with a graph-backed implementation on the same consequential workflow. If both achieve equivalent semantic error, abstention, and audit performance, the broader graph is not justified.
Counterargument 2: Formal semantic models become stale, expensive, and politically contested.
Weight: Strong. Accepted.
Structured meaning can be wrong. Metrics change, organizations reorganize, policies acquire exceptions, and business units disagree. A machine-readable definition can become more dangerous when systems treat it as unquestionable authority.
Falsification test: Introduce controlled definition, relationship, and effective-date changes. Measure whether the semantic layer detects staleness, preserves provenance, routes ownership, and prevents obsolete rules from authorizing or validating actions.
Counterargument 3: Modern models can reason over messy, unstructured enterprise information.
Weight: Moderate. Accepted for many low-consequence tasks.
A capable model may infer the right meaning from documents or conversational context without a formal semantic layer. The brief does not dispute that. Its claim is bounded to consequential or cross-system actions where silent ambiguity creates material risk.
Falsification test: Run the same ambiguous-action cases with unstructured context only and with governed semantic context. If the unstructured approach matches the semantic approach across wrong-action, clarification, and abstention measures, formalization may not add value for that workflow.
Counterargument 4: The direct evidence is too bounded to justify enterprise conclusions.
Weight: Strong. Accepted in part.
The central diagnostic source is a recent preprint. The executable ontology source is a proof of principle. The peer-reviewed KG evidence covers reasoning benchmarks rather than enterprise policy action. These sources establish mechanisms and available control patterns, not prevalence or realized outcomes.
Falsification test: Replicate the hidden-state and semantic-constraint results across multiple models, agent frameworks, enterprise-like data domains, and real policy conditions. Failure to reproduce the mechanism outside the original settings would lower confidence in its general enterprise relevance.
Confidence and Limitations
Confidence by dimension, and the material limits of the analysis.
| Dimension | Confidence | Rationale |
|---|---|---|
| Missing context can hide policy-relevant conditions from agents | High | Directly supported by the central diagnostic source |
| Authorization differs from semantic interpretation | High as conceptual distinction | Authorization model and hidden-state evidence address different questions |
| Executable semantics can constrain or validate action | Moderate to high | Demonstrated through a standard, policy language, and proof-of-principle implementation |
| KG grounding can improve bounded reasoning reliability | Moderate | Peer-reviewed benchmark evidence, but not enterprise outcome evidence |
| General enterprise relevance | Moderate | Mechanism maps plausibly to cross-system actions; prevalence remains unknown |
| Current production prevalence | Low / unknown | No prevalence study in the source set |
| Universal knowledge-graph requirement | Rejected | Evidence supports multiple implementation forms and proportional controls |
| 12–24 month semantic-context-contract forecast | Moderate | Converging architecture and standards signals; adoption path uncertain |
Material limitations
- The central policy-invisibility source is a recent preprint.
- The ontology-to-tools evidence is a proof of principle in a scientific extraction setting.
- The peer-reviewed KG paper does not test enterprise organizational policy.
- SHACL specifies validation capability but does not establish agent reliability outcomes.
- Cedar documentation establishes documented authorization behavior only.
- Forrester supplies architecture framing and prior art, not independent technical proof.
- The source set does not measure production incidence, business impact, legal obligation, or market penetration.
- Machine-readable meaning can itself be stale, incomplete, disputed, or incorrectly governed.
- The enterprise examples in the brief are illustrative applications of the mechanism, not documented incidents.
A Reproducible Enterprise Test
A sandboxed semantic-action-integrity test that uses no production data or live systems.
An enterprise can test semantic action integrity without using production employee data or modifying live systems.
Test setup
- Select one consequential sandbox workflow. Examples include mandatory training assignment, access review, internal-opportunity ranking, skills-profile updates, or an HR-service workflow.
- Create synthetic but structurally realistic records across two or more mock systems.
- Hold identity and authorization constant. The test agent should have valid credentials and permission to perform the action in every trial.
- Introduce controlled semantic ambiguities:
- two entities with similar or reused identifiers;
- one metric name with different definitions;
- relationships with current and future effective dates;
- business-unit or jurisdiction exceptions;
- a policy attribute unavailable in one system;
- a status whose administrative and substantive meanings differ.
- Capture the proposed action, supporting context, applied definitions, validation result, clarification behavior, and final write decision.
- Use reversible mock writes only.
Compare these operating modes
- permissions and raw records only;
- permissions plus unstructured documentation;
- typed schema and governed metric contract;
- contextual policy engine;
- relationship or graph validation;
- schema- or ontology-constrained tool interface;
- explicit clarification or abstention when context is incomplete.
Core measures
- authorized-but-semantically-wrong action rate;
- correct entity-resolution rate;
- correct metric-definition selection rate;
- temporal and relationship-state error rate;
- policy-context omission rate;
- clarification and abstention precision;
- false-block rate;
- time and maintenance cost per control mode;
- percentage of actions with reconstructable semantic provenance.
Pass condition
A workflow should not be considered semantically governed merely because every action is authorized and logged.
The stronger pass condition is:
- the agent identifies the governed entity, definition, relationship, effective state, and relevant policy context before action;
- the proposed write passes structural and policy validation;
- incomplete or conflicting meaning causes clarification, abstention, or escalation;
- the audit record preserves which semantic contract justified the action.
If a simpler typed-schema or policy solution meets that condition, a broader knowledge graph is unnecessary. If no mode materially reduces semantic wrong actions, the proposed meaning layer has not demonstrated operational value.
Methodology
How sources were reviewed and what was deliberately excluded.
The audit reviewed six public sources across five independent domains and separated direct research, peer-reviewed benchmark evidence, technical standards, first-party documented behavior, and analyst prior art. Two load-bearing claims were verified against exact source passages before drafting. Each public claim was evaluated against the source role and its stated limitation. Claims were excluded when the evidence did not support production prevalence, universal reliability, legal mandate, vendor-specific outcomes, or quantified business impact. Analytic synthesis and the 12–24 month forecast were evaluated separately from source-level facts. The public brief and this packet use the same evidence boundaries.
Sources
Numbered to match the citations in Brief №015 and in this packet.
- Policy-Invisible Violations in LLM-Based Agents, arXiv, 2026.
- R2-KG: General-Purpose Dual-Agent Framework for Reliable Reasoning on Knowledge Graphs, Findings of IJCNLP-AACL, 2025.
- Ontology-to-tools compilation for executable semantic constraint enforcement in LLM agents, arXiv, 2026.
- World Wide Web Consortium — Shapes Constraint Language (SHACL), W3C Recommendation.
- Cedar Policy — How Cedar authorization works.
- Forrester — Build Meaning Before Machines: Why Semantics, Ontologies, and Knowledge Graphs Matter for Agentic AI.