AI audit trail & evidence

Architecture direction

If AI acted, the business should be able to reconstruct what happened.

AI audit is not a demand to expose private model reasoning. It is a requirement to preserve the operational evidence around material work: who initiated it, which agent acted, what context was used, which tools were called, what policy decided, who approved and what action was actually released.

That is the evidence model behind the TEMRIK AI control plane: organisational authority should remain reconstructable even when the underlying model changes.

Identity · context · models · tools · policy · approval · action · outcome

Operational evidence record

01

Actor

User, agent, service identity and tenant.

02

Goal

The authorised task or business objective.

03

Context

Sources retrieved, records referenced and evidence identifiers.

04

Model

Provider, model identifier and relevant runtime metadata.

05

Tools

Tool requested, scoped arguments, result status and duration.

06

Policy

Allow, deny, restrict, escalate or request-more-information result.

07

Delegation

Which agent or service handed work to another actor.

08

Approval

Who approved a consequential release, and when.

09

Action

What was actually sent, changed, created, released or executed.

10

Outcome

Success, failure, exception, reversal or later business result.

Observability does not require private model reasoning.

It requires operational evidence.

The audit question

You do not need to know every thought.You need to know every material action.

Traditional application logs often tell engineers that something failed. Enterprise AI evidence needs a wider question: can the organisation prove the path from intent to action?

Who initiated the work?
Which tenant and identity boundary applied?
What business goal was authorised?
Which records were retrieved?
Which model/provider processed the task?
Which tool calls were requested?
What policy result was produced?
Was work delegated to another agent?
Who approved the consequential step?
What action reached the external system?
What outcome followed?
Can the evidence be correlated later?

Reconstructable action

A useful audit trail follows the work, not just the model call.

The evidence boundary should continue across retrieval, reasoning interfaces, tools, policy, human authority and the final business system. The model response is one event inside a larger operational run.

Event 01

Goal

A bounded task enters the system with actor and tenant context.

Event 02

Context retrieved

Permission-aware sources are selected and referenced.

Event 03

Model called

The chosen provider/model processes the defined task.

Event 04

Tool requested

The agent requests a capability with structured arguments.

Event 05

Policy result

A control layer allows, restricts, blocks or escalates.

Event 06

Human approval

A named authority approves when the action threshold requires it.

Event 07

Action

The permitted business action is released to the target system.

Event 08

Outcome

Result, error, exception and relevant evidence are correlated.

Material-action timeline

01

GOAL

02

CONTEXT RETRIEVED

03

MODEL CALLED

04

TOOL REQUESTED

05

POLICY RESULT

06

HUMAN APPROVAL

07

ACTION

08

OUTCOME

On mobile this sequence remains a vertical narrative: each event stands alone and the evidence chain still reads in execution order without relying on animation or horizontal scrolling.

The vocabulary matters

Logging, tracing, audit, evidence, monitoring and evaluation are related.They are not the same thing.

Treating every telemetry stream as an audit record produces either too much data or too little proof. Each layer serves a different operating purpose.

Logging

Discrete records of events or state, usually optimized for search and diagnosis.

Tracing

A correlated sequence of spans showing how one request moved across models, tools, agents and services.

Monitoring

Ongoing visibility into health, latency, errors, volume and operational conditions.

Audit

A business-control record designed to demonstrate who did what, under which authority, and with what result.

Evidence

The records, references and approvals needed to support a material claim or action.

Evaluation

A structured assessment of quality, safety, policy compliance or task performance.

What to capture

Record enough to explain the business event.Not everything because storage is cheap.

A defensible evidence design begins with purpose. Decide which events matter, how long they matter, which identifiers link them, who can inspect them and which fields should be minimized or redacted.

Actor

User, agent, service identity and tenant.

Goal

The authorised task or business objective.

Context

Sources retrieved, records referenced and evidence identifiers.

Model

Provider, model identifier and relevant runtime metadata.

Tools

Tool requested, scoped arguments, result status and duration.

Policy

Allow, deny, restrict, escalate or request-more-information result.

Delegation

Which agent or service handed work to another actor.

Approval

Who approved a consequential release, and when.

Action

What was actually sent, changed, created, released or executed.

Outcome

Success, failure, exception, reversal or later business result.

Privacy and telemetry

Evidence principleProvider dependent

What should never be logged carelessly?

Trace systems can become one of the richest datasets in the organisation. That makes collection policy, redaction, retention, access control and deletion part of the architecture—not an afterthought.

Passwords and secrets

Credentials, API keys, tokens and private keys should not become ordinary telemetry.

Unnecessary sensitive data

Collect only what is proportionate to the operational purpose and retention need.

Private chain-of-thought

Audit should capture material actions and evidence, not demand hidden model reasoning.

Unrestricted raw payloads

Full prompts, files, tool inputs and outputs can create a second sensitive-data store.

Security material without controls

Trace stores need their own identity, access, encryption, retention and deletion rules.

Private reasoning is not the audit target.

For enterprise control, the more reliable audit question is whether the organisation can reconstruct the observable operations: inputs selected, external evidence referenced, tools requested, policy decisions, approvals, actions and outcomes. This avoids making hidden chain-of-thought a governance dependency.

Observability architecture

Agent observability should cross provider boundaries.

OpenTelemetry provides a widely used foundation for correlated traces, metrics and logs, while major agent platforms are adding their own observability layers. The enterprise design problem is to preserve a useful business record even when individual providers expose different telemetry.

Execution telemetry

Trace IDs, spans, latency, status, model and tool operations.

Business evidence

Source references, policy outcomes, approvals, released actions and business results.

Control evidence

Identity, tenant, permission, delegation and exception decisions made outside the model.

TEMRIK architecture direction

01

Correlation

Carry durable run, actor, tenant and workflow identifiers.

02

Policy evidence

Record the allow / deny / escalate decision independently of the prompt.

03

Authority

Bind human approvals to the material action they release.

04

Outcome

Correlate the final system response or business result back to the run.

Human authority

Approval should be evidence, not just a button click.

When a workflow crosses an action ceiling, the audit record should be able to identify what the person saw, which action was proposed, who had authority, when they approved it and what was ultimately released.

See how this fits into TEMRIK Human Control and the wider agentic AI architecture.

1

Proposed action

What the agent wants the business to release.

2

Evidence bundle

The sources and context relevant to the decision.

3

Policy state

Why the request requires approval rather than automatic execution.

4

Decision owner

The named role or person with authority over the consequence.

5

Decision

Approve, reject, amend, request more information or escalate.

6

Released action

The exact action that reached the external system.

Security boundary

Audit data is security-sensitive data.

Observability can reveal user inputs, retrieved records, tool arguments, operational metadata and system structure. The evidence platform therefore needs its own access model and should follow the same least-privilege discipline as the workflow it observes.

The related AI security architecture explains why data boundaries, model choice, permissions and actions should be controlled outside the model itself.

Access

Limit evidence access by tenant, role and operational need.

Retention

Keep evidence for a defined purpose and period rather than indefinitely by default.

Integrity

Protect high-value audit events from silent alteration or deletion.

Redaction

Remove or mask sensitive fields that are unnecessary for the evidence purpose.

Separation

Do not make the same compromised agent the sole authority over its own audit record.

Export

Design for defensible review without uncontrolled bulk extraction of sensitive telemetry.

How it fits together

Auditability is the memory of controlled AI.

TEMRIK's operating approach is to place identity, company context, policy, playbooks, tools and human authority around model capability. A reconstructable evidence trail is how those controls become reviewable after the event.

Control loop

01

Identity establishes the actor.

02

Policy establishes the boundary.

03

Tracing establishes the execution path.

04

Evidence establishes what supported the action.

05

Human approval establishes consequential authority.

06

Audit establishes what happened afterwards.

Start with one workflow

Map the evidence trail for one AI workflow.

Identify the actor, data, model, tool, policy, approval point, released action and outcome that would need to be reconstructed if the workflow were challenged six months later.

Agents need rules as well as traces. Get the free 27 Rules of Peace field guide, or return to the TEMRIK platform.