Between actions and evidence · Observe · Correlate

Portable telemetry and evidence
for AI-agent runtimes

Normalize the governance facts around an agent run: policy decisions, approvals, actions, classified data flow, usage, cost, and evidence, without replacing your OpenTelemetry stack, policy engine, framework, or UI.

Alpha contract 0.1.0-alpha.4 · Python and TypeScript · MIT licensed
Where it fits

One correlation layer across the trust chain

AgenTrust Telemetry sits between the actions and evidence steps of the AgenTrust chain. It records runtime facts in a backend-neutral contract, correlates them with the application's existing OpenTelemetry context, and can finalize a complete durable evidence set into a TRACE record.

01Agent ManifestAgent step: declared identity and authority
02cMCP / cA2AActions step: policy and delegated actions
03TelemetryCorrelated governance facts
04TRACEEvidence step: verifiable evidence

Agent applications can already emit model and tool traces. Governance facts are the ones that stay trapped: in policy-engine decision logs, approval databases, cost modules, and proprietary dashboards. This gives those facts one privacy-conscious contract and correlates them with the trace the application already produces.

The contract

Six event families

Each family is a JSON Schema 2020-12 contract with valid and invalid fixtures. An event reports a fact observed elsewhere; this project never makes the authorization decision itself.

FamilyWhat it records
Policy decisionAllow, deny, challenge, error, enforcement mode, policy identity and timing
Approval lifecycleRequested through terminal decision and execution outcome, bound to an action digest
Action executionResolved tool, MCP, A2A, file, HTTP, and database attempts
Data flowClassified source-to-destination metadata, without payload capture
UsagePer-call and per-run token and cost facts, with explicit cost provenance
Evidence lifecycleRun checkpoints, completeness, and optional TRACE finalization status

Privacy boundary: the metadata-only profile prohibits raw prompts, outputs, source code, tool arguments or results, credentials, and authorization tokens. Propagated identifiers are correlation data, not proof of identity or authorization.

Use it

Keep the stack you already operate

The reference SDK uses the caller's current OpenTelemetry span. It does not install a provider, processor, exporter, collector, or global propagator. Standard OpenTelemetry fields take precedence over AgenTrust extensions.

Before you start: Python 3.11+, Git, and Bash on macOS, Linux, or Windows with WSL. Commands run from the repository root. No telemetry backend is needed for the first example.

1

Install the Python SDK

Install from the source checkout so the SDK and runnable examples use the same revision.

Terminal
git clone https://github.com/agentrust-io/agentrust-telemetry
cd agentrust-telemetry
python3 -m venv .venv
source .venv/bin/activate
python -m pip install -e ".[otel]"
2

Run a complete event example

Run the included synthetic policy event. It defines every required field, validates the event, and prints it through a local log emitter.

Terminal
python examples/manual_governance.py
3

Carry correlation across process boundaries

Integration sketch: this block assumes your application already defines tracer and event. For a complete runnable version, use python examples/governed_workflow.py after installing ".[test]" as shown below.

For synchronous cross-process agent calls, propagate the caller's W3C context and the durable AgenTrust identifiers, then use the extracted context as the receiving span's remote parent. For asynchronous queue handoffs, start a new trace with links=[remote.link()] instead: that preserves causality without representing queued work as a synchronous child span.

Python
from agentrust_telemetry import extract_context, inject_context

carrier = {}
inject_context(carrier, run_id="run-123", agent_id="planner")

remote = extract_context(carrier)
with tracer.start_as_current_span("worker", context=remote.otel_context):
    event.update(remote.event_fields(agent_id="worker"))

Propagated metadata is untrusted input. It does not establish agent identity or authorization.

4

Or use the TypeScript SDK

Use a separate terminal at the repository root, with Node.js and npm installed. Return to the root for the Python examples below.

The pre-alpha Node package validates the same fixtures and the same privacy profile as Python, and preserves nanosecond wire timestamps as decimal strings. It installs no OTel provider, exporter, or global propagator either.

Terminal · packages/typescript
cd packages/typescript
npm ci
npm run check

Expected first result: a printed policy event and an EmitResult with accepted=True and log_emitted=True. A missing span is expected without an active OpenTelemetry span. This synthetic event does not enforce a policy.

Two runnable examples

Both ship in the repository. The second walks the complete governance, OpenTelemetry, durable-evidence, and TRACE reference workflow.

Terminal
python examples/manual_governance.py

python -m pip install -e ".[test]"
python examples/governed_workflow.py
Adapters

Map what your policy engine already emits

You should not have to hand-build envelopes. The event factory owns specification version, producer identity, event UUID, timestamp, correlation fields, and both schema and privacy validation, so adapters only translate a source record. All of these exist in both Python and TypeScript, and TypeScript takes plain source records with no runtime dependency on the upstream package.

Cedar

Accepts the final Allow or Deny response plus determining policy IDs and caller-normalized diagnostic codes. Evaluation errors are recorded without rewriting the final decision, matching Cedar's skip-on-error semantics. Free-form error messages are rejected so content cannot leak through a diagnostic.

OPA

Accepts one decision-log event. Booleans map to allow or deny, an absent result maps to not-applicable, and structured results require an explicit mapper because OPA permits any JSON value. Input and result documents are never copied, and a trusted bundle digest stays required because a revision is not a content digest.

AGT bridge

Implements the batch sink shape used by Agent OS without making AGT a core dependency. The whole batch is normalized before the first event is emitted, so a malformed later event cannot cause partial delivery. Durable evidence callbacks must stay idempotent by event ID.

AGT's action-bound approval protocol maps as a linked sequence: the policy decision becomes a challenge, the approval request references that normalized policy event, and only a resolution becomes approved or rejected. Read the adapter reference.

Verify it

Don't take the contract on trust

The repository ships valid and invalid conformance fixtures and an independent Python runner. The invalid fixtures are the interesting half: raw content, a numeric timestamp, a zero trace ID, an action without a digest, usage without a measurement, an aggregate cost without a rollup. Each one must be rejected.

Terminal
python -m pip install -r conformance/requirements.txt
python conformance/runner/validate.py
python -m unittest discover -s tests -v

What the validator does not check. It verifies schema conformance and the metadata-only privacy invariant. It does not yet validate OTLP projection or TRACE mapping, so a passing run is evidence about the event contract, not about the whole pipeline.

Contract principles

The rules the schemas enforce

Known limits

What is not true yet

Stated here rather than in a footnote, because a governance evidence layer that oversells itself is worse than none.

The contract is alpha. Version 0.1.0-alpha.4 carries no stable SDK API and no compatibility guarantee. It may change incompatibly.

The SDK validates declared metadata; it cannot prove the declaration is true. A producer can report an inaccurate policy decision, identity, classification, token count, or cost. TRACE can bind signed evidence, and verified hardware attestation can establish runtime provenance within its trust policy. Neither automatically establishes the truth of every reported fact.

Operational delivery is not audit evidence. OpenTelemetry spans may be sampled or dropped. Complete evidence must be accumulated durably before any TRACE claim is made, and TRACE finalization here is software-only: it requires explicitly complete evidence plus trusted caller configuration.

Evidence memory mode is not durable. Callback mode defines acknowledgement and retry behaviour, but the adopter owns storage, idempotency, and recovery.

Interoperability has been exercised locally against OpenTelemetry Python, not across collectors, backends, or languages. Action telemetry also records resolved attempts only; it does not expose in-flight lifecycle transitions.

Boundaries

What it deliberately does not become

It is not a telemetry collector, storage service, dashboard, or SaaS backend. It is not a policy engine or an approval workflow. It is not an agent framework, not a model pricing catalog, and not auto-instrumentation duplicating OpenTelemetry GenAI or OpenInference.

Reference

The documents to implement against