Audit evidence

Logs are not evidence

A log line proves a request happened. An auditor, a customer's security team and your own board want something stricter: who ran it, under what authority, which model answered, what it cost β€” and whether the record can still be trusted months later. Most AI stacks quietly fail that second question.

The gap, concretely

  • Cost lives in the provider invoice, the decision lives in a code review, the identity lives nowhere β€” so no single record answers a question end to end.
  • Preparing for a review means an engineer exporting CSVs and stitching them together by hand.
  • "Show me every AI decision from March in this product" is currently a two-week project.
  • Your logs can be edited by anyone with database access, and nothing would show it.

What we actually have

Every capability on this page carries its real status: shipped code, work we do inside a pilot, or a design we have not built yet.

Cost, model and actor on one record

Live today

Our own router writes a record per call: who ran it, the source and operation, provider and model, tokens, latency and estimated cost in USD. That is a real table with real rows β€” it is the shape everything else builds on.

A published field list you can check

Live today

The evidence-pack structure is written out field by field on a public page, with each field marked as existing or design. Take it to your auditor before you take a call with us.

Scope instead of sampling

Pilot scope

Evidence is always about a scope: this system, this quarter, these decisions. Agreeing that scope and filling it is what a pilot does.

Tamper-evidence and packaged export

Design intent

Hash-chaining records so a silent edit breaks the chain, and exporting a scoped pack β€” both are designed and neither is built. No download button here will pretend otherwise.

How to check us

1. Read the field list

Live today

The proof page states every field, with its status. If a field your auditor needs is marked design, you have learned that in two minutes instead of two calls.

2. Bring your reviewer's actual question

Live today

Not "do you have audit logs" but the exact question you keep failing to answer. We map it to fields, or we tell you we cannot.

3. Fill one scope in a pilot

Pilot scope

One system, one period, records that survive being read by someone who was not there.

What exists today

  • Per-call records with actor, source, operation, provider, model, tokens, latency and estimated cost β€” in our own router, backed by tests.
  • A public, field-by-field evidence-pack structure, each field labelled existing or design.
  • The free External AI Exposure Check, which finds AI use nobody registered β€” the input to any honest inventory.

What we do not claim

  • We do not have tamper-evident storage. Hash-chained records are design intent, not a shipped feature.
  • There is no evidence-pack download today, and we will not fake one.
  • Team, project and policy-result fields do not exist yet; a record today carries actor, source, operation, provider, model, tokens, latency and cost.
  • We hold no SOC 2 attestation and publish no customer names.
  • Records only ever cover traffic that is instrumented. Anything running around it leaves nothing β€” no vendor can honestly say otherwise.

Questions people actually ask

How is this different from our observability tool?

Observability is built for debugging: high-volume, sampled, short retention, freely mutable. Evidence is built for review: scoped, complete inside that scope, retained on your schedule, and ideally hard to edit silently. We are honest that the "hard to edit" part is still design.

What is actually in a record right now?

Actor, source, operation, provider, model, input/output tokens, latency and estimated cost in USD, with a timestamp. The proof page lists it in full alongside the fields that do not exist yet.

Do you filter PII out of prompts?

No. Today there is no PII or secret filter in the path β€” the only output check looks for refusals and absurd lengths. Treat prompt content handling as pilot-scoping work, not a shipped guardrail.

Can records live in our own infrastructure?

That is pilot scoping, not a switch we flip. The honest answer depends on your residency constraints, and we would rather discuss it than promise a deployment mode we have not built.

Who decides this is enough?

Your auditor or counsel. Our job is to make sure that when they ask, the answer is a query against a defined structure rather than an archaeology project.

Send us the question your reviewer keeps asking

We will map it to the fields that answer it β€” or tell you plainly which part of it we cannot cover today.

Prefer to check us first? The record structure is written out field by field on the evidence page, and the external check shows its full report without an email.

Request an evidence walkthrough

Email or Telegram is enough. A person replies; there is no automated sales sequence.

Also useful

Other problems we cover