Fictional workspace. Real automation pattern. No customer data is used. Run the demo →

Resolve routine support questions without inventing company policy.

Provenance reads approved documents, drafts a cited answer, and knows when to involve your operations team — visibly, with the reasoning shown at every step.

Open the guided demoSee reliability evidence

No signup. No login. Live pipeline, not scripted.

51
indexed passages
43
eval cases
3
responsible outcomes

From policy update to accountable action.

The system automates the safe portion of support work and preserves human review where business judgment is required.

1
Approved policies

Operations controls the source material.

2
Index updated

Changed documents become searchable.

3
Ticket arrives

Email, chat, or helpdesk request.

4
Evidence retrieved

Relevant passages come back ranked.

5
Claims verified

Every statement must be supported.

6
Reply or route

Send safely, or involve staff.

7
Decision logged

Evidence and outcome stay auditable.

Cited answers

Every material claim maps to an approved policy passage.

Human judgment

Unsupported requests are routed to a person, not guessed at.

Safety gates

Unsafe instructions are blocked before generation, not after.

Audit history

Every automated decision leaves a real, persisted record.

Measure time saved without hiding the tradeoffs.

Adjust the assumptions for your own operation. These are not measured customer results — this demonstrates the calculation.

Illustrative calculator

Adjust the assumptions for your own operation. These are not measured customer results — this is a demonstration of the calculation, seeded with the fictional example from the product plan.

Tickets eligible for automatic resolution
270

per month

Staff hours potentially returned
41.9

per month

Human-review rate (target)
18%

of tickets not auto-resolved

Estimated handling-cost reduction
$1,465

per month, before automation cost

Estimated automation cost
$1.35

per month · ~$0.005/resolution

Estimated payback period
6.8 mo

on the one-time implementation cost

A scorecard, published as-is.

43 evaluation cases test routine answers, unsupported questions, and adversarial prompts against the same pipeline the demo runs.

100%Accuracy across 43 cases
0%False refusal rate
0%Fabrication rate
3Test buckets: answerable, unanswerable, adversarial
Test groupCasesAccuracyFabrication
Answerable21100.0%0.0%
Unanswerable13100.0%0.0%
Adversarial9100.0%0.0%
Overall43100.0%0.0%

Live scorecard from the committed eval suite — shown transparently, not reframed as production performance.

Known limitation

A clean scorecard is a dev-set number, not held-out proof.

This 43-case set was iterated against directly — two real fabrication bugs were found and fixed during development. A perfect score on the set used to find and fix those bugs isn't evidence the fix generalizes.

  • Next: build a held-out eval set never tuned against
  • Next: calibrate thresholds on that set, not this one
  • Until then: treat auto-send as demo-grade, not production-grade

More than a chat box over documents.

A portfolio case study covering the full decision path: controlled knowledge, retrieval, generation, verification, routing, and measurement.

SYSTEM / 01
Grounded retrieval

Policy documents are chunked, indexed, ranked, and shown beside every proposed response.

SYSTEM / 02
Claim verification

Generated statements are checked against retrieved passages before an outcome is chosen.

SYSTEM / 03
Decision routing

Evidence sufficiency determines reply, escalate, or block — before generation, not after.

SYSTEM / 04
Evaluation harness

Answerable, unanswerable, and adversarial cases track accuracy, fabrication, and latency.

Read the full architecture walkthrough →

Ready to see a refusal happen live?

This prototype shows how policy-heavy businesses can reduce repetitive work while keeping evidence, oversight, and failure modes visible.

Replay the demo