Finality Assurance™ Standards

Know what work you can rely on. Know why. Know when that changes.

A person, a system or an AI agent says the work is done. Before your business acts, you need to know what was required, what actually happened and whether the evidence still applies. Finality Assurance Standards (FAS), Expound’s first product, figures that out from the evidence instead of taking anyone’s word for it. It is available and licensable now. Start with the decision that matters to your business.

The Model Router · in development · built on Finality Assurance Standards

Today’s routers pick
a model. The work
needs a route.

A route is the model or the person, the tools, the evidence to be gathered, the checks to be run, the stopping rules, the budget, and when to escalate. The Expound Model Router selects the whole way a job gets done. It rules out every route that cannot prove what the job requires before anything is ranked, then hands the result to a kernel that figures out what you may rely on.

Early access

Be first on the router.

Sign up and we will tell you the moment it ships. In the meantime we will send the routing argument, the comparison and the measured results.

Top: a request goes to a router that ranks models on cost, latency and a quality score, picks one model, and the answer is taken as done. Bottom: a typed call resolves what this call must prove, routes that cannot meet it are excluded before ranking, admitted routes each carry a model or a person plus tools, evidence, checks, stopping rules and a budget, the route runs and returns evidence, and the Finality Kernel computes what may be relied on, or holds and names the failing condition. RECORDED BENCHMARK · THE SAME 36 TASKS, EVERY OUTCOME CORRECT IN BOTH ARMS A third fewer model calls. None of them to a frontier model. Each block is one model call spent to reach the same 36 correct outcomes. Fewer is better. Conventional routerstrongest fixed baseline 36 36 calls, every one to a frontier model Expound Model Routersame tasks, same correct answers 24 12 calls never made · zero frontier calls 12 of 36 outcomes reached by an exact, rule-based substitute, with no model call at all. The other 24 went to non-frontier models. A CONVENTIONAL ROUTER Requesta prompt Routerranks models on cost · latency ·a quality score Model Bthe pick Model A Model C Answerarrives, and stops “Done”assumed from the answer arriving THE EXPOUND MODEL ROUTER · BUILT ON FINALITY ASSURANCE STANDARDS Typed callwhat is being asked,for what consequence What this callmust proveevidence set, resolved per callfrom a governed profile Admissibilitybefore rankingroutes that cannot meet itare excluded, not down-weighted Route: model + toolsevidence · checks · stop rules · budget Route: deterministic computationexact, no model call at all Route: a personrules-selected; its own evidence contract Route: excludedcannot produce the required evidence Rank admittedroutes, run onelearning only ranks;it never widens the set returns evidence, not a claim Finality Kernelcomputes what may be relied on,by whom, for what, now Verdict + licensea named party may act,for a named decision Or: HOLDnothing admissible, or evidence short:it names the failing condition
What a conventional router decides, and what the Expound Model Router decides. The Expound Model Router never chooses what evidence counts and never creates Accepted Work; the Finality Kernel does that, and the router can be told no.

What is different, in six ways

Six structural differences.

Routers on the market rank models. This one selects routes, and it is built on a standard that computes whether the result may be relied on. The differences are structural, not tuning.

It selects a route, not a model

Two routes can use the same model and still be different routes, because the evidence, the checks and the stopping rules attached to them differ, and so does the full cost of reaching an outcome. The route selected is tied to the route that actually ran.

Eligibility is decided before anything is ranked

For each job, the evidence it must produce is looked up from an approved profile. Routes that cannot produce it are ruled out before anything is ranked. No score buys its way around a hard requirement.

It holds instead of guessing

When no eligible route remains, the router stops and names exactly what was missing: something you can act on, not a quiet best-of-a-bad-set.

A person is a route

Handing the job to a person is a route like any other, with its own evidence requirements, chosen by the same rules, not bolted on when automation fails.

Learning that cannot promote itself

The learning half, Constrained Policy Reinforcement Learning, only ranks routes that have already been ruled eligible. It cannot decide eligibility, change evidence requirements, weaken obligations, widen what may be claimed, or promote itself. It can be told no.

Built on the standard that computes reliance

The route returns evidence, never a claim. The Finality Kernel computes what a named party may rely on, for what, now, and recomputes when the basis moves. That kernel is Finality Assurance Standards, available and licensable today.

Harvey-ball comparison of today’s model routers against GMRS plus CP-RL across eleven dimensionsHarvey-ball comparison of today’s model routers against GMRS plus CP-RL across eleven dimensions
What each layer keeps track of. A router selects a model; a governed route selects the whole path, and when nothing is eligible it holds and names the condition that failed.
24 vs 36model calls to reach the same outcomes as the strongest fixed baseline on a 36-task set: a third fewer, with every outcome correct. The battery saturated: it licenses no parity claim, and it does not show lower total cost.
12 of 36outcomes reached by an exact, rule-based substitute, with no model call at all. The other 24 calls went to non-frontier models.
0wrong or fabricated answers accepted for reliance under computed acceptance, across 2,520 graded answers from six language models. The models still produced them; none got through.

The core it is built on

The router picks the route. FAS decides what you may rely on.

The Model Router is built on the product at the top of this page. Finality Assurance Standards computes whether consequential work is complete and what may happen next, from evidence a second party can re-check, and recomputes it when something underneath changes.

One standard and four products: Finality Assurance Standards itself, the Model Router, the Mechanistic Conflation Engine, and Work Management. Much of this page is about the question the standard answers: when is an agent done?

The question

Who in your organization may rely on an agent’s result, for what, right now?

There are two ways to answer when is an agent done. The first: the agent stops and hands something back, and you take that as done. The second: the agent has produced evidence that the intended change actually happened, someone with the standing to accept it has done so, and a named person may act on the result for a named purpose, now.

The intended change is the thing that matters, not the agent’s report that it made it, and not a call returning without an error. If the job was to revoke someone’s access, done means the account can no longer sign in.

Agents do not answer this question at all. An agent returns an answer and stops. Sometimes a field gets written, sometimes nothing is written, and either way nothing obliges the result to still be true tomorrow. Today an agent has no obligation to produce any evidence at all.

This is easy to miss, because the loop looks complete. The agent was authorized. It ran. It reported the effect verified. Nothing is broken; it ended exactly where it was built to end. What it was not built to do is stay interested. The result travels outward and the governance stays behind: a result is produced once and relied upon many times, by parties the producer never listed, for purposes the producer never knew.

FAS gives you the second answer, and keeps giving it as the world moves.

An answer arriving≠Done

An agent does not assert that done is done. It returns an answer, and done is what you assume from the answer arriving. Nobody can rely on an assumption.

Five questions, usually collapsed into one word

A status field is memory. It is not a measurement.

Where a system does record that work is finished, it is a flag on a row: approved, complete, done. The flag records that some process once reached a conclusion. It does not check whether the conclusion still holds. A status field is a promise that a human will notice when it stops being true. And with an agent, often there is no field at all: the answer being there is the whole of it.

It can go stale while the record stays byte-identical. It can hide uncertainty, because one label absorbs a whole distribution. It is usually written by the party with the most reason to say yes. And when the underlying evidence is corrected, the flag does not update itself.

What you actually need to knowWhat FAS computes separately
Was this allowed?Admission
Did it happen?Execution, and Occurrence, separately
Did the intended result follow?Outcome support
What may we say about it?Canonicality (which record wins), and Assertion (what may be claimed)
May this person act on it now?Reliance, and Applicability

Eight decisions, each kept separate. A process can run without its effect occurring. An effect can occur without authorization. An authorized, verified effect can still carry no license for a particular person to act on it. FAS never lets one stand in for another.

A call returning≠The change actually happening

Confidence≠Permission to rely

Provenance tells you how a result was made. Confidence tells you how sure the producer felt. Neither tells you whether you may act on it, for this decision, today. FAS does.

Evidence before done. Reliance after.

Two things every workflow is missing, on either side of “done.”

Evidence comes first, because done is computed from it. Done means two things: acceptance (someone with standing took the result as satisfying what it was for) and finality (that acceptance is settled, not provisional).

Both are computed from the evidence attached to the result, item by item, not summarized. Evidence is not bolted on after the work. It is what makes done determinable at all.

Reliance comes after, and the question is live twice. Right now: an agent hands you a result. May you act on it, for which decision, at what strength? The same result can be sound enough to draft a summary and not sound enough to move money.

And later, as things change: the documents it drew on, the model version, an upstream result. FAS is the computation in your stack that answers what anyone may rely on the result for, now.

One worked example, run twice

Same invoice. Same recommendation. Same starting evidence.

Two ways to decide. The ordinary workflow marks it payable. FAS checks the evidence against what this kind of payment actually requires, and computes whether Finance may act.

Stages 0–1 · 1:46Full recording
Why FAS is different, and the whole demonstration in one screen.
Ordinary workflowPAYABLE

The false pass

An invoice arrives and an agent recommends paying it. The purchase order matches. The vendor, the amount and the currency match. Finance has approved, and a purchasing email says the goods arrived. The workflow marks it payable.

By writing payable, the agent has stated the invoice is payable, and the statement leaves out a required record: an authorized confirmation from Warehouse Receiving. The purchasing email is useful information and it is not the record this decision requires. Nothing computed acceptance. Nothing computed whether Finance may act.

The omission is invisible, because the field says payable and the field is all there is. This is the ordinary failure mode of a confident agent: the answer sounds right and the missing piece never surfaces.

Stages 2–3 · 2:17Full recording
The ordinary workflow runs and marks the invoice PAYABLE. The omitted receiving proof never surfaces.
Under FASHOLD

Naming the exact proof that is absent

The list of what must be proved for this kind of work, its obligation roster, names six things that must be proved before payment. Five are supported. The sixth, the Warehouse Receiving confirmation, is not.

So the result is hold, naming the missing proof, with no payment action authorized. Nothing here concludes the goods never arrived. FAS refuses to rely on the payable claim while the required evidence is missing.

Accepted Work 0
Stage 4 · 1:26Full recording
FAS resolves the profile, finds five of six obligations supported, and returns HOLD naming the missing proof.
Independent checkHOLD

Who watches the watcher

A separate read-only verifier examines the same primary evidence by its own path and also returns hold. It can narrow a result, hold it, or refuse it, and it cannot raise a hold to a pass. A verifier able to turn a hold into a pass would be a second producer, and its agreement would be worth no more than the first one’s.

The Four-Corners Test™ runs alongside and asks a different question: can a second party reconstruct the determination from what is attached to it and nothing else? It passes. That does not mean the invoice passed; it means the reasoning for holding it can be rebuilt by someone who was not there.

Stage 5 · 1:04Full recording
The independent verifier returns HOLD by its own path. The Four-Corners Test passes.
Evidence suppliedPASS

A license to act, not the action itself

Warehouse Receiving supplies the authorized confirmation. The signature is checked, so is the issuer’s authority to give it, so is the invoice it applies to, and so is whether it is still current. All six requirements are now supported.

The result becomes a pass and a reliance license, a scoped permission to act: Finance may mark this invoice payable, for this payment attempt, while that evidence stays current. No payment was made. What was authorized is the next controlled action, not the action itself.

Accepted Work 1
Stage 6 · 0:54Full recording
The receiving record is admitted, all six obligations are met, and the result becomes PASS with Accepted Work 1.
Basis movesWITHDRAWN

History preserved, permission withdrawn

An authorized revocation record is admitted, and everything recomputes. The earlier pass stays in the record, and so does its unit of Accepted Work. Nothing erases what was true before.

What changes is what may be relied upon now: the receiving evidence is no longer current, so that accepted work leaves the set currently eligible to be relied upon, the payment permission is withdrawn, and further payment action on that evidence is blocked.

Three things are kept apart: what was accepted historically, what may be relied upon now, and what may be done next. The ordinary workflow holds only the first.

The recomputation produces no new Accepted Work
Stage 7 · 1:12Full recording
A revocation is admitted. Permission is withdrawn; the earlier pass stays in the record.
Flow diagram of the invoice example: PAYABLE in the ordinary workflow; HOLD, HOLD, PASS, WITHDRAWN under FASFlow diagram of the invoice example: PAYABLE in the ordinary workflow; HOLD, HOLD, PASS, WITHDRAWN under FAS
The same invoice under two decision methods. PAYABLE is the ordinary workflow's false pass; HOLD, HOLD, PASS and WITHDRAWN are the FAS path, with Accepted Work counted only at the evidence-cleared PASS.

The invoice demonstration is a working example that makes the computation visible. It runs on the FAS standards family and uses selected components on one invoice profile. PAYABLE, HOLD and PASS are its on-screen labels; the standard’s own verdict values are BLOCKED, UNKNOWN, DEGRADED, VERIFIED. An invoice is a convenient example because everything in it already has a name. The architecture is not invoice-specific: the same computation governs a release, a trade, a batch, or a benefits determination.

How this compares

The common approaches answer a different question.

Running the work several times and picking the best, having a model judge the output, an evaluation suite, traces and observability, a policy engine, provenance: each is useful, and each is a tool for the side that produced the work. FAS is built to answer a further question: what a second party may act on, from the evidence, and when that stops being true.

Harvey-ball comparison of best-of-N, model-as-judge, eval suites, traces, policy engines and provenance against FAS across seven questionsHarvey-ball comparison of best-of-N, model-as-judge, eval suites, traces, policy engines and provenance against FAS across seven questions
What each approach answers. A filled circle means the approach answers the question for a named party, from evidence, and keeps the answer current. An empty circle means it does not. This compares what each approach is built to answer, not how each performed on a shared workload.

Where do you stand?

Score one decision. Find your band.

Seven capabilities every organization needs before it can rely on machine-produced work. Each starts with the question you would actually ask, then lays out what most teams do today, what best practice looks like, how Expound delivers it, and what to measure. Then a thirteen-question self-assessment tells you your maturity band and what to work on first.

0

Absent

The capability does not exist. Nobody could answer the question.

1

Asserted

Somebody declares it. The declaration is the record, and it is not checked.

2

Documented

It is written down and reviewable, but assembled after the fact, by the party being examined.

3

Evidenced

Evidence exists at the moment of change, is typed, and is separable from whoever produced the work.

4

Computed

The answer is derived from that evidence on demand, cannot be written by anyone, and is withdrawn when its basis fails.

The economics

A cost no budget line carries.

$0.8–1.4 trilliona year, modeled, worldwide, on verification, attestation, review, reconciliation and audit-evidence work, performed by tens of millions of people and tracked by almost nobody as a single cost.

When machines produce work faster than people can check it, an organization has two ways to keep its decisions defensible: refuse to hand consequential work to machines, and give up the speed; or accept results on the producer’s say-so, and take on silent risk at machine speed.

Most organizations pay to avoid choosing, with reviewers, reconciliation and reconstruction. Reruns, approval chains, audit reconstruction, retries, second reviews and rework are the actual trust layer of modern operations.

Added up, they are a reliance tax: paid continuously, at expert rates, for something the organization still does not end up with, because none of that spend makes the next decision defensible on its own.

Expert review is the scarce input. When production multiplies and review does not, one of three things happens: review becomes spot-checking, spot-checking becomes sign-off, or the backlog becomes the product. All three turn the acceptance decision into an assumption.

Produced≠Accepted

What we measured

Zero wrong or fabricated answers accepted for reliance, across 2,520 graded answers.

Across 2,520 graded answers from six language models, the ordinary accept-on-claim approach accepted seven wrong or fabricated answers. Under computed acceptance, none was accepted for reliance. The models still produced them; the evidence requirement stopped them. Evidence became a real condition of getting in, and on a fixed 36-task set governed routing reached the same outcomes as the strongest static comparator with a third fewer model calls. On a separate live campaign it cost outcome quality, and that result is published beside this one.

Three panels: 7 versus 0 false accepts; 0 of 270 evidenced answers refused against 90 of 90 unevidenced refused; 24 versus 36 model callsThree panels: 7 versus 0 false accepts; 0 of 270 evidenced answers refused against 90 of 90 unevidenced refused; 24 versus 36 model calls
What we measured: decision quality, evidence-gated admission, and route selection.