Home · Writing · Design

Evidence-First UX: Showing Claims with Their Provenance

TLDR

  1. A wrong balance figure in a customer letter shows why generated claims need visible, per-claim provenance instead of a single buried sources link.
  2. The following is an illustrative composite , not a disclosed client incident. Its balances, sample sizes, timings, remediation effort and interface telemetry are chosen to make the control mechanics testable.
  3. The scenario starts with a case handler clearing eleven arrears letters before lunch. An assisted-summary tool combines account history, recent payments and an outstanding balance into a paragraph ready for a hardship-letter template.
  4. One further point on this data model: extraction confidence and generation confidence are not the same number and should never be collapsed into one.
  5. The worked acceptance case starts with 3 percent footer-link use across a six-week baseline and tests whether at least one relevant inline marker is inspected on materially more cases during an eight-week shadow period.
Figure 1Servicing system record to click to verifyCausal and control schematic
Servicing system record to click to verify10 declared states connected by 11 authored relations. The figure supports the section The provenance data model underneath the interface. L0L1L2L3L4 01
Servicing System Record
02
Retrieval Step
03
Provenance Tagging
04
Synthesis Step
05
Tagged Claim Store
06
Claim Renderer
07
Inline Citation Marker
08
Confidence Badge
09
Human Reviewer
10
Click To Verify
Reading. The authored topology makes 11 declared relations across 10 states inspectable. Read it as the control structure for “The provenance data model underneath the interface”, not as measured performance. Schematic derived from the paper's authored topology; no measured quantities.
On this page

Worked scenario: a wrong number in a customer letter

The following is an illustrative composite, not a disclosed client incident. Its balances, sample sizes, timings, remediation effort and interface telemetry are chosen to make the control mechanics testable. They must not be read as measured institutional results.

The scenario starts with a case handler clearing eleven arrears letters before lunch. An assisted-summary tool combines account history, recent payments and an outstanding balance into a paragraph ready for a hardship-letter template. It states a balance of 4,212 dollars. The handler copies it, sends the letter and moves on.

The actual outstanding balance was 2,940 dollars. The agent had aggregated two sub-accounts under the same customer ID, one of which had been closed and settled four months earlier, and summed both balances into a single figure. Nothing in the interface distinguished that summed number from the ordinary sourced fields around it. The customer's name, the account number, and the last payment date were all pulled directly from the servicing system of record and displayed with the same font weight, the same colour, and the same confidence as the aggregated balance, which was in fact a derived value built by the agent from two records it had silently decided belonged together.

In the composite, the error is not caught at generation or by the handler because nothing indicates that one number carries different risk from the rest. A quality sample eleven days later finds one of forty letters does not reconcile to the servicing ledger. That discrepancy triggers a wider trace across every case sharing the same sub-account structure.

The modelled trace finds 340 letters involving multiple sub-accounts; 87 contain a balance error above fifty dollars. The operating hypothesis sends all 340 to review and assumes 87 corrections, roughly 96 person-hours across quality, legal and operations, an operational-risk assessment and a jurisdiction-specific notification decision. These figures show how to construct an exposure model. They do not assert that a regulator reviewed or declined notification for a real institution.

The scenario deliberately leaves reputational cost unpriced. Complaint handling, customer correction, notification and governance effort depend on jurisdiction and harm. The important point is narrower: an apparently small rendering defect can create a larger control and remediation burden when unsupported arithmetic is styled like a sourced fact.

The agent's underlying logic for aggregating sub-accounts was arguably defensible in some contexts and wrong in this one. That is not unusual for generated content. What made it dangerous was that the interface gave the case handler no way to tell, at a glance, which parts of the paragraph she was reading were retrieved facts and which parts were the agent's own arithmetic. Everything in the letter looked equally certain. That is the failure mode this article is about, and it is a design failure before it is anything else.

Automation bias as a ux problem

Automation bias, the tendency to over-trust automated output, is usually treated as a training and culture problem. Teams run “trust but verify” workshops and add a source-checking step to the procedure. That instruction competes with caseload pressure every day. The interface is therefore part of the control: it shapes what reviewers notice and how much effort verification requires.

Consider what the composite interface teaches. It never asks the handler to distinguish a retrieved fact from a computed one because it does not distinguish them visually. The worked baseline assumes a footer-level “show sources” link is used on 3 percent of viewed summaries. That percentage is a test input, not production telemetry. It establishes the question a real pilot must answer: does moving provenance to the claim materially increase meaningful verification rather than superficial clicks?

This is the core claim of evidence-first design: automation bias is not simply a property of the human reading the output, it is a joint property of the human and the interface, and the interface is the variable you can actually change at scale. You cannot retrain every case handler's risk instincts through a policy document, particularly under a lunchtime caseload, but you can change what the screen shows them, and the screen change persists whether or not anyone remembers the training.

The mechanism is straightforward. Fluency, meaning grammatically smooth, confidently worded, well-structured text, is processed by a reader as a proxy for correctness, largely because in ordinary human writing the two are correlated: a person who writes clearly and confidently about a topic they know usually does know it. Generated text breaks that correlation. It is fluent regardless of whether the underlying claim is a solid retrieval from a system of record or an unsupported guess dressed in confident prose. When an interface presents all generated text at one visual register, it imports the fluency-as-correctness heuristic wholesale, and applies it evenly across sourced fact and pure inference. The reader has no cue telling them where the correlation between fluency and correctness has actually held and where it has not.

A single "show sources" link at the foot of a document does nothing to break this pattern, because it asks the reader to do additional work, opening a panel, scanning a list of documents, matching document titles back to specific sentences, entirely on their own initiative, with no visual signal that any particular sentence needs it more than any other. Under caseload pressure, additional undirected work is the first thing to be dropped. The fix is not to add more sourcing information behind a link. The fix is to change where the burden sits: instead of the reader having to go looking for provenance, provenance has to come to the reader, attached to the specific claim it supports, at the moment they are reading that claim.

What evidence-first ux actually means

Evidence-first UX is a design stance with one governing rule: no generated claim that could materially affect a decision is displayed without a visible, low-friction path back to whatever supports it, rendered at the granularity of the claim itself rather than the document it came from. That single sentence carries three separate design obligations, and it is worth taking them apart, because each one fails independently if you skip it.

The first obligation is claim-level granularity. A document-level citation, "this summary was generated using documents A, B, and C", satisfies almost no reviewer need, because it does not say which of the twelve sentences came from which document, or whether a given sentence is a paraphrase, an aggregation, or an unsupported inference dressed as fact. A reviewer checking a single disputed number needs to jump straight to the paragraph and page that supports that number, not be handed a reading list.

The second obligation is low friction. If checking a claim requires more than a glance and, at most, one click or hover, reviewers under caseload pressure will do it selectively. The scenario's 3 percent baseline is a deliberately weak operating hypothesis, not a judgement about diligence. Evidence has to sit close enough to the claim that checking it costs almost nothing.

The third obligation is honest differentiation. Not every claim in a generated document is the same kind of claim. Some are direct retrievals, a name, a date, a field value pulled straight from a system of record with no transformation. Some are aggregated or derived, a sum, an average, a "most recent of three dates", something computed from more than one source value by the agent itself. Some are pure model inference or judgement, a characterisation like "the customer has shown a pattern of late payment" or a recommendation like "this account should be flagged for enhanced review", where no single source record contains that sentence and the agent is exercising judgement over the evidence rather than reporting it.

An interface that renders all three the same way is, in effect, asserting a confidence level for the third category that it has not earned. Evidence-first design requires these three categories to look different from each other, consistently, everywhere they appear, so that a reviewer's eye is drawn to exactly the claims that need scrutiny and reassured, correctly, about the ones that do not.

None of this is purely visual. It works only if the pipeline produces the metadata the interface needs: every material claim must carry its source identifiers, transformations and timestamps forward from retrieval. A UI team cannot bolt honest provenance onto a pipeline that discarded those links. A cosmetic layer that guesses which document probably supported a finished sentence is worse than no citation because it manufactures evidence.

The provenance data model underneath the interface

Before any interface pattern can be trusted, the generation pipeline has to carry a data model rich enough to support it. This is the part of the work that is invisible to the end user and is also the part that most naive builds skip, because it adds engineering effort with no immediately visible payoff until the first incident makes the payoff obvious. Every claim that reaches the rendering layer needs, at minimum, the following attached to it.

A source document identifier, pointing to the system of record the claim was drawn from, whether a servicing platform record, a scanned correspondence item, a call transcript, or a policy document. A location within that document fine enough to jump to, a page number, a paragraph anchor, a field name, or a transcript timestamp, not merely the document's title. An extraction confidence score from the retrieval step, distinct from the model's own confidence in phrasing, because a badly scanned page can produce a low-confidence extraction even when the model's paraphrase of it reads fluent and assured. A generation timestamp, so a reviewer can tell whether the claim reflects data as of this morning or data cached from three days ago.

A provenance type flag, marking the claim as direct retrieval, aggregated, or model inference with no direct source. And, where aggregated, a list of the contributing source identifiers rather than a single one, so a reviewer can judge whether the combination was appropriate, precisely the check that would have caught the sub-account aggregation error before it reached a customer.

This is more metadata than many generation pipelines produce, and carrying it through retrieval and synthesis is real engineering work. The alternative is fluent output with no honest way to distinguish sourced, derived and inferred claims. The data model has to begin at retrieval, not after a summary has become plain text, because composition usually destroys the fine-grained links to individual source fields.

Notice the loop back at the bottom of that diagram. A provenance model that only flows forward, from source to rendered claim, gives you a display feature. A provenance model that closes the loop, letting the reviewer's verification click land back on the actual source record rather than a static snippet copied at generation time, gives you something closer to an audit trail, and it is the difference that matters when a regulator asks how a specific number in a customer letter was produced six months after the letter was sent. The static snippet answers "what did the model say it used". The closed loop answers "what was actually in the system of record", which is the question that actually gets asked in a post-incident review.

One further point on this data model: extraction confidence and generation confidence are not the same number and should never be collapsed into one. A page can be extracted with high confidence, clean text, unambiguous field boundaries, and the model can still misuse that clean extraction in a way that produces a wrong sentence. Conversely a poorly scanned page can be extracted with genuinely low confidence and the model can still, by chance, produce a correct sentence from it. A confidence badge that blends these two into a single number is throwing away exactly the information a reviewer needs to know which failure mode they are looking at.

Concrete ui patterns for evidence-first design

With the data model in place, the interface work becomes tractable. There are four patterns that, together, do most of the work, and each addresses a different part of the automation bias problem described earlier.

Inline citation markers at claim level

Instead of a single reference list at the end of a document, each sentence or clause that makes a checkable claim carries a small marker immediately after it, visually distinct from ordinary punctuation, something like a superscript numeral or a small coloured tick, placed at the point of the specific number or fact rather than the end of the paragraph. The marker is the visual anchor that makes claim-level granularity real rather than theoretical. Critically, the marker only appears on claims that actually have a traceable source. A sentence with no marker is not an oversight, it is itself informative, telling the reader that this particular sentence carries no direct source and should be read as the model's own construction.

The absence of a marker has to be a deliberate signal, not a gap that looks the same as a forgotten citation, which means the rendering layer needs a third state, marked, unmarked-because-unsupported, and unmarked-because-not-a-claim (a connective sentence with no factual content), and these three need to be visually distinguishable from each other too, or the "no marker" signal collapses back into ambiguity.

Hover or click to reveal source snippets

The marker itself does little if checking it still requires leaving the document. On hover, or one tap for touch interfaces, it should surface the supporting excerpt, record name, precise location and relevant timestamp. Moving from a footer link to a claim-level reveal is the main friction hypothesis in this design. Test it with task completion, source inspection and reviewer amendments rather than assuming a click proves verification.

Visual distinction between sourced, derived, and inferred

This pattern would expose the opening scenario directly. A field retrieved from a system of record uses the ordinary body style with a source marker. A sum, average or computed status receives a distinct treatment, such as a subtle tint plus an icon rather than colour alone, and its marker shows every contributing record and the operation performed: “summed from two open sub-accounts,” not merely the result.

Text that is model inference or judgement, a characterisation, a recommendation, a risk assessment with no single source record behind it, renders with a third, more clearly flagged treatment, commonly a border or icon explicitly labelled as model judgement, precisely so that a reviewer's eye is drawn to it as the category needing the most scrutiny rather than the least.

Figure 2Generated claim to reviewer attentionCausal and control schematic
Generated claim to reviewer attention6 declared states connected by 4 authored relations. The figure supports the section Visual distinction between sourced, derived, and inferred. L0L1 01
Generated Claim
02
Provenance Type
03
Sourced Style Plus Marker
04
Derived Style Plus Sources List
05
Judgement Flag Plus Badge
06
Reviewer Attention
Reading. The authored topology makes 4 declared relations across 6 states inspectable. Read it as the control structure for “Visual distinction between sourced, derived, and inferred”, not as measured performance. Schematic derived from the paper's authored topology; no measured quantities.

Confidence badges that are calibrated rather than decorative

A badge showing "92% confidence" next to a claim is worse than useless if that number does not correspond to anything measurable, because it manufactures false reassurance, which is arguably a worse outcome than no badge at all, since it actively suggests scrutiny has already been done. A calibrated badge has to be backed by an actual measurement process, typically built by sampling a population of claims at each displayed confidence band and checking, against ground truth, what fraction were actually correct.

If claims badged "high confidence" are wrong 1 time in 50 and claims badged "medium confidence" are wrong 1 time in 8, and a reviewer relying on the badge to decide where to spend their limited attention gets that ratio roughly right in practice, the badge is doing real work. If the bands are set once at launch from a vendor's default output and never checked against ground truth again, the badge is decoration, and it should be labelled honestly as such or removed, because a reviewer who has learned to trust an uncalibrated badge is worse off than one who trusts nothing and checks everything.

Pattern Trust calibration Reviewer friction Error catch rate
No citation, fluent text only Poor, uniform over-trust None, but no verification path exists Very low, caught only by downstream audit
Document-level citation list Weak, list is rarely consulted High, requires manual matching to claims Low
Citation on every sentence (spam) Weak, signal buried by volume Low to open, high to actually use Low to moderate, reviewers skim past
Claim-level inline marker with hover reveal Strong, differentiated by claim Very low High
Claim-level marker plus sourced/derived/inferred styling Strongest, matches display to actual certainty Low Highest
Confidence badge, uncalibrated Actively harmful, manufactures false trust Low but misleading Low, reviewers defer to the badge
Confidence badge, calibrated and monitored Strong, directs attention correctly Low High

Failure modes of naive attempts

Most teams that attempt evidence-first design fail on one of three specific technical points, and it is worth being precise about each, because the failures are easy to walk into with good intentions.

The first is citation spam. Marking every sentence, including connective prose, destroys the marker's value as a signal. In the worked credit-memo scenario, citation coverage reaches 100 percent while click-through falls below 2 percent within three weeks. The figures illustrate reviewer desensitization rather than report a client measurement. The fix is selective provenance for checkable facts, numbers, dates and statuses, supported by a generation pipeline that distinguishes factual claims from connective prose.

The second is false precision, a citation marker that points to a real document and a real page but does not actually support the specific claim next to it. This happens most often with aggregated claims, where the retrieval step correctly identifies two source records as relevant, and the synthesis step performs an operation on them, a sum, an average, a "most recent of", that the citation link does not represent.

A marker that points to "Document A, page 3" when the actual claim is "the sum of the value on Document A page 3 and the value on Document B page 1" is technically citing something, but it is not honestly representing what was done, and a reviewer who clicks through, sees a real number on a real page, and concludes the claim is verified, has been given false reassurance by a citation that looks rigorous and is not.

This is precisely what would make the opening scenario worse under a cosmetic citation layer: the handler clicks a marker, sees one real sub-account balance and concludes that the sum is verified after checking only one input. Provenance for a derived claim has to represent the derivation, not merely one input.

The third is latency. Provenance lookups that resolve to a live system of record add request time to a path that previously rendered static text. The worked performance budget assumes 340 milliseconds per uncached hover and six or seven markers inspected in quick succession. Those figures are stress-test inputs, not a bank result. The acceptance test should measure p50 and p95 hover latency on the intended network and device estate.

One design response pre-fetches approved snippets at render time so the hover hits a local cache. That trade accepts some unused reads in exchange for predictable interaction latency. It works only if the cached snippet is versioned and invalidated when the underlying record changes before the reviewer acts. The cache and freshness control belong in the feature estimate, not in post-launch remediation.

A worked example: the complaint response draft

Consider a worked complaint-response draft for a customer who disputed three transactions and asked why the account moved to a higher fee tier. The following table compares a fluent-text baseline with the proposed evidence-first rendering; all dates, balances and calibration figures are illustrative.

Claim In Draft Old fluent-text rendering Evidence-first rendering
"Your account was moved to Tier 2 on 14 March 2026" Plain sentence, no visual distinction from surrounding text Sourced marker linking to the tier-change event record, timestamp 14 March 2026 09:12, shown on hover with the exact system field
"Your average monthly balance over the prior quarter was 1,840 dollars" Plain sentence, identical weight to the date above Derived-claim tint plus marker showing three contributing monthly balance records and the averaging operation applied, so a reviewer can check the three inputs individually
"The three disputed transactions do not appear to be linked to a common merchant category" Plain sentence, reads as a simple fact Model-judgement flag, distinct border and icon, no source marker, with a short note that this is the model's characterisation of the transaction data shown above it, inviting the reviewer to check the transaction list themselves before relying on the characterisation
"Based on your payment history, we recommend enrolling in the fee waiver programme" Plain sentence, same visual register as the transaction date Model-judgement flag with a calibrated confidence badge reading moderate, based on a sampled accuracy check showing recommendations at this badge level match the outcome a human case handler would independently reach roughly 74 percent of the time, an explicit invitation to review rather than paste

The pattern across all four rows is the same. Under the old rendering, a case handler skimming this draft before sending it has no reason to spend more attention on the average balance calculation or the recommendation than on the tier-change date, because nothing on the screen tells them these are different kinds of claim carrying different risk. Under the evidence-first rendering, the two claims that most need a human check, the derived average and the model's recommendation, are the two claims visually distinguished as needing it, while the two directly sourced claims are visually reassured as such, which lets the reviewer's limited attention go where it is actually needed rather than spread evenly, or worse, not spent at all.

Measuring what changed

None of this is worth building unless it changes reviewer behaviour and downstream error rates. Treat the earlier numbers as a measurement design, not a before-and-after claim. A pilot should establish a baseline for footer-source use, meaningful source inspection, reconciliation defects, p50 and p95 rendering latency, review time and the rate at which reviewers amend a generated claim.

The worked acceptance case starts with 3 percent footer-link use across a six-week baseline and tests whether at least one relevant inline marker is inspected on materially more cases during an eight-week shadow period. Do not set 61 percent as a universal target. Segment the result by claim type and consequence, and distinguish a hover from evidence that changed or confirmed a decision.

For quality, stratify sampled letters into directly sourced, derived and model-judgement claims. Report discrepancies with confidence intervals and absolute counts. A change from one in forty to one in 260 would be material in the illustrative worksheet, but a real release decision must account for population mix, sample size and severity. A correct date and an incorrect hardship recommendation do not carry the same loss.

For performance, measure document-load and marker-open latency before and after caching under representative network conditions. The scenario uses 340 milliseconds per uncached hover and 40 milliseconds of added document-load time as test thresholds, not observations. For operations, expect review time to rise if reviewers are genuinely examining evidence. Record the increase alongside avoided defects and escalation quality instead of presenting evidence-first design as a pure productivity gain.

Treat every claim as an inspectable object

A paragraph is a poor unit for provenance. It may mix a quoted fact, a calculation, an inference and an unsupported transition. The interface should retain those differences even when the prose reads smoothly. One fluent sentence can contain several evidence obligations.

Figure 3Source records to audit and feedbackCausal and control schematic
Source records to audit and feedback8 declared states connected by 7 authored relations. The figure supports the section Treat every claim as an inspectable object. L0L1L2L3L4 01
Source records
02
Typed evidence objects
03
Claim construction
04
Claim-source edges
05
Rendered answer
06
Reviewer inspection
07
Challenge or approval
08
Audit and feedback
Reading. The authored topology makes 7 declared relations across 8 states inspectable. Read it as the control structure for “Treat every claim as an inspectable object”, not as measured performance. Schematic derived from the paper's authored topology; no measured quantities.
Claim type Evidence shown Interaction Failure state
Direct fact Exact source passage and record identity Open in context Source missing or superseded
Derived value Inputs, formula and rounding rule Recalculate Input or transformation unavailable
Policy interpretation Controlling provision and scope Compare rule to case Applicability unresolved
Model judgement Relevant evidence and uncertainty Challenge or seek second review No independent support
Omission claim Search scope and exclusions Inspect coverage Corpus or query incomplete
Do not attach one citation to a sentence when it supports only one clause. Split the claim or show clause-level evidence. Citation density is not the goal; support accuracy is.

The interface must expose evidence quality

All sources are not equal. An approved current policy should not look like an old draft. A system record should not look like a model inference. Visual treatment should encode authority, freshness and transformation.

Figure 4Claim selected to label model judgementCausal and control schematic
Claim selected to label model judgement10 declared states connected by 7 authored relations. The figure supports the section The interface must expose evidence quality. L0L1 01
Claim selected
02
Directly sourced?
03
Source current and authoritative?
04
Yes
05
Show verified source state
06
Show warning and alternative
07
No
08
Derived from inspectable inputs?
09
Show calculation lineage
10
Label model judgement
Reading. The authored topology makes 7 declared relations across 10 states inspectable. Read it as the control structure for “The interface must expose evidence quality”, not as measured performance. Schematic derived from the paper's authored topology; no measured quantities.

High authority · high freshness

Present as the primary evidence, with direct context and scope.

High authority · low freshness

Show an expiry warning. Require confirmation before material action.

Low authority · high freshness

Use as a lead, not a controlling source. Seek corroboration.

Low authority · low freshness

Exclude from the default evidence panel or mark it clearly as historical context.

Verification must survive handoff and audit

Evidence links should resolve after the case closes. They also need access control. A reviewer may be allowed to see a derived claim but not every underlying record.

Figure 5Reviewer opens claim to durable audit packageCausal and control schematic
Reviewer opens claim to durable audit package7 declared states connected by 6 authored relations. The figure supports the section Verification must survive handoff and audit. L0L1L2L3L4 01
Reviewer opens claim
02
Identity and purpose check
03
Resolve versioned source
04
Render authorised context
05
Record inspection event
06
Decision and reason
07
Durable audit package
Reading. The authored topology makes 6 declared relations across 7 states inspectable. Read it as the control structure for “Verification must survive handoff and audit”, not as measured performance. Schematic derived from the paper's authored topology; no measured quantities.

A broken citation is an operational defect. A citation to an inaccessible source is not reviewer evidence. A current claim tied to an old snapshot needs an explicit warning. Derived values need inputs and method. Unmarked model judgement must never inherit the visual authority of a sourced fact. Absence of evidence is itself a state the interface should name.

The W3C PROV-O recommendation supplies a formal provenance vocabulary. The ICO guidance on AI explanations separates explanation needs by audience and context. WCAG 2.2 is relevant because evidence must remain perceivable and operable. The NIST AI RMF places transparency inside a wider risk process. The EU AI Act includes record-keeping, transparency and oversight obligations for covered systems.

The practical standard is demanding but clear. A material claim should reveal what supports it, how it was produced, whether the source still governs and who accepted the residual uncertainty.

Notes for practitioners

Build the provenance tagging into the retrieval and synthesis steps first, before any interface work starts. If the agent cannot tell you, at generation time, which source records fed a specific sentence, no amount of front-end polish will produce honest citations later. Retrofitting provenance onto an already-composed paragraph is close to impossible to do accurately, since the fine-grained links are usually lost in the act of writing the sentence.

Treat “aggregated or derived” as a first-class category from day one, not a special case of “sourced.” Aggregation is a distinct failure surface, and naive citation layers often collapse it into one source pointer, recreating the false-precision problem described earlier.

Measure citation opens and meaningful evidence inspection from the first week, but do not turn either into a vanity metric. Compare the rate with the workflow's own baseline and risk-tier target. A low rate may indicate friction, irrelevant markers or a low-variance case mix; investigate the cause before declaring the interface successful or failed.

Never ship a confidence badge before you have a sampling process to check it against outcomes, and re-run that sampling monthly. A badge that has drifted out of calibration is actively worse than no badge, because reviewers who have learned to trust it will keep trusting it after it has stopped meaning anything.

Budget for the caching layer at the same time you budget for the citation feature itself. Provenance that resolves live against a system of record is the version worth having, but it is also the version that will add visible latency if you do not pre-fetch and cache it, and a stuttering interface will get blamed on the whole redesign, not on the specific caching gap that caused it.

Expect review time per document to go up, and say so plainly to whoever is sponsoring the work. The value of evidence-first design is not that it makes reviewers faster, it is that it makes their trust proportionate to what actually deserves it, and that trade is worth defending on its own terms rather than dressed up as a productivity gain it does not deliver.

Finally, keep the unmarked state deliberate. A sentence with no citation marker should read, to a trained reviewer, as "the model's own construction, treat with appropriate scepticism", never as "the citation was forgotten here". If your interface cannot make that distinction reliably, visually, every time, the absence of a marker carries no information, and you have quietly rebuilt the original fluent-text problem underneath a layer of markers that only sometimes mean something.