Three attractive ideas, one testable hypothesis

In this fictional scene set in 2024, a committee receives three generative AI proposals: draft support replies, search internal procedures and prepare commercial summaries. The demonstrations produce convincing text. The committee must choose where to run a bounded pilot. Every person, volume, cost, timing and result is invented for teaching purposes. None describes a CYTIZEN client engagement. The question is which task offers verifiable value with workable controls and acceptable total cost.

A use case is a defined task within a flow, with a user, input, output and decision about that output. “An assistant for everyone” is too broad to evaluate value or bound risk. This retrospective discussion uses primary sources checked in October 2026. NIST's AI RMF, published in 2023, and its generative AI profile, published on 26 July 2024, are dated references. Current developments described on their pages are not projected backwards onto every decision made in 2024.

Define the task before selecting the model

For the fictional procedure search, the user is a support agent, the input is a question about an approved procedure and the output is a suggested answer with verifiable passages. The decision remains human. The pilot cannot modify business records or automatically send customer replies. This allows a concrete benefit to be tested: finding usable information faster without increasing answer errors. Fluent text is not evidence that the underlying work has been completed correctly.

Describe the current method and at least one simpler alternative. Better conventional search, a form or clearer procedures may offer greater value. The presence of text does not establish a need for AI. Identify ambiguity, input diversity and how an output can be checked. A task whose answers are difficult to verify may cost more to operate than its demonstration suggests. Make that cost visible during selection, before integration creates pressure to continue.

Build a cost, control and value matrix

Gross value may include saved time, better quality or additional capacity. Net value subtracts data preparation, integration, licences, model calls, human review, support and maintenance. Theoretical minutes do not automatically reduce expenditure. The owner explains how released capacity will be used. Measure tasks completed correctly rather than answers generated. A quick response requiring lengthy verification may reduce value or shift work to more expensive specialists. Include that effect instead of counting only generation time.

Control cost depends on verifiability. A response about a short approved procedure can be checked against a passage. A summary of conflicting sources needs more judgement. An irreversible action needs explicit authority and validation. A human in the flow is not automatically an effective control: reviewers need time, accessible evidence and authority to reject. Rushed or habitual approval can allow errors through while creating an appearance of accountability. Test review behaviour as well as model output.

Compare plausible total cost, workable control and operational value. A combined score may help ranking, but must not compensate for a disqualifying condition. An attractive time-saving claim cannot make a use case acceptable when data access is unauthorised, ownership is absent or important outputs cannot be checked. Resolve these conditions before ranking. A less impressive demonstration may be a better first pilot because its value hypothesis and controls are more credible.

A completed selection matrix

This table is entirely fictional. Costs are illustrative pilot estimates, not market prices. Decisions depend on each organisation's data, error consequences and resources.

Use case and valueIllustrative total costConcrete controlDecision and stopping condition
Procedure search; faster verified answers12,000 euros for preparation and testing, including reviewApproved corpus, cited passages, agent validationBounded pilot; stop if critical answers cannot be verified
Automatic support replies; less drafting18,000 euros; review cost unresolvedExternal sending and unauthorised commitment riskDefer automatic sending; test human-approved drafts only
Commercial summary; meeting preparation8,000 euros; source reconciliation requiredVerify amounts and dates against source documentsTest after access clarification; stop if review consumes the gain

The supporting register names a value owner, control owner and technical lead. It records scope, observation period, exclusions and replacement option. Keep cost assumptions open to revision. An attractive low-volume task may become expensive when long inputs multiply calls, sources change or experts must handle exceptions. Request sensitivity to volume and review time rather than a single supposedly certain figure. Include costs of maintaining the corpus and answering disputed outputs after the pilot ends.

Specify evidence and stopping criteria first

The test set represents actual tasks and includes difficult cases: absent information, contradictory documents, out-of-scope questions and malicious instructions embedded in a source. Competent reviewers validate reference answers. Assess correctness, evidence quality, usefulness, appropriate refusal and total task time separately. A favourable average can hide a material error, so examine critical categories. Do not improve the apparent result by removing tasks where retrieval or verification fails. Record what the trial can and cannot establish.

In the scenario, procedure search is first tested on controlled copies without actions in a business system. Users compare their current method with assistance on comparable task categories. They record time to a verified answer, including source reading and correction. A fast response with untraceable evidence counts as unresolved. Test situations where the user should seek expertise rather than accept plausible prose. Successful abstention may be more valuable than a complete-looking answer that cannot be justified.

Stopping criteria combine safety and value. Exposure outside authorised data scope triggers suspension and investigation. Inability to verify critical outputs prevents expansion. Review workload exceeding expected benefit leads to scope revision or termination. Define these local rules before seeing results and record them. This prevents a pilot becoming permanent because the team likes the demonstration or has already invested in integration. Reviewers should know who can stop the pilot and how ordinary work resumes.

The exit decision is continuation within tested scope, modification and retest, or termination with an alternative. A successful pilot does not establish suitability for all documents, languages, volumes or users. Record what was tested and what remains unknown. Expansion requires fresh assessment of data and controls. A new model or supplier update does not remove the need to verify the use case in its operating context. Keep the pilot's conclusions bounded by its actual evidence.

Prepare operation during selection

The service needs to identify which model, corpus and rule versions produced an output. Choose useful records according to evidence needs and data constraints rather than keeping everything indefinitely. A corpus owner updates approved procedures and withdraws obsolete versions. Support staff need a route for disputed answers. A pilot working on stable documents can deteriorate when operating rules change unless these responsibilities are assigned and funded. Include this maintenance in the value calculation.

Access controls cover both source documents and responses. An assistant should not disclose information a user cannot consult directly. Test different permissions and access changes. An output remains a proposal unless the flow explicitly grants other authority. If the system later modifies a record or sends a message, evaluate that capability separately, including limits, validation evidence and suspension. Moving from advice to action materially changes the required controls, even when the underlying model stays the same.

Select according to error consequences

Manufacturing document assistance requires approved sources and competent review for operating instructions. Financial-service summaries need traceable figures and must not become implicit authorisations. Public-service answers should remain explainable, with access to correction or review. Professional-service drafts may reduce preparation work if checking does not consume the gain and data stays within authorised scope. These examples are design considerations, not observed engagements or universal rules. Their purpose is to connect concrete value with control capability and cost.

The strongest first pilot can answer a useful question and support a credible stopping decision. A use case deserves a trial when the organisation knows what it wants to learn, which evidence it will collect and what it will do if the expected value fails to appear. That decision discipline matters more than the persuasiveness of the initial demonstration.

Primary sources and chronology

References checked in October 2026. AI RMF 1.0 was released on 26 January 2023 and the generative AI profile on 26 July 2024. They provide voluntary risk-management references. The matrix, costs and stopping rules above are independent teaching examples, not regulatory thresholds.