The pilot starts with the decision it must inform
This guide revisits 2024 AI pilots from a 2026 perspective. The scenario and values are illustrative and do not represent CYTIZEN client results. A pilot is neither a sales demonstration nor an irreversible first rollout phase. It should resolve a specific uncertainty: does a use improve a journey enough to justify its cost and controls? The possibility of concluding no is a condition of a good pilot.
The launch record states a few decisions: task, users, admitted data, comparison, success criteria, spending cap and stopping owner. A team unable to explain how it will evaluate results is not ready to test, even if it already has model access. An impressive response is not evidence of a usable service.
Illustrative scenario: finding the right quality procedure
Inès, the quality lead, wants teams to find approved procedures more easily. Hugo, the applications project manager, proposes an assistant answering questions from an authorised corpus. Its demonstration produces a clear response within seconds. Inès asks what happens when two procedures look similar but one belongs to another site.
The team restricts the trial to document retrieval. The assistant creates no approved instruction, changes no source documents and authorises no operation. For each answer, the user must open the document, check its version and identify the site. The journey's outcome is not generated text but a relevant source found and verified.
The pilot uses twelve participants and two hundred prepared questions. Those counts are teaching choices. Eighty questions cover frequent requests, sixty cover ambiguity, forty cover missing or obsolete documents and twenty test attempts to use an unauthorised source. A confident answer to a question without an admissible source is a failure; explicit abstention can be the correct result.
A completed mandate before the first test
| Decision | Example mandate |
|---|---|
| Scope | Search approved procedures for site A in French |
| Exclusions | Regulatory advice, new instructions, quality decisions and unapproved documents |
| Data | Identified corpus and versions frozen for evaluation; no sensitive individual case files |
| Comparison | Current search and existing document engine on equivalent categories |
| Criteria | Correct source, correct version, appropriate abstention and time to verification |
| Decision | Inès accepts business results; Hugo confirms supportability; sponsor approves expansion |
| Limit | Illustrative €15,000 budget; no additional access or site without review |
Testers need to know the mandate. They cannot quietly widen the corpus to improve an answer or use excluded real data because the trial is internal. Record every change, which may require another comparison. Without this discipline, evaluation becomes inseparable from continuous alteration of the scope.
The business owner describes acceptable outcomes while the technical team explains limitations. A commitment to no errors is unrealistic. Controls and restrictions can, however, make particular actions impossible and some outcomes verifiable. The launch meeting should distinguish these enforceable properties from goals that will only be assessed during the trial.
Prepare an evaluation set and reference answers
Document specialists establish expected sources, allowed alternatives and no-answer cases for each question. A second reviewer examines ambiguity. The set should represent journey difficulties rather than only questions used in sponsor demonstrations. Keep a subset out of configuration work to reduce optimism from repeatedly seeing the same examples.
A completed row could specify Q-084: revision conditions for a site A procedure; admissible sources P-17 version 4 and appendix A; prohibited source P-17 for site B; an acceptable result cites version 4 and opens the document; an unacceptable result combines the two sites. The tester can judge this without assuming persuasive text is correct.
The reference answer may itself be disputed. If two experts disagree, resolve the document issue before counting the tool's response as right or wrong. A pilot can reveal a contradictory corpus. That is useful but distinct from improved AI performance, and its remediation belongs with the document owner.
Measure the whole journey
The correct-source rate divides answers identifying a relevant approved source by questions for which such a source exists. Appropriate abstention applies to questions lacking an authorised source. Combining them into one rate can hide a tendency to answer at any cost. Also measure wrong versions and permission violations, whose consequences differ from awkward wording.
The clock starts when the user reads the question and stops after source verification. Compare tasks of similar complexity between current search and the assistant. Generation time is only one component. An answer generated in five seconds but requiring five minutes of checking may perform worse than a less impressive document engine.
An illustrative result has 133 correct answers among 140 questions with a source, or 95%. Across forty missing or obsolete-document cases, thirty-six abstentions are appropriate. The twenty permission cases produce no unauthorised access. Those figures describe one sample rather than universal accuracy or absolute security.
Keep task-level results available. A good median can hide a particularly difficult group, such as questions involving site-specific annexes. If those cases affect an important user population, the expansion decision needs their results separately instead of relying on the entire sample's average.
Decide using errors rather than the average alone
The committee examines errors by consequence. An overly long answer needs a different correction from another site's procedure presented as applicable. A permission violation blocks expansion even when average time savings are high. The document owner determines which errors require changes to the corpus, retrieval or interface.
The mandate may require a median net saving of at least one minute, correct sources on at least 95% of source-bearing questions, appropriate abstention on at least 90% of no-source cases and no observed prohibited disclosure. These teaching thresholds explain a decision; actual values must reflect risk and sample size. No observed prohibited event remains accompanied by technical permissions and monitoring.
If quality is acceptable but net time does not improve, retain a narrower use or stop. If time is good but document versions are uncertain, correct that issue before expansion. One aggregate score cannot distinguish those situations. A decision should state the acceptable boundary, unresolved uncertainty and proof required to move beyond it.
A short sequence without hidden commitment
The first week establishes the task, data and reference answers. The next prepares the environment and checks access. Two trial weeks observe users and collect results. A final review chooses stopping, a bounded correction or expansion. This five-week sequence fits the example document pilot; it is not a mandatory duration for every AI project.
A change register identifies adjustments between trials. If configuration changes halfway through a set, do not merge results as if only one version had been assessed. Retest a stable subset and identify the new version in the decision record. Retain enough information to explain which configuration produced each material error.
A budget cap is insufficient if business hours are uncounted. Inès reserves specialist time to prepare and review cases. Hugo plans configuration, incidents and trial shutdown. The provider explains consumption limits and result export. Include the cost of closing the pilot if no continuation is authorised, including removal of trial access and controlled treatment of retained data.
What must transfer if the trial becomes a service
A service needs corpus ownership, permission administration, monitoring, support and a withdrawal procedure. Each new approved procedure version must replace the previous version in used sources. Reliable answers against a frozen corpus do not demonstrate that this maintenance exists.
Support needs to identify the question, system version and retrieved sources without unnecessarily duplicating sensitive information. The business must report an incorrect answer and track its treatment. A model or retrieval update triggers proportionate testing using relevant reference cases. A provider change announcement is therefore an operating input rather than merely a procurement email.
The example's exit decision may permit limited expansion at site A while preserving mandatory source verification and excluding generated instructions. Outstanding gaps remain named and dated. A successful demonstration does not become tacit authority to expand across all sites and languages. Extension should also verify that local document ownership and support capacity are available.
Adapt the pilot to the decision involved
In legal work, judge contract summaries on material omissions, sources and checking cost rather than fluency alone. In procurement, clause retrieval should not turn into an automatic supplier commitment. In industrial quality, identifying the approved version and site often matters more than producing a long response.
A trial processing personal data adds appropriate analysis and information with competent owners. Internal testing is not a general exemption. Guidance issued after 2024 must be distinguished from information available during the historical year. The retrospective perspective can recommend a stronger current practice without pretending it was a settled requirement in every earlier pilot.
For finance, define whether the assistant explains an existing control or proposes an accounting treatment. Those tasks need different reviewers and consequence assessments. A pilot should not broaden from document retrieval to professional judgement merely because the same interface can answer both questions.
Sources, method and limitations
Primary sources consulted on 4 October 2026: NIST AI RMF Playbook on measurement; CNIL on data protection in system design; CNIL on AI-development security. These inform evaluation and design. Consultation in 2026 does not imply that every recommendation was available in 2024.
The method starts with a verifiable task, compares an alternative and retains errors and limitations. Cases, thresholds and amounts are teaching examples. Sample size, data conditions and validation requirements need adaptation to actual use. The guide provides a mandate and evaluation examples without certifying a system or guaranteeing that future errors cannot occur.