The service begins where the demonstration ends

This illustrative scenario describes neither a CYTIZEN engagement nor a client outcome. Claire runs operations for an industrial group. A team demonstrates an assistant that finds instructions in a large document collection. Its answer is convincing and its citation looks precise. Marc, who runs IT services, asks what happens when a procedure changes, an employee leaves, or the answer contradicts an approved document. The prototype cannot yet answer those questions.

Scaling does not mean giving more users access to the prototype. It means turning an application into a service with a bounded purpose, authorised data, an accountable owner, verifiable quality and a withdrawal mechanism. Management can fund imperfect learning, but it needs to understand its exposure before funding an operational dependency. Maturity is the ability to explain a bad answer and prevent its recurrence, rather than the attractiveness of a good demonstration.

An inventory that supports decisions

Claire rejects an inventory consisting only of tool names. It contains one row for each business use: finding procedures, drafting contractual clauses or extracting invoice data. For each row, the process owner records users, decisions influenced, data consumed, interfaces, people affected and the fallback. A single platform can support uses with very different levels of criticality. A platform-wide classification would obscure that difference.

The document assistant's record is completed as follows: owner, quality management; users, authorised site personnel; output, an answer accompanied by source passages; permitted decision, locating information; prohibited decision, approving a quality deviation; data, approved procedures from the designated repository; fallback, the existing document search. It also identifies who can change the model, publish a corpus and suspend the service. A record that assigns every responsibility to the project team is not ready for operations.

Regulatory classification follows this inventory rather than replacing it. The European Commission publishes a differentiated AI Act timetable, updated following the AI Omnibus that entered into force in July 2026. The organisation's role, use and applicable regime require assessment by the competent functions. Permission to run a pilot demonstrates neither legal compliance nor readiness to make decisions autonomously. This article proposes a working organisation, not individual legal classification.

Evaluate errors that have consequences

Marc builds an authorised, de-identified evaluation set from questions actually encountered in the business. Subject-matter experts prepare expected answers before reviewing the system's responses. The sample includes obsolete documents, contradictory procedures, unanswerable questions, version changes and requests about another site. Difficult cases remain in the set rather than being removed to improve the average. Development data helps tune the service; a separate holdout set supports the release decision.

Three measures remain separate. The grounded-answer rate is the number of responses whose useful claims are supported by correct sources, divided by evaluated responses. The appropriate-abstention rate is correctly declined cases lacking evidence divided by evaluated cases requiring abstention. The critical-error rate is evaluated outputs capable of triggering dangerous or prohibited action divided by all evaluated outputs, using a business-approved definition. Rates are undefined when their denominator is zero; the corresponding counts remain visible. An overall satisfaction score must not compensate for a critical error with many easy answers.

Before testing, management sets illustrative decision rules: unauthorised document access triggers suspension; unsupported operational instructions block deployment; other errors are assessed by frequency and impact. These are scenario choices, not universal standards. The release record states sample size, populations not covered and uncertainty. No errors in a small sample do not provide a safety guarantee.

Access rights are not negotiated with the model

Authorisation must apply to documents before they reach the model. Giving the assistant a confidential document and then instructing it not to disclose it places the control at the wrong point. Marc checks propagation of rights through the repository, index, retrieval and user session. Tests include account revocation, group changes and document deletion, including caches, secondary indexes and logs.

Agents able to act introduce another risk layer. A technical identity with administrator rights is not acceptable merely because its usual tasks are modest. Each permitted tool needs a scope, time-bounded rights and logging, with explicit human approval for sensitive actions. Here, the assistant can prepare a change request but cannot approve a procedure or alter access rights. A malicious instruction inside a document must not expand its mandate.

Operate the service and count its full cost

A useful response costs more than a model call. Claire adds inference, indexing, storage, monitoring, evaluation, support and human review, then divides the total by requests completed to an acceptable quality level. Rejected attempts and corrections remain in the numerator; excluding them would hide the cost of errors. Fixed and variable charges are separated to model both low activity and demand peaks.

The benefit is not the theoretical time saved per question. A faster answer may require longer checking. The chosen measure compares the complete workflow, including retrieval and verification, on comparable tasks. The business then states what it will do with recovered capacity: reduce a backlog, strengthen controls or avoid external spending. Without that decision, saved minutes do not automatically become budget savings.

Support receives an escalation path for access incidents, unsupported answers, outages and cost drift. An operator can disable a source or revert to the previous version without waiting for the prototype developer. Model or corpus changes trigger proportionate evaluation against the reference set. History must connect a disputed answer to its configuration without unnecessarily retaining sensitive user data.

In practice: a useful disagreement

Claire wants to open the service to several countries. Marc offers two options: include all documents immediately, or open one workflow with controlled rights and versions. Quality management points out that local procedures sometimes have identical titles but different instructions. Management chooses the narrower option and requires site and version to appear with every answer. It reduces the pilot's promise but makes the actual delivery clearer.

The decision records the authorised population, corpus, outstanding tests and person entitled to suspend use. Expansion depends on resolving rights discrepancies and appointing a local owner. The committee does not ask for a green status; it asks which dependency still prevents the organisation from taking responsibility for the service. That question distinguishes demonstration governance from operational governance.

What genuinely changes by sector

In regulated manufacturing, the assistant must distinguish approved documents from drafts and preserve quality management's authority. In a legal team, a suggested clause remains a proposal; matter-level rights and approval traceability matter more than fluent drafting. In IT support, the decisive issue is action scope: reading a knowledge base, opening a ticket and restarting an application are different risks. Each variation changes the tests and responsibilities, not merely the article's terminology.

A small business without its own operating capacity may buy a service with explicit limitations. A larger organisation may provide a common platform while business teams retain responsibility for each use. The central decision remains: who will own quality, consequences and withdrawal when the pilot team leaves?

Sources, method and limitations

Primary sources checked on 4 October 2026. NIST provides a voluntary risk-management framework, not certification of this service. The proposed organisation, illustrative thresholds and scenario are educational analysis requiring adaptation to the context, contracts and applicable law.