Two hours to make a crisis governable
At 08:12 in this illustrative 2020 scenario, Claire receives three messages: warehouse teams cannot print shipping documents, Finance cannot see the morning entries, and a supplier mentions a network intervention. Julien, the operations lead, suggests a restart. Nadia, the logistics lead, wants to use paper. The director asks for a recovery time. Nobody is lying; each person sees a different part of the problem. Establishing a crisis cell starts by combining those fragments and giving someone authority to decide under uncertainty.
The aim of the first 120 minutes is not to fix every outage within two hours. It is to establish credible scope, command, an authorised minimum service, controlled changes and regular communication. The timings below are exercise guidelines, to be adapted to human safety, production constraints and existing mandates. An industrial emergency may require stopping an operation before the crisis meeting even begins.
0–15 minutes: declare, protect and assemble
Claire opens one coordination channel. Julien leads technical interventions, Nadia represents business operations, and Marc keeps the incident log. The sponsor explicitly gives Claire authority to coordinate decisions until relieved. Each participant supplies a callback number and deputy. If the normal messaging platform depends on the failed service, the conference number and contact list must be available elsewhere. Essential decision makers attend; observers receive updates rather than permanent invitations.
Marc records the initial fact: printing has failed since 08:05 on two terminals at site A; site B has not been checked. This does not become an organisation-wide outage without evidence. Julien freezes nonessential changes and preserves relevant logs. A suspected cyber incident requires security input; restarting or reconnecting equipment is not automatically appropriate. Nadia suspends shipments whose traceability cannot be guaranteed and identifies already loaded stock. The cell must protect the loading dock as well as the server.
15–30 minutes: separate facts, hypotheses and decisions
Start with three columns. Fact: two printers return the same error. Hypothesis: a shared network dependency. Decision: test from a third terminal without changing configuration. Every hypothesis needs a test that could disprove it. Julien states the duration and risk of each test. Claire prevents simultaneous interventions that would make their effects indistinguishable. The supplier identifies the precise change, its time and the person able to reverse it. Routine maintenance is not an adequate technical explanation.
Nadia establishes the business window: the carrier leaves at 10:00, some batches can wait, and another requires specific storage conditions. Acceptable interruption is defined by business flow, not just by application. Claire frames the first alternative: wait for recovery or ship a limited subset through an authorised paper procedure. The business owner retains authority to accept degraded traceability; IT operations cannot decide this alone.
30–60 minutes: choose a controlled fallback
Paper is not a magic solution. Nadia checks the last reliable numbering sequence, mandatory fields, signatures and subsequent reconciliation with the ERP. She limits the fallback to ten shipments with known status. Marc reserves identifiers and records who holds the register. Corrections retain their original information and history. A shipment with an uncertain batch or order status remains suspended, even when its truck is waiting.
Julien prepares a configuration rollback, describing what it will restore, what it will not and how duplicate transactions will be detected. Claire authorises it only after the technical owner confirms readiness and a business participant is available for the recovery test. The sponsor accepts the cost of delaying the suspended flow. The decision records the rejected option, accepted risk and trigger for reconsideration; it is more precise than restart the service.
60–120 minutes: validate service and hand over
At 09:25, printing works on the test terminal. Claire declines to announce resolution until one order completes preparation, printing, shipment and reconciliation. Nadia runs this journey with Julien. They also identify transactions submitted during the outage: which completed, failed or remain queued? A reachable application may still contain a transaction backlog capable of producing a second incident later that day.
By the end of two hours, the group has either stabilised the situation or agreed an explicit continuation plan. The incoming lead receives confirmed facts, open hypotheses, prohibited actions, temporary permissions and the next communication deadline. The cell does not disappear when the director leaves the call. Named owners follow reconciliation and technical investigation. Everybody knows who can declare normal operation and who can recall the cell if the discrepancy returns.
In practice: a completed incident log
| Time | Fact or decision | Owner and evidence |
|---|---|---|
| 08:15 | Impact confirmed on two terminals at site A; site B still unverified. Nonessential changes frozen. | Marc retains exact errors and timestamps. Julien checks site B before 08:25. |
| 08:32 | Paper fallback limited to orders with established status and identifiers. No uncertain batch leaves the dock. | Nadia signs the authorised list and maintains the document register; Finance receives the affected identifiers. |
| 08:48 | Configuration rollback authorised. No concurrent intervention. Next update at 09:10. | Julien records configuration before and after. The supplier supplies its change log. |
| 09:30 | End-to-end transaction succeeds; paper shipment reconciliation remains open. | Nadia confirms shipment and Finance checks postings. Marc records restored service under observation. |
| 10:00 | Handover accepted; reconciliation and two pending orders require follow-up. | Claire gives Thomas the mandate, unknowns and recall thresholds; he confirms availability. |
A message employees can act on
An actionable update states: since 08:05, site A cannot print some shipping documents. Logistics is checking loaded orders; do not recreate an order to bypass the error. Authorised shipments use the temporary register maintained by Nadia. The next update will be issued at 09:10 even if the cause remains unknown. This message explains the impact, prevents a dangerous workaround and sets the next checkpoint without inventing a resolution time.
Support receives different instructions: capture terminal identifier, time, operation and exact error, linking tickets to the incident while retaining local differences. Management receives impact, alternatives and accepted risk. A supplier receives relevant tests and authorisations. Sending the same lengthy message to everyone creates noise and can expose unnecessary technical information. Marc retains issued versions so the cell does not discover conflicting promises afterwards.
Sector differences: deciding what can continue
In manufacturing or pharma, IT continuity remains subordinate to safety, batch traceability and authorised procedures. Manual processing does not replace a required validation. Quality participates in affected business flows, not every network test. The cell distinguishes operations that can continue, records that can be captured and products that can be released.
In banking or insurance, retrying a payment or case requires establishing whether it already completed. A simple retry can create a duplicate. The cell assigns capacity to reconciling status and external references. In services and public administration, a fallback may provide limited telephone assistance. Teams must state which procedures remain possible, protect collected information and reconcile subsequent entry for completeness.
Prepare the next exercise without writing another manual
The next exercise removes the normal crisis channel, makes the primary decision maker unavailable and delays the supplier response. Measure time to appoint command, establish scope, validate the fallback register and complete a handover. Every defect needs an owner and a verifiable correction. Participants are not judged on guessing the hidden technical fault. The arrangement is judged on producing a safe decision with the information actually available.
Sources and method
Primary sources checked on 4 October 2026. The year identifies the period being examined; this retrospective was written in 2026. Later documents provide present-day comparisons, not knowledge attributed to that period. People, scenarios and numerical examples are illustrative teaching material, not results of a CYTIZEN engagement.