The failure nobody sees inside their own scope

In this illustrative 2020 scenario, Nora runs an order-processing service. The ERP works, the portal is available and the service desk meets its targets. Yet factory orders no longer become shipments. Karim, the application lead, shows green indicators; Lucie, the factory lead, shows stationary pallets. Both are right. Between their scopes sit a service identity, an interface, an on-call team and a printer whose owner is absent from the meeting.

A useful dependency map does not depict the entire information system. It explains how a business outcome can stop while individual components remain available. In 2020, unavailable teams and suppliers, combined with emergency remote access, made this particularly visible. Written in 2026, this retrospective does not attribute later products or obligations to 2020. Its purpose is programme leadership and coordination, not a claim of technical expertise in every component.

Start with a transaction rather than an application list

Nora selects an order in the exercise environment and records its identifier and creation, approval, transfer and picking times. For each transition, she asks which actor or component must respond before the next step starts. An arrow from ERP to warehouse is insufficient. Transfer uses an account, channel, format, schedule and recovery mechanism. Apparently automatic interfaces deserve scrutiny because unusual conditions often require human intervention.

Karim adds identity, name resolution, network, certificates, scheduling, storage and backup. Lucie adds the shared terminal, scanner, printer, shift coverage and exception-handling knowledge. They distinguish normal processing dependencies from recovery dependencies. Backup is not used for every order, but matters when data needs restoration. Supplier assistance may be optional during normal service and indispensable when a certificate expires or an interface becomes stuck.

A completed register exposes the real failure point

Step and dependencyPossible interruptionOwner, fallback and check
Order approval: interface service identity.Account locked or secret expired; the ERP remains accessible but messages stop.Karim appoints an identity owner. Replay only unacknowledged messages, retaining identifiers and checking for duplicates.
Transfer: message queue and network between data centre and site.Connection is up but messages are delayed; displayed business status does not prove receipt.The interface team records dispatch, acknowledgement and queue depth. Lucie allows picking only after receipt is established.
Picking: terminal, scanner, printer and trained operator.Equipment works but the replacement operator cannot handle exceptions.The factory provides a spare terminal and a procedure tested by the next shift. Paper requires reserved numbering and reconciliation.
Recovery: usable backup and restoration capability.Backup exists but access rights or reading tools are unavailable.Julien restores an isolated sample and compares records. The supplier supplies prerequisites; business acceptance remains internal.
Supplier support: portal, contact and contractual authority.The portal uses the failed identity or the main contact is absent.Procurement and operations retain an alternative channel and deputy. Test that assistance can be requested independently of the failed service.

In practice: three hidden dependencies in one exercise

Nora simulates an unavailable interface account within an authorised environment. The dashboard initially still shows its last known value. Karim adds freshness: green at 07:30 does not establish availability at 08:15. The account used to inspect the queue then turns out to rely on the same identity service. Finally, only one authorised specialist can replay messages, and that specialist works in another time zone. The technical diagram was correct; its recovery assumptions were not.

An exercise is not an unannounced production outage. It has defined scope, a window, participants, an immediate stop condition and a known recovery route. Dangerous scenarios may require a procedural walkthrough rather than live simulation. Evidence must still be labelled accurately: procedure review, restoration on a copy and full business-flow exercise demonstrate different levels of preparedness. The aim is to avoid false confidence, not to stage impressive failures.

Prioritise impact and concentration

Prioritise dependencies whose loss stops an essential flow before fallback can begin, and dependencies supporting several flows simultaneously. Nora compares her map with invoicing. Both use the same directory and the same interface specialist. Apparently autonomous services therefore share a concentration of expertise. Adding a supplier does not reduce exposure if it uses the same infrastructure or individual.

The register separates tolerated business interruption, detection time, mobilisation time, technical restoration and reconciliation. Restoring a server in twenty minutes may still mean two hours of business interruption. Figures are measured or explicitly labelled estimates, with date and conditions. Unknowns remain unknown; entering zero to make the table reassuring destroys the decision value of the map.

Build a fallback that does not recreate the dependency

A second site is not independent when it shares the same network, administrative identity and unavailable team. A daily export helps only if someone can read it, establish its age and reconcile changes made afterwards. Nora asks each owner what remains usable when the dependency disappears. That question distinguishes a technical copy from operational continuity.

Degraded operation has a maximum volume, duration, excluded categories and exit procedure. In the scenario, the factory handles only previously approved orders, retains original identifiers and excludes price or batch changes during fallback. Before normal processing resumes, a named team compares temporary operations with the ERP. The decision maker accepts delayed and excluded orders. Promising to do everything differently without limiting throughput merely transfers the outage to employees.

Sector differences depend on the business outcome

In manufacturing and pharma, the map covers central services supporting local equipment and dependencies needed to demonstrate batch conformity. A working controller is insufficient if required records are lost. Maintenance windows and operator handovers become continuity inputs. Quality and safety define authorised conditions for degraded processing.

In banking and insurance, an operation may depend on an external provider, data source and later reconciliation. Separate accepting a request, executing it and confirming its result to the customer. Unique identifiers and external status protect against duplicates. In services and public administration, contact-centre capacity or a callback channel may be more urgent than a secondary application. Check fallback accessibility with affected users rather than assuming an available telephone number is enough.

Maintain the map without creating another inventory

Keep dependencies needed for decisions on selected business flows. Owners review identity, site, supplier and on-call changes when those changes occur. Track assumptions unsupported by tests and give business acceptance criteria a named owner. Link to existing technical references instead of copying their details. Each review must decide whether to strengthen a dependency, test fallback, temporarily accept a risk or reduce service. An inventory alone provides no continuity.

Sources and method

Primary sources checked on 4 October 2026. The year identifies the period being examined; this retrospective was written in 2026. Later documents provide present-day comparisons, not knowledge attributed to that period. People, scenarios and numerical examples are illustrative teaching material, not results of a CYTIZEN engagement.