At Karolinska Institutet, Cleveland Fertility Center and Pacific Fertility Center, sophisticated laboratories were monitoring environments designed to protect irreplaceable biological material. Thousands of samples were still lost. Not because nobody was monitoring, but because monitoring alone was not enough.
These incidents expose a fundamental weakness in the way critical laboratory environments are often protected. Sensors measure. Alarms alert. Dashboards display. But when something goes wrong, none of that guarantees that the right person will know, understand, take responsibility and act in time.
The difference between monitoring and operational assurance can be measured in what happens next.

Karolinska Institutet: the failure started before the temperature changed

In December 2023, Karolinska Institutet suffered a catastrophic failure of the cryogenic storage infrastructure at its Neo facility in Stockholm. Approximately 47,100 biological samples, accumulated over decades, were ultimately lost.
The investigation is revealing precisely because the incident did not begin with a freezer suddenly becoming too warm. During planned maintenance an oxygen level alarm was triggered. That closed the valve on the bulk liquid nitrogen tank and interrupted the supply to the cryogenic storage tanks. The supply was never restored.
The individual tanks held enough reserve to keep going for a while, so their liquid nitrogen levels declined gradually. Only later did the conditions inside the tanks deteriorate. The sequence was effectively:
Maintenance event → LN₂ supply interrupted → reserve depleted → conditions deteriorated → alarms generated → intervention failed → samples lost.
The post mortem found a chain of technical and organisational shortcomings rather than one isolated mistake. Responsibilities were unclear. Communication and information sharing were inadequate. Alarm testing was insufficient. Contact details were out of date. Earlier incidents had not been documented properly, and the organisation had no effective 24/7 response capability over the Christmas period.
The alarms were not simply absent. Signals were generated. Messages were sent. What the organisation lacked was a reliable mechanism for turning those signals into a verified response.
The lesson is profound: an alarm is not a control system. It is one link in a control system.

Cleveland Fertility Center: when the safety net was switched off

In March 2018, University Hospitals Fertility Center in Cleveland suffered a cryogenic storage incident affecting more than 4,000 eggs and embryos belonging to roughly 950 patients.
The storage tank had a remote alarm intended to notify staff when conditions deteriorated. The alarm had been switched off. The hospital could not establish when it had been disabled, or by whom. The tank also had problems with its liquid nitrogen refill system. When storage conditions deteriorated, the expected remote warning simply did not provide the protection it was designed to provide.
Line drawing of a laboratory monitoring unit with its remote alarm output switched off and its network cable disconnected
This is a different weakness from Karolinska. The question was not whether an alarm existed, but whether the organisation knew that the alarm itself was available, connected and operational.
A monitoring system that can silently become unavailable creates something more dangerous than no monitoring at all: false confidence.

Pacific Fertility Center: the human became the last line of defence

On the same day as the Cleveland incident, Pacific Fertility Center in San Francisco suffered a cryogenic storage failure in a tank holding approximately 2,500 embryos and 1,500 eggs.
This case teaches the opposite lesson. An embryologist discovered during a physical walkthrough that the liquid nitrogen level in one tank had fallen dangerously low. The team responded immediately and transferred the material to another tank. Human inspection became the critical safety barrier.
The subsequent investigations examined the failure of the tank and the reliability and redundancy of its monitoring and alarm systems. Detailed reporting highlighted an issue that faces every fertility clinic and cryogenic facility: no single safeguard should be assumed to be infallible.
The lesson is not that human inspection is unnecessary. It is that human inspection should not have to be the final defence against a failure an independent monitoring system ought to detect.

Three incidents, three failure modes, one problem

Technically, the three cases were very different.
  • At Karolinska, the supply feeding the controlled environment was interrupted and the organisation failed to recover it in time.
  • At Cleveland, the remote alarm system itself was unavailable.
  • At Pacific Fertility, human observation caught a deteriorating condition that automated protection apparently failed to surface reliably.
The underlying problem was remarkably similar. The organisation could not reliably close the loop between a change in the critical environment and an effective intervention.
That loop needs far more than a sensor. It needs:
Line diagram of the assurance chain from sense to assure, with a broken link between alert and escalate
And every link matters:
  • a valve can be closed
  • a tank can lose nitrogen
  • a sensor can fail
  • a network connection can disappear
  • an alarm can be disabled
  • a phone number can be out of date, an email can land in a junk folder, a maintenance intervention can change the state of a system
  • the responsible person can be unavailable
  • an earlier alarm can create alarm fatigue
None of these events is necessarily catastrophic on its own. The catastrophe happens when the organisation cannot detect, understand and respond to the chain of events before the remaining safety margin is exhausted.

The limits of conventional environmental monitoring

Traditional environmental monitoring rests on a simple proposition: measure the environment and raise an alarm when a threshold is exceeded.
That model remains essential, but critical environments demand more. Consider Karolinska. Monitoring the temperature of the cryogenic tanks alone would have detected the problem only after the LN₂ supply had already been interrupted and the reserve had begun to disappear.
A more resilient approach also asks:
  • Is the LN₂ supply functioning?
  • Is the replenishment system operating?
  • Has a maintenance intervention changed the state of the system?
  • Is the expected refill actually happening?
  • Is the alarm system itself operational?
  • Are the right people receiving the alerts?
  • Has anyone acknowledged the event?
  • What happens if nobody responds?
  • Is there 24/7 escalation?
  • Can the organisation prove what happened?
This is the difference between measuring an environment and assuring an environment.

From environmental monitoring to operational assurance

Operational assurance starts from a different question. Not "what is the temperature?", but "can we demonstrate that this critical environment remains under control?"
That requires an independent layer connecting infrastructure, data, technology and people. It means continuously understanding not only the environmental condition, but also the health of the system responsible for protecting that condition. In practice:
  • recognising abnormal patterns
  • detecting loss of communication
  • verifying that alarms are operational
  • knowing who owns a critical event
  • escalating when an alert is not acknowledged
  • keeping an auditable record of what happened and how the organisation responded
  • recognising that the most important warning may arrive before the final environmental parameter crosses its critical threshold
The objective is not to promise that equipment will never fail. Equipment will fail. The objective is to make sure that equipment failure does not silently become asset failure.

Where XiltriX Resilience is different

XiltriX Resilience was designed around this principle. Rather than treating environmental monitoring as a collection of sensors, alarms and dashboards, it creates an independent operational assurance layer around critical environments, connecting three essential elements.
Line drawing of the closed assurance loop: laboratory infrastructure, a monitoring station and a 24/7 operations desk, with a verified response returning to the equipment
Infrastructure. The physical systems and assets that create and maintain the required environment.
Intelligence. The continuous collection, validation and interpretation of environmental and operational data, identifying not only deviations but patterns, anomalies and failures in the monitoring chain itself.
Care. The people, processes and accountability that turn a detected issue into an appropriate response rather than another unanswered alarm.
Together they form a closed assurance loop:

Infrastructure → Intelligence → Alert → Escalation → Care → Verified response

This is fundamentally different from installing more sensors. The lesson from Karolinska, Cleveland and Pacific Fertility is not that laboratories need more data. They need confidence that the data, the infrastructure and the response system can be trusted when it matters most.

Ordinary events, irreversible consequences

The three incidents in this paper were not caused by one extraordinary failure. They emerged from ordinary events:
  • a maintenance intervention
  • a closed valve
  • a depleted reserve
  • a disabled alarm
  • a failed detection mechanism
  • an incomplete response
That is precisely why they matter. Critical laboratory environments do not fail only when something spectacular happens. They fail when small deviations accumulate without being detected, understood or acted upon.
Operational assurance is about making those invisible chains visible and interrupting them before they become irreversible.
That is what XiltriX Resilience is about. Not simply telling you when the environment has changed, but providing continuous confidence that the environment, the monitoring system and the response process are all working together to protect what cannot be replaced.
Nothing left to chance.

Meet us at WOTS 2026

XiltriX will be at WOTS, World of Industry, Technology & Science, 22-25 September 2026 at Jaarbeurs Utrecht, booth 11B006. Come and see what operational resilience means in practice and how XiltriX Resilience delivers asset and process assurance.