Get Full Government Meeting Transcripts, Videos, & Alerts Forever!
Get email alerts on the It Incident topic
No spam. Unsubscribe anytime.
DPH describes Zuckerberg San Francisco General data-center overheating incident; no patient data lost
Summary
Acting CIO Winona Miedolovich told the commission a chiller failure and multiple alert failures caused partial server outages on March 2 at a Zuckerberg San Francisco General data-center; clinical systems were down between ~2.5 and 5.5 hours depending on the application, there was no confirmed data loss, and DPH will complete an after-action review with recommendations within 90 days.
Get email alerts on the It Incident topic
No spam. Unsubscribe anytime.
The Department of Public Health reported to the Health Commission that on March 2 a data-center room overheated at a facility supporting Zuckerberg San Francisco General, causing intermittent server outages. Acting chief information officer Winona Miedolovich said the problem stemmed from a chiller that failed to fail over to backup units and from a cascade of alert failures.
Miedolovich described four alert layers; she said the first three failed and only the fourth produced an email that prompted a manager to contact on-call staff. "The first 3 alerts failed," she said, and later explained the data center had cooled after the secondary chiller was manually activated; systems returned online between roughly 08:30 and 11:30 depending on the application.
When asked by commissioners whether any data had been lost, Miedolovich said, "We did not lose data but there is what we think is damaged because we've had failures on boards and systems within the service center within the data center." She added the department is assessing hardware stress and will rely on maintenance agreements and vendor repairs where needed.
Clinical systems experienced limited downtime; Miedolovich said some clinical interfaces were unavailable for between two-and-a-half and five-and-a-half hours, and one interface was not verified functional until the afternoon. The hospital activated incident command and used Epic downtime procedures and local backups so that critical medication and patient-care data remained accessible via manual charting and local systems.
DPH and Facilities are updating alarm-notification procedures, will retrain staff, broaden IT alerts to include facilities, and plan an after-action report by April with recommendations within 90 days. The department noted prior work to refresh aging infrastructure and said Epic is fully redundant while other systems need prioritized redundancy planning.
Next steps: Facilities and IT will determine why the chiller failed to fail over, implement updated alerting and monitoring, retrain staff, and deliver recommendations from the after-action review within the stated timeline.
