New York's 911 Failure Didn't Look Like a Failure. That's the Problem.
Seven hours, roughly 1,700 affected calls, and a degraded-service failure that delivery and answer-state metrics can miss.

Update, August 26, 2026: On August 24 the mayor said the city had identified an early-morning Motorola software update as the event that created the problem. He also said a significant number of calls were rerouted to another part of the 911 system, and that fewer than 400 unique calls were not connected, a more precise formulation than the earlier estimate. The public record does not establish who initiated, approved, scheduled, or executed the change. There is still no published technical root cause, nothing on what the update changed, which component failed, or how it produced the loss of audio. The disclosure tells us what initiated the incident at a high level. It does not tell us how that update produced the audio failure, or what New York's monitoring saw. I have corrected the sections below that this supersedes. The rerouting disclosure also narrows one part of my argument: calls did move elsewhere in the system that morning. The city has not said whether that rerouting was automatic, what triggered it, when it began, or whether it was routine load balancing rather than a response to the failure.
The phrase in New York City's statement I keep going back to is "unable to hear one another."
Between 3:13 a.m. and roughly 10:19 a.m. Tuesday morning, the city says about 1,700 of the 5,659 calls into 911 were impacted, and that on certain affected calls, callers and call takers could not hear each other. The mayor later said fewer than 400 unique calls were not connected. Certain affected calls were routed through a communications router at Public Safety Answering Center II in the Bronx. The city has not publicly identified that component or described its function beyond that. The mayor has since said the city identified a Motorola software update as the triggering change, but the city has not published a technical root-cause analysis. The investigation is open.
One precision point, because it is the whole article. The public record establishes three things: roughly 1,700 impacted calls, fewer than 400 unique calls that were not connected, and audio failure on certain calls. The city has not published how those categories overlap. Impacted is not the same as connected without audio, and I am not going to pretend I know how many of each there were.
But the connected-without-audio category is the one worth an article, because it is the one delivery and answer-state metrics can miss. A call that arrives, rings, and gets answered is not a failure to anything measuring whether calls arrive, ring, and get answered.
Two things up front. I am not claiming New York's dashboards showed the system as healthy. The city has not disclosed what its monitoring showed, and the narrower point I am making is that delivery and answer-state metrics alone cannot establish that a call carried usable two-way media. And this is one practitioner reading public reporting, not guidance from anyone. Every center's environment differs enough that what follows is meant as questions worth asking, not steps worth taking.
Why this kind of failure hides
For an IP voice call, signaling establishes and controls the session, and media carries the audio. Signaling sets the call up, routes it, rings the destination, and marks it answered. They are negotiated together during setup but travel as separate streams, often over separate paths, and one can be intact while the other is dead.
That is a class of failure, not a diagnosis. No audio with good signaling is consistent with a media path problem, and also with a gateway or session border controller fault, a codec or transcoding problem, or something in the console audio itself. The city has identified the initiating change but has not publicly identified the failure mechanism. I also do not know what they were monitoring, which alarms fired, or how much of those seven hours was detection versus isolation versus repair. So the question I can put to you is not what their monitoring did. It is what yours would do. If a metric only establishes that a call was delivered and answered, that metric by itself cannot establish that the call carried usable two-way media. That is not a hard problem to spot. It is outside what the measurement covers.
Now the part I keep coming back to, because it would be true almost anywhere. To the person in the headset, an individual no-audio call can present as a silent or open-line call rather than an obvious system fault. Someone who dialed and cannot speak, a pocket dial, a caller hiding and unable to make noise. Silent-call handling is standard practice, and where I went looking it is a required written procedure.
That procedure is correct. It is also, in this failure mode, a mechanism that could absorb the symptom, making individual instances look routine until an aggregate pattern emerges. I am describing a risk, not New York's chronology. But it is the rare case where a correct procedure can work against detection, because many silent calls have plausible non-technical explanations and the fault only shows up in aggregate.
So who in your building is watching that aggregate at three in the morning?
The rule says what must happen. It doesn't say how you'd know.
One correction first. PSAC II is sometimes described as the city's backup center. The city's own design documents describe a more active architecture: a parallel operation to the Brooklyn center, load-balanced with it, and a redundant hot site. It provides backup capacity too, but calling it the city's backup center is incomplete, and it was carrying live traffic when this started.
Redundancy only helps with a degraded failure if something can recognize a condition worth failing away from.
So I went and pulled a state's PSAP minimum standards. Yours will differ in the details, and that is the point. The one I read requires systems and processes tested to provide automatic immediate rerouting should a PSAP become unable to receive and process requests for emergency assistance.
That is a requirement about outcome. It does not specify the machine-detectable condition that has to stand in for unable to receive and process. Whatever mechanism implements that reroute requirement has to operationalize what unable to receive and process means, and in a partial media failure that implementation detail is everything. If the configured condition is calls stopped arriving, or the center stopped answering, nothing fires until one of those things actually happens, however degraded service already is.
The continuity requirement does name PSAP system failures, so it is not blind to technical failure. It requires a plan for maintaining mission-critical call-taking and dispatch during those failures, and separately addresses evacuation, transfer to the backup site, and overflow. Those minimum standards do not expressly require media-path or voice-quality monitoring, or define how a center should detect partial audio degradation. And the staffing standard, which requires staffing adequate to answer ninety percent of incoming 911 requests within ten seconds, measures when the request was answered, not whether the call worked.
Here is what surprised me. The field's own guidance already has this right. NENA's i3 architecture treats two-way real-time media as part of a human-initiated call, not as an optional success criterion checked after signaling completes. Its separate Managing and Monitoring NG9-1-1 information document goes further, and explicitly discusses monitoring streaming media quality against thresholds. But NENA describes that one as guidance and best practices rather than requirements or specifications. That is what an information document is for, and no knock on it. When I checked that state's binding equipment requirements, the NENA technical standard incorporated there by name is the NG911 GIS data model. That mandates a data model for location and routing data. I found no equivalent requirement for media-path or voice-quality monitoring.
I am not singling that state out, and the useful move is to go read your own rather than take my word for what it says.
That is the gap. Not ignorance. Translation.
The federal picture rhymes. A modernized FCC NG911 reliability framework took effect eight days before this incident, but its obligations run to covered providers in the 911 delivery network and expressly not to a PSAP or 911 authority to the extent it is providing those capabilities itself. One of its IP-network monitoring benchmarks centers on geographically distributed automatic disruption detection and alarming. And the framework became effective eight days before the incident, although most of its new NG911 compliance obligations were not yet due. Meanwhile New York's own utility regulator told the Commission that prior widespread 911 outages in the state took significantly longer to identify and understand because the networks involved fell outside the reliability rules.
A state regulator, on the record, saying the hard part was knowing.
What I'd be asking, by role
Dispatch floor. Is anyone treating silent and open-line volume as a signal about the system rather than a category in a monthly report? Nothing here argues for abandoning normal silent-call handling. The question is whether there is a threshold above baseline that triggers a technical escalation, and whether telecommunicators have a low-friction way to say "something upstream feels off" before they are certain, with no penalty for being wrong. That asymmetry usually favors reporting it.
Agency leaders. Does your continuity plan have a degraded-service play? Not the evacuation play. The one covering your center staffed, powered, reachable, answering calls, and not delivering service. And who can order a rollback at 3:40 in the morning? If the honest answer is "the director, once somebody wakes him up," that adds avoidable recovery time that no amount of redundancy will lower.
IT and technical staff. Two questions for whoever owns your call delivery path, asked rather than accused, because the answers may be good ones. Does anything being monitored distinguish a call that was delivered from a call that had two-way audio? Delivered-call counts, trunk status, and console state can all keep reporting healthy through some media failures, so some signal has to represent media health, or aggregate the symptom. And what specifically triggers the automatic reroute? If the condition is "the center stopped answering," ask out loud what happens when the center keeps answering.
Procurement. If your reliability language is written around binary availability, ask what the contract says about degraded service that never meets its definition of an outage, and who has to tell you before a change is pushed into your call path. That second one is a contract question, not a regulatory one. FCC rules since April 2025 require originating service providers and covered 911 service providers to notify potentially affected PSAPs within thirty minutes of discovering an outage that potentially affects a 911 special facility, but that is outage notification, not advance warning.
The city has identified the triggering change. It has not published the technical failure mechanism, the affected component, or a completed analysis. None of this needs to wait for that.
The question worth carrying onto your own floor this week is narrower than anything New York has to answer. If a call came in right now, connected, got answered, and carried no usable two-way audio, what in your building would know?
The specifics of Tuesday belong to New York. The shape of it does not.
I write here in a personal capacity. Nothing above represents the position of any employer, agency, or professional association I am affiliated with, and nothing in it should be read as guidance from any of them.





