Skip to main content

Command Palette

Search for a command to run...

New York's 911 Failure Didn't Look Like a Failure. That's the Problem.

Seven hours, roughly 1,700 affected calls, and a degraded-service failure that ordinary availability and answer-time metrics can miss.

Updated
9 min readView as Markdown
New York's 911 Failure Didn't Look Like a Failure. That's the Problem.
K
I'm a systems engineer in a public safety environment. The systems I keep running are the ones people reach after something has already gone wrong, which sets a different bar than most IT work — the failure isn't a bad quarter, it's a call that doesn't connect. My work sits in three places: the network and virtualization layer under emergency communications, the security posture around it, and the practical question of where AI belongs in systems that can't afford to be confidently wrong. What I keep returning to are the failures that don't announce themselves — degraded service that still passes an availability check, a vendor dependency nobody inventoried, a monitor measuring the wrong layer entirely. I run my own infrastructure for the same reason: local inference stacks, self-hosted services, hardware I own end to end. It's the only way to see how something actually fails rather than how the documentation says it will. I write that up as field notes at blog.theknowngood.com, and maintain a reference dataset on AI model evaluation at theknowngood.com. Licensed amateur radio operator, KO6JKE. Troubleshooting an RF path and debugging a network stack are closer than they look, and both matter most when the usual channels are down.

The phrase in New York City's statement I keep going back to is "unable to hear one another."

Between 3:13 a.m. and roughly 10:19 a.m. Tuesday morning, the city says about 1,700 of the 5,659 calls into 911 were impacted, and that on certain affected calls, callers and call takers could not hear each other. After backing out alarm traffic and duplicates, City Hall estimates up to 400 New Yorkers may not have connected with a call taker at all. The affected calls were routed through a communications router at Public Safety Answering Center II in the Bronx. The city has not publicly identified that component or described its function beyond that. No cause has been publicly confirmed. The investigation is open.

One precision point, because it is the whole article. Impacted is not connected without audio. The city's own 400-person estimate says some of that traffic never reached anyone. What the record supports is a range: roughly 1,700 calls impacted during a seven-hour window, and inside that, some unknown number in which callers and call takers connected but could not hear one another.

That second category is the one ordinary delivery and answer-time metrics can miss. A call that arrives, rings, and gets answered is not a failure to anything measuring whether calls arrive, ring, and get answered.

One thing up front. This is one practitioner reading public reporting, not guidance from anyone. Every center's environment differs enough that what follows is meant as questions worth asking, not steps worth taking.


Why this kind of failure hides

An IP call is two things at once. Signaling sets up the call, routes it, rings the destination, and marks it answered. Media is the actual audio. They are negotiated together during setup but travel as separate streams, often over separate paths, and one can be intact while the other is dead.

That is a class of failure, not a diagnosis. No audio with good signaling is consistent with a media path problem, and also with a gateway or session border controller fault, a codec or transcoding problem, or something in the console audio itself. I do not know which one New York had, and neither does anyone outside the investigation. I also do not know what they were monitoring, which alarms fired, or how much of those seven hours was detection versus isolation versus repair. So the question I can put to you is not what their monitoring did. It is what yours would do. If your dashboard defines a healthy call as a delivered call, a media failure is not something you would struggle to spot. It is something that measurement cannot see at all.

Now the part I keep coming back to, because it would be true almost anywhere. To the person in the headset, a call with no audio does not present as a system fault. It presents as a silent call. Open line. Someone who dialed and cannot speak, a pocket dial, a caller hiding and unable to make noise. Silent-call handling is standard practice, and where I went looking it is a required written procedure.

That procedure is correct. It is also, in this failure mode, a mechanism that could absorb the symptom, converting a systemic fault into a routine, individually-explainable task nobody has reason to escalate. I am describing a risk, not New York's chronology. But it is the rare case where a correct procedure works against detection, because every silent call has a good non-technical explanation and the fault only exists in aggregate.

So who in your building is watching that aggregate at three in the morning?


The rule says what must happen. It doesn't say how you'd know.

One correction first. Several outlets called PSAC II the city's backup center. The city's own design documents describe PSAC II as a parallel operation to the Brooklyn center and a redundant hot site working with it. It provides backup capacity too, but calling it the city's backup center is incomplete, and it was carrying live traffic when this started.

Redundancy answers "what happens if this stops?" It has much less to say about "what happens if this keeps running and keeps being wrong?"

So I went and pulled a state's PSAP minimum standards. Yours will differ in the details, and that is the point. The one I read requires systems and processes tested to provide automatic immediate rerouting should a PSAP become unable to receive and process requests for emergency assistance.

That is a requirement about outcome. It does not specify the machine-detectable condition that has to stand in for unable to receive and process. Somebody in your building, or at your vendor, has already configured what that means in practice, and in a partial media failure that implementation detail is everything. If the configured condition is calls stopped arriving, or the center stopped answering, nothing fires until one of those things actually happens, however degraded service already is.

The continuity requirement does name PSAP system failures, so it is not blind to technical failure. It requires a plan for maintaining mission-critical call-taking and dispatch during those failures, and separately addresses evacuation, transfer to the backup site, and overflow. Those minimum standards do not expressly require media-path or voice-quality monitoring, or define how a center should detect partial audio degradation. And the answer-time standard, ninety percent within ten seconds, measures the beginning of a call, not whether the call worked.

Here is what surprised me. The field's own guidance already has this right. NENA's Managing and Monitoring NG9-1-1 information document treats two-way real-time media as part of the call itself, not as an optional success criterion checked after signaling completes. But that is an information document, which NENA describes as a source for the voluntary use of communication centers rather than an operational directive. That is what an information document is for, and no knock on it. When I checked that state's binding equipment requirements, the NENA technical standard incorporated there by name is the NG911 GIS data model. That mandates a data model for location and routing data. I found no equivalent requirement for media-path or voice-quality monitoring.

I am not singling that state out. I would expect the same shape in most of them, which is why the useful move is to go read your own rather than take my word for what it says.

That is the gap. Not ignorance. Translation.

The federal picture rhymes. A modernized FCC NG911 reliability framework took effect eight days before this incident, but its obligations run to covered providers in the 911 delivery network and expressly not to a PSAP or 911 authority to the extent it is providing those capabilities itself. One of its IP-network monitoring benchmarks centers on geographically distributed automatic disruption detection and alarming. And several core provisions do not require compliance until eighteen months after a future public notice. Meanwhile New York's own utility regulator told the Commission that widespread 911 outages in the state took significantly longer to identify and understand because the networks involved fell outside the reliability rules.

A state regulator, on the record, saying the hard part was knowing.


What I'd be asking, by role

Dispatch floor. Is anyone treating silent and open-line volume as a signal about the system rather than a category in a monthly report? The silent-call procedure itself should stay exactly as written. The question is whether there is a threshold above baseline that triggers a technical escalation, and whether telecommunicators have a low-friction way to say "something upstream feels off" before they are certain, with no penalty for being wrong. A false positive costs a phone call.

Agency leaders. Does your continuity plan have a degraded-service play? Not the evacuation play. The one covering your center staffed, powered, reachable, answering calls, and not delivering service. And who can order a rollback at 3:40 in the morning? If the honest answer is "the director, once somebody wakes him up," that sets a floor on recovery time that no amount of redundancy will lower.

IT and technical staff. Two questions for whoever owns your call delivery path, asked rather than accused, because the answers may be good ones. Does anything being monitored distinguish a call that was delivered from a call that had two-way audio? Delivered-call counts, trunk status, and console state can all keep reporting healthy through some media failures, so something has to measure media health, not just call delivery. And what specifically triggers the automatic reroute? If the condition is "the center stopped answering," ask out loud what happens when the center keeps answering.

Procurement. If your reliability language is written around binary availability, ask what the contract says about degraded service that never meets its definition of an outage, and who has to tell you before a change is pushed into your call path. That second one is a contract question, not a regulatory one. FCC rules since April 2025 require originating service providers and covered 911 service providers to notify potentially affected PSAPs within thirty minutes of discovering a qualifying outage, but that is outage notification, not advance warning.

No verified root cause is public yet, and an early confident causal claim deserves scrutiny until the investigation is documented. None of this needs to wait for it.

The question worth carrying onto your own floor this week is narrower than anything New York has to answer. If a call came in right now, connected, got answered, and carried no usable two-way audio, what in your building would know?

The specifics of Tuesday belong to New York. The shape of it does not.


I write here in a personal capacity. Nothing above represents the position of any employer, agency, or professional association I am affiliated with, and nothing in it should be read as guidance from any of them.