Honestly. Of all the places to deploy first I’m happy it’s that one.
The government politicians who approved this are frequent flyers out of there, so they’re at least getting their lines out on the line before the rest of the country if it proves ill managed or conceived.
Over the weekend, I got linked to https://admiralcloudberg.medium.com/reaping-the-whirlwind-in..., which is an exhaustive look at the causes of the first airspace collision in the US in several decades. After reading that, my conclusion is that, if the software being rolled out is what I think it is, it would be a borderline impeachable offense to not first trying it out in the DC airspace.
The fundamental problem with DCA, and one of the two main causes of the crash [1], is that they are required by political pressure to operate at a higher operational tempo than they can safely operate at. DCA has essentially 1½ usable runaways--large planes can only use the larger runway, and given the capacity restrictions, airlines have been pushing to use fewer small planes at the airport. This shift in plane size means the effective safe slot capacity has gone down, but people still keep citing the same number as justification for safe numbers, and the politicization of the issue has shut down everyone who complained that the actual traffic just couldn't be safely handled.
If DCA's theoretical slot capacity (36 landings and takeoffs each per hour) is to be reached one just one runway, you have about 100 seconds to go from plane 1 touchdown; exit the runway, letting plane 2 on to take off; plane 2's wheel leaving the runway, clearing plane 3 touchdown. That's doable, but with essentially 0 margin for error. Alternating the runways used for each landing would give you closer to 30s of margin, but if only 10-20% of the planes can use the alternate runway, you can't divert enough planes to use the alternate runway.
Now DCA and the FAA aren't stupid enough to actually schedule 36 landing slots every hour, but the problem is that because of a few various factors, what was scheduled as, say, 30 slots for an hour (giving 120 seconds between landings, probably sufficient margin) ends up being 12 slots used in the first half-hour and 18 slots used in the second half-hour, which means a lot of the actual operation ends up having no safety margin even though on paper you have sufficient margin.
One of the ways you can rectify that is to assign slots further in advance so that you don't get the bunching. Effectively saying "oh, if you leave right now, you'll arrive at 3:30 with three other planes, but if I hold you for 10 minutes, I can push you into a less busy arrival time." Doing this requires good, accurate prediction of the actual flight travel times, and my understanding is that this is what the new software is meant to provide.
[1] The other main cause is essentially that the DoD's aviation practices in the area is a giant clusterfuck that endangers lives, and unfortunately that sentence is not relegated to the past tense.
Local people know where they are because once upon a time a D8 drove through pulling a big orange spool of fiber conduit and passers by chatted about how another fiber was going in. Also around here the contractors seeded the disturbed ground with a wildflower mix (presumably to discourage weeds) so you can also look for straps of wildflowers.
And of course there are marker posts that say "buried fiber optic cable do not dig".
Once you get out farther from civilization these signs become less common and there are more situations where they can get removed.
Most cheap hired labor doesn't realize you shouldn't dig up the market posts and throw them in the dumpster on construction sites. Add a little bit of rain and suddenly the contractor showing up with an excavator that was told everything was properly marked, and can see markings in other places but not where they are digging, and disaster occurs.
Also soils aren't static, I've been on sites where the midpoint of a buried telecommunications cable had drifted over 8 feet from where the markers were a few hundred yards apart. Just looking at the markers and assuming a straight line isn't a safe bet. Looking at the fence rows to the left and right of the cable and you could see a matching bow in the fence rows.
Fiber was all installed long after we knew locating it was important. I suspect it all had a tracer wire installed with it. Of course sometimes the tracer wire can break.
Before fiber a lot of things were put into the ground with no thought how you would find it again.
Indeed, tracer wire is usually bonded to a (sacrificial) magnesium anode, but it’s a crapshoot whether the tracer wire is still intact.
I sometimes hire directional boring and excavation contractors and if there’s any doubt as to where an electrical conduit or natural gas pipe is, I opt for the hydrovac truck to minimize risk.
Always bring a length of fiber optic cable when you're in the wilderness. If you get lost or stranded, bury the cable. Within a few hours, a backhoe will stop by to dig it up, and the crew will rescue you.
I don't quite understand this sort of thing happening. Wasn't the whole point of the internet to be a self-healing network where we route around severed cables, etc?
Or is it that these ATC networks are their own air-gapped network with less redundancy? That just doesn't add up. Or maybe there was only one line going to the ATC, with no multiple "ISPs" like a datacenter would have?
* Yes, TRACON have their own dedicated links. You can look into ASTERIX and STARS to learn some of the cursed ways the data processing and dataflow work.
* There is supposed to be a primary and a secondary link, in this case the primary failed and the fail-over also failed. It's unclear from the reporting if they were damaged in the same incident or if the failover was not tested or monitored adequately.
The Philadelphia TRACON site has been notoriously unreliable and was supposedly improved in 2025, it's also unclear if these issue actually could stem from that implementation.
I’ve had enough double-redundant links fail - and I mean pretty thoroughly double-redundant, different provider, different media, different physical path - that I’m quite surprised something like the air traffic control system only has two links.
One thing that needs to be reviewed is if they are dual links or redundant links.
In any system where both sides of redundancy always carry traffic then loss of one link can cause congestion failure if any link fails. This is a very common means of failure in electrical networks that requires load shedding. Well, you can't load shed air traffic.
A more complex system that I like, but comes with it's own set of constraints and implementation issues is a system where both lines carry all the traffic at all times. This way the default state of the system is always working and your first failure isn't invisibly critical.
But this is very hard as we see in TCP when things get out of order and high latency creeps in. You have to manage a lot more state at the data level.
Contractually dual-path, or actually dual-path? My understanding is that there's enough infrastructure horse-trading going on behind the scenes that it's very difficulty to be certain that two circuits between points A and B don't share the same infrastructure somewhere in between.
First big oops of this form that I remember:
"In December 1986, the ARPANET had 7 dedicated trunk lines between NY and Boston, except that they all went through the same conduit -- which was accidentally cut by a backhoe. "
I'm talking about data not able to leave the building, not even data getting lost on the way.
One example was a site that had fiber and coax, from different companies. They might have shared a pipe at some point along their length, hard to say. But both connections went down at the same time for digital reasons, not physical. The providers had simultaneous unrelated backend router problems and the site lost contact to both gateways.
This was with a nice SD-WAN system that routed everything dynamically across both pipes to address latency and errors, but if the packets get dropped at the first hop in both networks, there's not much you can do!
The saving grace in that case was the third redundant connection, a 4G cell modem, with a lot less bandwidth but able to keep the critical transactions going. These days I'd certainly want Starlink on the roof as well.
Yup. Seven different paths comes with seven different sets of equipment, and seven different rights-of-way that have to be negotiated, purchased, leased, whatever.
Reporting is saying that the backup was cut a while ago and they only discovered it when the primary failed. Apparently nobody was ping testing the backup link.
True. But what if there are more people competent to set up & maintain secure tunnels through the internet than there are people competent to set up & maintain dedicated fiber links?
But what if the actual bottleneck isn't the size of the various cohorts of people with the necessary expertise? What if the backup line goes untested because bureaucrats inflict some irrational set of conditions? If so, then changing media won't help, because the organization will impede competent operation of whatever scheme you offer.
Yes, but low-functioning bureaucracies are also the most prone to outsourcing. They can't manage to get anything done internally, so they hire outside orgs (hopefully higher functioning) to perform necessary basic services. Often late in the game, when (figuratively) a VIP visit could reveal that their bureaucracy is too broken to keep its office bathrooms clean.
In this case, that outsourcing could easily look like "call local ISP's, get connections, set up tunnel".
Similar events happened in Europe disguised as thieves stealing fiber optic. This does not have any sense economically, as the value in the market is zero so... either is an honest accident and is cleared in a few days, or is sabotage
This used to happen more often in the past, when copper thieves mistook fiber for conductor. There were cases where critical systems were shut down due to fiber optic theft, and the thieves were caught burning the sheathing to expose the... glass.
>disguised as thieves stealing fiber optic. This does not have any sense economically, as the value in the market is zero
Do you honestly think that crackheads think that far in advance?
I've seen fiberoptic cables stolen from 2 (city) jobsites in the last 5 years, once by tweakers later caught trying to sell them as scrap copper and the second thief was never caught.
This happened even though the spools had big signs on them saying "Fiber Optic Cable - NO COPPER".
In my country if you want to dig into the ground more than a few meters you'd need to get a permit. Officially everything that is buried is registered but every once in a while they find an old pipe that nobody knows what it is for.
Yes that was the 811 I referred to - a US compliance and permitting hotline. They send any utilities mapped in your digging area to come mark their buried lines.
That's very uncharacteristic, especially given all those contingency requirements (backup policies, failovers, testing schedules etc.) imposed on corporations/companies in the wake of 9/11.
Well I think it’s technically pretty similar to 5G, but it doesn’t use existing 5G towers. It’s this company in Denver that sets up receivers on top of certain buildings and distributes using a line of sight connection. What’s weird is that we plug our modem into the phone jack in the wall
Not my area of expertise though so I could be leaving out info
Edit: Starry is the company, I believe they were just purchased by Verizon
Besides what the others have said, speaking from a neteng perspective, sometimes backup lines (and the core infrastructure in general) are engineered in such a way that it can't easily be tested properly without taking other things down, or manually rolling a truck specifically to test it in isolation with extra equipment.
Not saying that's what is going on here, just that it's possible.
I think this the same system that was the topic of a major contractor dispute a few months ago, between Verizon (incumbent fiber contractor) and SpaceX (lobbying to replace it).
As a backup for when ground communications are cut? As opposed to having no backup in that situation? Yeah, maybe we should, no matter who controls the company.
You'd hope Putin is not so stupid as to start committing acts of war on US territory. Helping the Iranians against the US seems to have gone mostly unnoticed but I'm pretty sure that if such an act was traced back to russia there would be a different tune played soon after.
It’s being gong on for years in Europe. Nobody’s starting a war because they cut a fibre. Or burn down a telephone exchange. Or blow up a plane. Or kill a policeman, even if it can be attributed.
So of course if the FAA is going to put AI into the watchtowers, they're obviously going to use local AI and not rely on some random Amazon cloud infra in some middle eastern country.
right. of course, no ones going to cut corners on the safe use of AI and secure infrastructure in this administration.
Building multiple diverse paths and monitoring for fiber cuts is not hard. Having 2 fiber paths is insufficient for even only moderately important workloads at a tech company. Overlapping fiber cuts happen. For something with significant economic and safety impact this is just crazy.
Sometimes the level of incompetence / lack of care in organizations like this astounds me. I understand issues like this can be complicated and systemic but it honestly makes me think very poorly of the technologists building these systems in government.
That sounds pretty plausible. But you still have to have enough internal competency to oversee a contractor, and ask the right questions/build proper requirements and verify they are meeting them.
I've told this story so many times I've worn the corners off it. BT was giving a presentation to MSN WAN OPS about the new datacenter buildout in London for us. They kept going on and on about physical security, man traps, and that our fiber left the building on each side and didn't get close to each other for so many km away.
BT guy ends that part of the presentation with "the IRA will really have to get shit together to take you off the net"
This was during the troubles. Same planet, different world.
The UK ATC system was down again for the second time this month. I'm no conspiracy theorist but hard to believe those two + this are 'accidents'. I believe the Netherlands has also had several issues recently.
Wasn't the UK ATC system the one where the whole thing core dumped, followed by the backup ATC system core dumping, when someone filed a flight plan with confusing same-name waypoints?
With low-enough engineering quality, you don't actually need saboteurs.
Finding out your backup fiber has been cut, aparently some time
ago, only when you try to use it as the backup isn’t an issue of “aging infrastructure.” It’s an issue of letting completely incompetent people run your tech.
It happens. There use to be a joke during the first big DC build out phase that went like if you're ever going into the wilderness take a 1ft length of fiber optic cable with you. If you get lost bury it and a back hoe operator will appear and sever it within an hour. You can get a ride back with them.
>A man stranded in the bush in northern Saskatchewan was rescued last week after chopping down four power poles — knocking out electricity to surrounding communities. [...]
>But he had an axe and he knew SaskPower would have to check the downed line, so he went to work.
I remember a story from an acquaintance working in construction. In one case they had to dug up right where a fiber optic ran, so they had cleared it with its operator and waited for their go-ahead. And waited. And waited, until the boss said something to the effect of 'start digging, we have a schedule to keep'...
back in the day, i worked for Energis (#3 wholesale telco in the UK) … we had about 2.2m active subscribers and ran a fibre backbone.
one day the backbone (E1 bidirectional ring iirc) was broken north of london … my boss drove out to find that travellers had entered the field, opened the inspection cover in the ground, stood a pole for their clothes line and filled with cement
very savvy, he didn’t let on that we had about £1m revenue per day close to collapse and convinced them to hang their clothes somewhere else (perhaps a few quid changed hands)
101 comments
[ 0.29 ms ] story [ 35.1 ms ] threadhttps://www.airwaysmag.com/new-post/faa-smart-first-deployme...
Good luck to all of us.
The government politicians who approved this are frequent flyers out of there, so they’re at least getting their lines out on the line before the rest of the country if it proves ill managed or conceived.
It’s basically dogfooding.
The fundamental problem with DCA, and one of the two main causes of the crash [1], is that they are required by political pressure to operate at a higher operational tempo than they can safely operate at. DCA has essentially 1½ usable runaways--large planes can only use the larger runway, and given the capacity restrictions, airlines have been pushing to use fewer small planes at the airport. This shift in plane size means the effective safe slot capacity has gone down, but people still keep citing the same number as justification for safe numbers, and the politicization of the issue has shut down everyone who complained that the actual traffic just couldn't be safely handled.
If DCA's theoretical slot capacity (36 landings and takeoffs each per hour) is to be reached one just one runway, you have about 100 seconds to go from plane 1 touchdown; exit the runway, letting plane 2 on to take off; plane 2's wheel leaving the runway, clearing plane 3 touchdown. That's doable, but with essentially 0 margin for error. Alternating the runways used for each landing would give you closer to 30s of margin, but if only 10-20% of the planes can use the alternate runway, you can't divert enough planes to use the alternate runway.
Now DCA and the FAA aren't stupid enough to actually schedule 36 landing slots every hour, but the problem is that because of a few various factors, what was scheduled as, say, 30 slots for an hour (giving 120 seconds between landings, probably sufficient margin) ends up being 12 slots used in the first half-hour and 18 slots used in the second half-hour, which means a lot of the actual operation ends up having no safety margin even though on paper you have sufficient margin.
One of the ways you can rectify that is to assign slots further in advance so that you don't get the bunching. Effectively saying "oh, if you leave right now, you'll arrive at 3:30 with three other planes, but if I hold you for 10 minutes, I can push you into a less busy arrival time." Doing this requires good, accurate prediction of the actual flight travel times, and my understanding is that this is what the new software is meant to provide.
[1] The other main cause is essentially that the DoD's aviation practices in the area is a giant clusterfuck that endangers lives, and unfortunately that sentence is not relegated to the past tense.
(The actual easiest method is to install the conduit or direct burial cable with tracer wire)
And of course there are marker posts that say "buried fiber optic cable do not dig".
Most cheap hired labor doesn't realize you shouldn't dig up the market posts and throw them in the dumpster on construction sites. Add a little bit of rain and suddenly the contractor showing up with an excavator that was told everything was properly marked, and can see markings in other places but not where they are digging, and disaster occurs.
Also soils aren't static, I've been on sites where the midpoint of a buried telecommunications cable had drifted over 8 feet from where the markers were a few hundred yards apart. Just looking at the markers and assuming a straight line isn't a safe bet. Looking at the fence rows to the left and right of the cable and you could see a matching bow in the fence rows.
Before fiber a lot of things were put into the ground with no thought how you would find it again.
I sometimes hire directional boring and excavation contractors and if there’s any doubt as to where an electrical conduit or natural gas pipe is, I opt for the hydrovac truck to minimize risk.
Or is it that these ATC networks are their own air-gapped network with less redundancy? That just doesn't add up. Or maybe there was only one line going to the ATC, with no multiple "ISPs" like a datacenter would have?
* There is supposed to be a primary and a secondary link, in this case the primary failed and the fail-over also failed. It's unclear from the reporting if they were damaged in the same incident or if the failover was not tested or monitored adequately.
The Philadelphia TRACON site has been notoriously unreliable and was supposedly improved in 2025, it's also unclear if these issue actually could stem from that implementation.
In any system where both sides of redundancy always carry traffic then loss of one link can cause congestion failure if any link fails. This is a very common means of failure in electrical networks that requires load shedding. Well, you can't load shed air traffic.
A more complex system that I like, but comes with it's own set of constraints and implementation issues is a system where both lines carry all the traffic at all times. This way the default state of the system is always working and your first failure isn't invisibly critical.
But this is very hard as we see in TCP when things get out of order and high latency creeps in. You have to manage a lot more state at the data level.
First big oops of this form that I remember:
"In December 1986, the ARPANET had 7 dedicated trunk lines between NY and Boston, except that they all went through the same conduit -- which was accidentally cut by a backhoe. "
https://www.csl.sri.com/~neumann/insiderisks06.html
One example was a site that had fiber and coax, from different companies. They might have shared a pipe at some point along their length, hard to say. But both connections went down at the same time for digital reasons, not physical. The providers had simultaneous unrelated backend router problems and the site lost contact to both gateways.
This was with a nice SD-WAN system that routed everything dynamically across both pipes to address latency and errors, but if the packets get dropped at the first hop in both networks, there's not much you can do!
The saving grace in that case was the third redundant connection, a 4G cell modem, with a lot less bandwidth but able to keep the critical transactions going. These days I'd certainly want Starlink on the roof as well.
That seems like an extremely foolhardy thing to do.
In this case, that outsourcing could easily look like "call local ISP's, get connections, set up tunnel".
One of the more popular means of diffusing responsibility.
Pretty grim that a life critical system wasn't designed to report that the backup fibre was unserviceable until they attempted to switch over to it.
I wonder how long it was down? Days, weeks, months?
Similar events happened in Europe disguised as thieves stealing fiber optic. This does not have any sense economically, as the value in the market is zero so... either is an honest accident and is cleared in a few days, or is sabotage
Do you honestly think that crackheads think that far in advance?
I've seen fiberoptic cables stolen from 2 (city) jobsites in the last 5 years, once by tweakers later caught trying to sell them as scrap copper and the second thief was never caught. This happened even though the spools had big signs on them saying "Fiber Optic Cable - NO COPPER".
Recently, I tried calling 811 before digging in my yard. The webpage was broken and the hotline kept me on hold forever. I gave up. Small wonder.
https://blackhydrovac.com/underground-utility-strikes-learn-...
https://cybersquirrel1.com/
Plenty of people don’t bother and YOLO it, and usually it’s fine - until it isn’t.
The punchline is that it's about 2000 ft from the local CO, where (I believe) half the town's lines terminate.
I am paying, $1000, $1800 & $1900 for the same service at 3 different location (20 mins from each other).
Two locations, I also have old coax lines that are still active, but not paying for it.
When I bought two businesses, I learned that they were paying for a dedicated fiber but using coax service.
At one of the location, we had fiber, and paying for backup coax and wireless. But if you turn off fiber box, it wont fail over to either one.
I wouldn't surprise it was down for weeks and no one bothered about it.
I thought about what I believe is your way until https://www.tomshardware.com/opinion/t-mobile-home-internet-... and https://www.tomshardware.com/news/t-mobile-misleads-home-int... .
Not my area of expertise though so I could be leaving out info
Edit: Starry is the company, I believe they were just purchased by Verizon
Not saying that's what is going on here, just that it's possible.
https://news.ycombinator.com/item?id=43205435 ("Starlink to take over $2.4B contract to overhaul air traffic control comms (theverge.com)")
Salami tactics.
right. of course, no ones going to cut corners on the safe use of AI and secure infrastructure in this administration.
Sometimes the level of incompetence / lack of care in organizations like this astounds me. I understand issues like this can be complicated and systemic but it honestly makes me think very poorly of the technologists building these systems in government.
BT guy ends that part of the presentation with "the IRA will really have to get shit together to take you off the net"
This was during the troubles. Same planet, different world.
Remove Elon Thiel Trump Vance ASAP
With low-enough engineering quality, you don't actually need saboteurs.
https://www.cnn.com/2025/02/25/business/musk-faa-starlink-co...
Press X for doubt.
[1] https://www.npr.org/2026/09/21/nx-s1-5976816/faa-ai-manage-a...
>A man stranded in the bush in northern Saskatchewan was rescued last week after chopping down four power poles — knocking out electricity to surrounding communities. [...]
>But he had an axe and he knew SaskPower would have to check the downed line, so he went to work.
one day the backbone (E1 bidirectional ring iirc) was broken north of london … my boss drove out to find that travellers had entered the field, opened the inspection cover in the ground, stood a pole for their clothes line and filled with cement
very savvy, he didn’t let on that we had about £1m revenue per day close to collapse and convinced them to hang their clothes somewhere else (perhaps a few quid changed hands)
Just sayin
In places like these, it's not the best contractor that wins, but the one that supports genocide against Palestinians.