ITnews

Telstra sowed seeds for its July mobile outage in 2020. Six years on they sprouted.


When Telstra engineers embarked on a simple server chassis replacement exercise in 2020, they could barely have contemplated the complex chain of missteps and complacency that would, six years later, lead to a national outage.



A technical audit of the July 8 outage which disrupted triple zero calls found that the server timing apparatus critical to Telstra’s mobile network was, at the time, reliable and “fit-for-purpose”.

However, after corner-cutting, skimping on costs, critical software maintenance lapses, documentation failures and missed red flags, by mid-2026 it was drifting from that.

It was morphing into the ticking time bomb that eventually took out its entire mobile network, could have cost lives, caused thousands-of-dollars in compensation claims and impacted public transport operators that thought the carrier would never let them down.

In addition, it now leaves Telstra facing potentially eyewatering penalties measured in the tens of millions of dollars.

Media were given just over an hour to prepare questions for chief executive Vicki Brady ahead of a teleconferenced press scrum the carrier held alongside the release of the findings of the technical autopsy of the outage prepared by Technology Audit Partners (TAP).

The TAP investigation

During the call, Brady admitted that the carrier had not recognised how critical its network timing apparatus was for its mobile network operations.

“What the investigation makes clear is why that was able to happen. At its core, this came down to the network timing system not being given the priority that a critical network capability requires. Modern networks are complex, but complexity is not an excuse,” Brady said.

For technical audiences the TAP report reads like a slowly unfolding disaster film across a timeline replete with incidents of haphazard decision-making, shortcuts and that other popular cinematic trope, alarms that baffle rather than demystify.

However, in this saga, rather than being ignored, the boffins are cast as well-meaning but nevertheless unwitting accomplices to the villainy.

And if, as any good film director will tell you, timing is important to any well-crafted popcorn-muncher, then it was doubly so in this story as it swung between engineering horror and comedy romp.

Timing and synchronisation are critical to any network. They keep data moving at the right time and in the right order, they are crucial for security mechanisms to work correctly and they’re critical for real-time applications.

That’s why network operators have evolved a robust network timing protocol (NTP) anchored to reference clocks in a cascading hierarchy of strata or “stratums”, with the highest reference stratum being zero.

Stratum zero clocks are usually linked to high precision atomic clocks, many of which are carried on satellites that can be referenced by global positioning system (GPS) hardware that can be installed in servers.

The seeds

According to the TAP report, in 2010 Telstra’s hierarchy was in good shape. It had two stratum 2 NTP servers in Melbourne and Sydney referencing stratum 1 sources provided by the National Measurement Institute (NMI).

The pair provided reliable times to three lower tier stratum 3 NTP servers in Perth, Sydney and Melbourne.

However, in 2020 Telstra started to make critical changes to its core timing configuration in order to upgrade an ageing model chassis used to house its network timing hardware. 

The upgrade introduced a replacement chassis with an operational limitation. Unlike their predecessors, the new chassis wouldn’t allow coexistence of stratum 2 and stratum 3 servers in a scenario where the former could provide time references for the latter.

Or as the TAP report described it, “the stratum 2 NTP server could not provide time to a stratum 3 NTP server in the same chassis”.

The TAP report strongly suggests that the server chassis were replaced in both its Sydney and Melbourne stratum 2 NTP servers.

The report found the Melbourne NTP stratum 3 server was thereon relying on the external  stratum 2 NTP server in Sydney NTP and vice-versa – the stratum 3 server in the Sydney chassis relied on the Melbourne stratum 2 time server.

Only the Perth stratum 3 server “in a different chassis, at a different location” continued to have stratum 2 redundancy, the TAP report stated.

Telstra was unable to confirm that the chassis were replaced at both the Melbourne and Sydney NTP server locations.

Nevertheless, after the chassis upgrade Telstra engineers introduced a change to the NTP configuration that would be one of the most decisive during the outage six years later.

Where once the servers had been in a client-server configuration, meaning they were constrained to seeking time from the twinned stratum 2 servers, the engineers switched it to a “peering” configuration.

The TAP report stated that this meant that that the server in the NTP timing chassis could acquire time from “any NTP source that it could find and select to use any of these sources that was providing the most stable and accurate time”.

“This introduced uncertainty in the time sources being used in mobile core timing hierarchy,” the TAP report stated.

“We have no evidence that the personnel involved with the mobile core timing architecture changes which were made in 2020 were aware of or investigated this risk in the architecture,” TAP added in its report.

TAP’s investigation also found that the new NTP timing chassis brought with it the option of GPS for a backup time source but that “it was never configured for automatic failover”.

Missed signs

Clear signs of trouble soon followed. The TAP report found that Telstra’s Melbourne stratum 3 time server started to exhibit “looping” behaviours during periods when it couldn’t reach its Sydney stratum 2 server for a reliable time.

The published version of the report did not explain the looping behaviour in detail other than to say that when the Melbourne NTP server lost connectivity with Sydney, it was able to become a client of servers lower in the ranking hierarchy.

TAP found that the Telstra “personnel involved did not recognise the meaning of Melbourne NTP at stratum 5” but didn’t investigate deeper.

“This anomaly was observed but not noted and no problem ticket was created for further investigation,” TAP stated in its report.

Alarms

By October 2025 NTP alarms were going off informing Telstra engineers working on the Melbourne stratum 3 server that it was losing connectivity with its only stratum 2 time source in Sydney.

To work around the problem, they fitted a GPS card to the server, letting it calculate time independently via satellites rather than another source on the network.

The move had the effect of promoting it to be the network’s most senior trusted time source, the equivalent of a ground truth that every other, less senior server would now defer to and copy.

Though now in an arguably fragile state, the configuration might have worked for a time if not for the final fumble in the unfolding folly – one which Telstra had the opportunity to correct years, if not decades, earlier.

The GPS card’s firmware hadn’t been updated to correct for a well-known flaw in older equipment: the internal counter GPS satellites use to track the date can only count up to 1024 weeks, just under 19.7 years, before it runs out and starts again from zero.

An up-to-date software version would have corrected for that reset automatically and the card’s manufacturer – originally Symmetricom which was later acquired by Microchip Technology – had been issuing warnings about it since the year 2000.

The most recent warning that Telstra could have heeded was issued in January 2025.

The card was not supposed to be in active use in the first place, so nobody at Telstra had prioritised applying the fix.

July 8

The problem slept until the wee hours of the morning of July 8 when a technician moved to replace the physical chassis housing the Melbourne NTP server to fix a power supply fault.

When the replacement unit powered up, the GPS card reset, hitting the 1024-week limit and calculated the date as November 2006.

As it was now the most senior time source in the hierarchy, other servers started following its lead taking on the incorrect date. The problem spread across Telstra’s mobile network as the error radiated outward to other NTP servers, causing disconnections and call failures, including to triple zero emergency services.

TAP found on the night of the outage two “primary” NTP engineers involved in the chassis power maintenance exercise had been mandatorily stood down.

The TAP report didn’t go into detail as to why the Melbourne NTP server originally started losing connectivity with the Sydney server, it had been relying on and necessitating the GPS card. However, the consultancy said it should have been investigated further.

“When Melbourne NTP was connected to the GPS card a ‘Network at Risk’ ticket should have been raised to find the real root cause of the frequent loss of the Sydney source and restore the NTP server back [to] the intended design but this did not happen,” its report stated.

‘Budgetary prioritisation’

TAP cautiously backed a position put forward by Telstra engineers that budgetary constraints impacted their decisions.

Citing this in connection with a missed opportunity on Telstra’s part to recognise NTP as a “Sovereign Function”, TAP said that budgetary priorities were evident in its analysis of the outage.

“While these reports are not based on documentation and evidence, it is our opinion that the circumstances we found in this outage analysis carry signs of these reports based on what has happened,” it stated.

TAP also found that Telstra’s networking division was short on expertise and resources to take “full ownership” of the NTP capability and recommended a review of staffing in the engineering division.

Not a broken design but more resources thrown at engineers, alarms improved and external experts sought

Brady, who took a personal $600,000 hit to her salary due to the outage, acknowledged the telco’s failures throughout the timing network’s design change but rejected characterisations of it as broken.

“I would not characterise [the design change] as broken. It was a change in design. It was a different design to what we had had in place from 2010. We made that change in 2020.

“What’s clear is once that design change was made, it was not well documented, and as a result, through that period, as architecture evolved, as changes happened, we didn’t have that end-to-end ownership and oversight of our timing system approach inside the network,” Brady said.

Brady said that the carrier had introduced additional alarms and measures to monitors its network time servers, and tightened up its lab testing regime.

She also confirmed that the carrier was giving its engineering teams additional resources, including input from external experts.

Additional reporting by Juha Saarinen



Source link