UK RACK 1 (Status: Post Incident Monitoring) (Investigating) Critical

Affecting System - Dell Chassis Issues + Aqua Node

  • 13/07/2026 01:00
  • Last Updated 16/07/2026 11:23

INCIDENT REPORT:

How did this happen
The QEMU VNC on the affected node (Aqua) was bound to 0.0.0.0 on legacy VMs - this was the infection vector. VMs provisioned by BoxLayer natively were correctly bound to 127.0.0.1 and were not infected.

What the attack was
Based on everything observed during the incident, the attack follows a well-documented pattern of opportunistic Linux VPS compromise for botnet expansion:

  • Primary objective: Find and infect more victims. The malware was scanning the entire internet for open VNC servers (port 5900-5999) to repeat the same attack on new targets. The VM infrastructure was being used as a scanning platform, not attacked as a target.
  • Secondary effect: The combined scanning traffic from a number of VMs simultaneously generated enough unique connection flows to overflow the Intel ixgbe NIC's Flow Director table, deadlocking the TX queue and killing external connectivity. The traffic seen from VMs services was a mass port scanning on systems not expecting to burst all instances with a flood of traffic.

Was any data leaked
Based on all information we have, with the traffic (TX/RX logging) sending out using ports consistent with port scanning for further attack vectors we do not see any data leakage from the node or customers VMs. All traffic was SYN-only packets (no data payload) targeting ports 5900-5999, and no sustained TCP sessions completing (no SYN-ACK-ACK handshakes).

The attackers components were configured to:

  • Scan for open VNC ports
  • Attempt to spread to new victims
  • Maintain persistence via systemd

As always, with any attack, of any sort, we always recommend changing your VM password and sanity check your service. This is also a good time to remind customers to always keep a copy of your critical data on your local computer or cloud storage devices in the case of a major incident.

What is the fix and how is it being prevented
The fix was implemented as soon as we found the VNC security gap, we configured the impacted VMs VNC configuration to be correct and secured them. Also, as hinted above, BoxLayer was built to be as secure as possible - it is how we designed it, we are extending our security auditing tools within BoxLayer to be able to detect such security gaps, report them and plug them quickly, this will also include a range of customer facing and internal facing advisory information.

So what is next
With these sorts of incidents that thankfully are rare, we treat extremely seriously, so we will continue to monitor all traffic, the node and chassis itself, to identify any issues or remaining gaps that need resolving. At the present time, we have found no other gaps within the security of our network and services have remained very stable (CPU/RAM/Network) since we resolved the issues. We are also working with remaining customers that were impacted to bring up backups or sanity check their services by their request.


AQUA NODE UPDATE: The node has been working normally for the past 24 hours and no inbound or outbound attacks has been seen, we continue to resolve issues on the small number of VPSs that continue to have issues.


AQUA NODE UPDATE: VPS customers that are impacted by the insecure OS's installed on their VPS's are being contacted by our technical teams to confirm permission to restore from available backups. Once restored our team will work with customers to secure their VPS's. Please note, it is down to customers to keep their VPS instances updated and secured but in this case we are happy to action the updates.


Summary of previous updates:

Our team detected issues within one of our Dell Chassis located in RACK 1 of our UK Coventry data centre, this caused a flood of unusual traffic similar to a DoS which was overloading the switches. We were able to make the flow of network activity stable to find the root node causing the issue was the Aqua KVM Node. It took some time to narrow down where and why this was happening due to the unusual signs the logs and data was giving us. It caused a up/down services, constantly through a number of days with periods of uptime ~9 hours but then randomly failed again.


NOTE: Due to a formatting issue in the server status form, previous line updates have been cleared. We will backfill where possible and summarise the incident.