TP-Link shipped my access point EOL out of the box. I had to write a kernel exploit to fix it.

TL;DR: My TP-Link EAP770 v1 corrupts downloads over Wi-Fi and has never received a firmware update. Disabling its forwarding offload fixes the problem, but TP-Link's shell cannot change that setting. I spent a weekend with an LLM putting together a kernel exploit so I could write two zeros into sysfs.

Long downloads kept breaking#

In 2024 I bought a shiny new Wi-Fi 7 access point, a TP-Link EAP770. There were occasional problems, but usually something would briefly stop working, reconnect, and carry on.

Rsync transfers over Wi-Fi started failing on files larger than a few gigabytes. SSH was reporting MAC errors, so I spent a while debugging rsync and SSH, then wrote a wrapper around rsync with retries and moved on. I trusted TCP to deliver the right bytes because that's what the fucking thing is for. I captured traffic with Wireshark and found zero bad checksums and zero retransmissions. Eventually I routed traffic through Tailscale inside my own home network, and that worked well enough that I left it alone.

About a year later, I found that the same 128 MiB file from the Nix binary cache on Fastly would fail over HTTP/2 and download over HTTP/1.1. git clones also failed, and HTTPS connections died with bad record mac, but the failures were random and pretty rare. I thought this was probably connected to the earlier SSH problems, but I still had no idea what was happening.

Everything worked over Ethernet, which I was using most of the time, so I knew Wi-Fi was involved. I went through the available settings, enabled Wi-Fi 7, disabled it, and suspected the MediaTek adapter in my Framework 16 enough to buy a Qualcomm replacement. That seemed a bit better, but it did not solve the problem. Turning GRO off on the client also seemed to help.

Zero firmware updates#

The device is an EAP770(EU) v1.0 with a Qualcomm IPQ9574 running Linux 5.4.164 and firmware 1.0.0 Build 20230605 Rel. 71641. That is the firmware it shipped with, and apparently the last time TP-Link remembered this thing even existed :)

I bought an early Wi-Fi 7 AP with no Wi-Fi Alliance certification, which was stupid of me. PLZ DON'T BUY STUFF UNTIL STANDARDS ARE FIXED AND CERTIFICATIONS ARE MADE.

The downloads were already breaking when I contacted support in June 2025, and the website only had firmware for v2. They could not find v1 in their database and asked for photos of the device and box before sending me this:

The EAP770 Version 1 is an End-of-Life product. While it may still be available through some retailers, it is no longer supported with firmware updates.

NICE.

Tracking down the corruption#

In October 2026, fifteen months after that exchange, I started noticing that the connection dropped around 11:00 every day and went back to debugging it with Opus 5.5.

I captured a failing HTTPS download with tcpdump and found a 1440-byte TCP segment with 64 wrong bytes in a row and a valid TCP checksum. TcpInCsumErrors stayed at zero. TCP accepted the segment and handed it to TLS, which rejected the record with bad record mac.

The damaged record would not decrypt, so I switched to raw TCP between a wired machine and the Wi-Fi laptop. The corruption only appeared with several TCP streams running at once, in the LAN-to-Wi-Fi direction. The test found three bad 64-byte blocks in 8 GiB, each starting with zeros.

Two writes that need root#

Reading the driver source suggested disabling the access point's forwarding offload, SFE/PPE, with these two settings:

echo 0 > /sys/sfe/l2_feature
echo 0 > /sys/sfe/ppe_rfs_feature

Both sysfs files are writable only by root, but TP-Link's SSH shell runs as uid 1 with no capabilities. There is no GUI setting for this either.

There is also no v1 firmware image to patch and flash. The v2 firmware is for different hardware. I tried the sibling EAP773's EAP773(EU)_V1_1.0.14, and the device rejected it. Newer v2 firmware releases are encrypted too. I could have bought a flash clip and tried programming it directly, with the risk of bricking it, but I did not have the fucking clip.

I was not going to spend another $200 replacing an AP because TP-Link could not be bothered to maintain the one they had already sold me. It is sitting in my house, I know which setting I need to change, and their software will not let me change it.

I gave Opus 4.8 guidance on the primitives to use, then had it work through the attempts until the exploit was stable. The exploit had to leave the AP running because the vendor's boot script enables the broken offload again after a reboot. The overflow also leaves the 4 KiB slab in a state where a second attempt panics, which means one attempt per boot.

The chain uses CVE-2022-0185 and the pipe-buffer repoint technique known from Dirty Pipe. CVE-2022-0185 is an integer underflow in legacy_parse_param(), reachable from an unprivileged user namespace, that lets a write overflow a 4 KiB slab object. The pipe-buffer manipulation turns that overflow into an arbitrary write. There were already public write-ups and proof-of-concept code for these; getting them to work on this particular kernel without crashing the AP took the time.

Exploit details

Without a v1 GPL source drop, we used the nearest published source, for the EAP783, whose kernel config differs from the AP's /proc/config.gz by three cosmetic lines. We built a QEMU harness with that config, got the chain working there, and then moved to the real device.

The kernel has no KASLR, so addresses stay the same between boots, but kallsyms is zeroed and we still had to find the addresses at runtime. The QEMU build loaded the kernel at physical 0x40080000, but the device loaded it at 0x42080000, putting the calculations off by 0x2000000. The correct address was in a build config variable in the EAP783 source. Even after correcting the base, the device's .text, .rodata and .data had each moved by a different amount relative to the QEMU build, so adding one offset to everything did not work.

The exploit:

  1. Unshare into a user namespace so that fsopen is reachable and the underflow can be triggered.
  2. Defragment the 4 KiB slab with a few hundred same-size allocations, then groom it to put the overflow source, a victim key, a SysV message and pipe rings next to each other.
  3. Overflow into the victim key's metadata and enlarge its length field. Reading the key back now leaks neighbouring objects, including a pipe's buf[0].ops, which contains the address of anon_pipe_buf_ops.
  4. Repoint one pipe's buffer at another pipe's ring page, then forge an entry in that second ring using the leaked ops pointer. pipe_buf_can_merge() checks that the ops pointer matches, so a merge write through the forged entry can reach an arbitrary physical page, including a page of the kernel image.
  5. Scan .data to find modprobe_path. Reading kernel-image pages through a pipe normally calls put_page() on the reserved page when the buffer is consumed, causing Bad page map and a panic. But pipe_read only calls ->confirm and ->release, and does not check that the ops table is real. A forged pipe_buf_operations with both entries pointing to generic_pipe_buf_confirm, which returns zero without doing anything else, lets the scan read the pages without releasing them. It can then walk .data in batches until it finds the string.
  6. Overwrite modprobe_path and execve a file with magic that matches no binary format. The kernel calls request_module(), which runs the replacement path as global uid 0.
  7. Keep a child alive holding the forged pipe. Letting the process exit would tear down the pipe and panic on the reserved page.

The overflow writes C strings, so the forged page address cannot contain a NUL byte. A candidate was usable about one time in five. Grooming two candidate rings and choosing the usable one at runtime was what finally made the attempt dependable.

Getting this working took about seventeen hours of token gambling with Opus 4.8, including reading the device's binaries in radare2, checking the enabled protections, and working through the crashes. With root, I could finally write both settings and test the change. The AP stayed up, and disabling the offload stopped the corruption.

A DMA cache coherency bug#

Qualcomm's qca-nss-dp driver reuses packet buffers without cleaning all the dirty CPU cache lines before handing them back to the NIC. On this hardware, DMA writes directly to DRAM without keeping the CPU caches in sync. A dirty cache line left over from the previous use of a buffer can be written back after the NIC has received a new packet, overwriting part of that packet in memory.

On x86, DMA is cache-coherent, so the same missing cache maintenance would not corrupt anything. The hardware keeps the caches in sync for you. On this chip, the driver has to do it, and it doesn't. While looking through the kernel sources, I found similar DMA cache maintenance problems in several other drivers too.

Driver details

SFE's fast transmit path cleans skb_headlen(skb) bytes from the current skb->data, covering the linear part of the outgoing packet. It then sets skb->fast_recycled = 1 and returns the buffer to a per-CPU recycle list without clearing its contents. The recycler zeroes the sk_buff struct up to tail, restores fast_recycled, and resets skb->data to skb->head + NET_SKB_PAD.

When that buffer is reused for receive, the refill path sees the flag and skips invalidation:

/* ... if the packet is fast transmitted and hence fast recycled,
 * we can be assured that invalidate was already done at the
 * time of previous transmit */
if (unlikely(!skb->fast_recycled)) { /* <--- bug: skips cache maintenance on recycled buffers */
    dmac_inv_range_no_dsb((void *)skb->data,
          (void *)(skb->data + rx_alloc_size - EDMA_RX_SKB_HEADROOM - NET_IP_ALIGN));
}
skb->fast_recycled = 0;

The receive region is 1856 bytes, or 29 cache lines, starting at head + 192. The previous transmit cleaned only two or three lines from a different offset. The flag causes receive to skip cache maintenance for the whole region based on a transmit that covered a different part of the allocation.

Dirty lines outside that cleaned range can be written back after the NIC's DMA, replacing 64 bytes of the received packet at a time. Normal cache eviction can cause the write-back. On arm64, invalidating an unaligned range can also do it: __dma_inv_area uses dc civac, clean and invalidate, on the partial edge lines. Once the dirty data has overwritten the packet in DRAM, invalidating the CPU cache cannot recover the received bytes.

The NIC checks the original packet before that overwrite, and the AP's receive path marks it CHECKSUM_UNNECESSARY using the hardware status bits. SFE/PPE then recalculates the L3/L4 checksums for the accelerated flow over the already corrupted payload, so the client gets a fresh TCP checksum that is valid for the garbage. TCP accepts it; the end-to-end MACs in TLS and SSH catch the damage, reporting bad record mac and corrupted MAC. The exact sequence inside PPE is hidden in proprietary hardware and driver code, but the packet capture shows that the checksum was calculated after the corruption.

Buffer recycling is enabled by CONFIG_SKB_RECYCLER, a compile-time option with no runtime switch. Disabling the offload stops buffers taking the fast transmit path, so the driver no longer sets fast_recycled and receive invalidates the full buffer before handing it to the NIC.

A download to a Wi-Fi client enters the AP through the wired receive path. An upload does not take that path. Traffic between two wired machines goes through the switch without ever passing through the AP.

The code is in the public openwrt/qca-nss-dp tree. The skipped invalidation is at edma_rx.c:484, the headlen-only clean is at edma_tx.c:490, and the flag is set at edma_tx.c:742. This is Qualcomm's out-of-tree NSS code.

		/*
		 * Invalidate skb->data
		 * A73 flush operation does an invalidate operation as well.
		 * If the packet is fast transmitted and hence fast recycled,
		 * we can be assured that invalidate was already done at the
		 * time of previous transmit
		 */
		if (unlikely(!skb->fast_recycled)) {
			dmac_inv_range_no_dsb((void *)skb->data,
					      (void *)(skb->data + rx_alloc_size -
					      EDMA_RX_SKB_HEADROOM -
					      NET_IP_ALIGN));

		}
		skb->fast_recycled = 0;
	buff_addr = (dma_addr_t)virt_to_phys(skb->data);
	EDMA_TXDESC_BUFFER_ADDR_SET(txd, buff_addr);

#ifdef EDMA_40BIT_SUPPORT
	EDMA_TXDESC_BUFFER_ADDR_HI_SET(txd, buff_addr);
#endif

	dmac_clean_range_no_dsb((void *)skb->data, (void *)(skb->data + buf_len));
		/*
		 * We set fast_recycled flag if packet has taken the
		 * SFE fast transmit path, so that any Rx DMA driver
		 * that allocates this packet later can avoid
		 * an invalidate operation
		 */
		if (likely(skb->fast_xmit) && likely(skb->is_from_recycler)) {
			skb->fast_recycled = 1;
		}

Before and after#

I ran the same raw TCP test between the same two machines with the offload enabled and disabled:

OffloadDirectionDataSpeedBad 64-byte blocks
onLAN -> Wi-Fi8 GiB867 Mbit/s3
offLAN -> Wi-Fi8 GiB800 Mbit/s0

Disabling the offload cost about 8% of the download speed in this test. I have to give up performance to get correct data because TP-Link will not let me properly fix my own fucking device.

With the offload on, that was one bad block per 2.7 GiB on average, and one is enough to kill a TLS connection. At full speed I could not finish a large download: a NixOS ISO failed at 739 MB on one attempt and 1.73 GB on another. Limiting the same download to 10 MB/s let it finish. Uploads from Wi-Fi to the LAN were already clean before the fix.

The NOT_SUPPORT_BRIDGE counter in /sys/sfe/exceptions increases after disabling the offload, as bridged flows take the normal path instead of being accelerated.

The vendor script /etc/rc.d/rc.ipq runs echo 1 > /sys/sfe/ppe_rfs_feature at boot, so I have a watcher on my NAS that polls the AP and reapplies the fix if the offload has been enabled again. It checks both switches afterwards. If the attempt fails, it reboots the AP through TP-Link's management API and tries on the next boot, with a hard cap on the number of reboots. The AP reboots about once a year, so most of the time the watcher has nothing to do.

I still need to send a pull request to qca-nss-dp so vendors still shipping firmware updates can pick up the fix. My own AP will keep running the workaround.

Opus 5.5 found the bug and then stopped talking about it#

When I moved from diagnosing the corruption to applying the fix, Opus 5.5 started cutting sessions off roughly every other message. Mentioning something like "local privilege escalation" is enough, and I have had a session cut while typing. I have the Cyber Verification Program enabled on my account, which has not helped. Opus 4.8 worked through the exploitation without that problem.

The funny part is that 5.5's guards also kill the session when I just ask it to read this article. The work is already done, the AP is mine, and apparently even reading about it is too much.