In the Linux kernel, the following vulnerability has been resolved:
bpf: tcp: Fix use-after-free in bpfitertcpestablishedbatch()
reqskqueuehashreq() publishes a TCPNEWSYNRECV requestsock onto the ehash chain, drops the bucket lock, and only afterwards sets rskrefcnt to 3.
Lockless readers such as __inetlookupestablished() handle this with refcountincnotzero(), but bpfitertcpestablishedbatch() uses plain sockhold() while holding the bucket lock, on the assumption that the lock guarantees skrefcnt > 0. That assumption does not hold for requestsock:
CPU 0 CPU 1 ----- ----- tcpconnrequest() reqskqueuehashreq() inetehashinsert(req) spinlock(bucket) _sknullsaddnodercu(req) // rskrefcnt == 0 spinunlock(bucket) bpfitertcpestablishedbatch() spinlock(bucket) sockhold(req) <-- addition on 0 spinunlock(bucket) refcountset(&req->rskrefcnt, 3) // clobbers saturated value
which surfaces as:
refcountt: addition on 0; use-after-free. WARNING: lib/refcount.c:25 at refcountwarnsaturate+0x48/0x90, CPU#1 Call Trace: bpfitertcpestablishedbatch+0x14e/0x170 bpfitertcpbatch+0x53/0x200 bpfitertcpseqnext+0x27/0x70 bpfseqread+0x107/0x410 vfs_read+0xb9/0x380
The iterator's stolen reference is lost when the publishing CPU's refcount_set() overwrites the count, leaving the socket one reference short. When the last legitimate owner drops its reference the reqsk is freed while still reachable, leading to use-after-free.
This reproduces in seconds with tcp_syncookies=0, a handful of threads doing connect()/close() to a local listener while others read an iter/tcp link in a tight loop.
Use refcountincnot_zero() and skip the socket on failure. A skipped socket is still part of the bucket, so keep counting it in expected. The reallocations are sized from expected, and a request sock whose refcount gets published while the lock is held across the last realloc must already have room.
A skipped socket is counted in expected but never batched, so endsk can be short of expected on a batch that is actually complete. Decide completeness by whether the walk left any socket behind instead. The WARN after the locked realloc checks the same, replacing an endsk == expected check that could not hold on that path since commit cdec67a489d4 ("bpf: tcp: Make sure iter->batch always contains a full bucket snapshot").
If every matching socket in a bucket is mid-init (refcount 0), end_sk stays 0. Advance to the next bucket rather than returning a batch entry that was never filled this round.