Introduction
In 2010, a company called Spread Networks spent about 300 million dollars digging a straighter fiber optic cable between Chicago and New York. The goal was to shave 3 milliseconds off the round trip. High-frequency trading (HFT) firms happily paid for access, because in their world, whoever reacts to a price change first wins the trade. Everyone else gets nothing. (The story opens Michael Lewis’ book Flash Boys, a good read if you want the finance side of this article.)
Since then, the race has moved from milliseconds to microseconds, and sometimes nanoseconds. At that scale, an unexpected bottleneck appears: the operating system itself. The Linux network stack, a marvel of engineering that powers most of the Internet, costs tens of microseconds per packet. For a web server, this is negligible. For a trading system, it is an eternity.
This article explains where those microseconds actually go, and how HFT systems get rid of them with a technique called kernel bypass. We will not talk about finance beyond this introduction: trading is only our motivation, and everything below is system software.
By the end of this article, you will:
- understand the journey of a packet through the Linux kernel, and where time is lost along the way;
- understand the three key ideas behind kernel bypass: polling, userspace packet access, and zero-copy;
- read through a minimal
AF_XDPreceiver and see, line by line, what it removes compared to a standard UDP socket.
Prerequisites: you should be comfortable with C, basic Linux usage, and know what a socket is. No prior knowledge of kernel internals, network drivers or finance is required.
Life of a packet in the Linux kernel
To understand what kernel bypass removes, we first need to see what the normal path looks like. Let’s follow a UDP packet from the wire to your application.
|
|
Figure 1: the standard receive path. A packet goes from the wire to the NIC, is copied to memory by DMA, triggers a hardware interrupt, is processed by the kernel stack (softirq, sk_buff, protocol layers), lands in the socket queue, and finally reaches the application through a syscall and a copy — with two context switches along the way.
Step 1: the NIC receives the frame. The network interface card (NIC) checks the frame and copies it into a receive ring buffer in main memory, using DMA (Direct Memory Access). No CPU involved so far: this part is fast.
Step 2: the interrupt. The NIC raises a hardware interrupt (IRQ) to tell the CPU that data has arrived. An interrupt is the hardware’s way of demanding attention: the CPU stops whatever it was doing, saves its state, and jumps to a handler function registered by the driver. (For the rest of this article, “IRQ” stands for such a hardware interrupt.) This detour is not free: classic OS benchmarks such as lmbench measure context switch costs in the microsecond range, and the interrupt also pollutes the CPU caches: whatever your application had loaded is partially evicted.
Step 3: softirq processing and the sk_buff.
Handlers must stay short, so modern kernels defer most of the work to a
softirq, a mechanism that runs deferred interrupt work outside the
hardware interrupt handler (this is where NAPI, the kernel’s polling-based
packet processing framework, runs).
There, the kernel allocates a sk_buff, the data structure that represents
a packet in the kernel, and starts pushing it through the protocol stack:
Ethernet, then IP (routing decision, netfilter firewall hooks, checksums),
then UDP.
Each layer means more function calls, more checks, and more cache misses.
Step 4: the socket queue.
The payload finally lands in the receive queue of the destination socket.
If your application was blocked in recvfrom(), the kernel wakes it up.
Waking up a process means a scheduler decision and yet another context
switch.
Step 5: the syscall and the copy.
Your recvfrom() call is a system call: the CPU switches from user mode to
kernel mode and back.
On the way, the packet payload is copied from kernel memory into your
application’s buffer.
That is the second copy of the data since it left the wire.
Add all of this up and a single UDP receive costs on the order of tens of microseconds between the wire and your code, depending on hardware and kernel version. The exact number matters less than the structure of the cost:
- interrupts interrupt (!) your application and trash its caches;
- context switches between user and kernel mode cost time on every syscall;
- copies burn memory bandwidth;
- generality has a price: the kernel checks things your application may not care about (firewall rules, routing, multiple protocols), because it must work for every use case.
None of these steps is a bug. They are the price of a stack that is safe, general and shared between all processes. HFT systems simply cannot afford that price.
The bypass approach
Kernel bypass attacks each cost we just identified. The idea is radical: take the kernel out of the data path entirely, and let the application talk to the NIC directly.
|
|
Figure 2: standard path versus bypass path. On the left, the application sits on top of the kernel stack and reaches the NIC through syscalls, interrupts and copies. On the right, the application busy-polls a memory area shared directly with the NIC, and the kernel is out of the data path.
Polling instead of interrupts. Rather than being notified when a packet arrives, the application runs an infinite loop asking “is there a packet? is there a packet?”. This sounds wasteful, and it is: one CPU core runs at 100% forever. But it removes interrupts, context switches, and scheduler wakeups from the critical path. When a packet arrives, the code that handles it is already running, with warm caches. Latency becomes not only lower but, crucially for trading, predictable.
The NIC mapped into userspace. Frameworks like DPDK (Data Plane Development Kit) ship userspace drivers, called Poll Mode Drivers. The NIC’s ring buffers are mapped directly into the application’s address space. Received frames appear in memory that the application can read without any syscall.
Zero-copy. Since the application reads the NIC’s buffers directly, the packet is not copied from kernel space to user space. The DMA transfer from the NIC is the only copy that ever happens.
What you give up. There is no kernel in the path anymore, so there is also no TCP/IP stack. DPDK hands you raw Ethernet frames; parsing IP and UDP is now your problem (or your library’s). This is one reason bypass code is harder to write and maintain than socket code.
Note
Why the hands-on below uses AF_XDP and not DPDK.
DPDK is the reference tool in the industry, but it requires dedicating a
NIC to a userspace driver and setting up hugepages, which is impractical
on a student laptop or a VM.
AF_XDP is a socket family built into Linux (since 4.18) that achieves
the same three ideas: polling, direct access to the NIC’s frames through
a shared memory area, and zero-copy on supported drivers.
Purists will point out that AF_XDP is technically a fast path through
the kernel rather than a full bypass: normal kernel drivers are still
used, and an eBPF program (a small verified program that userspace can
load into the kernel) decides which packets go to the fast path.
The mental model, and the performance behavior, are close enough for our
purposes.
Hands-on: a tale of two receivers
Time to make this concrete.
We will look at two small UDP receivers, one with a standard socket and one
with AF_XDP, and compare what each iteration of their receive loop
actually costs.
If you want to build and run the AF_XDP version yourself, the
xdp-project tutorial provides a complete working setup that the
code below is modeled on.
The baseline: a standard UDP socket. Nothing exotic here.
|
|
(Error handling for socket() and bind() is elided for brevity; the
check on recvfrom() matters here because the loop must survive
interrupted or failed reads, as documented in man 2 recvfrom.)
Every iteration of this loop is a syscall, a wakeup, and a copy, as seen in the previous section.
The AF_XDP receiver.
An AF_XDP socket (called an XSK) works differently.
The application allocates a chunk of memory called the UMEM, splits it into
frames, and shares it with the kernel.
Four rings (fill, completion, RX, TX) are used to pass frame descriptors
back and forth, but the packet data itself never moves: it is written once
into the UMEM by the driver, and read in place by the application.
Setting up the UMEM and the rings takes some boilerplate (about a hundred
lines with libxdp; the tutorial linked above walks through all of it),
but the receive loop itself is short:
|
|
Note the comment: handle_packet() receives a raw Ethernet frame.
Our code has to skip the Ethernet and IP headers itself to find the UDP
payload.
The kernel used to do that for us.
What each loop iteration costs. Put side by side, the two receivers differ on every axis identified in the “life of a packet” section:
| Per received packet | UDP socket | AF_XDP receiver |
|---|---|---|
| Syscalls | 1 (recvfrom) |
0 on the happy path |
| Copies of the payload | 2 (DMA + to user) | 1 (DMA only) |
| Interrupt + wakeup | yes | no (busy-polling) |
| Protocol parsing | done by the kernel | done by you |
| CPU usage while idle | ~0% | 100% of one core |
The last two rows are the price tag, and they lead directly to the next section.
Tip
If you benchmark this yourself, pin the receiver to an isolated core
(isolcpus kernel parameter), set the CPU frequency governor to
performance, and report the median and the 99th percentile rather than
the average: for a trading system, the tail is what hurts, since an
occasional slow packet is a missed trade.
Without those precautions, scheduler noise and frequency scaling will be
larger than the effect you are measuring.
Published results consistently show the same shape.
Cloudflare measured the kernel path’s limits in
“How to receive a million packets per second”: a naive UDP
receiver caps out around 200-370 kpps on one core, and reaching 1 Mpps
already requires a multi-queue NIC, SO_REUSEPORT across several
processes and careful CPU pinning — their best case, with queues and
threads perfectly aligned on a single NUMA node, was 1.4 Mpps.
Those limits are why they turned to bypass techniques for
DDoS mitigation, and why exchanges’ microbursts overwhelm standard
stacks.
On the bypass side, the DPDK performance reports published
for each release routinely show tens of millions of packets per second on
a single core, which puts the per-packet budget at well under a
microsecond.
More important than the median gain: the variance collapses, because
nothing interrupts the polling core anymore.
Predictability is the real product.
Tradeoffs: when bypass is not worth it
If kernel bypass were free, everyone would use it. Here is what it actually costs.
A core burning at 100%. The polling loop consumes an entire CPU core doing mostly nothing. On a trading server with dozens of cores, dedicating a few to networking is a bargain. On a shared machine, a laptop, or a cloud instance billed for CPU time, it is a terrible deal.
You lose the kernel’s tooling.
Packets on the bypass path are invisible to the tools you rely on daily:
tcpdump sees nothing, iptables/nftables rules do not apply, standard
statistics in /proc/net stop counting.
Debugging a production issue at 3 a.m. without tcpdump builds character.
(AF_XDP is gentler here than DPDK, precisely because it stays inside the
kernel’s driver model.)
You inherit the kernel’s job. No stack means no TCP, no fragmentation handling, no ARP, no ICMP. Trading systems either use UDP-based protocols, buy a userspace TCP stack, or implement the minimum they need. That is a lot of subtle, security-sensitive code the kernel used to provide for free, hardened by thirty years of production use.
Operational complexity. Bypass setups pin threads to isolated cores, disable CPU frequency scaling and deep sleep states, and tune IRQ affinity for everything else. The configuration becomes part of your correctness.
The rule of thumb: if your latency budget is measured in milliseconds, keep the kernel; it is doing more for you than you think. Kernel bypass makes sense when microseconds are the product, as in trading, or when packet rates overwhelm the kernel entirely, as in DDoS mitigation at companies like Cloudflare.
Conclusion
The Linux network stack pays for its generality with interrupts, context
switches and copies, adding tens of microseconds to every packet.
High-frequency trading systems remove that cost by taking the kernel out of
the data path: they poll instead of sleeping, read the NIC’s memory
directly, and never copy a packet.
And the idea is no longer reserved to trading firms with exotic hardware:
with AF_XDP, the same three principles are one socket call away on any
recent Linux kernel.
But the kernel path is not slow because it is badly written; it is slow because it is general, safe and shared. Bypass trades those qualities for speed. Like most interesting decisions in system programming, it is not about finding the best solution, but about knowing precisely what you are giving up.
Going further
- Carl Cook, “When a Microsecond Is an Eternity” (CppCon 2017): what the rest of an HFT system looks like, beyond networking.
- Cloudflare, “How to receive a million packets per second”: the measurements behind the kernel-path ceiling cited in this article.
- Cloudflare, “Why we use the Linux kernel’s TCP stack”: a thoughtful defense of not bypassing, from people who partially do.
- The xdp-project tutorial: the best hands-on resource to go deeper into XDP and AF_XDP.
- AF_XDP kernel documentation: the reference for UMEM and ring semantics.
- DPDK documentation: the industry-standard framework, if you want the full bypass experience.