Overview

I put off learning eBPF for two years because every introduction started with "it's a virtual machine in the kernel" and then got into verifier details. The mental model that finally made it click: eBPF lets you attach small programs to kernel events and run them safely, without writing a kernel module or patching the kernel.

That's it. What makes it interesting is that the events are everywhere — every syscall, every network packet, every function call in the kernel — and the programs run at native speed.

Why this matters in practice

Before eBPFWith eBPF
Kernel modules for tracingLoad programs at runtime, no reboot
Rebuild and reinstall on kernel updateIndependent of kernel version (mostly)
A crash takes down the machineVerifier prevents unsafe programs from loading
iptables rules for firewallingXDP processing at the driver level
Sampling profilers with 10ms resolutionPer-event observation at nanosecond scale

The most visible consumer is Cilium, which replaced kube-proxy in many Kubernetes clusters. The most visible developer tool is probably bpftrace, which turns one-liners into kernel tracing programs.

The architecture in one paragraph

You write a program in a restricted C dialect. The compiler (clang) produces eBPF bytecode. You submit the bytecode to the kernel via the bpf() syscall. The kernel verifier walks the program, rejects anything unsafe (unbounded loops, out-of-bounds memory access, uninitialized reads), and if it passes, JIT-compiles it to native instructions. The program attaches to a hook — a kprobe on a function, a tracepoint, an XDP hook on a network interface — and runs whenever the hook fires.

The verifier is the piece that makes this safe enough to use in production. It's also the piece that makes eBPF programming frustrating, because it rejects a lot of reasonable-looking code. More on that later.

Your first bpftrace one-liner

bpftrace is the fastest way to see what eBPF can do. Install it (apt install bpftrace), then:

# Count syscalls by process name
sudo bpftrace -e 'tracepoint:syscalls:sys_enter_* { @[comm] = count(); }'

Run it for ten seconds, Ctrl+C, and you get a table of every process on the system ranked by syscall count. No agents, no restarts, no configuration.

# Show open() calls over 4KB in real time
sudo bpftrace -e '
tracepoint:syscalls:sys_enter_openat {
  printf("%s opened %s\n", comm, str(args->filename));
}'

Things that used to require strace on a specific PID now work system-wide, with no measurable overhead on the traced processes. This is what sells people on eBPF.

Common tracing scenarios

Questionbpftrace one-liner
Which files does process X open?bpftrace -e 'tracepoint:syscalls:sys_enter_openat /pid == PID/ { printf("%s\n", str(args->filename)); }'
What's the slowest syscall?bpftrace -e 'tracepoint:raw_syscalls:sys_enter { @start[tid] = nsecs; } tracepoint:raw_syscalls:sys_exit /@start[tid]/ { @time[comm] = sum(nsecs - @start[tid]); delete(@start[tid]); }'
Which DNS queries are slow?Attach to udp_sendmsg, filter by port 53
TCP retransmits by destinationtracepoint:tcp:tcp_retransmit_skb { @[args->saddr, args->daddr] = count(); }

I've used the DNS one to find a resolver that took 3 seconds on first lookup, and the TCP retransmit one to confirm that a client's "network is slow" complaint was actually packet loss on their end. Both took about two minutes to write.

Writing a real eBPF program in C

bpftrace is for exploration. For anything you're going to run repeatedly, you write a proper eBPF program in C and load it with a library like libbpf or the Go library cilium/ebpf.

// trace_open.c
#include <linux/bpf.h>
#include <bpf/bpf_helpers.h>

struct event {
    __u32 pid;
    __u32 uid;
    char comm[16];
    char filename[256];
};

struct {
    __uint(type, BPF_MAP_TYPE_RINGBUF);
    __uint(max_entries, 256 * 1024);
} events SEC(".maps");

SEC("tracepoint/syscalls/sys_enter_openat")
int trace_openat(struct trace_event_raw_sys_enter *ctx) {
    struct event *e;
    __u64 pid_tgid = bpf_get_current_pid_tgid();

    e = bpf_ringbuf_reserve(&events, sizeof(*e), 0);
    if (!e) return 0;

    e->pid = pid_tgid >> 32;
    e->uid = bpf_get_current_uid_gid();
    bpf_get_current_comm(&e->comm, sizeof(e->comm));
    bpf_probe_read_user_str(&e->filename, sizeof(e->filename),
                             (void *)ctx->args[1]);

    bpf_ringbuf_submit(e, 0);
    return 0;
}

char LICENSE[] SEC("license") = "GPL";

A few things worth pointing out:

  • The ring buffer is the modern way to send data from kernel to userspace. It's faster than the older perf buffer and has fewer ordering issues.
  • bpf_probe_read_user_str is required because the filename pointer points to userspace memory. Accessing it directly would crash the program, and the verifier won't let you try.
  • The GPL license declaration is required for programs using GPL-only helpers. Without it, the kernel refuses to load the program.

The verifier is strict

Every eBPF developer hits this in the first week. The verifier rejects programs that could, in theory, do something unsafe. Common rejections:

ErrorCauseFix
invalid mem accessReading a pointer the verifier can't prove is validUse a probe helper or bounds check
back-edge from insn X to YUnbounded loop (older kernels)Bounded loops with #pragma unroll or kernel 5.3+
R1 type=map_value expected=fpPassing a map value as a stack pointerCopy to a local variable first
program too largeOver 1M instructions (older kernels)Split into multiple programs
cannot call GPL only functionMissing license declarationAdd char LICENSE[] SEC("license") = "GPL";

The verifier's error messages have gotten much better in recent kernels. It now shows you the exact instruction that failed and the state of each register, which is what makes debugging tolerable.

If you're starting out, use a CO-RE (Compile Once, Run Everywhere) approach with libbpf. It lets you write one program that works across kernel versions by using BTF type information instead of hardcoded offsets.

XDP: the networking use case

XDP (eXpress Data Path) is eBPF attached at the network driver level, before the kernel allocates an sk_buff. It runs on every incoming packet, and it's fast enough for DDoS mitigation at 10Gbps per core.

SEC("xdp")
int xdp_drop_udp(struct xdp_md *ctx) {
    void *data = (void *)(long)ctx->data;
    void *data_end = (void *)(long)ctx->data_end;
    struct ethhdr *eth = data;

    if ((void *)(eth + 1) > data_end) return XDP_PASS;
    if (eth->h_proto != bpf_htons(ETH_P_IP)) return XDP_PASS;

    struct iphdr *ip = (void *)(eth + 1);
    if ((void *)(ip + 1) > data_end) return XDP_PASS;
    if (ip->protocol == IPPROTO_UDP && ip->dport == bpf_htons(1234)) {
        return XDP_DROP;
    }
    return XDP_PASS;
}

Return XDP_DROP and the packet is gone before the kernel networking stack sees it. This is the same technique Cloudflare uses to handle volumetric attacks.

Attaching:

sudo ip link set dev eth0 xdp obj xdp_drop.o sec xdp

# Check
ip link show eth0 | grep xdp

# Detach
sudo ip link set dev eth0 xdp off

Tools built on eBPF

ToolPurpose
CiliumKubernetes networking, replacing kube-proxy
FalcoRuntime security — alerts on suspicious syscalls
PixieAuto-instrumented observability for k8s
ParcaContinuous profiling without sampling gaps
bpftraceOne-liners for kernel tracing
bccOlder toolkit, still useful for specific tools

I've run Falco in production and it caught a container attempting to write to /etc/passwd — something no userspace monitoring would have seen in real time without enormous overhead.

When eBPF is the wrong tool

  • You need to trace userspace only. dtrace or perf handle that without kernel-level complexity.
  • You're on an old kernel. Anything before 4.9 has limited support; CO-RE needs 5.2+ for BTF. RHEL 7 and older are painful.
  • You need portability across architectures. eBPF works on x86_64, arm64, and a few others, but not everything.
  • Your use case is already covered. Don't write an XDP load balancer when nginx exists; don't write syscall tracing when Falco has 200 community rules.

The learning curve is real, but the entry point with bpftrace is shallow. Ten minutes of one-liners will teach you more about what's happening on your system than a week of reading documentation.