Overview
I put off learning eBPF for two years because every introduction started with "it's a virtual machine in the kernel" and then got into verifier details. The mental model that finally made it click: eBPF lets you attach small programs to kernel events and run them safely, without writing a kernel module or patching the kernel.
That's it. What makes it interesting is that the events are everywhere — every syscall, every network packet, every function call in the kernel — and the programs run at native speed.
Why this matters in practice
| Before eBPF | With eBPF |
|---|---|
| Kernel modules for tracing | Load programs at runtime, no reboot |
| Rebuild and reinstall on kernel update | Independent of kernel version (mostly) |
| A crash takes down the machine | Verifier prevents unsafe programs from loading |
| iptables rules for firewalling | XDP processing at the driver level |
| Sampling profilers with 10ms resolution | Per-event observation at nanosecond scale |
The most visible consumer is Cilium, which replaced kube-proxy in many Kubernetes clusters. The most visible developer tool is probably bpftrace, which turns one-liners into kernel tracing programs.
The architecture in one paragraph
You write a program in a restricted C dialect. The compiler (clang) produces eBPF bytecode. You submit the bytecode to the kernel via the bpf() syscall. The kernel verifier walks the program, rejects anything unsafe (unbounded loops, out-of-bounds memory access, uninitialized reads), and if it passes, JIT-compiles it to native instructions. The program attaches to a hook — a kprobe on a function, a tracepoint, an XDP hook on a network interface — and runs whenever the hook fires.
The verifier is the piece that makes this safe enough to use in production. It's also the piece that makes eBPF programming frustrating, because it rejects a lot of reasonable-looking code. More on that later.
Your first bpftrace one-liner
bpftrace is the fastest way to see what eBPF can do. Install it (apt install bpftrace), then:
# Count syscalls by process name
sudo bpftrace -e 'tracepoint:syscalls:sys_enter_* { @[comm] = count(); }'
Run it for ten seconds, Ctrl+C, and you get a table of every process on the system ranked by syscall count. No agents, no restarts, no configuration.
# Show open() calls over 4KB in real time
sudo bpftrace -e '
tracepoint:syscalls:sys_enter_openat {
printf("%s opened %s\n", comm, str(args->filename));
}'
Things that used to require strace on a specific PID now work system-wide, with no measurable overhead on the traced processes. This is what sells people on eBPF.
Common tracing scenarios
| Question | bpftrace one-liner |
|---|---|
| Which files does process X open? | bpftrace -e 'tracepoint:syscalls:sys_enter_openat /pid == PID/ { printf("%s\n", str(args->filename)); }' |
| What's the slowest syscall? | bpftrace -e 'tracepoint:raw_syscalls:sys_enter { @start[tid] = nsecs; } tracepoint:raw_syscalls:sys_exit /@start[tid]/ { @time[comm] = sum(nsecs - @start[tid]); delete(@start[tid]); }' |
| Which DNS queries are slow? | Attach to udp_sendmsg, filter by port 53 |
| TCP retransmits by destination | tracepoint:tcp:tcp_retransmit_skb { @[args->saddr, args->daddr] = count(); } |
I've used the DNS one to find a resolver that took 3 seconds on first lookup, and the TCP retransmit one to confirm that a client's "network is slow" complaint was actually packet loss on their end. Both took about two minutes to write.
Writing a real eBPF program in C
bpftrace is for exploration. For anything you're going to run repeatedly, you write a proper eBPF program in C and load it with a library like libbpf or the Go library cilium/ebpf.
// trace_open.c
#include <linux/bpf.h>
#include <bpf/bpf_helpers.h>
struct event {
__u32 pid;
__u32 uid;
char comm[16];
char filename[256];
};
struct {
__uint(type, BPF_MAP_TYPE_RINGBUF);
__uint(max_entries, 256 * 1024);
} events SEC(".maps");
SEC("tracepoint/syscalls/sys_enter_openat")
int trace_openat(struct trace_event_raw_sys_enter *ctx) {
struct event *e;
__u64 pid_tgid = bpf_get_current_pid_tgid();
e = bpf_ringbuf_reserve(&events, sizeof(*e), 0);
if (!e) return 0;
e->pid = pid_tgid >> 32;
e->uid = bpf_get_current_uid_gid();
bpf_get_current_comm(&e->comm, sizeof(e->comm));
bpf_probe_read_user_str(&e->filename, sizeof(e->filename),
(void *)ctx->args[1]);
bpf_ringbuf_submit(e, 0);
return 0;
}
char LICENSE[] SEC("license") = "GPL";
A few things worth pointing out:
- The ring buffer is the modern way to send data from kernel to userspace. It's faster than the older perf buffer and has fewer ordering issues.
bpf_probe_read_user_stris required because the filename pointer points to userspace memory. Accessing it directly would crash the program, and the verifier won't let you try.- The GPL license declaration is required for programs using GPL-only helpers. Without it, the kernel refuses to load the program.
The verifier is strict
Every eBPF developer hits this in the first week. The verifier rejects programs that could, in theory, do something unsafe. Common rejections:
| Error | Cause | Fix |
|---|---|---|
invalid mem access | Reading a pointer the verifier can't prove is valid | Use a probe helper or bounds check |
back-edge from insn X to Y | Unbounded loop (older kernels) | Bounded loops with #pragma unroll or kernel 5.3+ |
R1 type=map_value expected=fp | Passing a map value as a stack pointer | Copy to a local variable first |
program too large | Over 1M instructions (older kernels) | Split into multiple programs |
cannot call GPL only function | Missing license declaration | Add char LICENSE[] SEC("license") = "GPL"; |
The verifier's error messages have gotten much better in recent kernels. It now shows you the exact instruction that failed and the state of each register, which is what makes debugging tolerable.
If you're starting out, use a CO-RE (Compile Once, Run Everywhere) approach with libbpf. It lets you write one program that works across kernel versions by using BTF type information instead of hardcoded offsets.
XDP: the networking use case
XDP (eXpress Data Path) is eBPF attached at the network driver level, before the kernel allocates an sk_buff. It runs on every incoming packet, and it's fast enough for DDoS mitigation at 10Gbps per core.
SEC("xdp")
int xdp_drop_udp(struct xdp_md *ctx) {
void *data = (void *)(long)ctx->data;
void *data_end = (void *)(long)ctx->data_end;
struct ethhdr *eth = data;
if ((void *)(eth + 1) > data_end) return XDP_PASS;
if (eth->h_proto != bpf_htons(ETH_P_IP)) return XDP_PASS;
struct iphdr *ip = (void *)(eth + 1);
if ((void *)(ip + 1) > data_end) return XDP_PASS;
if (ip->protocol == IPPROTO_UDP && ip->dport == bpf_htons(1234)) {
return XDP_DROP;
}
return XDP_PASS;
}
Return XDP_DROP and the packet is gone before the kernel networking stack sees it. This is the same technique Cloudflare uses to handle volumetric attacks.
Attaching:
sudo ip link set dev eth0 xdp obj xdp_drop.o sec xdp
# Check
ip link show eth0 | grep xdp
# Detach
sudo ip link set dev eth0 xdp off
Tools built on eBPF
| Tool | Purpose |
|---|---|
| Cilium | Kubernetes networking, replacing kube-proxy |
| Falco | Runtime security — alerts on suspicious syscalls |
| Pixie | Auto-instrumented observability for k8s |
| Parca | Continuous profiling without sampling gaps |
| bpftrace | One-liners for kernel tracing |
| bcc | Older toolkit, still useful for specific tools |
I've run Falco in production and it caught a container attempting to write to /etc/passwd — something no userspace monitoring would have seen in real time without enormous overhead.
When eBPF is the wrong tool
- You need to trace userspace only. dtrace or perf handle that without kernel-level complexity.
- You're on an old kernel. Anything before 4.9 has limited support; CO-RE needs 5.2+ for BTF. RHEL 7 and older are painful.
- You need portability across architectures. eBPF works on x86_64, arm64, and a few others, but not everything.
- Your use case is already covered. Don't write an XDP load balancer when nginx exists; don't write syscall tracing when Falco has 200 community rules.
The learning curve is real, but the entry point with bpftrace is shallow. Ten minutes of one-liners will teach you more about what's happening on your system than a week of reading documentation.
