What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

eBPF turns kernel analysis into programmable, live instrumentation. A userspace loader can place a verified eBPF program on a kernel or userspace hook, filter and aggregate events in the kernel, then deliver only useful data through maps or event buffers. That makes eBPF valuable for tracing system calls, scheduling, I/O, networking, memory, security decisions, and application-to-kernel behavior—without changing kernel source or loading a traditional kernel module.

The important qualification is that eBPF is a measurement layer, not a magic “kernel visibility” switch. Your conclusion is only as good as the event selected, the timing measured, the data retained, and the interpretation checked against another signal.

What eBPF contributes to kernel analysis

Classic BPF began as a packet-filtering virtual machine. Extended BPF (eBPF) adds a richer instruction set, maps, helper functions, multiple program types, and many attachment points. The kernel accepts a program through the bpf() system call after the verifier checks its control flow, pointer bounds, stack initialization, alignment, reference handling, and allowed helper use. The subsystem is documented at kernel.org and in the broader BPF documentation.

A typical analysis has this shape:

  1. Turn an operational question into a measurable event or function.
  2. Choose the most stable hook that exposes the needed context.
  3. Attach a small program and filter early.
  4. Aggregate counters, distributions, or stacks in BPF maps.
  5. Send selected records through a ring buffer or perf buffer.
  6. Cross-check the result with another source before claiming cause.

eBPF observes live execution. It does not provide unlimited historical state, and a successful attachment does not prove that the chosen event represents the user-visible problem. A syscall entry, for example, is not a successful completion; a frequently called function is not necessarily the bottleneck.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How the eBPF data path works

The practical lifecycle is:

Source or script
   ↓
Compiler or bpftrace
   ↓
ELF/BPF object
   ↓
libbpf or another loader
   ↓
Verifier
   ↓
Attach to hook
   ↓
Maps and event buffer
   ↓
Userspace analysis

Programs, helpers, maps, and links

An eBPF program runs at a defined hook and can call approved helpers—for example, to read context, obtain timestamps, look up map values, or emit an event. Maps hold counters, histograms, configuration, and state shared with userspace or other BPF programs. A link represents an attachment and can be inspected and detached independently of the program object.

Map choice affects memory, concurrency, and lookup cost. Per-CPU maps reduce contention for hot counters but require userspace to combine values from each CPU. Ring buffers provide efficient, ordered event delivery; perf buffers remain widely supported. Histograms preserve a latency distribution instead of hiding tail behavior behind one average. Stack traces are useful for attribution but depend on symbols, unwinding support, frame pointers, and kernel configuration.

The verifier is a safety gate, not a correctness proof

The verifier simulates possible paths and tracks register types, pointer ranges, stack state, map pointers, and references before loading a program. It can reject an unchecked nullable map lookup, an out-of-bounds packet read, an invalid stack offset, or an unsupported helper call; see the verifier documentation. Acceptance means the program satisfies safety constraints. It does not mean that the event is semantically correct, that your timing is end-to-end latency, or that the workload will be unaffected by a poorly chosen high-frequency probe.

Choose the hook before choosing the command

Event selection is usually harder than writing a one-liner. Establish whether you need entry, return, waiting, failure, completion, or a sampled stack, then prefer the most stable interface available.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Hook Best use Strength Main risk
Tracepoint Defined kernel events and syscall tracing Static interface, generally more stable than kprobes Fields may omit internal detail
Raw tracepoint Lower-overhead tracepoint access Direct access to raw arguments More dependent on raw layout and program type
kprobe Dynamic kernel-function entry tracing Broad reach Names, arguments, and semantics can change
kretprobe Return values and completion timing Useful for errors and duration Return context may not retain original arguments
fentry/fexit BTF-enabled function tracing Typed arguments and low-overhead trampolines Requires suitable kernel features and BTF
Perf/profile CPU and hardware/software sampling Efficient statistical profiling Does not capture every event
Uprobe/uretprobe User-process functions Correlates application and kernel behavior Symbols, ASLR, inlining, and ABI issues
USDT Application-provided user events Better semantic stability than arbitrary uprobes The application must provide probes
LSM Security decisions and enforcement Can observe or restrict security actions Requires careful privilege and policy design
XDP/tc Early packet-path analysis Very early, efficient packet visibility Not equivalent to socket-layer observations

Tracepoints

Use a tracepoint when the event exists and its fields answer your question. Tracepoints are statically defined and generally more resilient across kernel upgrades than probes tied to internal functions. Their behavior and definitions are described in the kernel tracing documentation and bpftrace’s language reference.

kprobes and kretprobes

Use a kprobe when no suitable tracepoint exists and you specifically need an internal function. Validate the symbol and signature on every target kernel. Optimization, inlining, architecture, configuration, module scope, and restrictions can make a function unavailable. A kretprobe is useful for return values and elapsed time, but the return context alone may not contain the entry arguments.

fentry and fexit

fentry/fexit use BTF-derived argument types and eBPF trampolines. They are often preferable to kprobes for typed function instrumentation when the running kernel supports the attachment type and exposes usable BTF. They still do not make function semantics stable across all releases.

Sampling, uprobes, and USDT

Choose profile or hardware sampling for “where is CPU time going?” Choose event tracing for counts, errors, state transitions, or individual operations. Uprobes observe binary functions; USDT probes are application-defined semantic events and are usually a better contract when available.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Check the environment first

Run these checks on the machine that will actually be traced:

uname -a
cat /etc/os-release

test -r /sys/kernel/btf/vmlinux && echo "BTF available" || echo "BTF unavailable"
sudo bpftool feature probe
mount | grep -E 'tracefs|debugfs' || true

BTF is commonly exposed at /sys/kernel/btf/vmlinux. libbpf uses it for type and field relocations, and you can generate a header with:

sudo bpftool btf dump file /sys/kernel/btf/vmlinux format c > vmlinux.h

BTF is not guaranteed on every distribution or kernel build. Without it, tooling may require kernel headers, manually supplied structures, or another probe type. Kernel configuration, architecture, vendor backports, tracefs/debugfs availability, and loaded modules all affect what can be attached.

Privilege is equally deployment-specific. Linux introduced more granular capabilities, including CAP_BPF and CAP_PERFMON, in Linux 5.8; networking programs may need CAP_NET_ADMIN, while older systems can rely on broader legacy paths. See the Linux eBPF capability reference. A container’s root user may still lack host capabilities, namespace access, tracefs, symbols, or process-memory access. Locked-down kernels, LSM policy, seccomp, and cloud-provider restrictions can block loading or attachment.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

First investigation with bpftrace

bpftrace is a high-level language for rapid exploration, one-liners, aggregations, and histograms. Discover available probes instead of guessing names:

sudo bpftrace -l 'tracepoint:syscalls:*open*'
sudo bpftrace -l 'tracepoint:sched:*'
sudo bpftrace -l 'kprobe:*vfs*'
sudo bpftrace -lv 'tracepoint:syscalls:sys_enter_openat'
sudo bpftrace -lv 'fentry:tcp_reset'

Count open-attempts by process name:

sudo bpftrace -e '
tracepoint:syscalls:sys_enter_openat
{
  @[comm] = count();
}'

Press Ctrl-C to print the aggregate. This is a count, not a rate; divide by a measured interval if you need events per second. Process names are not unique, so production analysis should normally key by PID, UID, cgroup, executable path, or a combination.

To inspect arguments and add process identity:

sudo bpftrace -e '
tracepoint:syscalls:sys_enter_openat
{
  printf("%-6d %-16s %sn", pid, comm, str(args.filename));
}'

The args syntax follows current bpftrace documentation; older releases used different forms in some examples, so check the installed version. Add an early filter rather than printing every process:

sudo bpftrace -e '
tracepoint:syscalls:sys_enter_openat
/pid == 1234/
{
  @[str(args.filename)] = count();
}'

Measure a function’s duration carefully

sudo bpftrace -e '
kprobe:vfs_read
{
  @start[tid] = nsecs;
}

kretprobe:vfs_read
/@start[tid]/
{
  @latency_us = hist((nsecs - @start[tid]) / 1000);
  delete(@start[tid]);
}'

This histogram measures time between entry and return for vfs_read. It can include nested calls, scheduling, and blocking, so it is not automatically application-visible read latency. Recursive paths, missing returns, and reuse of a thread ID can leave incorrect state; production code needs bounded state and cleanup logic. A kretprobe may also be unavailable for the target.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Profile CPU hotspots without tracing every call

sudo bpftrace -e '
profile:hz:99
{
  @[kstack] = count();
}'

This samples kernel stacks at 99 Hz. Sampling is usually safer for broad hotspot discovery; event tracing is better for specific state transitions and counts. Reduce stack frequency or use a hardware PMU workflow with perf when the workload is especially hot.

Inspect loaded programs and BTF with bpftool

bpftool is the low-level inspection utility for BPF objects and capabilities:

bpftool version
bpftool help
sudo bpftool feature probe
sudo bpftool prog show
sudo bpftool map show
sudo bpftool link show
sudo bpftool btf show

Use these commands to confirm that the program loaded, the expected link exists, maps are present, and the running kernel exposes BTF. In a deployed application, ring-buffer or perf-buffer readers should consume events continuously; debug printing is not a production transport.

From an exploratory script to production tooling

Use bpftrace for discovery, then move to libbpf when the tool must run repeatedly or across a fleet. libbpf handles object opening, map creation, relocation, verification/loading, attachment, and teardown. BPF skeletons provide generated lifecycle code, while CO-RE (Compile Once—Run Everywhere) records relocation information and lets libbpf adapt types and fields using target-kernel BTF.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

CO-RE improves portability, but type portability is not semantic portability. It cannot restore a removed function, create a missing hook, provide an unavailable helper or program type, compensate for absent BTF, or prove that a field means the same thing after a semantic change. Test across actual distribution kernels, architectures, configurations, and vendor backports. Account for lost events, map cardinality, cleanup, feature probing, capability grants, and shutdown behavior.

BCC remains useful when an existing Python- or Lua-based diagnostic tool already solves the problem. Its runtime compilation and header compatibility requirements can complicate broad deployment; libbpf with CO-RE is generally the stronger foundation for a new portable binary.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Common failures and recovery paths

No probes found

sudo bpftrace -l 'tracepoint:*'
sudo bpftrace -l 'kprobe:*'
sudo bpftool feature probe
test -r /sys/kernel/btf/vmlinux

The event or function may not exist, a module may be unloaded, tracing may be disabled, tracefs may be unavailable, the tool’s naming convention may differ, or permissions may block enumeration. Search tracepoints first, inspect trace-event definitions, verify the exact kernel build and architecture, then try supported fentry/fexit or kprobe fallbacks. perf list, /proc/kallsyms, and kernel source can validate a target.

Cannot attach to a kprobe

The function may be inlined, optimized away, static, renamed, untraceable, or restricted. Try the corresponding tracepoint, an fentry hook, a caller or callee, or a module-qualified symbol. Confirm availability with bpftrace -l and bpftool feature probe.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Verifier rejection

Typical causes are uninitialized stack reads, unchecked nullable map results, missing packet bounds checks, invalid pointer arithmetic, misalignment, leaked references, unsupported helpers, and excessive state complexity. Check the verifier log rather than guessing. A safe map access has this shape:

value = bpf_map_lookup_elem(&map, &key);
if (!value)
    return 0;

/* Access value only after the NULL check. */

For packet data, establish data_end and verify that the complete header lies within it before dereferencing. Reduce complexity with bounded loops, smaller state, less stack use, tail calls, and userspace interpretation.

The program loads but output is empty

Confirm that the event occurs, remove restrictive filters, check the program and link with bpftool, verify that the buffer reader is running, and ensure the process or cgroup namespace is the one generating activity. Replace event output with a simple counter to separate attachment problems from transport problems.

Overhead or dropped events

Symptoms include CPU growth, scheduler perturbation, lock contention, memory growth, and ring-buffer loss. Filter by PID, cgroup, UID, device, namespace, or operation at the hook; aggregate in kernel; use per-CPU maps where suitable; sample broad profiles; avoid printf() on hot paths; reduce stack capture; and measure the workload with and without the probe.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Technically correct, analytically wrong

  • A function duration can include blocking and nested work.
  • Syscall entry does not prove completion or success.
  • A kprobe can observe internal calls unrelated to the user-visible operation.
  • Process names merge different PIDs.
  • Incomplete stacks, lost events, and sampling bias can change conclusions.
  • Kernel CPU time is not the same as wall-clock request latency.

Validate an important conclusion with an independent signal such as perf, ftrace, application latency, /proc, block statistics, scheduler data, or network counters.

When another tool is the better choice

Question or constraint Often preferable Why
Hardware PMU or mature statistical profiling perf Established sampling and hardware-event workflows
Existing kernel event streams ftrace or trace-cmd Small deployment surface and standardized trace formats
One process’s system-call behavior strace Simple per-process syscall visibility without BPF privileges
Existing scripts and runtime compilation SystemTap or BCC May already fit the team’s tooling
Current counters and state /proc and /sys Lower complexity for data already exported by the kernel
Historical or business context Application metrics and tracing eBPF alone observes live kernel activity, not retained history or intent

eBPF complements these tools rather than universally replacing them. Commercial platforms can add fleet deployment, dashboards, retention, alerting, Kubernetes integration, security policy, or continuous profiling, but they do not remove kernel-version, privilege, cardinality, overhead, and interpretation constraints. Native tools remain the sensible choice for a one-off host investigation.

A concise decision guide

  • Need a defined, relatively stable event? Start with a tracepoint.
  • Need an internal function with no tracepoint? Use a kprobe, accepting version maintenance.
  • Need typed function arguments and BTF is available? Try fentry/fexit.
  • Need CPU hotspots? Sample with a profile event or perf.
  • Need counts, errors, or transitions? Trace events and aggregate them.
  • Need a repeatedly deployed application? Build with libbpf and CO-RE, with feature probing and lost-event accounting.
  • Need rapid exploration? Use bpftrace.
  • Need an existing diagnostic? Check BCC.
  • Need historical state or user intent? Combine eBPF with application, kernel, or system telemetry.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.