Overview
Metrics tell you a service is slow. Traces tell you which requests are slow. Profiling tells you which function is slow — the actual line of code that's consuming CPU or allocating memory.
Traditionally, profiling was something you did in development with a profiler attached, or in production with a lot of ceremony and a performance hit. Continuous profiling makes it always-on, low-overhead, and just another data source alongside metrics and logs.
What continuous profiling actually is
A profiler samples the call stack at intervals — typically every 10ms — and aggregates the samples over time. The result is a picture of where CPU time is actually being spent, across the whole fleet.
| Traditional profiling | Continuous profiling |
|---|---|
| Manual, on-demand | Always running |
| One process at a time | Entire fleet |
| High overhead | 1-3% CPU |
| Snapshot in time | Historical trends |
| Dev environment | Production |
The always-on part is what changes things. When a user reports that things feel slow, you don't need to reproduce it. You look at the last hour of profiles and see what changed.
The overhead question
The "1-3% CPU" claim is real for sampling profilers, and the reason is straightforward: sampling doesn't instrument every function. It interrupts the process periodically and records the current stack. The cost is proportional to sample rate, not to how much code runs.
At 100Hz (10ms intervals), the profiler wakes up 100 times per second per core. That's negligible compared to the work the process is doing. At 1000Hz, it starts to matter.
This is different from tracing, which can have significant overhead because it instruments every operation. If you've ever enabled detailed Database tracing in production and watched throughput drop by 30%, you've seen the difference.
Reading a flame graph
Flame graphs are the standard visualization. They're confusing the first time and obvious the second.
[main]
[handleRequest] [backgroundWorker]
[parseJSON] [dbQuery] [render] [processBatch]
[read] [decode] [exec] [template] [compute] [write]
The rules:
- Width = time. A function that occupies 30% of the width was on the stack for 30% of samples.
- Height = call depth. The stack grows upward; leaf functions are at the top.
- Left to right has no meaning beyond the order of the call stack. It's not a timeline.
The thing to look for: wide boxes near the top of the graph. Those are functions that are themselves doing a lot of work, not delegating. A wide box that has children is a parent that's waiting — the interesting work is one level down.
A function that's wide and has no children at the top of the graph is a leaf — that's where the CPU time actually is. This is where you look for optimization opportunities.
Setting up pprof for Go
import (
"net/http"
_ "net/http/pprof"
)
func main() {
go func() {
http.ListenAndServe("localhost:6060", nil)
}()
// your application
}
That's it. Import the package, start an HTTP server, and you have a full profiling endpoint at localhost:6060/debug/pprof/.
# CPU profile for 30 seconds
go tool pprof http://localhost:6060/debug/pprof/profile?seconds=30
# Heap profile
go tool pprof http://localhost:6060/debug/pprof/heap
# Goroutine dump
go tool pprof http://localhost:6060/debug/pprof/goroutine
# Block profile (why are goroutines blocked?)
go tool pprof http://localhost:6060/debug/pprof/block
# Mutex contention
go tool pprof http://localhost:6060/debug/pprof/mutex
Each of these opens an interactive prompt. The commands worth knowing:
# Top functions by CPU
(pprof) top
# Top with cumulative time (includes children)
(pprof) top -cum
# Open a flame graph in the browser
(pprof) web
# Show source for a specific function
(pprof) list processBatch
top -cum is the one I use first. It shows the call tree sorted by cumulative time — how much time a function and everything below it consumed. If a top-level handler is 40% of CPU, that's the place to look.
Then list functionName shows the source with per-line timings, which is where the actual optimization target becomes obvious.
Continuous profiling with Pyroscope or Parca
pprof is on-demand. For continuous, you need an agent that scrapes profiles periodically and a backend to store them.
# Docker Compose for a Go app with Pyroscope
services:
pyroscope:
image: grafana/pyroscope:latest
ports:
- "4040:4040"
volumes:
- pyroscope-data:/var/lib/pyroscope
myapp:
build: .
environment:
PYROSCOPE_SERVER_ADDRESS: http://pyroscope:4040
PYROSCOPE_APPLICATION_NAME: myapp
depends_on:
- pyroscope
volumes:
pyroscope-data:
// In your Go app
import "GitHub.com/grafana/pyroscope-go"
func main() {
pyroscope.Start(pyroscope.Config{
ApplicationName: "myapp",
ServerAddress: "http://pyroscope:4040",
ProfileTypes: []pyroscope.ProfileType{
pyroscope.ProfileCPU,
pyroscope.ProfileAllocObjects,
pyroscope.ProfileAllocSpace,
pyroscope.ProfileInuseObjects,
pyroscope.ProfileInuseSpace,
},
})
// your application
}
The result: every 10 seconds, a profile is scraped, tagged with service name, version, and environment, and stored. You can then query "show me CPU profiles for myapp in production over the last 24 hours, grouped by version" and compare.
The killer feature: comparing versions
Here's the scenario that continuous profiling is built for.
You deploy version 2.3.0 at 10am. At 10:15, response times for one endpoint are up 40%. Metrics confirm it. Traces show it's happening somewhere in a specific code path. Profiling shows that parsePayload went from 5% of CPU to 45%.
Before continuous profiling, this comparison took hours of manual profiling in a dev environment that might not even reproduce. With a continuous profiler, it's a diff view: "here's the profile for v2.2.0, here's v2.3.0, here's what changed."
The comparison is by function name and by stack. A function that didn't exist in v2.2.0 shows as new. A function that was 5% and is now 45% shows as a regression.
This is how you catch performance regressions before they become incidents, and it's the feature that justifies the setup cost.
Memory profiling: the other half
CPU profiles show where cycles go. Memory profiles show where allocations come from, which is often more actionable for fixing slowdowns because allocation triggers garbage collection, and GC pauses show up as latency.
# Heap profile — currently allocated memory
go tool pprof http://localhost:6060/debug/pprof/heap
# Allocation profile — total bytes allocated over time
go tool pprof http://localhost:6060/debug/pprof/allocs
The distinction matters. A heap profile shows what's currently held. An allocation profile shows what was ever allocated, which is where you find functions that allocate and immediately release — invisible in the heap profile but expensive because of GC pressure.
(pprof) top -cum
flat flat% sum% cum cum%
0.02s 0.5% 0.5% 1.85s 45.2% encoding/json.Marshal
0.45s 11.0% 11.5% 0.45s 11.0% runtime.mallocgc
...
A function that shows 0% flat CPU but 45% cumulative under json.Marshal is a function that's spending all its time serializing. The fix is often to marshal once instead of per-item, or to use a faster encoder.
What to profile in production
| Profile type | Finds | Overhead |
|---|---|---|
| CPU | Where cycles go | 1-3% |
| Heap (in-use) | What's held in memory | Low |
| Allocs | Allocation hot spots | 1-3% |
| Goroutine | Leaks, blocked goroutines | Negligible |
| Block | Channel and mutex blocking | Moderate |
| Mutex | Lock contention | Moderate |
I run CPU, heap, and goroutine profiles always. Block and mutex profiles are opt-in — they add measurable overhead and they're only useful when you're specifically debugging a contention problem.
Labels: the feature that makes profiles actionable
pprof.Do(ctx, pprof.Labels(
"endpoint", "/api/search",
"tenant", tenantID,
"version", buildVersion,
), func(ctx context.Context) {
handleSearch(ctx)
})
Labels attach metadata to profile samples. In the profiler's UI, you can then filter by label: "show me CPU usage for tenant=acme" or "compare endpoint=/api/search across versions."
Without labels, a profile is an aggregate across all requests. With labels, you can slice it. This is what turns profiling from "the service is slow" into "this specific customer's requests are slow because of this specific code path."
What profiling tells you that nothing else does
Metrics: the p99 latency for /api/search is 800ms.
Traces: the slow requests spend 700ms in a database query.
Profiling: the query itself is fast, but the JSON serialization of the result set is calling reflection 40,000 times per request, and that's where the 700ms is.
The last one is only visible in a profiler. The database shows as fast. The trace shows the time attributed to the database call, but the actual CPU cost is in code that runs after the query returns.
This is a real example from a project I worked on. We spent a week optimizing queries before profiling showed the problem was in serialization.
When continuous profiling is worth it
- You have a latency or CPU problem you can't explain. Profiling answers questions that metrics and traces can't.
- Your traffic is expensive. Shaving 20% off CPU on a fleet of 200 instances saves real money.
- You have performance regressions you can't catch in staging. Continuous profiling catches them in production, where the workload is real.
- You have many services. The comparison and aggregation features become more valuable as the fleet grows.
For a small app with occasional performance issues, on-demand pprof is enough. Continuous profiling's value is in the always-on data and the historical comparison, which only matters when you have a fleet to manage and regressions to catch.
Getting started without the infrastructure
If you're not ready to run a continuous profiler, start with pprof. Import the package, expose the endpoint on an internal port, and next time something is slow, grab a profile.
# Grab a 30-second CPU profile and open it locally
go tool pprof -http=:8080 http://prod-server:6060/debug/pprof/profile?seconds=30
That single command gives you a flame graph in your browser for a production process. It's the fastest way to go from "something is slow" to "this function is the problem," and it requires about two lines of code in your application.
Continuous profiling is the next step, not the first one. Get value from on-demand pprof first, then decide if you need the always-on version.
