When a Service Slows Down, Where Should You Look?
An Introduction to Profiling and Tracing with Go
It is 8:50 a.m. and you have ten minutes to complete an online form before starting work. Once you have overcome your fear of paperwork, you dive into a page full of increasingly incomprehensible text fields and wonder why on earth you chose to sacrifice those precious minutes to boredom.
At 8:59, you finally submit the form — and nothing happens. You question your life choices, click the submit button again — just in case — and curse the endless waste of time. After several seconds, a message finally appears: your surname cannot contain accents. Data validation is legitimate, but what is more frustrating than a slow system without knowing why it is slow?
Now take the perspective of the developer behind that platform. In software engineering, the possible culprits are numerous and enough to make anyone sweat: a bug, an algorithm that works well for two users but collapses at ten, a slow network call, or excessive memory allocations. You can follow an intuition and take a guess. But without looking at what is actually happening, it is easy to spend a day optimising code that does not explain the observed slowdown.
“It is a capital mistake to theorize before one has data.”
— Sherlock Holmes in “A Scandal in Bohemia”, The Adventures of Sherlock Holmes , Arthur Conan Doyle.

Like Sherlock, I prefer to start with facts before forming a hypothesis. Having used Go for nearly a decade, I particularly appreciate the tools that come with it: pprof — which we will see below — lets us inspect what a process is doing in detail and put real numbers behind a feeling.
For a web service in production, I often find it more practical to start with tracing. It lets me follow real requests through several layers and identify where they spend time. Comparing several traces then helps me spot a recurring slowdown.
In this article, I will give an overview of these two approaches: the questions they answer, the first things to look at, and their limits. Let us see where to look before deciding what to change.
What are we talking about?¶
Before opening a tool, let us pause on two questions: what does my program do while it runs? and what happens to a request between its arrival and its response?
Profiling: what does my program do while it runs?¶
A profile is a measurement collected over a period of time, or a state captured at a specific instant. Depending on the question, it can show:
- CPU time consumed by functions;
- where the program allocates or retains memory;
- goroutines present at a given moment.
You cannot observe everything with a single profile, but each view helps attribute a cost or period of waiting to a specific part of the program.

Tracing: what happens to a request between its arrival and response?¶
A trace follows an operation — for example, an HTTP request — from its arrival to its response. It consists of spans: each span represents an instrumented step, with a start, an end, and therefore a duration. It can show time spent processing a request in one service, then time spent calling a database or another service.
When a request crosses several services, its trace context must travel with those calls to preserve the full journey. Tools such as Jaeger and Datadog APM can then display it. Tracing shows where a request spends time; on its own, it does not show which functions consume the most CPU inside a service.

Now that these two views are in place, let us start with the most local one: what happens inside a program.
Profiling: looking inside a program¶
Imagine a traffic jam. Should you add a lane, redesign an on-ramp, or remove a bottleneck that slows every car down? It depends on where traffic is actually blocked: making cars faster on an already clear stretch will not remove the jam.
The same caution applies to a Go service that slows down under load. Before changing a function that looks expensive, I collect a profile in a situation that resembles the observed problem. pprof can do this on a running process or during tests; here, let us start with an HTTP service.
1. Expose pprof¶
The net/http/pprof package exposes diagnostic routes out of the box. Run them on a local port, separate from the business server:
import (
"log"
"net/http"
_ "net/http/pprof"
)
func startPprof() {
go func() {
log.Printf("pprof: %v", http.ListenAndServe("127.0.0.1:6060", nil))
}()
}Calling startPprof() at startup gives a dedicated HTTP server access to the diagnostic routes. Importing net/http/pprof registers them on the default HTTP mux; all that remains is starting a regular HTTP server with http.ListenAndServe. Here, the server listens locally only, so those routes are not exposed to public traffic.
2. Capture under representative load¶
While the slow scenario is happening, capture thirty seconds of CPU time, for example:
go tool pprof 'http://localhost:6060/debug/pprof/profile?seconds=30'3. Read the result¶
You can analyse a profile in the CLI or in pprof’s web interface.
At the CLI prompt, top displays the functions most present in the samples. flat is time attributed directly to a function; cum also includes time spent in its calls.

To explore calls visually, start the web interface and open the displayed address:
go tool pprof -http=:8081 'http://localhost:6060/debug/pprof/profile?seconds=30'The graph shows cumulative CPU time and makes the functions taking the largest share easy to spot. A function at the top of the CPU ranking is not automatically something to fix, but a lead to investigate. Serialising complex structures as JSON or XML can, for example, be a significant cost while still being required to serve the request. The question is then whether that cost can be reduced, avoided, or accepted.

4. Choose the right profile¶
heaphelps identify allocations or objects that remain in memory;goroutineshows call stacks when they accumulate or appear blocked;- a CPU profile does not show time spent waiting for a network call.
Profiling adds some overhead of its own, which varies by profile type and duration. Before using it broadly in production, measure its impact and account for that disturbance in the analysis.
Similar tools in other languages¶
- Java: Java Flight Recorder
- Python:
cProfile - Rust:
cargo flamegraph - Node.js:
--cpu-prof - PHP: Xdebug Profiler
Tracing: following an operation’s journey¶
Tracing is not limited to microservices. It follows an operation — for example, an HTTP request — through the instrumented parts of a monolith, its database calls, and, when needed, other services. It answers a simple question: where was time spent?
1. Export traces¶
For this example, I will use Jaeger as the backend that stores traces. The application sends its traces in OTLP, an open standard that can export traces to compatible backends.
At application startup, create a TracerProvider, attach an exporter, and sample — here, 10% — of traces:
exporter, err := otlptracehttp.New(ctx)
if err != nil {
return err
}
tp := trace.NewTracerProvider(
trace.WithBatcher(exporter),
trace.WithSampler(trace.ParentBased(trace.TraceIDRatioBased(0.1))),
)
otel.SetTracerProvider(tp)
defer tp.Shutdown(ctx)Configure the exporter to send traces to Jaeger’s OTLP collector. Locally, the HTTP exporter can use http://localhost:4318 as its base address; traces are then sent to /v1/traces.
In production, sampling avoids retaining every request, which can take significant storage. At 10%, you keep one trace out of ten on average, but adjust that number to the real need.
2. Instrument entry points and dependencies¶
Now that the application can export traces, instrument the application code. For an HTTP server, the simplest approach is to create a root span for every request with HTTP middleware. Here is the setup with the chi router:
r.Use(telemetry.HTTPHandler)
r.Use(telemetry.RouteAttributes)HTTPHandler uses otelhttp.NewHandler. The second middleware groups requests by route pattern — for example, GET /v1/profiles/{id} — so it does not create a different value for every identifier.
Instrumentation libraries then do much of the work: otelsql for my database (MySQL here), redisotel for Redis, and otelhttp.NewTransport for HTTP calls to internal or external services. Trace context is passed to outgoing HTTP calls, so their spans join the same trace, provided the remote service is instrumented too.
3. Read a trace¶
In Jaeger, open a slow request and follow its spans from start to response. A long step tells you where to look, not necessarily the cause: it might contain computation, wait for a database, or call a remote dependency. Compare several traces before drawing a conclusion.

Here is one example from my experience. At Molotov, this instrumentation repeatedly helped me investigate slowdowns. I identified one call that was made twice: once in the API gateway, then again in the specialised service that manages recordings (bookmarks). The trace showed where to look; reading the code confirmed that one call could be removed, saving a few hundred milliseconds.
A trace is not a benchmark and does not replace a pprof profile. It locates the slowdown along an operation’s journey; a profile then helps explain what happens inside the affected process.
Useful resources¶
- OpenTelemetry instrumentation in Go
- Jaeger
- Tempo , an OTLP-compatible alternative
go tool trace, for observing the Go runtime rather than distributed tracing
Where should you start?¶
Start with the observed symptom: it points to the first measurement to take.
| Symptom | Start with |
|---|---|
| The service consumes a lot of CPU | A CPU profile with pprof |
| Memory use or allocations grow | A heap profile with pprof |
| Goroutines accumulate or appear blocked | A goroutine profile with pprof |
| A request is slow, with no clearly identified step | A tracing tool such as Jaeger or Datadog APM |
The first tool does not end the investigation. A trace can point to a call or process to inspect with pprof; a CPU profile showing little CPU activity can instead indicate that you should follow dependencies. After a fix, measure the same symptom again.
What I take away¶
The bookmarks case at Molotov illustrates the approach well: measurement found a duplicated step, then removed a few hundred milliseconds without complex optimisation.
pprof and tracing give two complementary views of a performance problem: the first examines what happens inside a Go process; the second follows an operation across instrumented steps. In both cases, numbers help you ask a better question. To judge a change, compare measurements obtained under similar conditions.
I have recently installed Tempo and configured it in Grafana to experiment with it in the NuCorder stack. It will probably be the subject of a future article once I have had time to get familiar with it.
Need help diagnosing a slow Go service?¶
If your service is slow and you are unsure where the time goes, I can help you choose the right measurements, find the bottleneck, and verify that a fix makes a difference.