Let’s set sail aboard a beautiful three-masted ship, bound for a promising land. The destination is far away, and the voyage is full of uncertainty!

For centuries, sailors have relied on instruments such as compasses, maps, and barometers to reduce that uncertainty. Each puts a number to something that can be hard to judge: speed, heading, or depth. On their own, these instruments provide limited information; together, they help the crew understand the situation and adjust course when needed.

A three-masted sailing ship crosses a blue sea, with a compass and a map on board.
AI-generated illustration (GPT Image).

Cabestan Technologies takes its name from nautical vocabulary. But the expertise I’m sharing today concerns another area that also needs to stay on course: the backend. When a service slows down or returns errors, users may be the ones to tell us — sometimes too late. Instrumenting the service lets us follow its behavior and spot changes that deserve attention sooner.

In this article, we’ll look at what backend metrics are for and how to expose a few of them from a Go service. Then we’ll see how to view them in Grafana on a dashboard. I’ll leave out the infrastructure setup to keep the article focused on metrics.

What is a metric?

A metric is a numerical measurement of how a service is running. By following it over time, we can see whether things are changing. For example, we can count incoming requests, measure how long they take, or track memory use.

Here are a few useful terms for reading these measurements:

  • Metric: a measured value, such as the number of requests received.
  • KPI (Key Performance Indicator): a metric chosen to track a goal. Response time is a metric; the share of requests completed in under 300 milliseconds can be a KPI for tracking the user experience.
  • Label: information used to distinguish measurements, such as the HTTP method (GET, POST) or response code (200, 500). Labels should have a limited number of possible values. A user ID or request ID would create almost as many time series as requests, increasing the amount of data to store.

Not every metric is a KPI. To choose which ones to track, start with the questions you want to answer: Is the service available? Are responses getting slower? Are errors increasing?

A note: in production, metrics help us track how a service changes and spot unusual behavior. During an incident, seeing latency and errors rise at the same time gives us a lead to investigate.

Metrics alone may not reveal the cause: higher latency does not tell us whether requests are waiting on a database, a remote service, or internal processing. Logs and tracing can help us investigate those details. We also need to read charts in context: a traffic spike, for example, is not necessarily unusual.

Now let’s look at the main metric types and what they can measure.

The main metric types

Three types of metrics are especially useful for tracking a service: counters, gauges, and histograms. Which one to use depends on the question you want to answer.

Counter: count what happened

A counter (Counter) adds up events over time: each request handled makes it go up. Its raw value does not decrease while the service is running. If the service restarts, the counter starts from zero again.

Dashboards often show the rate calculated from a counter — for example, requests per second — instead of its raw total. This calculation accounts for resets, so the rate does not suddenly drop because the service restarted.

Someone adds tokens to a growing stack next to an abacus.
AI-generated illustration (GPT Image).

Gauge: track a value that can go up or down

A gauge (Gauge) reports a value at a given moment, and that value can go up or down. It can measure the number of requests in progress, memory use, or open connections. If the value keeps rising without falling, that may point to a buildup. After a restart, the previous value is not automatically preserved: the application reports what it measures or initializes at that point.

A dial and a slider illustrate a value that can change.
AI-generated illustration (GPT Image).

Histogram: observe a distribution

A histogram (Histogram) groups observations into buckets. For HTTP requests, it groups their durations into time ranges. This lets us estimate, for example, how long it takes for 95% of requests to finish.

The buckets should match the durations we want to observe. A histogram does not keep every individual duration; it records a distribution from which percentiles can be estimated.

Tokens are sorted into compartments to represent durations grouped into buckets.
AI-generated illustration (GPT Image).

In practice, a counter can track the number of requests, a gauge the number of connected users, and a histogram how long requests take. Together, these measurements give us several views of an HTTP server. We still need to choose labels that help distinguish them without creating too many time series.

From application to dashboard: the tools

There are several ways to collect, store, and display metrics. I use Prometheus and Grafana , the stack I usually work with. Other tools and combinations are available; the general idea is the same.

The metrics flow has three steps:

  1. The backend service measures its activity and exposes the values on an HTTP endpoint, often /metrics.
  2. Prometheus regularly queries that endpoint — this is called scraping —, stores the measurements over time, and lets us query them with PromQL.
  3. Grafana connects to Prometheus and turns those measurements into charts and dashboards.

This article uses pull collection: Prometheus regularly queries the service. Push is also possible: the service is responsible for sending its metrics to a server, usually through an intermediary such as the Pushgateway . This is especially useful for very short-lived jobs that may finish before Prometheus can query them, or that do not expose an HTTP server.

For NuCorder , I use VictoriaMetrics instead of Prometheus to collect and store metrics. It integrates with the Prometheus ecosystem, and Grafana can query it as a data source. The flow stays the same: the service exposes metrics, VictoriaMetrics collects them, and Grafana displays them.

Which metrics should you choose for an HTTP service?

To get started, there is no need to measure everything the service does. I would begin with a few concrete questions:

  • How many requests are coming in? We can track their number and, if useful, break them down by HTTP method or route.
  • How many are failing? The response code helps us spot errors, such as 5xx responses.
  • How long do responses take? Tracking their duration helps us see whether requests are getting slower, especially the slowest ones.

These three measurements provide an initial view of the service’s health. If traffic stays steady while responses slow down, a dependency or the service itself may be struggling. If errors increase after a deployment, that may point to a regression.

Once we have those first results, we can add business metrics to investigate more specific questions. For example, counting login attempts and separating successful attempts from failed ones helps us track login failures over time.

A note: labels let us filter measurements by method, response code, or route. Avoid values that change with every request: use the route pattern /users/{id} instead of a path containing the actual ID, and do not add a user ID, email address, or search parameter. Every distinct value creates a new time series, and the number can grow quickly.

Instrumenting an HTTP endpoint in Go

Go is the language I use every day, so that’s what I’ll use in this example. To instrument the service, I use the official Prometheus Go client , which provides metric types and tools to expose them over HTTP.

Prometheus clients in other languages

The same idea works in other languages with a suitable library:

The Prometheus documentation lists more client libraries.

Here is a deliberately small example. Middleware wraps the handler, records the response code and duration after it runs, then counts requests by method, route, and code:

go
package main

import (
	"fmt"
	"log"
	"net/http"
	"strconv"
	"time"

	"github.com/prometheus/client_golang/prometheus"
	"github.com/prometheus/client_golang/prometheus/promhttp"
)

type responseRecorder struct {
	http.ResponseWriter
	status      int
	wroteHeader bool
}

// responseRecorder stores the status code returned by the handler.
func (r *responseRecorder) WriteHeader(status int) {
	if r.wroteHeader {
		return
	}
	r.status = status
	r.wroteHeader = true
	r.ResponseWriter.WriteHeader(status)
}

func (r *responseRecorder) Write(p []byte) (int, error) {
	if !r.wroteHeader {
		r.WriteHeader(http.StatusOK)
	}
	return r.ResponseWriter.Write(p)
}

func instrumentHTTP(requests *prometheus.CounterVec, duration *prometheus.HistogramVec, next http.Handler) http.Handler {
	return http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
		start := time.Now()
		rec := &responseRecorder{ResponseWriter: w, status: http.StatusOK}
		next.ServeHTTP(rec, r)

		// ServeMux provides a stable pattern, unlike a URL containing IDs.
		route := r.Pattern
		code := strconv.Itoa(rec.status)
		requests.WithLabelValues(r.Method, route, code).Inc()
		duration.WithLabelValues(r.Method, route).Observe(time.Since(start).Seconds())
	})
}

func main() {
	requests := prometheus.NewCounterVec(prometheus.CounterOpts{
		Name: "http_requests_total",
		Help: "Total number of HTTP requests.",
	}, []string{"method", "route", "code"})

	duration := prometheus.NewHistogramVec(prometheus.HistogramOpts{
		Name:    "http_request_duration_seconds",
		Help:    "HTTP request duration in seconds.",
		Buckets: prometheus.DefBuckets,
	}, []string{"method", "route"})

	// Register the metrics so the /metrics handler can expose them.
	prometheus.MustRegister(requests, duration)

	app := http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
		fmt.Fprintln(w, "Hello!")
	})

	mux := http.NewServeMux()
	mux.Handle("/hello", instrumentHTTP(requests, duration, app))
	// Keep /metrics outside the middleware so scrapes are not counted.
	mux.Handle("/metrics", promhttp.Handler())

	log.Println("Listening on http://localhost:8080")
	log.Fatal(http.ListenAndServe(":8080", mux))
}

To try the example, add the github.com/prometheus/client_golang dependency, start the server, and open /hello. The measurements are available at /metrics in a format Prometheus can read.

The metrics exposed are:

  • http_requests_total counts requests by method, route pattern, and response code.
  • http_request_duration_seconds tracks their duration by method and route pattern.
  • The route pattern — for example, /users/{id} — groups requests for different users instead of using each full URL.

A note: retrieving the route pattern with r.Pattern requires Go 1.22 or later.

Viewing metrics in Grafana

To visualize the metrics, add Prometheus as a data source in Grafana. Then create a dashboard with three charts: request rate, error rate, and 95th-percentile latency.

Grafana dashboard showing request rate, the share of 5xx responses, and 95th-percentile latency.
Grafana dashboard built from the Go example’s HTTP metrics.

Request rate

To show the number of requests per second, calculate the rate of increase of the counter over a time window, then add the series together:

promql
sum(rate(http_requests_total[$__rate_interval]))

This gives the overall rate. To break it down by HTTP code, group the series by the code label:

promql
sum by (code) (rate(http_requests_total[$__rate_interval]))

Error rate

We can compare the rate of 5xx responses with the total request rate:

promql
sum(rate(http_requests_total{code=~"5.."}[$__rate_interval]))
/
sum(rate(http_requests_total[$__rate_interval]))

The result is a ratio that Grafana can display as a percentage. If no requests were received during the time window, the error rate is undefined.

95th-percentile latency

The histogram provides cumulative counts for each duration bucket. We can aggregate them and ask for an estimate of the 95th percentile:

promql
histogram_quantile(
  0.95,
  sum by (le) (rate(http_request_duration_seconds_bucket[$__rate_interval]))
)

What does p95 mean? At a given point on the chart, if p95 is 300 milliseconds, that means about 95% of requests in the calculation window completed in under 300 milliseconds. The remaining 5% took longer.

P95 complements the average, which can hide slowdowns affecting only the slowest requests.

These three charts help us spot changes in a service’s behavior; we still need to find out what caused them.

What metrics cannot show on their own

Metrics provide an overview of a service over a period of time. They help us track trends and spot anomalies, but they do not necessarily tell us what happened to a specific request.

A dashboard can show that p95 has increased and that there are more 5xx responses, but it will not automatically tell us which request failed or which internal operation took a long time. Logs provide detailed events with context; tracing follows a request across components; profiling helps us examine the work done inside a process. These tools answer different questions and can complement each other during an investigation.

We can also define alerts based on metrics, for example when the error rate stays high or latency exceeds a threshold for several minutes. An alert should correspond to a situation that actually calls for action: too many sensitive thresholds can create noise and make notifications less useful. The rules and how alerts are routed depend on the environment, which is beyond the scope of this article.

What I take away

Backend metrics provide concrete signals about a service’s traffic, errors, latency, and resource use. With Prometheus — or VictoriaMetrics in my case — and Grafana, a few well-chosen measurements, shown with their units and context, are enough to track changes in the service. By instrumenting a few Go handlers, we can move from a hunch — “the API seems slow” — to specific questions: since when, for what share of requests, and where should we look next? Metrics do not steer the service for us, but they help us know when it is time to correct course.

Need help getting a clearer picture of your backend?

If you want to choose useful metrics for your service or set up a first dashboard, feel free to get in touch.

Let’s talk