Skip to main content
Back to blog

Traceroute: How to Read the Path Your Data Takes to a Server

Technical

Traceroute: How to Read the Path Your Data Takes to a Server article illustration

Something is slow, so you run a traceroute. Halfway down the output one hop reports 180ms while every hop after it sits at 30ms, and it looks like you have found the culprit.

You almost certainly have not. That is the classic traceroute misreading, and understanding why it happens is most of what makes the rest of the output worth reading.

What the tool actually does

Traceroute does not ask the network for a route. It works a cruder trick, using a field in every IP packet that was put there for something else.

That field is the time to live. RFC 1812, the standards-track document setting out what an IPv4 router has to do, describes it as “a timer limiting the lifetime of a datagram” measured in seconds, and then explains why nobody treats it that way. “Each router (or other module) that handles a packet MUST decrement the TTL by at least one, even if the elapsed time was much less than a second. Since this is very often the case, the TTL is effectively a hop count limit on how far a datagram can propagate through the Internet.” The same section says what happens when the count runs out: “If the TTL is reduced to zero (or less), the packet MUST be discarded, and … the router MUST send an ICMP Time Exceeded message, Code 0 (TTL Exceeded in Transit) message to the source.”

That error message is the entire mechanism. Send a packet with a time to live of one and the first router discards it and reports back, naming itself in the process. Send one with a time to live of two and the second router does the same. Keep incrementing and each router along the route announces itself in turn. Cisco’s documentation describes the sequence and notes the ceiling: probes increment “up to the maximum specified hop count. This is 30 by default.”

The traditional Unix method uses UDP aimed at a port nothing is likely to be using, starting at 33434, so the destination answers with a port unreachable message and the trace knows to stop. Windows tracert sends ICMP echo requests instead, so the destination finishes the trace with an ordinary echo reply. Three probes are sent at each step by default, which is where the three timings on each line come from.

Each of those three numbers is a separate attempt

A traceroute line is not a measurement of a link. It is three independent probes that happened to expire at the same router.

Your machine notes when it sent the probe, notes when the reply arrived, and subtracts. Routers along the way do no timing of their own. So the figure you see is the sum of three things: the time for your probe to reach that router, the time for that router to produce an ICMP message about it, and the time for that message to travel back to you.

Only the first and third reflect the network your traffic actually uses. The middle one is work the router does solely because you asked, and it is the part that misleads people. A 2016 presentation on reading traceroute output, created by Richard Steenbergen and published by ARIN, puts the distinction plainly: the generation step “is an artificially imposed constraint which affects ONLY traceroute packets, not real network traffic”.

Why a slow hop in the middle can mean nothing

Modern routers move ordinary traffic in hardware. That is the fast path, and it is what your actual packets use. Producing an ICMP Time Exceeded message is not ordinary traffic, so it falls to a general purpose processor handling exceptions, and that processor is modest by comparison. The same ARIN presentation is blunt about the priority it gets: “ICMP Generation is NOT a priority for the router.”

A router can be forwarding traffic through itself perfectly while taking its time to answer a probe aimed at itself, and traceroute reports the slow answer.

There is one test that settles it. Look at the hops after the slow one. If the following hops drop back to a normal figure, the packets clearly passed through fine and the spike was that router answering slowly rather than forwarding slowly. The same presentation states the rule both ways round: “If there is an actual forwarding issue, the loss or latency will persist across ALL future hops as well”, and “Latency spikes in the middle of a traceroute mean absolutely nothing if they do not continue forward.”

For what the numbers mean once you have a figure you trust, what counts as a good ping covers round-trip time, jitter and the distance floor you cannot get under.

What the stars mean, and what they do not

An asterisk means no reply arrived before the timeout. It does not mean the packet was lost.

Plenty of routers simply decline to answer. Microsoft’s documentation for tracert says so directly: “some routers don’t return time Exceeded messages for packets with expired TTL values and are invisible to the tracert command. In this case, a row of asterisks (*) is displayed for that hop.” A silent hop sitting between two hops that answer normally is a router with a policy, not a fault.

A star on the final line often has a different explanation again. Cisco’s routers rate limit ICMP unreachable messages, the type that ends a UDP trace, “to one packet per 500 ms (as a protection for Denial of Service (DoS) attacks)”, and Cisco notes that this “limitation does not affect other packets like ICMP echo requests or ICMP time exceeded messages”. Its own worked example shows the destination answering the first and third probes while the middle one comes back as a star, because the second reply was suppressed rather than lost.

Stars running from some point all the way to the end are a different case. The traceroute manual explains why, noting that traditional methods “can not be always applicable, because of widespread use of firewalls” which “filter the ‘unlikely’ UDP ports, or even ICMP echoes”, so the “whole tracerouting will just stop at such a firewall”. That is why the tool offers ICMP and TCP modes, described in the manual as intended “to bypass firewalls”.

Reading the addresses and names

The address on each line is one specific interface, not the router as a whole. Microsoft describes which one: “The near/side interface is the interface of the router that is closest to the sending host in the path.” A large router has many addresses, and only the one facing you appears.

Where a name appears beside the address, it is a reverse DNS record the operator chose to publish, and operators often encode a city, an interface type and a role into it. The same ARIN presentation warns that this “may not always be up to date”, and that some networks are “surprisingly bad at keeping DNS data updated”, which is one of several reasons reverse DNS tells you less than it appears to.

If you want to know whose network a hop belongs to rather than guessing from its name, look the address up and read the network identifier that comes back. What that number means, and why it describes a routing policy rather than an owner, is the subject of what an ASN is.

The first line or two are your own equipment, showing private addresses because that is what your home network runs on, and how one address serves every device explains where those come from.

The things it cannot show you

The return path is invisible. Every figure in the output includes a journey back to you that traceroute never displays, and the route home can differ from the route out. A hop that looks slow may be answering over a completely different path from the one carrying your traffic.

The path may not be a single path. Networks spread traffic across parallel links, and each probe is an independent trial that may take a different one. Sometimes you can see this, because a hop shows more than one address. Often you cannot, and the column is a blend of two routes rather than a picture of one.

Some hops are hidden by design. RFC 3443, published in January 2003, sets out how the time to live is handled inside MPLS networks, and describes a mode of operation called the Pipe Model introduced “to support the practice of configuring MPLS LSPs such that packets transiting the LSP see the tunnel as a single hop regardless of the number of intermediary label switch routers”. Where that is in use, a provider’s whole core can appear as one line. Nothing is broken; the hop count is simply not a count of machines.

Using it for what it is good at

Traceroute answers structural questions well and timing questions badly.

It will tell you where your traffic leaves your provider, whether it is going somewhere geographically sensible, and at what point it stops getting through. What it will not give you is a reliable latency measurement for any single hop, and that is what a ping test is for, run against the endpoint you actually care about rather than against the machinery in between.

So read it in that spirit. Check the route looks reasonable, note where it stops, and when a hop looks slow, check what comes after it before believing the number. A traceroute that clears itself in thirty seconds has still told you something, and it beats a guess dressed up as a diagnosis.