SMS deliverability differs from email deliverability in one decisive way: the confirmation a carrier returns is a network receipt, not a handset confirmation, and the route the message traveled decides how much that receipt is worth. A sender that treats "delivered" as a binary fact builds its program on a number it cannot defend. This post walks what delivery receipts actually report, how to read benchmark ranges by vertical and route class without inventing percentages, where receipts come from before the handset ever confirms, and the tenant-owned levers in Devotel Orbit that turn receipt quality into a closed measurement loop — detection, alerting, and route selection you control.
Delivery receipts vs. handset delivery — what carriers actually report
An SMS moves through a chain of carriers and aggregators before it reaches a handset, and a delivery receipt (DLR) is a status returned by one hop in that chain. When an aggregator or a downstream carrier marks a message DELIVERED, it has received a delivery acknowledgement from the next hop down — which may itself be another aggregator — and passes that claim up to you. The handset is the source that matters; everything between your sending platform and the handset is an intermediary relaying a claim.
That distinction matters because receipt semantics define what a delivery benchmark can honestly claim. A direct carrier interconnect usually returns the handset's own delivery acknowledgement, so DELIVERED on that route class is close to a true handset confirmation. An aggregating route returns the downstream party's claim of delivery, which may be several hops from the handset, and a route where the receipt is unverifiable — a gray route — is where departures between "delivered" and "actually on the handset" concentrate. Treat every receipt as a claim from the hop that answered, then measure the quality of that hop before you treat the receipt as evidence.
Benchmark ranges by vertical and route class
You benchmark a program against something, and the first rule is that a benchmark only means what its source supports. Orbit publishes its own medians for customers where the measurement loop closes, and carriers publish statistics in their regulatory filings and network disclosures; a post that quotes a percentage without one of those two sources is fabricating it, and that is the failure this post names explicitly so no one reads an invented range as authoritative.
Two variables organize every benchmark honestly. Vertical: one-time passwords and transactional alerts draw fast, high-confirmation traffic — users request them and immediately engage — while marketing and promotional traffic draws opt-out and complaint friction that makes delivery weaker. Route class: a direct carrier interconnect and an aggregating downstream answer different receipt semantics, as above, so their medians are not comparable and a blended number across route classes is meaningless. Compare OTP to OTP and marketing to marketing, direct interconnect routes to direct interconnect routes, and then the benchmark tells you something; compare across those boundaries and it tells you nothing. A universal benchmark range would have to be fabricated, so treat any single global range as a marketing artifact rather than a measurement.
Gray-route and convergence signals — where receipts come from before the handset confirms
A gray route is a route whose origin is not registered for the traffic class it carries — A2P traffic traveling as if it were person-to-person, often terminated through SIM boxes. It is the worst case for receipt reliability because a receipt that arrives on a gray route may come from any hop in that chain, or from nothing at all; carriers now price unregistered-origin surcharges above the registered tariff and screen blocked ranges before what they call "delivery," so a gray-route receipt may report a hop that never reached the handset. The canonical explainer on this class — SMS pumping and gray-route fraud: the operator explainer — walks the economics that make gray routes a billing trap as well as a measurement trap.
The signals that distinguish a trustworthy route from a gray route are the ones that let you infer what a receipt is worth before the handset ever confirms. Sender registration status, the per-hop latency pattern across your traffic, and the rate at which a route's receipts disagree with downstream confirmation are the three observable signals; a convergence toward one of them breaking pattern — a latency profile shifting, a sender losing registration, a confirmation rate departing from its baseline — is the fingerprint of a route drifting gray. Where receipts arrive, treat them as claims; where they do not, the route's silence is itself the signal you measure.
Tenant-owned levers in Orbit — detection, alerts, and A/B route selection
In Orbit, the quality loop above runs entirely on controls the tenant owns, per the platform's compliance posture that every control except the federal TCPA dialing window is tenant-configurable and defaults open. Three levers close it:
Anomaly detection. Orbit's Insights → Anomalies surface usage and delivery anomalies, and a route whose receipts dry up or whose confirmation rate departs from baseline raises on the detector before the gap costs you meaning. The detector classes separate wallet (spend) scope from webhook/inbound scope, so a delivery anomaly raises independently of a billing anomaly — you get the alarm on the quality signal, not only on the cost.
Sender-island alerts. A sender identity that tilts toward a route whose receipts cannot be trusted gets flagged in sender-island route-quality alerts — the alert names the affected sender and route so you can act before the drift is structural. Sender registration status feeds directly into this detection, which is why a dropped registration surfaces as a route-quality alert rather than a silent delivery gap.
A/B route selection. Where route classes differ — direct interconnect against aggregating downstream — you select the route per sender and per traffic class, then compare receipt quality across the A/B selection before committing traffic. That is the honest way to use the vertical/route-class frame from the benchmark discussion: choose the route class against the same vertical's confirmation baseline, and the selection is measurable rather than aspirational.
Sibling posts — the email-deliverability pair
The email side of deliverability differs in mechanism but shares the measurement discipline: receivers return claims, and the sender owns the loop. Devotel Orbit's email deliverability playbook covers SPF, DKIM, DMARC alignment and bounce handling on the email side, and transactional email sending identities covers the identity-and-DNS-verification half that protects deliverability across a shared account. SMS and email differ in what a "delivered" claim proves — SMS receipts come from the network hop rather than the handset, while email's authentication grid is cryptographically checkable — but both reduce to the same tenant-owned loop: know what the confirmation actually claims, measure where it departs, and control the route.
Frequently asked questions
What is the difference between a delivery receipt and a handset confirmation?
A delivery receipt is a status returned by one hop in the carrier chain — an aggregator or a downstream carrier acknowledging it accepted delivery responsibility. A handset confirmation comes from the device itself. On a direct carrier interconnect the two usually coincide; on an aggregating or gray route, the receipt can arrive while the handset never sees the message.
Why can't a post like this quote one universal SMS delivery benchmark?
A universal benchmark range would have to be fabricated — vertical (OTP vs. marketing) and route class (direct interconnect vs. aggregating) shift the number by structure, not by noise. Orbit cites its own medians for customers where the loop closes, and carriers publish statistics in their own disclosures; anything between those two sources with no citation is invented.
What is a gray route, and why does it matter for measurement?
A gray route carries A2P traffic dressed as ordinary person-to-person traffic, often terminating through SIM boxes, on an origin not registered for that traffic class. It matters for measurement because receipts on such a route may come from any hop in an unverifiable chain, so "delivered" cannot be read as a handset confirmation.
Which Orbit controls close the SMS quality loop?
Three tenant-owned controls: the anomaly detection surface flags when a route's receipts dry up or a confirmation rate departs baseline, sender-island route-quality alerts name the sender and route drifting toward untrustworthy receipts, and per-sender A/B route selection lets you compare route classes against the same vertical baseline before committing traffic.