Both terms describe the same idea — controlling how many API requests an application can make in a given window — but they operate at different points in the system, and the difference matters when you design messaging flows on Devotel Orbit. Conflating them is the root of the most common integration misdiagnosis: developers see delayed sends, assume a 429 somewhere, and tune a retry loop that was never the problem. This post defines each term against the machinery Orbit actually ships, so the next time traffic slows you know which mechanism fired and what to fix. If you want the operational walkthrough with per-endpoint budgets and a worked queue-drain, /blog/rate-limiting-vs-throttling-cpaas-apis picks up where this one leaves off.
Rate limiting: the client's budget
A rate limit is a quota the platform assigns to a client, expressed as requests per second, per minute, or per day. The client knows the ceiling up front and can design around it: batch sends, spread traffic, queue non-urgent requests. On Orbit, rate limits for messaging APIs are documented per endpoint in the rate limits guide, so a well-behaved client shapes its traffic deterministically and avoids rejections entirely.
The key property is contractual: a rate limit exists to keep shared infrastructure predictable for every tenant, not to punish a specific caller. What happens when you exceed it is the part developers get wrong. There are two mechanisms, and they are not interchangeable:
- Rejection-style limiting. Past a hard cap, the API answers 429 Too Many Requests with a
retry_aftervalue in the error body and a matchingRetry-Afterheader. The client backs off, waits, and tries again. Orbit's request-path limits behave this way — the per-endpoint per-minute budgets and the global 50-requests-per-second per-key ceiling all reject the excess. - Throttle-style limiting. On Orbit, messaging sender-velocity limits are shaped, not hard-rejected: a burst past the velocity hint is spread and paced around your configured limit instead of failing outright. The send takes longer; it does not come back as an error.
Reading a paced send as a rate-limit failure (or a 429 as a throttle) is the misdiagnosis this post exists to prevent.
Throttling: the platform's survival valve
Throttling is what happens when the platform deliberately slows, queues, or refuses traffic to protect itself, the carrier network, or the downstream channel. Unlike a contractual quota, throttling is often dynamic: it can kick in below your nominal quota when carrier routes congest, during seasonal peaks, or when fraud screens flag a suspicious sender pattern.
On messaging APIs specifically, throttling exists at multiple layers:
- The API gateway, shaping inbound request bursts against the per-key second ceiling.
- The sender-velocity layer, where Orbit paces message sends around each organization's configured throughput instead of rejecting a burst.
- The carrier side, downstream of the CPaaS entirely: US 10DLC throughput tiers cap messages per second per brand and route, and a newly registered number carries a daily send ceiling that ramps as it builds delivery history. Cross the ramp and the send is refused with
DAILY_CAP_EXCEEDEDorWARMING_QUOTA_EXCEEDED; send too fast inside the day andNUMBER_MPS_EXCEEDEDfires. To the API caller that shows up as a 429 — but it is carrier-side shaping, not your quota. The number warming caps page walks those gates. - The channel itself, e.g. WhatsApp quality-based tiering that throttles delivery for low-quality senders, or email ISPs throttling when complaint rates spike.
Note the asymmetry the FAQ returns to: carrier-side gates arrive as 429 responses even though throttling itself is the delay mechanism — which is exactly why "429 = my quota" is an unreliable diagnosis.
Why the distinction matters
When sends are rejected or delayed, the diagnosis differs by mechanism, and Orbit makes the mechanism readable: a 429 body's error code names which limiter family fired, and the right response differs by family.
| Symptom | Mechanism | Where to fix |
|---|---|---|
| 429 with a request-cap or throughput code | Quota exceeded (rejection-style) | Client: token bucket, bounded queue, honor retry_after |
| Sends paced slower than fired, no 429 | Sender-velocity throttling | Expected behavior — drain a queue instead of firing parallel |
429 with DAILY_CAP_EXCEEDED / WARMING_QUOTA_EXCEEDED | Carrier warming ramp | Spread traffic across warming numbers; wait for UTC reset |
| Delayed WhatsApp delivery, quality tier degraded | Channel-side throttling | Template quality and sender reputation, not retry logic |
If your quota is genuinely exceeded, fix the client: implement a token-bucket limiter, respect retry_after, retry idempotently. If the route is congested or a quality filter flagged the sender, fix that instead: check template quality, sender-ID status, and route health in the Orbit dashboard, and read the error code before treating any 429 as "wait a second and retry." The rate-limit and cooldown taxonomy maps every code to its family, and the troubleshooting flow walks the live decision tree.
A worked case: a campaign fires 3,000 SMS from twenty parallel workers and starts seeing 429s. Reading the code, half are the global request ceiling (queue discipline fixes those) and half are NUMBER_MPS_EXCEEDED on a warming number (a one-second backoff clears those, but the right fix is spreading the campaign over more senders). Two 429s, two different defects — treating both as "retry slower" would have stalled the campaign for hours without fixing either.
Definitions you can cite
- Rate limiting: a predetermined quota on the number of requests a client may make in a time window, enforced at the API boundary. On overflow the overflow is either rejected — a 429 with a
retry_afterit is safe to honor — or, for Orbit's sender-velocity limits, shaped into a paced stream. - Throttling: dynamic, sometimes quota-independent request shaping applied by the platform, the carrier, or the channel to protect infrastructure and delivery quality. It delays or paces rather than refusing outright, though its downstream gates (warming ramps, MPS hints) surface as 429s to the caller.
A healthy integration on Devotel Orbit does both:
- Stay inside your documented rate limits so the gateway never has to reject you.
- Make your retry logic throttling-tolerant — backoff on the server's
retry_afterhint, idempotency keys on every send — so transient throttling surfaces as a delay, not an error. The idempotency and safe retries guide walks that contract end to end.
Frequently asked questions
What is the difference between rate limiting and throttling?
Rate limiting is a published quota on request volume; exceeding it reliably signals the excess — on Orbit, either a 429 with a retry_after hint (request-path limits) or a paced send (sender-velocity limits). Throttling is dynamic traffic shaping that slows, queues, or paces rather than enforcing a fixed number — carrier ramps, congestion, and quality-based tiering all throttle without any quota of yours being crossed.
On Orbit, does exceeding my rate limit always return a 429?
No. Request-path limits — the per-endpoint per-minute budgets and the global 50-per-second per-key ceiling — reject with a 429. Sender-velocity limits are throttles: a burst past the velocity hint is spread around your configured limit, so sends go out slower instead of failing. And some 429s you receive (DAILY_CAP_EXCEEDED, WARMING_QUOTA_EXCEEDED) come from carrier-side gates downstream, not from your quota at all — the error code in the body tells them apart.
Should I retry every 429?
Only after reading the code. Throughput codes like NUMBER_MPS_EXCEEDED clear in about a second and are safe to retry on the retry_after hint. Warming codes point at the next UTC midnight — retrying every second burns your queue without clearing anything. Frequency-cap and per-recipient cooldown codes should not be retried against the same recipient inside the window at all. The troubleshooting flow has the full decision tree.
How do I stay throttling-tolerant?
Drain a bounded queue slightly below the published ceiling, send an Idempotency-Key on every attempt so retries cannot double-send, and watch the X-RateLimit-Remaining header so you slow the drain before the first rejection rather than after it.
For the queue-and-retry patterns behind this, see the idempotency and safe retries guide and the webhook retry guidance.