Don't aim for the threshold — aim well under it
Amazon's and Walmart's published thresholds (for example, late shipment rate under 4%) are the point at which enforcement action becomes likely, not a target to hover near. Because these metrics run on a trailing window, a single bad week can push a seller who's normally at 2% up toward the threshold fast, and it takes weeks of clean performance to bring the average back down. Sellers with healthy accounts generally run at half the stated threshold or better.
Late shipment rate
The percentage of orders shipped after the promised ship-by date. The most common causes are supplier delays that aren't reflected in your handling-time settings, and simply underestimating pick-pack-ship time during high order volume. Fix: pad handling time realistically rather than optimistically, and build order fulfillment into a daily (not every-other-day) routine.
Cancellation rate
The percentage of orders you cancel, most often because of a stockout. This is why inventory sync accuracy (see How to Prevent and Recover From Stockouts) is really an account-health issue, not just a lost-sale issue.
Valid tracking rate
The percentage of shipments with tracking that's both uploaded and actually scanned by the carrier as picked up. Common failure: uploading a tracking number the moment a label is printed, before the carrier has physically taken the package — if the package sits for a day before pickup, some marketplaces still count this against you if the scan doesn't happen quickly enough.
A worked example: doing the math on a bad week
Say a seller ships 300 orders in a rolling 30-day window and normally runs a 2% late shipment rate (6 late orders). A one-week carrier disruption causes 10 additional late shipments in that window. The trailing rate jumps to (6+10)/300 = 5.3%, above the common 4% threshold. Because the window is rolling, those 10 late orders stay in the calculation for the full window length (commonly 30 days on a rolling basis) even after the carrier issue is resolved — meaning the seller could see the metric stay elevated for weeks after service has actually returned to normal, simply because the bad days haven't aged out of the window yet. This is the most common reason sellers feel like "we fixed it but the dashboard doesn't show it" — the dashboard is telling the truth about the trailing average, just not about today specifically.
Reading the threshold math for cancellation and tracking rates
The same rolling-window logic applies to cancellation rate and valid tracking rate. A useful mental model: divide the platform's stated threshold in half, and treat that halved number as your actual operating target. If the published cancellation-rate threshold is around 2.5%, operate as if your real ceiling is roughly 1.25% — this gives you room to absorb a genuinely unavoidable bad week (a carrier-wide disruption, a supplier stockout you couldn't have forecast) without crossing into enforcement territory.
Edge cases worth knowing about
Marketplace-fulfilled vs. seller-fulfilled orders: on programs where the marketplace itself handles fulfillment (Amazon FBA, Walmart WFS), shipping-reliability metrics are generally excluded from your seller-fulfilled performance calculations, since the marketplace — not you — controls the ship time. Don't assume FBA/WFS orders are silently dragging down your seller-fulfilled metrics; verify in your dashboard's metric definitions which order types are actually included.
Multi-warehouse operations: if orders route to different warehouses with different actual capabilities, a single blended handling-time setting can be wrong for some fraction of your orders no matter what you set it to. Where the platform allows warehouse-specific or SKU-specific handling times, use them rather than one global setting.
Holiday and peak-season volume: a jump in order volume without a proportional jump in fulfillment capacity is one of the most common causes of a seasonal late-shipment spike. Build in extra handling-time buffer or temporary handling-time extensions (many platforms allow a temporary vacation/high-volume handling-time setting) rather than letting the metric absorb the hit.
Building in a buffer
A practical rule: set your internal ship-by deadline at least one day earlier than the marketplace's stated deadline, so ordinary variability (a slow morning, a printer jam, one missed pickup) doesn't turn into a metric hit.
Best practices
- Treat half the published threshold as your real operating ceiling, not the published number itself.
- Set handling time to reflect your actual pick-pack-ship capability under normal conditions, with an added buffer — not your best-case time.
- Use warehouse- or SKU-specific handling times where the platform supports them, instead of one global setting.
- Adjust handling-time settings proactively before a known high-volume period rather than reactively after a metric spike.
Troubleshooting
I fixed the root cause weeks ago but the metric hasn't recovered. Confirm the metric's exact trailing-window length in the platform's documentation — a 90-day window takes up to 90 days to fully "forget" a bad period, even with zero new defects.
My seller-fulfilled late shipment rate looks wrong. Check whether the report is accidentally including or excluding marketplace-fulfilled orders — this is one of the most common sources of confusion when reading a blended shipping report.
A single large batch of orders tanked my rate overnight. A short burst of volume beyond your normal capacity (a promotion, a viral moment) is a common and identifiable cause — review whether your handling-time setting or staffing needs a temporary adjustment before the next similar event.
FAQs
Is it better to have a slightly longer stated handling time or a fast one with occasional misses? A realistic, slightly longer handling time that you consistently meet is almost always better for account health than an aggressive one you occasionally miss — consistency is what these metrics reward.
Do late shipment rate and valid tracking rate use the same trailing window? Not necessarily — check each metric's specific window in your dashboard rather than assuming they match, since platforms sometimes apply different windows to different metrics.
Does a single very late shipment count more than several slightly late ones? Generally these metrics count each late shipment as one event regardless of by how much it was late, though platforms may separately flag chronic severe lateness as a distinct policy concern.