K-01
OTIF — On Time In Full
OTIF is the joint test: did the customer get what they ordered, when they were told, complete. It is also the most configurable number in supply chain reporting — six switch settings decide whether the same operation scores 95% or 60%. This entry defines the metric and makes every switch explicit.
There is a reference build of this metric: the Service & Delivery Performance dashboard in the showcase runs this definition over frozen, simulated data — not a live client system. Its Definitions view runs every switch on this page over one frozen order book — nine defensible OTIF numbers from the same orders.
Definitions below are current to the date above — verify against current SAP, Oracle, Microsoft, Infor, and Databricks documentation before you build.
Formula
Formula
In words
OTIF = eligible lines on time (arrival on or before the first-confirmed promise) and in full (first attempt) ÷ eligible lines
Where
- an eligible order line
- eligible order lines (not fully cancelled, promise date due)
- arrival date of line ℓ
- first-confirmed promise date
- quantity shipped on the first attempt
- ordered quantity
- 1 when the bracketed test holds, else 0
This is the practice default; every other OTIF in the wild is a setting of the six switches below.
What OTIF measures
On Time In Full is the share of eligible order lines that pass the date test and the quantity test jointly on the same line: the goods arrived on or before the first-confirmed promise date, and the first shipment carried the full ordered quantity. Every other reading of OTIF is a setting of six declared switches: grain, date basis, ship or arrival, first attempt or cumulative, tolerance, and exclusions.
One order line, three questions: did the goods move on time (against a date somebody committed to), in full (the complete quantity, not most of it), and does the line count at all (or was it canceled out of the denominator). A line passes OTIF only when it passes the date test and the quantity test together — jointly, on the same line, not as two separate percentages that get multiplied later.
Also answers to OTIF · DIFOT (delivered in full, on time) · OTIF-D — same joint test, different acronym communities.
The core thing to internalize: there is no standard OTIF. Every definition in the wild is a configuration of six decisions — grain, date basis, ship vs arrival, first attempt vs cumulative, tolerance, and exclusions. Change one switch and the same month of shipments moves by double digits. That makes the number meaningless in isolation and only comparable to itself, computed identically over time. It also makes “our OTIF is 94%” an incomplete sentence until the six settings are stated.
OTIF vs fill rate vs OTD
Three metrics get conflated under this banner, and they are not the same measure:
- Fill rate tests quantity only — shipped over ordered, usually cumulative, with no time test at all. A line that ships complete three weeks late has a 100% fill rate.
- OTD (on-time delivery) tests time only. A line that ships one unit of a hundred on the promised date is on time.
- OTIF tests both, jointly, at a declared grain. Under matching definitions it can never exceed either component: OTIF ≤ min(fill rate, OTD).
The trap. Dashboards that compute on_time_pct × in_full_pct are assuming the two failure modes are independent. They never are — late lines are disproportionately also short lines — so the product misstates the joint metric, and there is no correction factor that fixes it. Compute the joint pass flag per line and aggregate the flag. If a scorecard shows OTIF above either of its own components, the joint test was never computed.
And one more name collision. “Service level” is used in the wild for OTIF, for fill rate, and for cycle service level — the last of which is a different object entirely: an event count of replenishment cycles that ended without a stockout, not a ratio over ordered quantity. A bare “service level” in a contract or a planning system is not yet a metric; ask which construction it scores.
The six-switch switchboard
Every OTIF definition — yours, your customer’s, the one in the board deck — is a setting of these six switches. Each links to its own section below.
| Switch | Settings | Practice default |
|---|---|---|
| Grain | Order · Line · Unit / case-weighted | Line grain — compute there, then roll up to order or unit views from the same fact. |
| Date basis | Customer-requested · First-confirmed / original promise · Current promise · Must-arrive-by (MABD) | First-confirmed / original promise — and report requested-vs-confirmed gap separately. |
| Ship vs arrival | Ship basis · Arrival basis | Arrival basis wherever the customer measures arrival; keep a ship-basis cut for warehouse accountability. |
| First attempt vs cumulative | First attempt · Cumulative by cutoff · Eventually complete | First attempt, with the split-shipment rate reported alongside. |
| Tolerance | Zero-early / zero-late · Late-side grace (e.g. −3/+0 days) · Delivery window, early fails too | On or before the first-confirmed date — no late-side tolerance, early arrivals accepted — unless a contract or scorecard defines the window. |
| Exclusions | Cancels excluded · Returns measured separately · Rebooked dates honored | Exclude cancels, measure returns separately, and if rebooked dates are honored, publish the rebooking rate next to OTIF. |
The practice default, stated once: line grain, first-confirmed date, arrival basis where the customer measures arrival, first-attempt, arrival on or before the promise — no late-side tolerance, early accepted — unless contracted, cancels excluded and returns tracked separately. And one override that beats all of it: if a customer scores you, their scorecard is the definition — model their configuration first and argue with it second.
Switch 1 — Grain: order, line, or unit
The grain decides what one “pass” means. Ten orders, forty lines, five hundred units; four orders carry six problem lines totalling twenty units. The same month scores 60% at order grain (six of ten orders fully clean), 85% at line grain (thirty-four of forty lines pass), and 96% unit-weighted (four hundred eighty of five hundred units on passing lines). Nobody lied — the switch moved.
- lowers the scoreOrder — every line must pass for the order to pass — strictest, punishes big orders
- dependsLine — the modeling default; every ERP's sales table is line-grain
- raises the scoreUnit / case-weighted — what retail scorecards effectively use; smooths order-size mix
Practice default Line grain — compute there, then roll up to order or unit views from the same fact.
Order grain is the customer’s experience of the order — and it structurally punishes large orders, since one short line fails twenty clean ones. Unit or case weighting is what retail compliance programs effectively apply (charges scale with the non-compliant volume), and it smooths order-size mix out of trend lines. Line grain is where the data lives: the sales line is the grain of VBAP, F4211, SalesLine, and OOLINE alike — compute there and both other grains are roll-ups of the same fact.
Switch 2 — Date basis: which promise counts
“On time” against what date? Four candidates, and they tell different stories: the date the customer asked for, the date you first committed to, the date you currently promise (after every reschedule), and the date the customer requires goods to arrive by (MABD — theirs, not yours).
- lowers the scoreCustomer-requested — scores the ask, including asks you never agreed to
- dependsFirst-confirmed / original promise — scores the promise as first made — the honest default
- raises the scoreCurrent promise — every reschedule resets the clock — lateness gets laundered
- dependsMust-arrive-by (MABD) — customer-imposed; what retail compliance programs score
Practice default First-confirmed / original promise — and report requested-vs-confirmed gap separately.
JD Edwards is the clearest teaching example because it keeps all three of its dates side by side on the sales line: SDDRQJ (the customer’s requested date), SDOPDJ (the original promised date — JDE keeps it precisely for on-time measurement), and SDPPDJ (the current promise, updated as the order reschedules). Measure against SDPPDJ and every reschedule launders lateness: the order that slipped three times arrives “on time” against its third promise.
The other systems split the same idea across tables. SAP: VBAK.VDATU is the requested date; the first confirmed date lives on the VBEP schedule lines — the earliest schedule line with a confirmed quantity (BMENG > 0) carries the confirmed date in EDATU, with WMENG as the ordered quantity beside it. D365 makes the switch a column swap: SalesLine.ShippingDateRequested vs ShippingDateConfirmed on the ship side, ReceiptDateRequested (line) vs SalesTable.ReceiptDateConfirmed (header) on the arrival side. M3 splits by level: OOHEAD.OARLDT is the header-level requested date, OOLINE.OBDWDT the line’s planned delivery date.
Switch 3 — Ship date or arrival date
“On time” measured where? Ship basis compares the date goods left the dock; arrival basis compares the date they reached the customer. The gap between the two is transit time — and retail compliance fines live in that gap: a shipment that left on schedule and sat two days at a carrier hub fails the customer’s arrival test while passing your ship test.
- raises the scoreShip basis — what the warehouse controls; ignores transit entirely
- lowers the scoreArrival basis — what the customer experiences and scores; needs POD or delivery confirmation
Practice default Arrival basis wherever the customer measures arrival; keep a ship-basis cut for warehouse accountability.
Where each lives. Ship basis is well-recorded everywhere: SAP posts actual goods issue in LIKP.WADAT_IST; JDE writes SDADDJ at ship confirm; D365’s warehouse-true timestamp is WHSShipmentTable.ShipConfirmUTCDateTime; M3 issues stock against the delivery. Arrival basis is where ERPs get thin, because arrival happens off your systems:
- SAP: LIKP.
LFDATis the planned customer-arrival date — the promise side of an arrival test. The actual side needs proof-of-delivery data (POD confirmation or carrier events), which base SD does not capture for you. - JDE: F4201.
SHADLJis the header-level actual delivery date — note the grain mismatch against line-level measurement, and that it stays blank unless transportation or delivery confirmation actually populates it. The richer path is F4941 shipment routing steps, which carry scheduled vs actual delivery timestamps per leg, joined to order lines through F4942 and anchored on the F4215 shipment header. - D365: SalesTable.
ReceiptDateConfirmedcarries the promised arrival; actual arrival needs POD or carrier integration, same as SAP. - M3: deliveries ride the MHDISH / MHDISL structure; this reference’s curated field set carries no verified actual-departure or arrival timestamp, so whichever field you adopt as the ship event is a proxy — name it on the dashboard rather than implying a measured departure.
If the customer scores arrival and you can only measure ship, say so on the dashboard — an unlabeled ship-basis OTIF quietly overstates performance by the transit time.
Switch 4 — First attempt or cumulative
When a line ships in pieces, which piece gets judged? Splits are not an edge case — every ERP models them as first-class records: LIPS writes one row per delivery split against the same order line (VGBEL/VGPOS), JDE spawns suffix records (SDSFXO) and tracks the backordered remainder in SDSOBK, M3 subdivides lines with OBPOSX and writes one MHDISL row per delivery, and D365 posts multiple packing-slip lines against one sales line.
- lowers the scoreFirst attempt — the first shipment must complete the line — the discipline metric
- dependsCumulative by cutoff — the line must complete by a declared date — the honest middle
- raises the scoreEventually complete — no cutoff; backorders that ship months late still count
Practice default First attempt, with the split-shipment rate reported alongside.
First-attempt is the discipline metric — it asks whether the promise was kept as made. Eventually-complete is nearly information-free: with enough months, almost everything eventually ships. Cumulative-by-cutoff (“complete within the tolerance window”) is the honest middle for operational review. Whichever you pick, publish the split-shipment rate next to OTIF: the share of order lines that shipped in more than one piece is what tells a reader whether the gap between the first-attempt and cumulative readings is a rounding difference or the whole story.
Switch 5 — Tolerance windows
How close is close enough? A zero-early/zero-late window is the strictest reading of the promise. Grace windows (ship within three days early, zero days late) are common — and every grace day quietly raises the score, so the window belongs in the metric’s label, not in a footnote. Delivery-window programs flip the intuition: under a must-arrive-by regime, early is also a failure, because the receiving DC plans labor and dock capacity against the window.
- lowers the scoreZero-early / zero-late — on the promised date or fail — the strictest window
- raises the scoreLate-side grace (e.g. −3/+0 days) — every grace day quietly raises the score
- dependsDelivery window, early fails too — how MABD programs work — DCs plan capacity, so early is non-compliant
Practice default On or before the first-confirmed date — no late-side tolerance, early arrivals accepted — unless a contract or scorecard defines the window.
Two mechanical traps once windows get tight. Time of day: a date-grain comparison calls 11:59 PM on the due date on-time — fine, as long as it’s deliberate. Timezone: D365 stores datetimes in UTC, so WHSShipmentTable.ShipConfirmUTCDateTime must go through from_utc_timestamp() before its date is taken — otherwise every shipment confirmed after 4–5 PM Pacific lands on tomorrow’s date and fails a zero-window test it actually passed (see the UTC datetimes quirk).
Switch 6 — Exclusions
What leaves the denominator decides what the number can hide. Three populations to rule on:
- Canceled lines. Standard practice excludes them — a canceled ask is no longer a promise — but track the cancel rate beside OTIF, because cancellation is also how demand quietly exits a struggling order book. The markers: SAP VBAP.
ABGRU(rejection reason set = rejected), JDE’s canceled quantitySDSOCNwith its cancel dateSDCNDJ, D365SalesStatus = 4(Canceled), M3’sOBORST = '99'— completed without delivery. - Returns. A line that delivered on time, in full, and came back a week later passed OTIF — the fulfillment promise was kept; the product failed some other promise. Exclude returns from OTIF and measure the return rate as its own metric.
- Rebooked dates — the integrity switch. When a customer agrees to move a slipping order’s date, does the line get scored against the new date? Honor renegotiated dates and lateness converts into paperwork: service looks stable while promises quietly move. If the business decides to honor them (sometimes legitimately — the customer asked for the move), publish the rebooking rate next to OTIF so the movement is visible.
JDE hands you the forensic trail for that last switch: F42199 writes a ledger row every time an order line changes — including every promise-date edit — and F4209 records which orders sat on hold, who released them, and when. Date-change frequency is a queryable fact, not a suspicion.
Retail compliance programs
The reason none of this is academic: large North American retailers run OTIF as a compliance program with money attached. Walmart’s program is the best-documented public example — suppliers are scored against a must-arrive-by date (MABD) at the purchase-order level, with delivery windows measured in days, early arrivals counted as non-compliant alongside late ones, compliance targets in the high 90s, and charges assessed per non-compliant case as a percentage of the cost of goods. Other large retail and grocery programs follow the same shape with different constants.
Read that against the switchboard: a retail program is a customer-imposed configuration of all six switches — their grain (PO and case), their date basis (MABD), their measurement point (arrival), their tolerance (a window where early fails), their exclusions. That is why no standard OTIF exists: any supplier serving two scorecard customers is already running two definitions, plus their own. The architecture answer is in the landing pattern: store components, compute every configuration as a view.
How this lands: bronze → silver → gold
The switchboard has an architectural consequence: never store the ratio — store the components.
- Bronze: the source tables as extracted — all companies, dates still in their native encodings (DATS strings, Julian integers, numeric YYYYMMDD, UTC datetimes), CDC flags kept.
- Silver: one company, encodings resolved exactly once — zero-dates and sentinels become real NULLs, Julian and YYYYMMDD become DATE, UTC becomes local. Conformed order-line and delivery-line entities per source.
- Gold: one
fct_order_line_fulfillmentat order-line grain carrying the components: requested / first-confirmed / current promise dates, first-ship and last-ship (and, where captured, arrival) dates, ordered / shipped / canceled quantities, the delivery count, and a first-attempt-complete flag — plus the usual conformed customer, product, and date dimensions. No OTIF column anywhere.
Every OTIF variant then becomes a SELECT: your line grain, first-confirmed, zero-window view; the customer’s PO-grain MABD-window view; last year’s definition for continuity — all reading the same fact. When the definition argument arrives (it always arrives), switching definitions is a view change, not a pipeline re-engineering. That one fact serves fill rate and OTD on the way, since both are projections of the same components. This is the medallion / Kimball pattern the rest of the library builds toward — the same shape as the deliveries guide’s landing pattern, extended with the promise-date components OTIF needs. And it is the fact the service & delivery reference build is built on.
Once the fact exists, the switchboard has one more form: a Unity Catalog metric view, where each switch setting is a named expression with its definition attached, and Databricks Genie answers “what was OTIF last quarter” from the definition you declared rather than from whichever columns looked plausible.