Data Quality in Logistics: The Hidden Cost of Bad Data
Bad data does not just produce bad reports — it produces bad decisions. Here is what poor data quality actually costs logistics companies.
The Margin Number That Was Never Real
Data quality is the problem nobody wants to talk about in logistics. It is unglamorous, technically tedious, and easy to ignore — until the consequences become impossible to dismiss. A CFO reviews margin data showing healthy 14% average margins and approves the quarterly strategy. What the CFO does not know is that 22% of shipments have NULL values in the CarrierCost field, which means the margin calculation excludes the worst-performing shipments entirely. The actual margin is closer to 9%.
This is not a hypothetical scenario. It is the norm across the logistics industry, where data flows through multiple systems — TMS, ERP, carrier portals, email confirmations, manual spreadsheets — each with its own data standards, formats, and quality controls (or lack thereof).
Bad data does not announce itself. It quietly corrupts the analysis, the KPI, and the decision built on top of it. When margins are already thin, a run of decisions made on numbers that were wrong is the kind of mistake a mid-size forwarder does not always survive.
The Four Dimensions of Data Quality
Data quality in logistics manifests across four dimensions, each with distinct symptoms and consequences:
1. Completeness (NULL Rates)
The most common data quality issue is missing values. In a typical freight forwarding dataset, NULL rates vary dramatically by field:
- ShipperRevenue: Usually 2-5% NULL — relatively well-maintained because it directly affects invoicing.
- CarrierCost: Often 8-15% NULL — less consistently captured, especially for intermodal shipments with multiple cost components.
- DueDate: Frequently 15-30% NULL — surprisingly neglected despite being critical for on-time delivery measurement.
- CompletionDate: Can reach 40-60% NULL — the most commonly missing field, because it requires post-delivery data entry that often does not happen.
Consider the implications: if 60% of delivery dates are missing in your dataset, your on-time delivery KPI is based on only 40% of shipments. Is that 40% representative? Almost certainly not — the shipments with completed dates are likely the well-managed ones. The missing 60% probably includes a disproportionate share of problematic deliveries that nobody bothered to update. Your reported 92% OTD rate may actually be 78%.
2. Accuracy (Type Mismatches)
Data accuracy issues occur when values exist but are wrong — often due to data type mismatches or entry errors. Common examples in logistics data:
- Revenue stored as text instead of numeric, causing calculations to fail silently.
- Dates in mixed formats (DD/MM/YYYY vs MM/DD/YYYY) causing sorting and comparison errors.
- Carrier names with inconsistent spelling ("DHL Express", "DHL express", "DHL-Express") creating phantom carriers in analysis.
- Route codes mixing IATA airport codes with city names, making lane analysis unreliable.
3. Consistency (Duplicate Records)
Duplicate shipment records are pervasive when data flows from multiple sources. The same shipment may appear in the TMS export, the ERP invoice record, and a carrier confirmation — each with slightly different field values. Without deduplication, your shipment volume is inflated, your revenue is double-counted, and your per-shipment metrics are diluted.
A common pattern: a freight forwarder reports 12,000 monthly shipments based on their database. After deduplication, the actual count is 10,200 — 15% of records were duplicates. Every volume-based metric was wrong by 15%.
4. Timeliness (Stale Data)
Data that was accurate when entered becomes stale when not updated. Shipment status fields showing "In Transit" for shipments that were delivered weeks ago. Cost fields reflecting estimated rates rather than actual invoiced amounts. The gap between estimated and actual data creates a parallel reality where your analytics show one picture and your bank account tells a different story.
Proof, not a pilot
Put this to work on your own operational data.
No integration project. No black box.
Start a 90-Day Proof of ValueThe Real Cost: Bad Decisions from Bad Data
The cost of bad data is not the cost of fixing it — it is the cost of the decisions made while it was unfixed. The scenarios below are illustrative, but each maps to a failure mode that recurs across the industry:
- Carrier negotiation failure: A logistics company enters rate negotiations citing a carrier's 94% on-time rate from their data. The carrier's own data shows 88%. The discrepancy exists because 30% of late deliveries in the company's dataset had NULL completion dates and were excluded from the OTD calculation. The company loses negotiation credibility and pays higher rates.
- False profitability: A route appears profitable at 16% margin because 20% of shipments on that route have NULL CarrierCost fields. When the costs are reconciled, the actual margin is 4% — below the company's minimum threshold. The route was kept active for six months based on false data, costing approximately €45,000 in avoidable margin erosion.
- Missed SLA penalties: A customer contract includes SLA penalties for OTD below 90%. The company's dashboard shows 91% OTD, so no action is taken. But when the customer audits using their own data (which includes the shipments the company failed to track), the actual OTD is 84%. The resulting penalty and relationship damage far exceed what proactive data quality management would have cost.
Fixing It Without Boiling the Ocean
Data quality improvement in logistics follows a predictable path:
- Measure first: Before fixing anything, quantify the problem. Calculate NULL rates per field, identify duplicate percentages, and catalog type mismatches. You cannot improve what you have not measured.
- Prioritize by impact: Not all data quality issues are equal. Focus first on fields that directly affect financial calculations (Revenue, Cost) and customer-facing metrics (OTD). Fixing these fields first delivers the highest ROI.
- Automate detection: Manual data quality checks are unsustainable. Implement automated auditing that runs on every data upload, flagging NULL rates above threshold, detecting duplicates, and validating data types before the data enters your analytics pipeline.
- Set targets: Establish minimum data quality standards: 95% completeness on financial fields, 99% on shipment identifiers, zero tolerance for duplicate records. Track progress weekly.
Data quality is not a one-time project. It is a continuous discipline — one that separates companies making decisions on solid ground from those building strategy on sand.
Put this to work on your own operational data.
Start with one lane, one workflow, one decision. Measure impact. Expand when value is proven.
No integration project. No black box.
Written by
Berna Bulgurcu
Co-founder & CEO, Syntask
The Syntask team writes about operational decision intelligence for logistics — turning the data teams already have into prioritized, evidence-backed decisions.
Topics
- Data Quality
- Deep Dive
- Cost Reduction