Two Inspectors, Two Verdicts: Why Your QC Numbers Keep Moving
๐ Two visits, two verdicts, one lot โ and now nobody trusts either number. It happens on repeat orders more than anywhere else: the factory's own QC reports 1.
๐ Two visits, two verdicts, one lot โ and now nobody trusts either number. It happens on repeat orders more than anywhere else: the factory's own QC reports 1.1% and passes, your third-party visit reports 3.2% and fails, and the argument eats two weeks of the shipping window. Most of the time the goods are not the problem. The sampling plan and the test method are โ and both are fixable in writing, which is what turns an Amazon FBA inspection into a decision instead of a debate. It costs from $169 per man-day.
๐งฎ The same lot can pass on Tuesday and fail on Friday
Sampling is a lottery with known odds. At AQL 2.5, Level II, a 2,400-unit lot is sampled at n=125 with an accept number of 7: find 7 defects and it passes, find 8 and it fails. A lot that genuinely contains 3% defective units will return 7 or fewer defects in roughly half the samples drawn from it. Nothing changed at the factory; the sample changed. Shrink the sample to n=32 โ what an in-house check often uses โ and the accept number falls to 2, so the same 3% lot now fails most of the time. Before comparing two reports, compare the two sampling plans. If one was Level I and the other Level II, you are not looking at a quality trend, you are looking at arithmetic.
๐ Five sources of wobble that are not supplier quality
Sample size and level. Record n, the level and the code letter in every report header. Measurement method. "Powers on" and "runs 60 seconds under load" are different tests with different pass rates. Instrument and zero. A caliper 0.02 mm out, or a scale never zeroed, drags every reading in one direction. Test duration. A cold-start unit and a run-in unit are different products: bearings free up, capacitors form, firmware finishes booting. The crew. Lighting, shift and how many units the inspector has already handled that day all move the borderline visual calls.
๐งพ Make two reports comparable: four lines in the order
Write these into every inspection and every repeat visit. One: the same AQL and level โ AQL 2.5, Level II, single sampling, normal severity. Two: the same gauge and the same named measuring points on the part. Three: the same functional list in the same order with the same pass line. Four: one lot per line in your tracking sheet, normalised per 100 units so lots of different sizes read side by side. Reports that follow those four lines can be compared. Reports that do not are opinions with a logo on them.
๐งช The split-sample repeat test
The cheapest way to find out whether your verdicts repeat: keep the sample. Take 10 units from a lot, have one inspector measure and label them, then have a second inspector re-measure the same 10 units blind two hours later. Compare the two sets of readings. If the dimensions agree inside your tolerance and the borderline calls land the same way, your procedure repeats and any gap between visits is real information about the goods. If they do not agree, you have found a measurement problem worth more than the lot in front of you. We run this on request for importers building a monthly defect trend who need the numbers to mean something.
๐ซ What better method cannot fix
Repeatability does not rescue a bad lot, and it does not catch a supplier who changed material quietly. A genuinely borderline lot โ one sitting on the accept number โ will always be borderline, and the answer there is not a third opinion, it is a decision: rework, accept with a negotiated adjustment, or reject. Write down which one and who pays. The method only tells you which of those three you are actually choosing.
๐ The variance table we attach to the report
| Source of wobble | What it moves | What we record |
|---|---|---|
| Sample size and level | How many defects a lot can show and still pass | n, level, code letter, accept and reject numbers |
| Measurement method | Which failures exist at all | The named test, its duration and its pass line |
| Gauge calibration | Every dimension, in one direction | Instrument ID and the check-weight or zero reading |
| Run-in vs cold start | Motor, battery and firmware behaviour | Warm-up time before the functional sequence |
| Crew and lighting | Borderline visual calls | Shift, inspector initials and photos with the reading in frame |
๐ฐ A 2,600-unit order where the argument cost more than the visit
Case: 2,600 cordless electric screwdrivers at $14.80 FOB = $38,480. The factory's in-house QC reported 1.1% majors from a 32-unit check and pushed to ship. Our visit at AQL 2.5 Level II, n=125, found 9 majors โ 7.2% โ all in a single failure mode: the motor stalled under a 60-second load test that the in-house check never ran. The factory disputed the result, so we repeated the visit with both parties' gauges on the same bench and a fresh random sample: 2.6% majors, same failure mode. The extra man-days cost $338. What that settled was a written rework plan for 74 units and a shipped lot with a verified defect rate, instead of roughly 68 escapes at $39.99 retail plus removal fees plus a replacement order flown in. Both reports came back typically ~24 hours after the visit, so the shipping window survived. After 2,000+ inspections the pattern is boring and consistent: the disagreement is almost never about the goods, and it is almost always settled by writing down the method.
Put the four comparability lines in your order, keep the sample, and check the accept numbers on our AQL calculator before you argue about a percentage. Man-day rates are on the pricing page, and we will scope a repeat visit for you from the contact page.
โ FAQs
Why do two inspectors get different defect rates on the same lot?
Usually because they sampled differently, not because the lot changed. Different sample sizes carry different accept numbers, and a lot near the borderline will pass one sample and fail another by pure randomness. Write down n, the AQL level and the test method so the two reports can be compared at all.
What is a split-sample repeat test?
Keep 10 units from a lot, let one inspector measure and label them, then have a second inspector re-measure the same units blind a couple of hours later. Agreement inside tolerance means your procedure repeats; disagreement is a measurement problem you want to find before it corrupts a year of trend data.
Is a supplier's in-house QC report worthless?
No, but it is a different instrument. In-house checks are typically smaller samples with a narrower test list, which is useful for catching gross problems as they happen. It becomes misleading when it is quoted as a defect rate comparable to an independent AQL 2.5 Level II inspection.
How do I compare lots of different sizes in one trend?
One lot per line and normalise per 100 units: majors per 100 pieces, criticals as raw counts, and time-to-fix in days. That way a 500-unit trial run and a 10,000-unit production lot sit on the same chart without one skewing the other.
Does a repeat inspection cost extra?
Every visit is a man-day from $169, including a repeat visit after a disputed result. Framed against a shipping window, a $338 repeat with both gauges on one bench is usually cheaper than two weeks of argument followed by an unverified shipment.
Frequently asked questions
Why do two inspectors get different defect rates on the same lot?
Usually because they sampled differently, not because the lot changed. Different sample sizes carry different accept numbers, and a lot near the borderline will pass one sample and fail another by pure randomness. Write down n, the AQL level and the test method so the two reports can be compared at all.
What is a split-sample repeat test?
Keep 10 units from a lot, let one inspector measure and label them, then have a second inspector re-measure the same units blind a couple of hours later. Agreement inside tolerance means your procedure repeats; disagreement is a measurement problem you want to find before it corrupts a year of trend data.
Is a supplier's in-house QC report worthless?
No, but it is a different instrument. In-house checks are typically smaller samples with a narrower test list, which is useful for catching gross problems as they happen. It becomes misleading when it is quoted as a defect rate comparable to an independent AQL 2.5 Level II inspection.
How do I compare lots of different sizes in one trend?
One lot per line and normalise per 100 units: majors per 100 pieces, criticals as raw counts, and time-to-fix in days. That way a 500-unit trial run and a 10,000-unit production lot sit on the same chart without one skewing the other.
Does a repeat inspection cost extra?
Every visit is a man-day from $169, including a repeat visit after a disputed result. Framed against a shipping window, a $338 repeat with both gauges on one bench is usually cheaper than two weeks of argument followed by an unverified shipment.