Two Inspectors, Two Verdicts: Why Your QC Numbers Keep Moving

๐Ÿ“Š Two visits, two verdicts, one lot โ€” and now nobody trusts either number. It happens on repeat orders more than anywhere else: the factory's own QC reports 1.

๐Ÿ“Š Two visits, two verdicts, one lot โ€” and now nobody trusts either number. It happens on repeat orders more than anywhere else: the factory's own QC reports 1.1% and passes, your third-party visit reports 3.2% and fails, and the argument eats two weeks of the shipping window. Most of the time the goods are not the problem. The sampling plan and the test method are โ€” and both are fixable in writing, which is what turns an Amazon FBA inspection into a decision instead of a debate. It costs from $169 per man-day.

๐Ÿงฎ The same lot can pass on Tuesday and fail on Friday

Sampling is a lottery with known odds. At AQL 2.5, Level II, a 2,400-unit lot is sampled at n=125 with an accept number of 7: find 7 defects and it passes, find 8 and it fails. A lot that genuinely contains 3% defective units will return 7 or fewer defects in roughly half the samples drawn from it. Nothing changed at the factory; the sample changed. Shrink the sample to n=32 โ€” what an in-house check often uses โ€” and the accept number falls to 2, so the same 3% lot now fails most of the time. Before comparing two reports, compare the two sampling plans. If one was Level I and the other Level II, you are not looking at a quality trend, you are looking at arithmetic.

๐Ÿ” Five sources of wobble that are not supplier quality

Sample size and level. Record n, the level and the code letter in every report header. Measurement method. "Powers on" and "runs 60 seconds under load" are different tests with different pass rates. Instrument and zero. A caliper 0.02 mm out, or a scale never zeroed, drags every reading in one direction. Test duration. A cold-start unit and a run-in unit are different products: bearings free up, capacitors form, firmware finishes booting. The crew. Lighting, shift and how many units the inspector has already handled that day all move the borderline visual calls.

๐Ÿงพ Make two reports comparable: four lines in the order

Write these into every inspection and every repeat visit. One: the same AQL and level โ€” AQL 2.5, Level II, single sampling, normal severity. Two: the same gauge and the same named measuring points on the part. Three: the same functional list in the same order with the same pass line. Four: one lot per line in your tracking sheet, normalised per 100 units so lots of different sizes read side by side. Reports that follow those four lines can be compared. Reports that do not are opinions with a logo on them.

๐Ÿงช The split-sample repeat test

The cheapest way to find out whether your verdicts repeat: keep the sample. Take 10 units from a lot, have one inspector measure and label them, then have a second inspector re-measure the same 10 units blind two hours later. Compare the two sets of readings. If the dimensions agree inside your tolerance and the borderline calls land the same way, your procedure repeats and any gap between visits is real information about the goods. If they do not agree, you have found a measurement problem worth more than the lot in front of you. We run this on request for importers building a monthly defect trend who need the numbers to mean something.

๐Ÿšซ What better method cannot fix

Repeatability does not rescue a bad lot, and it does not catch a supplier who changed material quietly. A genuinely borderline lot โ€” one sitting on the accept number โ€” will always be borderline, and the answer there is not a third opinion, it is a decision: rework, accept with a negotiated adjustment, or reject. Write down which one and who pays. The method only tells you which of those three you are actually choosing.

๐Ÿ“‹ The variance table we attach to the report

Source of wobbleWhat it movesWhat we record
Sample size and levelHow many defects a lot can show and still passn, level, code letter, accept and reject numbers
Measurement methodWhich failures exist at allThe named test, its duration and its pass line
Gauge calibrationEvery dimension, in one directionInstrument ID and the check-weight or zero reading
Run-in vs cold startMotor, battery and firmware behaviourWarm-up time before the functional sequence
Crew and lightingBorderline visual callsShift, inspector initials and photos with the reading in frame

๐Ÿ’ฐ A 2,600-unit order where the argument cost more than the visit

Case: 2,600 cordless electric screwdrivers at $14.80 FOB = $38,480. The factory's in-house QC reported 1.1% majors from a 32-unit check and pushed to ship. Our visit at AQL 2.5 Level II, n=125, found 9 majors โ€” 7.2% โ€” all in a single failure mode: the motor stalled under a 60-second load test that the in-house check never ran. The factory disputed the result, so we repeated the visit with both parties' gauges on the same bench and a fresh random sample: 2.6% majors, same failure mode. The extra man-days cost $338. What that settled was a written rework plan for 74 units and a shipped lot with a verified defect rate, instead of roughly 68 escapes at $39.99 retail plus removal fees plus a replacement order flown in. Both reports came back typically ~24 hours after the visit, so the shipping window survived. After 2,000+ inspections the pattern is boring and consistent: the disagreement is almost never about the goods, and it is almost always settled by writing down the method.

Put the four comparability lines in your order, keep the sample, and check the accept numbers on our AQL calculator before you argue about a percentage. Man-day rates are on the pricing page, and we will scope a repeat visit for you from the contact page.

โ“ FAQs

Why do two inspectors get different defect rates on the same lot?

Usually because they sampled differently, not because the lot changed. Different sample sizes carry different accept numbers, and a lot near the borderline will pass one sample and fail another by pure randomness. Write down n, the AQL level and the test method so the two reports can be compared at all.

What is a split-sample repeat test?

Keep 10 units from a lot, let one inspector measure and label them, then have a second inspector re-measure the same units blind a couple of hours later. Agreement inside tolerance means your procedure repeats; disagreement is a measurement problem you want to find before it corrupts a year of trend data.

Is a supplier's in-house QC report worthless?

No, but it is a different instrument. In-house checks are typically smaller samples with a narrower test list, which is useful for catching gross problems as they happen. It becomes misleading when it is quoted as a defect rate comparable to an independent AQL 2.5 Level II inspection.

How do I compare lots of different sizes in one trend?

One lot per line and normalise per 100 units: majors per 100 pieces, criticals as raw counts, and time-to-fix in days. That way a 500-unit trial run and a 10,000-unit production lot sit on the same chart without one skewing the other.

Does a repeat inspection cost extra?

Every visit is a man-day from $169, including a repeat visit after a disputed result. Framed against a shipping window, a $338 repeat with both gauges on one bench is usually cheaper than two weeks of argument followed by an unverified shipment.

Frequently asked questions

Why do two inspectors get different defect rates on the same lot?

Usually because they sampled differently, not because the lot changed. Different sample sizes carry different accept numbers, and a lot near the borderline will pass one sample and fail another by pure randomness. Write down n, the AQL level and the test method so the two reports can be compared at all.

What is a split-sample repeat test?

Keep 10 units from a lot, let one inspector measure and label them, then have a second inspector re-measure the same units blind a couple of hours later. Agreement inside tolerance means your procedure repeats; disagreement is a measurement problem you want to find before it corrupts a year of trend data.

Is a supplier's in-house QC report worthless?

No, but it is a different instrument. In-house checks are typically smaller samples with a narrower test list, which is useful for catching gross problems as they happen. It becomes misleading when it is quoted as a defect rate comparable to an independent AQL 2.5 Level II inspection.

How do I compare lots of different sizes in one trend?

One lot per line and normalise per 100 units: majors per 100 pieces, criticals as raw counts, and time-to-fix in days. That way a 500-unit trial run and a 10,000-unit production lot sit on the same chart without one skewing the other.

Does a repeat inspection cost extra?

Every visit is a man-day from $169, including a repeat visit after a disputed result. Framed against a shipping window, a $338 repeat with both gauges on one bench is usually cheaper than two weeks of argument followed by an unverified shipment.