Gauge R&R explained: how to run, calculate and read a study

What gauge R&R measures, how to run a 10 × 3 × 3 study, a worked example with published AIAG data (ANOVA and average & range), how to read %GRR and ndc, and what to do when it fails.

Updated 17 primary sourcesHow we researchCite

On this page
  1. What gauge R&R measures
  2. Repeatability vs reproducibility
  3. How to run a gauge R&R study
  4. How to calculate gauge R&R: a worked example
  5. How to read the results: %GRR, %Contribution and %Tolerance
  6. What to do when a gauge R&R fails
  7. Gauge R&R vs calibration vs measurement uncertainty
  8. Other gauge studies
  9. Gauge R&R study checklist
  10. FAQ
  11. Sources

A gauge R&R (gage repeatability and reproducibility) study estimates how much of the variation you see in your measurements comes from the measurement system itself. It splits that variation into repeatability (the same operator measuring the same part several times) and reproducibility (different operators measuring the same parts), and compares the total with the variation between parts or with the tolerance. The usual study has 10 parts, 3 operators and 2 or 3 trials. Under the AIAG criteria, a %GRR below 10 % is generally acceptable, 10–30 % may be acceptable, above 30 % is not, and the study should resolve at least 5 distinct categories.

What gauge R&R measures

Every measurement you record is the part’s real value plus some measurement error. If the error is large compared with how much parts really differ, you can’t separate good parts from bad ones, and your SPC charts mostly show gauge noise. A gauge R&R study puts numbers on that error. In the average and range method used by the AIAG manual it reports five quantities, each as a standard deviation (AIAG MSA 4th ed., as reproduced by SPC for Excel, 2015):

Term Short name What it captures
Repeatability EV, equipment variation Spread when one operator measures the same part several times with the same gauge
Reproducibility AV, appraiser variation Spread between the operators’ averages (plus, in ANOVA, any operator × part interaction)
Gauge R&R GRR The measurement system’s total variation: GRR² = EV² + AV²
Part variation PV How much the parts in the study really differ
Total variation TV Everything you observed: TV² = GRR² + PV²

Variances add; standard deviations don’t. That one fact explains most of the confusion around gauge R&R percentages, and we’ll come back to it in the worked example.

The words come from metrology, with a twist. The international vocabulary (VIM) defines repeatability as precision under the same procedure, operators, measuring system, conditions and location over a short period, and reproducibility as precision across different locations, operators and measuring systems (VIM 2.20–2.25). Precision itself is the closeness of agreement between repeated measurements, usually expressed as a standard deviation, variance or coefficient of variation, and it is not the same thing as accuracy (VIM 2.15). In an automotive gauge R&R, “reproducibility” is narrower: usually only the operator changes (SPC for Excel, 2015; Minitab, methods and formulas).

Repeatability vs reproducibility

A quick example. One inspector measures the same shaft diameter three times with a bore gauge and gets 25.012, 25.018 and 25.009 mm: that scatter is repeatability. Three inspectors each measure it three times and their averages are 25.013, 25.021 and 25.006 mm: the difference between those averages is reproducibility.

They point to different fixes:

Repeatability (EV) Reproducibility (AV)
Question it answers Does the gauge give the same answer twice? Do different people get the same answer?
Typical causes Gauge wear or play, fixture, resolution, part form (an out-of-round shaft), environment Unclear method, different technique or force, training, where on the part each person measures
Where to look first Equipment and setup Procedure and people

How to run a gauge R&R study

The common design is a crossed study: every operator measures every part in every trial. The example in the AIAG manual uses 10 parts, 3 operators and 3 trials, 90 readings in total (SPC for Excel, 2015; Minitab example). Two or three trials and two or three operators are common: those are the cases covered by the published average and range constants (Ermer, 2006, table 1). The manual’s example is a design, not a rule; your customer may specify its own.

10 parts×3 operators×3 trials=90 readings
Part12345678910
Op. A
Op. B
Op. C

= one reading. Each cell is one operator measuring one part 3 times, never back to back.

  1. Trial 136810419527
  2. Trial 246271081359
  3. Trial 357410283691

Example random order for one operator. Each operator gets a fresh order in every trial, and nobody sees the part number or earlier readings.

A crossed study: 3 operators each measure 10 parts 3 times. In each trial round, every operator gets the parts in a new random order. Source: design of the AIAG MSA 4th ed. example (as published by SPC for Excel and Minitab); random orders generated by Acribi.
  1. Calibrate and check the gauge first

    Fix bias and calibration problems before you study precision (Ermer, 2006). A gauge R&R cannot see a gauge that reads consistently wrong: shifting every reading by the same amount leaves all the variance components unchanged.

  2. Choose parts that cover the process range

    Take parts from normal production that represent its expected variation, not 10 consecutive parts and not parts of different types. Number them where the operators can't see the number.

  3. Choose the operators

    Use people who actually do this measurement, with the gauge, fixture and procedure used in production.

  4. Randomize and measure blind

    Each operator measures all parts in a new random order in each trial, without seeing earlier readings. Record values to the full resolution of the gauge.

  5. Analyze

    Use ANOVA (preferred) or the average and range method, and look at the components, not only the headline %GRR.

  6. Decide and document

    Compare with the criteria your customer uses, and record parts, operators, gauge ID, date and conclusions.

    Fails: diagnose before you buy a new gauge

Two design choices matter more than any formula:

  • Part selection. If the parts barely differ, part variation is small, total variation is small, and %GRR (which divides by total variation) looks bad even with an excellent gauge. The reverse is also true: parts far outside the normal range make any gauge look good. This follows directly from how the percentages are calculated (Minitab, methods and formulas).
  • Randomization and blind measurement. Operators who remember a part’s earlier reading, or measure it twice in a row, make repeatability look better than it is (SPC for Excel, 2015).

In a calibration or test lab with few artifacts and operators, the NIST e-Handbook suggests more than two check standards, all operators if there are only a few (or a random sample of more than two from a large group), and measurements on more than two days if there is only one operator (NIST e-Handbook 2.4.2).

How to calculate gauge R&R: a worked example

We’ll use the data of the AIAG MSA example (4th edition) as published independently by SPC for Excel and Minitab: 10 parts, operators A, B and C, 3 trials, and a tolerance of 8 units. All results below were recomputed with our own code and match both publications to the last published decimal (apart from small rounding differences between the tables and the text of the SPC for Excel article itself). You can load the same data in the gauge R&R calculator.

Part A1 A2 A3 B1 B2 B3 C1 C2 C3
1 0.29 0.41 0.64 0.08 0.25 0.07 0.04 −0.11 −0.15
2 −0.56 −0.68 −0.58 −0.47 −1.22 −0.68 −1.38 −1.13 −0.96
3 1.34 1.17 1.27 1.19 0.94 1.34 0.88 1.09 0.67
4 0.47 0.50 0.64 0.01 1.03 0.20 0.14 0.20 0.11
5 −0.80 −0.92 −0.84 −0.56 −1.20 −1.28 −1.46 −1.07 −1.45
6 0.02 −0.11 −0.21 −0.20 0.22 0.06 −0.29 −0.67 −0.49
7 0.59 0.75 0.66 0.47 0.55 0.83 0.02 0.01 0.21
8 −0.31 −0.20 −0.17 −0.63 0.08 −0.34 −0.46 −0.56 −0.49
9 2.26 1.99 2.01 1.80 2.12 2.19 1.77 1.45 1.87
10 −1.36 −1.25 −1.31 −1.68 −1.62 −1.50 −1.49 −1.77 −2.16

Average and range method (by hand)

The average and range method (ARM) turns ranges into standard deviations with constants. In the AIAG 4th edition convention the constants give 1σ values: K1 = 0.5908 for 3 trials, K2 = 0.5231 for 3 operators and K3 = 0.3146 for 10 parts (SPC for Excel, 2015). They are simply 1/d2 and 1/d2* from the statistics of ranges of normal samples, which is why they change with the number of trials, operators and parts (Ermer, 2006).

  1. Average range of the 30 operator-part cells: R̿ = 0.3417 (operators: 0.184, 0.513, 0.328).
  2. Repeatability: EV = R̿ × K1 = 0.3417 × 0.5908 = 0.202.
  3. Operator averages 0.190, 0.068 and −0.254, so their range is X̄diff = 0.4447.
  4. Reproducibility: AV = √((X̄diff × K2)² − EV² / (parts × trials)) = √((0.4447 × 0.5231)² − 0.202² / 30) = 0.230. The subtraction removes the share of repeatability that leaks into the operator averages.
  5. Gauge R&R: GRR = √(EV² + AV²) = 0.306.
  6. Part variation: the part averages range over Rp = 3.511, so PV = Rp × K3 = 1.105.
  7. Total variation: TV = √(GRR² + PV²) = 1.146.
  8. Percentages: %EV = 17.61 %, %AV = 20.04 %, %GRR = 26.68 %, %PV = 96.4 %.

These match the published values (SPC for Excel, 2015, tables 2 and 7).

ANOVA method (preferred)

The analysis of variance (ANOVA) method uses all 90 readings instead of ranges, and it can estimate one more thing the ARM can’t: the operator × part interaction, which shows up when operators disagree on some parts but not on others (SPC for Excel, 2015; Ermer, 2006). It is the method we recommend and the default in our calculator.

For this data the interaction has F = 0.434 and p = 0.974, far above the usual 0.05 cutoff, so it is pooled into repeatability and the model is refitted without it (Minitab example). The 0.05 cutoff is a software setting, not a universal rule (Minitab lets you change it, and some tools keep the interaction in every study), so state the rule you used. Any variance component that comes out negative is set to zero (Minitab, methods and formulas; NIST e-Handbook 2.4.4).

The variance components then give:

% Contribution (of variance) — adds up to 100 %
0 %25 %50 %75 %100 %Repeatability (EV)3.39 %Reproducibility (AV)4.37 %Total gauge R&R7.76 %Part-to-part (PV)92.24 %
% Study variation (of standard deviation) — does not add up to 100 %
10 %30 %AIAG bands for GRR0 %25 %50 %75 %100 %Repeatability (EV)18.42 %Reproducibility (AV)20.90 %Total gauge R&R27.86 %Part-to-part (PV)96.04 %
Show the values as a table
SourceVarianceSD% Contribution% Study var% Tolerance
Repeatability (EV)0.039970.199933.39 %18.42 %14.99 %
Reproducibility (AV)0.051460.226844.37 %20.90 %17.01 %
Total gauge R&R0.091430.302377.76 %27.86 %22.68 %
Part-to-part (PV)1.086451.0423392.24 %96.04 %78.17 %
Total variation (TV)1.177881.08530100.00 %100.00 %81.40 %

ANOVA, interaction pooled (p = 0.974), study variation = 6 × SD, tolerance = 8. Number of distinct categories: 4.

Gauge R&R results for the AIAG example data (ANOVA). The same study reads 7.76 % as a share of variance and 27.86 % as a share of standard deviation. Bands apply to the total gauge R&R row. Source: data from the AIAG MSA 4th ed. example as published by Minitab and SPC for Excel; computed by Acribi (src/lib/grr.ts, validated against both).
Source Std. dev. % Contribution (variance) % Study variation (SD) % Tolerance (tol. = 8)
Repeatability 0.19993 3.39 % 18.42 % 14.99 %
Reproducibility 0.22684 4.37 % 20.90 % 17.01 %
Total gauge R&R 0.30237 7.76 % 27.86 % 22.68 %
Part-to-part 1.04233 92.24 % 96.04 % 78.17 %
Total variation 1.08530 100 % 100 % 81.40 %

Published values: Minitab example; %Contribution also in SPC for Excel (2015).

Same data, three different “answers”

Average and range ANOVA
%GRR of study variation (SD) 26.68 % 27.86 %
%GRR contribution (variance) 7.13 % 7.76 %
Number of distinct categories 5 4

Nothing here is wrong. The two methods estimate the same components in different ways, and 7.76 % and 27.86 % are the same result expressed on different scales: 0.2786² ≈ 0.0776. Judged on the standard-deviation percentages, as Minitab applies the AIAG guideline, the study is marginal by either method. SPC for Excel (2015) notes that the AIAG manual itself calls the ANOVA result for these data acceptable because the %GRR is below 10 %, which is exactly the mix of scales described below. The ndc difference shows how sensitive a truncated number can be: 5.09 becomes 5, and 4.86 becomes 4.

How to read the results: %GRR, %Contribution and %Tolerance

Each percentage divides the gauge R&R by something different, and each answers a different question:

Metric Formula Question it answers Thresholds
%Study Variation (%GRR, %TV) 100 × GRR / TV Can the gauge see the variation of this process? (process control, SPC) AIAG: < 10 % acceptable · 10–30 % may be acceptable · > 30 % unacceptable
%Tolerance (P/T) 100 × 6 × GRR / (USL − LSL) Can the gauge sort good parts from bad ones? (product inspection) The same 10 / 30 % bands are commonly applied
%Contribution 100 × GRR² / TV² What share of the observed variance is measurement? ≈ 1 % and 9 % are the squares of 10 % and 30 %

Formulas: Minitab, methods and formulas (6 × SD by default). Criteria: AIAG MSA 4th ed., paraphrased from the table reproduced by SPC for Excel (2015, citing p. 78) and from Minitab (interpreting the results, which applies the 10 % guideline to %Study Var and %Tolerance).

Three practical rules:

  • Use the 10 / 30 % bands on standard-deviation metrics, as Minitab does with %Study Var and %Tolerance (Minitab, interpreting the results). Applying them to %Contribution makes a gauge look far better than it is: 10 % of the variance is about 32 % of the standard deviation. That is the root of the claim, repeated online, that “10 % in ANOVA equals 30 % in average and range”: the methods agree; the columns differ.
  • Pick the denominator that matches the gauge’s job. A gauge used to accept parts at final inspection is judged against the tolerance. A gauge feeding an SPC chart has to see process variation, so %Study Variation matters. In the example, the gauge uses 22.68 % of the tolerance but 27.86 % of the process variation.
  • Check the multiplier. Older material and some software use 5.15σ (99 %) instead of 6σ for study variation and %Tolerance (Ermer, 2006; SAS/QC). %Study Variation doesn’t change; %Tolerance does.

Number of distinct categories (ndc)

The ndc estimates how many groups of parts the measurement system can reliably tell apart within the process variation:

ndc = 1.41 × PV / GRR, truncated to a whole number (1.41 ≈ √2; if the result is below 1, it is reported as 1, as Minitab does).

The AIAG asks for ndc ≥ 5 (Minitab example). In the example, ANOVA gives 1.41 × 1.04233 / 0.30237 = 4.86, so ndc = 4; the average and range method gives 1.41 × 1.105 / 0.306 = 5.09, so ndc = 5 (truncation rule: Minitab, number of distinct categories; the ndc of 4 is published in the Minitab example, the average and range value is our own calculation). An ndc of 1 means the gauge can’t distinguish parts at all within this process; ndc ≥ 5 corresponds to a %Study Variation of roughly 27 % or less.

What to do when a gauge R&R fails

Don’t start by buying a gauge. The components tell you where the problem is:

What the results show Likely cause What to try
Repeatability dominates Gauge play or wear, fixture, clamping, resolution too coarse, part form (for example, an out-of-round shaft measured at random positions), environment Fix or improve the fixture, define the measurement location, maintain or replace the gauge, check resolution
Reproducibility dominates Different technique, force or measurement location between operators; unclear procedure Write and train a precise method, add fixtures or stops that remove differences in technique; then repeat the study
Significant operator × part interaction Some operators struggle with some parts (size, shape, access) Look at those parts and that operator’s readings; often a method problem
Small part variation, high %GRR Parts don’t represent the process Re-select parts; judge against the tolerance if the process really is that tight, and document why
Result “too good” (around 1 %) Parts too different, readings not blind, or repeat readings without re-positioning the part Check the study design before celebrating

The NIST e-Handbook adds a useful warning: gauges can respond differently to the working environment (temperature, humidity, operators, wear), and a gauge study is how those biases come to light (NIST e-Handbook 2.4.5). Once you’ve changed something, repeat the study on the same parts if you can, so before and after are comparable.

Ready to try your own data? The gauge R&R calculator runs ANOVA and average and range side by side and shows each convention it uses.

Gauge R&R vs calibration vs measurement uncertainty

These three answer different questions, and you need all of them.

Calibration Gauge R&R Measurement uncertainty
Question Does the gauge read correctly against a reference? How much does the measurement vary in real use? How much doubt is there about a result, from every source?
Conditions Lab or controlled conditions, reference standards Production: your parts, operators, fixtures Any, as long as all significant sources are included
Detects Bias (error) at test points Repeatability, operator effects, interaction Combined effect of bias corrections, repeatability, environment, reference, resolution…
Result Errors with an uncertainty, on a certificate %GRR, %Tolerance, ndc U (expanded uncertainty) with k

A calibration tells you how the gauge relates to the true value in a near-ideal setting, but not whether it behaves the same way in the working environment. That is the job of the gauge study (NIST e-Handbook 2.4.5). Calibrate first, then study precision: a gauge R&R can’t see a gauge that is consistently off (Ermer, 2006).

Gauge studies also feed uncertainty budgets. The standard deviations from a well-run study can serve as Type A contributions (repeatability, operator effects), but if the measurement process drifts, the study no longer describes it; NIST suggests running a gauge study before and after the measurements it is meant to describe, so you can tell whether the process changed in between (NIST e-Handbook 2.4.6). A full budget then adds what a gauge R&R never sees, such as the calibration uncertainty of the gauge and temperature effects. See measurement uncertainty explained.

Other gauge studies

A crossed gauge R&R is the most common study, not the only one:

  • Type 1 study: one reference part of known value measured repeatedly, to check bias and repeatability against the tolerance before a full study. It’s part of the VDA 5 sequence and is a useful diagnostic when an R&R fails (Dietrich and Radeck, 2014).
  • No operator influence: for automated systems such as a CMM running a program, reproducibility between operators doesn’t apply; studies without the operator factor are used instead (Dietrich and Radeck, 2014).
  • Attribute studies: go/no-go gauges and visual inspection give pass/fail answers, not numbers, so they need an attribute agreement study instead, which checks whether appraisers agree with themselves, with each other and with a known standard (Minitab, attribute agreement analysis).
  • Destructive tests: when a part can be measured only once, operators can’t re-measure the same part, and the crossed design above doesn’t apply. A nested design, where each operator measures different parts, is the usual alternative (Minitab, nested gage R&R).

ASTM also publishes a general MSA guide, ASTM E2782-24, for measurement systems used on manufactured parts. For the broader picture of bias, linearity and stability, see measurement system analysis.

Gauge R&R study checklist

Print this and keep it with the study record.

Gauge R&R study checklist

Before the study
Parts and operators
Running it
Analysis and decision

FAQ

What is a good gauge R&R result?

Under the criteria of the AIAG MSA manual (4th edition), a %GRR below 10 % is generally acceptable, 10 % to 30 % may be acceptable depending on how important the measurement is and what improving it would cost (with the customer's approval), and above 30 % is unacceptable. Software such as Minitab applies them to ratios of standard deviations (%Study Variation or %Tolerance), not to the %Contribution column. The manual also asks for at least 5 distinct categories.

Is gauge R&R the same as MSA?

No. Measurement system analysis (MSA) is the whole set of studies that check a measurement system: bias, linearity, stability, attribute agreement and gauge R&R. Gauge R&R is the study of the measurement system's variation (repeatability and reproducibility), usually the one customers ask for first.

How many parts, operators and trials do I need?

The published AIAG example uses 10 parts, 3 operators and 3 trials (90 readings), and 10 × 3 × 2 or 10 × 3 × 3 is the common design. More parts improve the estimate of part variation; more operators and trials improve the estimates of reproducibility and repeatability. Check what your customer specifies.

Can I use more than 3 operators or 3 trials?

Yes. The ANOVA method works with any balanced study. Some average-and-range spreadsheets only include constants up to 3 operators or trials, which is a limitation of the spreadsheet, not of the method.

Can I use out-of-spec or prototype parts?

Use parts that represent the variation of your production process. If production parts don't vary enough, the %Study Variation will look bad even with a good gauge; in that case, judge the gauge against the tolerance (%Tolerance) or a known process variation, and document why. Parts from a different process or of a different type make the study meaningless.

My gauge R&R is 1 %. Is that too good to be true?

Not necessarily. Automated systems such as a CMM, with no operator influence and parts that vary a lot, can give very low values. Check that the parts were measured blind and in random order, that the trials were real repeat measurements (not repeated readings of one setup), and that the gauge resolution can actually see the variation.

Why do Excel and statistics software give different results for the same data?

Usually because of conventions: ANOVA versus average and range, whether the operator × part interaction is kept or pooled, a 6 or 5.15 multiplier for study variation, 1.41 or √2 for the number of distinct categories, and slightly different range constants. With the AIAG example data the %GRR is 26.68 % by average and range and 27.86 % by ANOVA.

Does a calibrated gauge still need a gauge R&R?

Yes, if your quality system or customer requires MSA. Calibration checks the gauge against a reference in controlled conditions. Gauge R&R checks how much the measurement varies in real use, with your operators, fixtures and parts.

Is it gage or gauge?

Both spellings are used for the same thing. "Gage" is common in US automotive and quality documents and software, while "gauge" is the general English spelling.

Sources

  1. NIST. NIST/SEMATECH e-Handbook of Statistical Methods, §2.4 Gauge R&R studies (2.4.1–2.4.6) — Levels of variation, study design for labs, bias, gauge studies in uncertainty budgets
  2. JCGM / BIPM. International Vocabulary of Metrology (VIM), JCGM 200:2012. 2012 — 2.15 precision, 2.20–2.21 repeatability, 2.24–2.25 reproducibility
  3. AIAG. Measurement Systems Analysis Reference Manual, 4th edition. 2010 — Paid; not reviewed directly. Acceptance criteria and example data cited through the secondary sources below
  4. BPI Consulting (SPC for Excel). Three Ways to Analyze a Gage R&R Study. 2015 — Reproduces the AIAG 4th ed. example data and results (average and range, ANOVA) and the AIAG acceptance table (manual p. 78)
  5. Minitab Support. Example of Crossed Gage R&R Study — Same AIAG data analyzed by ANOVA (tolerance 8); AIAG ndc ≥ 5
  6. Minitab Support. Methods and formulas for the gage R&R table (Crossed Gage R&R Study) — Variance components, negative values set to zero, %Contribution, %Study Var, %Tolerance, ndc
  7. Minitab Support. Number of distinct categories (Crossed Gage R&R Study, methods and formulas) — ndc truncated; values below 1 reported as 1
  8. Minitab Support. Interpret the key results for Crossed Gage R&R Study — AIAG 10 % guideline applied to %Study Var and %Tolerance
  9. Minitab Support. Attribute Agreement Analysis overview; Nested Gage R&R Study overview — Pass/fail and subjective ratings; nested design when operators cannot measure the same parts
  10. D. S. Ermer, Quality Progress (ASQ). Improved Gage R&R Measurement Studies; Appraiser Variation in Gage R&R Measurement. 2006 — Average and range formulas, d2 and d2* constants, 5.15σ convention, calibration before precision studies
  11. SAS Institute. SAS/QC User's Guide: Gage R&R, Average and Range Method — Study-variation multiplier of 4, 5.15 or 6
  12. IATF. IATF 16949:2016 Frequently Asked Questions (FAQ 6, clause 7.1.5.1.1). 2026
  13. J. Muelaner, engineering.com. VDA-5: Combining Uncertainty Evaluation with Gage Studies
  14. E. Dietrich, M. Radeck, Carl Hanser Verlag. Prüfprozesseignung nach VDA 5 und ISO 22514-7 (sample chapter). 2014
  15. Quality Magazine. The New VDA Volume 5: Obligation and Opportunity. 2023
  16. ISO. ISO 22514-7:2026, Capability of measurement processes (3rd edition; replaces the 2021 edition). 2026
  17. ASTM International. ASTM E2782-24, Standard Guide for Measurement Systems Analysis (MSA). 2024