Guides · Understanding fit statistics

Learn · Guides

Understanding fit statistics

The Rasch model predicts a probability for every single response. Fit statistics ask, item by item and person by person: did reality behave like the prediction? This is where flawed questions, miskeyed answers, and guessing get caught — and it’s the part of a Rasch report reviewers read first.

The intuition: surprise, averaged

Take one response. The model expected an 80% chance of success; the person failed. That’s surprising — a large residual. Fit statistics are averages of squared, standardized residuals: a mean-square (MNSQ) of 1.0 means exactly as much surprise as the model predicts for itself. Above 1.0 means more noise than expected (underfit); below 1.0 means less noise than expected (overfit) — responses suspiciously tidy, as when items duplicate each other.

Infit vs outfit

The two flavors differ only in weighting. Outfit is the plain average, so it’s dominated by far-away encounters — the expert who misses a trivial item, the novice who nails a hard one. It flags lucky guesses and careless slips. Infit weights each residual by its information, so it’s dominated by encounters near the item’s own difficulty — the responses that carry the most measurement. That makes infit the more serious alarm: an item with high infit is misbehaving exactly where it’s supposed to be doing its job. A practical reading order: scan infit first; use outfit to distinguish “broken item” from “a few weird encounters.”

The 0.5–1.5 folklore — useful, not sacred

Convention treats MNSQ between roughly 0.5 and 1.5 as productive for measurement, 1.5–2.0 as worth inspecting, and beyond 2.0 as degrading. Logit’s tables use those bands for its amber/red highlighting. But they are rules of thumb, not critical values: mean-squares tighten with sample size, so a 1.3 from n = 2,000 can be more damning than a 1.6 from n = 60. The standardized form (ZSTD) accounts for sample size but becomes hypersensitive in large samples — with enough data, everything “significantly” misfits. Read MNSQ for size of the problem and ZSTD for confidence it’s real, and let neither replace looking at the item.

What to do with a misfitting item

Misfit is a symptom, not a verdict. High outfit with fine infit → look for a handful of surprising respondents (data-entry errors, one confusing distractor). High infit → the item probably measures something else in the middle of the range: double-barreled wording, a second skill, a miskey. Very low mean-squares → redundancy; the item adds reliability on paper but little information. The remedies are editorial (rewrite, re-key, drop) — the statistics only point the flashlight. Logit’s ICC tab overlays the observed proportions on the model curve, which is usually the fastest way to see what a flagged item is doing wrong.

In Logit

The Fit tab plots every item as a bubble (measure × infit, sized by sample); the Items and Persons tables sort by any fit column; the PDF report’s Fit Flags section lists everything outside the conventional bands. See it live in your first analysis.