Learn · Guides
Rating scales & category structure
A Likert scale is a claim: that “disagree — neutral — agree” are three ordered steps along one trait, and that respondents actually use them that way. Category structure analysis is how you audit the claim. It regularly finds that a response option you paid for — in respondent attention and questionnaire length — is measuring nothing.
Thresholds: where the scale changes gear
In the Rating Scale Model (Andrich, 1978), each boundary between adjacent categories gets a threshold — the point on the ruler where a respondent becomes equally likely to pick the higher category as the lower one. With three categories there are two thresholds: below τ₁ the bottom category is the most probable response, between τ₁ and τ₂ the middle one is, above τ₂ the top one is. Each category should own a stretch of the ruler — that’s what makes it a rung worth having.
Play with the thresholds
These are category characteristic curves — the probability of each response option across the trait. Drag the thresholds together, then past each other, and watch the middle category’s territory shrink and vanish:
Each category owns a region of the scale: as a person’s position rises, the most probable response moves 0 → 1 → 2. The thresholds τ1 and τ2 mark the crossings. This is a healthy rating scale.
Disordered thresholds: the classic pathology
When τ₁ > τ₂, no region of the trait exists where the middle category is the most likely answer — its curve never rises above the others. Respondents aren’t treating it as a step between the extremes; they may read “neutral” as “no opinion,” “not applicable,” or “refuse.” The usual remedies are collapsing the category into a neighbor and re-running, or rewording it in the next revision. (One nuance: Andrich threshold disorder is a warning light, not a conviction — low-frequency categories can disorder thresholds even when the scale is working. Look at the category frequencies and average measures alongside.)
RSM or PCM?
The Rating Scale Model shares one set of thresholds across all items — right when every item uses the same response options and you want to talk about “the scale” as a single structure. The Partial Credit Model (Masters, 1982) gives every item its own thresholds — right when items differ in format, or when partial-credit scoring rubrics differ by item. PCM is more flexible; RSM is more parsimonious and its category diagnostics pool across items, so they’re stabler in small samples. If you’re unsure, run both in Logit and compare — with a well-behaved common scale they’ll largely agree.
Reading Logit’s category table
For each category Logit reports the observed count, the average measure of the people who chose it (these should rise monotonically with the category — a violation is a stronger red flag than threshold disorder), the category’s outfit, and the Andrich threshold with its standard error. The same structure drives the CCC plot, which looks exactly like the demo above but for your data. This is the analysis behind the Liking for Science walkthrough, where the textbook three-category scale calibrates to ordered, symmetric thresholds of ±0.83 logits.
In Logit
Choose RSM or PCM at analysis setup; the results include a Category Structure tab (table + CCC curves) and the PDF report carries the same table. Start with your first analysis.