In Parts I and II of this series, we established why the ALSFRS-R became the dominant outcome measure in ALS trials, and catalogued the methodological limitations that have accumulated against it. Most of those limitations — interrater variability, non-linearity, informative censoring — are problems of measurement precision or analytical misspecification. They introduce noise or bias into estimates of functional decline, but they do not challenge whether functional decline is a coherent quantity to measure in the first place.
Multidimensionality is different. It is not a problem of noise or model misspecification. It is a problem of construct validity — a challenge to the assumption that the ALSFRS-R total score is measuring a single coherent quantity at all. If that assumption fails, then the total score is not just noisy. It is, in a precise technical sense, uninterpretable.
This is the most consequential limitation in the scale’s profile, and the one that has received the least systematic attention in trial design and analysis. This article explains what multidimensionality means, what evidence we have for it in the ALSFRS-R, and why it matters — concretely — for the interpretation of recent clinical trial results.
What does unidimensionality mean, and why does it matter?#
A measurement scale is unidimensional if its items all reflect a single underlying construct — a latent variable — that fully explains the correlations among them. In a unidimensional scale, once you know a patient’s position on that latent construct, knowing their score on any individual item tells you nothing new about their scores on the other items. The items are, in psychometric terms, locally independent given the latent trait.
The classic test is confirmatory factor analysis (CFA): does a single underlying factor explain how patients score across all items? If yes, the scale is unidimensional and a total score is meaningful. If not — if the items cluster into separate groups that a single factor cannot account for — the scale is multidimensional, and summing those groups into a total score conflates distinct things.
Rasch analysis approaches the same question differently: it asks whether each item’s scores can be predicted from a single underlying trait, or whether items systematically deviate from that prediction in ways that suggest separate dimensions. Franchignoni and colleagues applied this to the ALSFRS-R in 2013 and found clear evidence of misfit — bulbar and respiratory items formed their own clusters that a single trait could not adequately capture.1
Why does this matter practically? Because if the ALSFRS-R is multidimensional, then the total score is a weighted average of distinct biological processes — bulbar neurodegeneration, limb motor neuron loss, respiratory muscle failure — that may respond differently to a given intervention. Aggregating them into a total score does not neutralise these differences. It obscures them.
If the ALSFRS-R is multidimensional, the total score is a weighted average of distinct biological processes that may respond differently to treatment. Aggregating them does not neutralise those differences. It obscures them.
Evidence for multidimensionality in the ALSFRS-R#
The psychometric case has been building for over a decade. The four-domain structure of the scale — bulbar, fine motor, gross motor, respiratory — was clinically motivated rather than empirically derived, and early factor analytic work suggested that this structure was at best approximate.
Multiple CFA studies have found that a single-factor model fits the ALSFRS-R poorly. A two-factor solution (motor versus bulbar/respiratory) typically fits better, and some analyses support a three- or four-factor structure corresponding roughly to the clinical domains.2 The respiratory subscale, in particular, consistently loads onto a factor that is partially distinct from the motor items — which is biologically unsurprising, since respiratory failure in ALS reflects diaphragmatic and intercostal denervation that can progress semi-independently of limb and bulbar involvement.
Importantly, the factors are not independent. Bulbar, motor, and respiratory decline are correlated — patients who decline rapidly in one domain tend to decline in others — which is why a total score captures some signal at all. But correlation is not identity. Two constructs can be meaningfully distinct and still co-vary. The relevant question is not whether the domains are correlated, but whether they respond differentially to a given treatment. And here the clinical trial evidence becomes instructive.
The clinical consequences: evidence from pridopidine#
The most instructive recent illustration comes from Regimen D of the HEALEY ALS Platform Trial, which evaluated pridopidine — a selective sigma-1 receptor agonist — against a shared placebo over 24 weeks.3 The primary endpoint, ALSFRS-R total score change at 24 weeks, was not met in the full analysis set. A straightforward negative result, on the face of it. But the domain-level picture was considerably more interesting. In the predefined subgroup of patients with definite ALS and early disease (under 18 months from symptom onset), pridopidine demonstrated numerically slower decline across the ALSFRS-R total score and across bulbar and respiratory subdomains specifically. In a subsequent exploratory analysis of fast progressors within that subgroup, pridopidine slowed ALSFRS-R total score decline by 32%, respiratory function decline by 62%, and dyspnoea decline by 88%.
The pattern is telling. The signal was domain-specific — concentrated in bulbar and respiratory function — and was invisible in the full analysis set total score. A drug with a potentially meaningful effect on specific biological processes registered as a null result because those processes were aggregated with domains where no effect was present.
The mechanism of obscuration is straightforward. Suppose a drug slows bulbar decline by 40% but has no effect on limb motor or respiratory progression. In a total score analysis, that 40% effect is diluted across all 12 items — three of which are bulbar. The effective signal-to-noise ratio is reduced by roughly a factor of four relative to what a bulbar-specific analysis would detect. If the bulbar effect is moderate to begin with, the total score analysis will return a null result.
This is not a hypothetical failure mode. It is a plausible explanation for some of the negative trial history in ALS — a field that has seen dozens of promising compounds fail to reach statistical significance on ALSFRS-R total score, with little systematic examination of whether domain-specific effects were present and masked.
A drug that slows bulbar decline by 40% but leaves limb and respiratory progression unchanged will likely return a null result on ALSFRS-R total score. The trial will be recorded as negative. The drug will not advance — not because it doesn’t work, but because we couldn’t see it working.
What constitutes a valid outcome measure: the normative framework#
The psychometric literature offers a clear normative framework for evaluating composite scales, and the ALSFRS-R fares poorly against it on the dimension that matters most for aggregation.
A total score is a valid summary statistic if and only if the scale is unidimensional — or sufficiently close to unidimensional that the departures do not materially distort inference. This is not a philosophical preference. It is a mathematical requirement. When you sum scores across items that load on different factors, you are performing an operation whose result depends on the arbitrary weighting implicit in the item structure, the number of items per domain, and the relative variances of the domain-specific trajectories. The resulting number has no clean interpretation in terms of any single biological process.
For the ALSFRS-R, the four domains contribute three items each to the total score — a design choice that was clinically motivated but effectively imposes equal weighting across domains regardless of their relative rates of change or clinical importance. A patient who loses all bulbar function while retaining full motor and respiratory function scores 36/48 — the same as a patient who has lost moderate function uniformly across all domains. These are not equivalent clinical states, and treating them as equivalent for the purposes of a treatment effect estimate introduces systematic bias whose direction depends on where the drug acts.
What the factor structure of the ALSFRS-R actually looks like#
The most methodologically rigorous examinations of the ALSFRS-R’s factor structure have used both exploratory and confirmatory approaches in large natural history datasets. The consistent finding is that a correlated three-factor model — broadly corresponding to bulbar, motor, and respiratory dimensions — fits substantially better than a single-factor model, with comparative fit indices (CFI) and root mean square error of approximation (RMSEA) values that clearly favour the multidimensional solution.2
The practical implication is that the ALSFRS-R contains at least three semi-independent clocks running simultaneously. In a given patient, any of those clocks may run faster or slower. In a given trial, a drug may affect one clock, two, or all three — and the effect sizes may differ substantially across domains.
Toward a solution#
Recognising the multidimensionality problem does not immediately resolve it. The more ambitious response is to replace the ALSFRS-R with a psychometrically sounder instrument — one designed from the outset to be unidimensional, with item selection guided by Rasch or IRT analysis. ROADS was built with exactly that goal — items selected and calibrated to reflect a single underlying construct. AIMS takes a different route, embracing the scale’s multidimensionality explicitly by structuring assessment into separate domains analysed independently. Both represent serious attempts to correct what the ALSFRS-R gets wrong, but both face the same practical barrier: new instruments lack the decades of natural history data that regulators expect from a primary endpoint. That validation takes years — years that patients in today’s trials do not have. The more immediately actionable response is to accept the scale’s multidimensional structure and model it explicitly — to keep using the ALSFRS-R but analyse it as what it actually is: a vector of domain trajectories, not a scalar. The statistical machinery for this already exists. The case for using it is, I would argue, already compelling. That is the focus of Part IV.
Franchignoni F, et al. Rasch analysis of the Revised Amyotrophic Lateral Sclerosis Functional Rating Scale. Eur J Phys Rehabil Med. 2013;49(4):461–470. ↩︎
Bakker LA, Schröder CD, van Es MA, et al. Assessment of the factorial validity and reliability of the ALSFRS-R: a revision of its measurement model. J Neurol. 2017;264(7):1413–1420. doi:10.1007/s00415-017-8538-4. ↩︎ ↩︎
Shefner J, et al. Pridopidine for the treatment of ALS — results from the Phase 2 Healey ALS Platform Trial. Neurology. 2024. doi:10.1212/WNL.0000000000206526. See also: Geva M, et al. Pridopidine treatment in ALS: subgroup analyses from the HEALEY ALS Platform trial. Amyotroph Lateral Scler Frontotemporal Degener. 2025. doi:10.1080/21678421.2025.2597935. ↩︎
