Showing posts with label part scores. Show all posts
Showing posts with label part scores. Show all posts

Sunday, November 24, 2024

#AAIDD's #IQ Part-Score Position (with reference to diagnosing #intellectual #disabiity [#ID]) Is at Variance With Other Authoritative Sources—Important for #schoolpsychologists



In a 2021 commentary regarding the most recent official AAIDD intellectual disability definition and classification manual (2021), I raised a concern regarding AAIDD’s position that only a full scale or global IQ score can be used for a Dx of ID.  No room was left for clinical judgement for n=1 unique cases.  Click here to download and read the complete article.  Below is some select text.

“AAIDD's [IQ] Part-Score Position Is at Variance With Other Authoritative Sources”

In AAIDD’s The Death Penalty and Intellectual Disability (Polloway, 2015), both McGrew (2015) and Watson (2015) suggest that [IQ] part scores can be used in special cases. (Note that these two chapters, although published in an AAIDD book, do not necessarily represent the official position of AAIDD.) The limited use of part scores is also described in the 2002 National Research Council book on ID and social security eligibility (see McGrew, 2015; Watson, 2015). The authoritative Diagnostic and Statistical Manual of Mental Disorder—Fifth Edition (DSM-5) manual implies that part scores may be necessary when it states that ‘‘highly discrepant subtest scores may make an overall IQ score invalid'' (American Psychiatric Association, 2013, p. 37). Finally, in the recent APA Handbook of Intellectual and Developmental Disabilities (Glidden, 2021), Floyd et al. (2021) state ‘‘in rare situations in which the repercussions of a false negative diagnostic decision would have undue or irreparable negative impact upon the client, a highly g-loaded part score (see McGrew, 2015a) might be selected to represent intellectual functioning'' (emphasis added; p. 412).

 In a unique n = 1 high-stakes setting, a psychologist may be ethically obligated to proffer an expert opinion whether the full-scale score is (or is not) the best indicator of general intelligence. There must be room for the judicious use of clinical judgment-based part scores. AAIDD's purple manual complicates rather than elucidates guidance for psychologists and the courts. In high-stakes settings, a psychologist may be hard pressed to explain that their proffered expert opinions are grounded in the AAIDD purple manual, but then explain why they disagree with the ‘‘just say no to part scores'' AAIDD position.”

Thursday, March 01, 2012

IAP101 Brief #12: Use of IQ component part scores as indicators of general intelligence in SLD and MR/ID diagnosis

   
            Historically the concept of general intelligence (g), as operationalized by intelligence test battery global full scale IQ scores, has been central to the definition and classification of individuals with a specific learning disability (SLD) as well as individuals with an intellectual disability (ID).  More recently, contemporary definitions and operational criteria have elevated intelligence test battery composite or part scores to a more prominent role in diagnosis and classification of SLD and more recently in ID.
            In the case of SLD, third-method consistency definitions prominently feature component or part scores in (a) the identification of consistency between low achievement and relevant cognitive abilities or processing disorders and (b) the requirement that an individual demonstrate relative cognitive and achievement strengths (see Flanagan, Fiorello & Ortiz, 2010).  The global IQ score is de-emphasized in the third-method SLD methods.
            In contrast, the 11th edition of the AAIDD Intellectual Disability: Definition, Classification, and Systems of Supports manual (AAIDD, 2010) placed general intelligence, and thus global composite IQ scores, as central to the definition of intellectual functioning.  This has not been without challenge.  For example, the AAIDD ID definition has been criticized for an over-reliance on the construct of general intelligence and for ignoring contemporary psychometric theoretical and empirical research that has converged on a multidimensional hierarchical model of intelligence (viz., Cattell-Horn-Carroll or CHC theory).
The potential constraints of the “ID-as-a-general-intelligence-disability” definition was anticipated by the Committee on Disability Determination for Mental Retardation, in its National Research Council report “Mental Retardation:  Determining Eligibility for Social Security Benefits” (Reschly, Meyers & Hartel, 2001).  This national committee of experts concluded that “during the next decade, even greater alignment of intelligence tests and the IQ scores derived from them and the Horn-Cattell and Carroll models is likely.  As a result, the future will almost certainly see greater reliance on part scores, such as IQ scores for Gc and Gf, in addition to the traditional composite IQ.  That is, the traditional composite IQ may not be dropped, but greater emphasis will be placed on part scores than has been the case in the past” (Reschly et al., 2002, p. 94).  The committee stated that “whenever the validity of one or more part scores (subtests, scales) is questioned, examiners must also question whether the test’s total score is appropriate for guiding diagnostic decision making.  The total test score is usually considered the best estimate of a client’s overall intellectual functioning.  However, there are instances in which, and individuals for whom, the total test score may not be the best representation of overall cognitive functioning.” (p. 106-107).
            The increased emphasis on intelligence test battery composite part scores in SLD and ID diagnosis and classification raises a number of measurement and conceptual issues (Reschly et al., 2002).  For example, what are statistically significant differences?  What is a meaningful difference?  What appropriate cognitive abilities should serve as proxies of general intelligence when the global IQ is questioned?  What should be the magnitude of the total test score? 
Appropriate cognitive abilities will only be the only issue discussed here.  This issue addresses  which component or part scores are more correlated with general intelligence (g)—that is, what component part scores are high g-loaders?  The traditional consensus has been that measures of Gc (crystallized intelligence; comprehension-knowledge) and Gf (fluid intelligence or reasoning) are the highest g-loading measures and constructs and are the most likely candidates for elevated status when diagnosing ID (Reschly et al., 2002).  Although not always stated explicitly, the third method consistency SLD definitions specify that an individual must demonstrate “at least an average level of general cognitive ability or intelligence” (Flanagan et al., 2010, p.745), a statement that implicitly suggests cognitive abilities and component scores with high g-ness.
Table 1 is intended to provide guidance when using component part scores in the diagnosis and classification of SLD and ID (click on images to enlarge and use the browser zoom feature  to view; it is recommended you click here to access a PDF copy of the table..and also zoom in on it).  Table 1 presents a summary of the comprehensive, nationally normed, individually administered intelligence batteries that possess satisfactory psychometric characteristics (i.e., national norm samples, adequate reliability and validity for the composite g-score) for use in the diagnosis of ID and SLD.



The Composite g-score column lists the global general intelligence score provided by each intelligence battery.  This score is the best estimate of a persons general intellectual ability, which currently is most relevant to the diagnosis of ID as per AAIDD.  All composite g-scores listed in Table 1 meet Jensens (1998) psychometric sampling error criteria as valid estimates of general intelligence.  As per Jensens number of tests criterion, all intelligence batteries g-composites are based on a minimum of nine tests that sample at least three primary cognitive ability domains.  As per Jensens variety of tests criterion (i.e., information content, skills and demands for a variety of mental operations), the batteries, when viewed from the perspective of CHC theory, vary in ability domain coveragefour (CAS, SB5), five (KABC-II, WISC-IV, WAIS-IV), six (DAS-II) and seven (WJ III) (Flanagan, Ortiz & Alfonso, 2007; Keith & Reynolds, 2010).   As recommended by Jensen (1998), the particular collection of tests used to estimate g should come as close as possible, with some limited number of tests, to being a representative sample of all types of mental tests, and the various kinds of test should be represented as equally as possible (p. 85).  Users should consult sources such as Flanagan et al. (2007) and Keith and Reynolds, 2010) to determine how each intelligence battery approximates Jensens optimal design criterion, the specific CHC domains measured, and the proportional representation of the CHC domains in each batteries composite g-score.
Also included in Table 1 are the component part scales provided by each battery (e.g., WAIS-IV Verbal Comprehension Index, Perceptual Reasoning Index, Working Memory Index, and Processing Speed Index), followed by their respective within-battery g-loadings.[1]  Examination of the g-ness of composite scores from existing batteries (see last three columns in Table 1) suggests the traditional assumption that measures of Gf and Gc are the best proxies of general intelligence may not hold across all intelligence batteries.[2] 
In the case of the SB5, all five composite part scores are very similar in g-loadings (h2 = .72 to .79).  No single SB5 composite part score appears better than the other SB5 scores for suggesting average general intelligence (when the global IQ score is not used for this purpose).  At the other extreme is the WJ III where the Fluid Reasoning, Comprehension-Knowledge, Long-term Storage and Retrieval cluster scores are the best g-proxies for part-score based interpretation within the WJ III.  The WJ III Visual Processing and Processing Speed clusters are not composite part scores that should be emphasized as indicators of general intelligence.  Across all batteries that include a processing speed component part score (DAS-II, WAIS-IV, WISC-IV, WJ III) the respective processing speed scale is always the weakest proxy for general intelligence and thus, would not be viewed as a good estimate of general intelligence. 
            It is also clear that one cannot assume that composites with similar sounding names of measured abilities should have similar relative g-ness status within different batteries.  For example, the Gv (visual-spatial or visual processing) clusters in the DAS-II (Spatial Ability), SB5 (Visual-Spatial Processing) are relatively strong g-measures within their respective battery, but the same cannot be said for the WJ III Visual Processing cluster.  Even more interesting are the differences in the WAIS-IV and WISC-IV relative g-loadings for similarly sounding index scores. 
For example, the Working Memory Index is the highest g-loading component part score (tied with Perceptual Reasoning Index) in the WAIS-IV but is only third (out of four) in the WISC-IV.   The Working Memory Index is comprised of the Digit Span and Arithmetic subtests in the WAIS-IV and the Digit Span and the Letter-Number Sequencing subtests in the WISC-IV.  The Arithmetic subtest has been reported to be a factorially complex test which may tap fluid intelligence (Gf-RQ—quantitative reasoning), quantitative knowledge (Gq), working memory (Gsm), and possible processing speed (Gs; Keith & Reynolds, 2010; Phelps, McGrew, Knopik & Ford, 2005).   The factorially complex characteristics of the Arithmetic subtest (which, in essence, makes it function like a mini-g proxy) would explain why the WAIS-IV Working Memory Index is a good proxy for g in the WAIS-IV but not in the WISC-IV. The WAIS-IV and WISC-IV Working Memory Index scales, although named the same, are not measuring identical constructs.

A critical caveat is that the g-loadings cannot be compared across different batteries.  g-loadings may change when the mixture of measures included in the analyses change.  Different "flavors" of g can result (Carroll, 1993; Jensen, 1998). The only way to compare the g-ness across batteries is with appropriately designed cross- or joint-battery analysis (e.g., WAIS-IV, SB5 and WJ III analyzed in a common sample).
The above within and across intelligence battery examples illustrates that those who use component part scores as an estimate of a person’s general intelligence must be aware of the composition and psychometric g-ness of the component scores within each intelligence battery.  Not all component part scores in different intelligence batteries are created equal (with regard to g-ness).  Also, not all similarly named factor-based composite scores may measure the same identical construct and may vary in degree of within battery g-ness.  This is not a new problem in the context of naming factors in factor analysis, and by extension, factor-based intelligence test composite scores, Cliff (1983) described this nominalistic fallacy in simple language—“if we name something, this does not mean we understand it” (p. 120). 




[1] As noted in the footnotes in Table 1, all composite score g-loadings were computed by Kevin McGrew by entering the smallest number (and largest age ranges covered) of the published correlation matrices within each intelligence batteries technical manual (note the exception for the WJ III) in order to obtain an average g-loading estimate.  It would have been possible to calculate and report these values for each age-differentiated correlation matrix for each intelligence battery.  However, the purpose of this table is to provide the best possible average value across the entire age-range of each intelligence battery.  Floyd and colleagues have published age-differentiated g-loadings for the DAS-II and WJ III.  Those values were not used as they are based on the use of the principal common factor analysis method, a method that  analyzes the reliable shared variance among tests.  Although principal factor and principal component loadings typically will order measures in the same relative position, the principal factor loadings typically will be lower.  Given that the imperfect manifest composite scale scores are those that are utilized in practice, and to also allow uniformity in the calculation of the g-loadings reported in Table 1, principal component analysis was used in this work. The same rationale was used for not using the latent factor loadings on a higher-order g-factor in SEM/CFA analysis of each test battery.  Loadings from CFA analyses represent the relations between the underlying theoretical ability constructs and g purged of measurement error.  Also, frequently the final CFA solutions reported in a batteries technical manual (or independent journal articles) allow tests to be factorially complex (load on more than one latent factor), a measurement model that does not resemble the real world reality of the manifest/observed composite scores used in practice.  Latent factor loadings on a higher-order g-factor will often differ significantly from principal component loadings based on the manifest measures, both in absolute magnitude and relative size (e.g., see high Ga loading on g in WJ III technical manual which is at variance with the manifest variable based Ga loading reported in Table 1) 
[2] The h2 values are the values that should be used to compare the relative amount of g-variance present in the component part scores within each intelligence battery.

Thursday, March 31, 2011

Why IQ composite scores often are higher or lower than the subtest scores: Awesome video explanation

This past week Dr. Joel Schneider and I released a paper called " 'Just say no' to averaging IQ subtest scores." The report generated considerable discussion on a number of professional listservs.

One small portion of the paper explained why composite/cluster scores from IQ tests often are higher (or lower) than the arithmetic mean of the tests that comprise the composite. This observation often baffles test users.

I would urge those who have ponder this question to read that section of the report. And THEN, be prepared to be blown away by an instructional video Joel posted at his blog where he leads you through a visual-graphic explanation of the phenomena. Don't be scared by the geometry or some of the terms. Just sit back and relax and now recognize, even if all the technical stuff is not your cup-of-tea, that there is an explanation for this score phenomena. And when colleagues ask, just refer them to Joel's blog.

It is brilliant and worth a view, even if you are not a quantitatively oriented thinker.

Below is a screen capture of the start [double click on icon to enlarge]



- iPost using BlogPress from my Kevin McGrew's iPad

Monday, March 28, 2011

Cognitive ability domain cohesion-why composite scores comprised of significantly different subtest scores are still valid

Some excellent discussion has been occurring on the NASP and CHC listservs in response to the "Just say no to averaging IQ subtest scores" blog post and report.

An issue/question that has surfaced (not for the first time) is why markedly discrepant subtest scores that form a composite can still be considered valid indicators of the construct domain. Often clinicians believe that if there is a significant and large discrepancy between tests within a composite, the total score should be considered invalid.

The issue is complex and was touched on briefly in our report and in the NASP and CHC threads by Joel Schneider. Here I mention just ONE concept for consideration.

Below is a 2-D MDS analysis of the WJ III Cog/Ach tests for subjects aged 6-18 in the norm sample. MDS also finds structure as does factor analysis. This 2D model is based on the analysis of the tests correlation matrix. What I think is a major value of MDS, and other spatial statistics, is that one can "see" the numerical relations between tests. Although the metrics are not identical, the visual-spatial map of the WJ III tests does, more-or-less, mirror the intercorrelations between tests. [Double click on image to enlarge]




So....take a look at the Gc, Grw, or Gq tests in this MDS map. All of these tests cluster closely together. Inspection of their intercorrelations finds high correlations among all measures. Conversely, look at the large amount of spatial territory covered by the WJ III Gv tests. Also look at the Ga tests (note that a red line is not connecting Auditory Attention, AA, down in the right-hand quadrant with the other Ga tests). Furthermore, even though most of the Gsm tests are relatively cohesive or tight, Memory for Sentences is further away from the other Gsm tests.

IMHO, these visual-spatial maps, which mirror intercorrelations, tell us than in humans, not all cognitive/ach domains include narrow abilities that are highly interrcorrrelated. I call it "ability domain cohesion." Clearly the different Gv abilities measured by the WJ III Gv tests indicate that the Gv domain is less cohesive (less tight) than the Gc or Grw domain. This does not suggest the tests are flawed..instead it tells us about the varying degrees of cohesiveness present in different ability domains.

Thus, for ability domains that are very very broad (in terms of domain cohesion--e.g., Gv and Ga in this MDS figure), wildly different test scores (e.g., between WJ III Spatial Relations, SR, and Picture Recognition, PR) may be valid and simply reflect that inherent lower cohesiveness (tightness) of these ability domains in human intelligence. Thus, if a person is significantly different in his/her respective Gv SR or PR scores, and these scores are providing valid indications of their relative standing on these measured abilities, then combining them together is appropriate and reflects a valid estimate of the Gv domain....which by nature is broad...and people will often display significant within-domain variability.

Bottom line. Composite scores produced by subtests that are markedly different are likely valid estimates of domains...it is just the nature of human intelligence that some of these domains are more tight or cohesive than others.
- iPost using BlogPress from my Kevin McGrew's iPad