Showing posts with label WISC-IV. Show all posts
Showing posts with label WISC-IV. Show all posts

Friday, November 22, 2024

Research Byte: Beyond Individual #Tests: Youths’ #Cognitive Abilities, Basic #Reading, and #Writing—relevant to #schoolpsychologists #CHC

An impressive multiple test-battery CHC theory cognitive and achievement (cross-battery) confirmatory factor analysis study based on a research design first conceptualized by Jack McArdle (planned missing data reference variable design) that finds that multiple broad CHC abilities are important in explaining reading and writing achievement above and beyond psychometric g.  Of course, the results would be different if a bi-factor model were run (see McGrew et al., 2023 for discussion of three major classes of cognitive-achievement CFA/SEM research designs).  

Click here to download/read this open access articles. 

This research group, IMHO, does some of the best CFA/SEM modeling in the assessment and school psychology literature.

Click on images to enlarge for easier viewing.



Beyond Individual Tests: Youths’ Cognitive Abilities, Basic Reading, and Writing 

by  1,* 1 2 and  3
1
Department of Educational Psychology, University of Connecticut, Storrs, CT 06268, USA
2
Department of Educational Psychology, University of Texas at Austin, Austin, TX 78712, USA
3
Department of Educational Psychology, University of Kansas, Lawrence, KS 66045, USA
*
Author to whom correspondence should be addressed. 
J. Intell. 202412(11), 120; https://doi.org/10.3390/jintelligence12110120
Abstract

Broadly, individuals’ cognitive abilities influence their academic skills, but the significance and strength of specific cognitive abilities varies across academic domains and may vary across age. Simultaneous analyses of data from many tests and cross-battery analyses can address inconsistent findings from prior studies by creating comprehensively defined constructs, which allow for greater generalizability of findings. The purpose of this study was to examine the cross-battery direct effects and developmental differences in youths’ cognitive abilities on their basic reading abilities, as well as the relations between their reading and writing achievement. Our sample included 3927 youth aged 6 to 18. Six intelligence tests (66 subtests) and three achievement tests (10 subtests) were analyzed. Youths’ general intelligence (g, large direct and indirect effects), verbal comprehension–knowledge (large direct effect), working memory (large direct effect), and learning efficiency (moderate direct effect) explained their basic reading skills. The influence of g and fluid reasoning were difficult to separate statistically. Most of the cognitive–basic reading relations were stable across age, except the influence of verbal comprehension–knowledge (Gc), which appeared to slightly increase with age. Youths’ basic reading had large influences on their written expression and spelling skills, and their spelling skills had a large influence on their written expression skills. The directionality of the effects most strongly supported the direct effects from the youths’ basic reading to their spelling skills, and not vice versa.

Thursday, March 23, 2017

Research Byte: The Predictive Validity of Four Intelligence Tests for School Grades: A Small Sample Longitudinal Study in Germany

The Predictive Validity of Four Intelligence Tests for School Grades: A Small Sample Longitudinal Study

  • Department of Psychology, University of Basel, Basel, Switzerland
Intelligence is considered the strongest single predictor of scholastic achievement. However, little is known regarding the predictive validity of well-established intelligence tests for school grades. We analyzed the predictive validity of four widely used intelligence tests in German-speaking countries: The Intelligence and Development Scales (IDS), the Reynolds Intellectual Assessment Scales (RIAS), the Snijders-Oomen Nonverbal Intelligence Test (SON-R 6-40), and the Wechsler Intelligence Scale for Children (WISC-IV), which were individually administered to 103 children (Mage = 9.17 years) enrolled in regular school. School grades were collected longitudinally after 3 years (averaged school grades, mathematics, and language) and were available for 54 children (Mage = 11.77 years). All four tests significantly predicted averaged school grades. Furthermore, the IDS and the RIAS predicted both mathematics and language, while the SON-R 6-40 predicted mathematics. The WISC-IV showed no significant association with longitudinal scholastic achievement when mathematics and language were analyzed separately. The results revealed the predictive validity of currently used intelligence tests for longitudinal scholastic achievement in German-speaking countries and support their use in psychological practice, in particular for predicting averaged school grades. However, this conclusion has to be considered as preliminary due to the small sample of children observed.

Tuesday, March 07, 2017

WISC-V CFA by Reynolds and Keith - a MUST read by two of the best intelligence test scholars I know

Available online 3 March 2017

Multi-group and hierarchical confirmatory factor analysis of the Wechsler Intelligence Scale for Children—Fifth Edition: What does it measure?


Highlights

WISC-V constructs are measured similarly across the 6–16-year age range.
g and five broad ability factors account for subtest covariances.
Our CFA findings diverged from EFA research.
g is measured strongly in the new 7 subtest FSIQ.

Abstract

The purpose of this research was to test the consistency in measurement of Wechsler Intelligence Scale for Children-Fifth Edition (WISC-V; Wechsler, 2014) constructs across the 6 through 16 age span and to understand the constructs measured by the WISC-V. First-order, higher-order, and bifactor confirmatory factor models were used. Results were compared with two recent studies using higher-order and bifactor exploratory factor analysis (Canivez, Watkins, & Dombrowski, 2015; Dombrowski, Canivez, Watkins, & Beaujean, 2015) and two using confirmatory factor analysis (Canivez, Watkins, & Dombrowski, 2016; Chen, Zhang, Raiford, Zhu, & Weiss, 2015). We found evidence of age-invariance for the constructs measured by the WISC-V. Further, both g and five distinct broad abilities (Verbal Comprehension, Visual Spatial Ability, Fluid Reasoning, Working Memory, and Processing Speed) were needed to explain the covariances among WISC-V subtests, although Fluid Reasoning was nearly equivalent to g. These findings were consistent whether a higher-order or a bifactor hierarchical model was used, but they were somewhat inconsistent with factor analyses from the prior studies. We found a correlation between Fluid Reasoning and Visual Spatial factors beyond a general factor (g) and that Arithmetic was primarily a direct indicator of g. Composite scores from the WISC-V correlated well with their corresponding underlying factors. For those concerned about the fewer numbers of subtests in the Full Scale IQ, the model implied relation between g and the FSIQ was very strong.

Sunday, December 07, 2014

The WJ IV Measurement of Auditory Processing (Ga): An online powerpoint slide show



The WJ IV Measurement of Auditory Processing (Ga) from Kevin McGrew

The WJ IV Cognitive and Oral Language include new measures of auditory processing (Ga) that are much more cognitively complex auditory measures of intelligence.  This short presentation provides an overview of the WJ IV Ga tests and presents evidence supporting the importance of Ga as a major component of human intelligence.

Thursday, July 31, 2014

WJ IV update: WJ IV g-scores (GIA,BIA,Gf-Gc composite) correlations with WISC-IV/WAIS-IV FS and GAI IQ scores

In the WJ IV technical manual (McGrew, LaForte, Schrank, 2014) concurrent validity results are presented for the WJ IV COG with the WISC-IV and WAIS-IV (click here for WJ IV COG overview and select correlation information from tech. manual).

A number of psychologists have asked about correlations between the primary WJ IV COG g-scores and the Wechsler General Ability Index (GAI).  They are not presented in the technical manual.  I have now computed those correlations, as well as a few others with the Wechsler GAI, and they are now part of the SlideShare at the link above and are also reported below.  Click on image to enlarge.


Monday, September 02, 2013

The Wechsler Arithmetic subtest measures quantitative reasoning...another study

A new study on the now "old" WISC-III which still provides insights into the debate regarding what the Wechsler Arithmetic subtest measures. Consistent with research I have coauthored and my analysis of other studies (click here to view), this new study is consistent with the classification of Arithmetic as primarily a measure of quantitative reasoning.

Click on images to enlarge.







- Posted using BlogPress from my iPad

Friday, May 17, 2013

Video tutorial: Estimating latent WISC-IV and WAIS-IV scores for individuals--Dr. Joel Schneider

Dr. Joel Schneider has done it again.  A brilliant video tutorial demonstrating how latent factor scores can be used, via Excel templates he provides, to interpret scores on the WISC-IV and WAIS-IV.  This is complex material but his beautiful visual video tutorial makes it easier to understand the complex constructs.  Dr. Schneider continues to push the envelope on psychometric based IQ test score interpretation.


Friday, September 07, 2012

Historical review of WISC to WISC-IV

This article is an open access article and thus I am providing a link to a copy.

The article helps users understand the various changes that have occurred in this series, changes which may help explain IQ score differences in individuals who may have taken various versions over time.

 

Saturday, April 28, 2012

Research byte: WISC-IV CHC-->math achievement study

A very interesting study that supports the cautions outlined by myself and Wendling (McGrew & Wendling, 2010) that the extant CHC cognitive-achievement relations research completed during the past 20 years consists primarily of studies conducted with the WJ-R and WJ-III (94 %) and that generalization of the COG-ACH conclusions to other instruments are not established and should only be undertaken with significant caution.
Click on images to enlarge




















Posted using BlogPress from Kevin McGrew's iPad
www.themindhub.com

Wednesday, March 28, 2012

Thursday, March 01, 2012

IAP101 Brief #12: Use of IQ component part scores as indicators of general intelligence in SLD and MR/ID diagnosis

   
            Historically the concept of general intelligence (g), as operationalized by intelligence test battery global full scale IQ scores, has been central to the definition and classification of individuals with a specific learning disability (SLD) as well as individuals with an intellectual disability (ID).  More recently, contemporary definitions and operational criteria have elevated intelligence test battery composite or part scores to a more prominent role in diagnosis and classification of SLD and more recently in ID.
            In the case of SLD, third-method consistency definitions prominently feature component or part scores in (a) the identification of consistency between low achievement and relevant cognitive abilities or processing disorders and (b) the requirement that an individual demonstrate relative cognitive and achievement strengths (see Flanagan, Fiorello & Ortiz, 2010).  The global IQ score is de-emphasized in the third-method SLD methods.
            In contrast, the 11th edition of the AAIDD Intellectual Disability: Definition, Classification, and Systems of Supports manual (AAIDD, 2010) placed general intelligence, and thus global composite IQ scores, as central to the definition of intellectual functioning.  This has not been without challenge.  For example, the AAIDD ID definition has been criticized for an over-reliance on the construct of general intelligence and for ignoring contemporary psychometric theoretical and empirical research that has converged on a multidimensional hierarchical model of intelligence (viz., Cattell-Horn-Carroll or CHC theory).
The potential constraints of the “ID-as-a-general-intelligence-disability” definition was anticipated by the Committee on Disability Determination for Mental Retardation, in its National Research Council report “Mental Retardation:  Determining Eligibility for Social Security Benefits” (Reschly, Meyers & Hartel, 2001).  This national committee of experts concluded that “during the next decade, even greater alignment of intelligence tests and the IQ scores derived from them and the Horn-Cattell and Carroll models is likely.  As a result, the future will almost certainly see greater reliance on part scores, such as IQ scores for Gc and Gf, in addition to the traditional composite IQ.  That is, the traditional composite IQ may not be dropped, but greater emphasis will be placed on part scores than has been the case in the past” (Reschly et al., 2002, p. 94).  The committee stated that “whenever the validity of one or more part scores (subtests, scales) is questioned, examiners must also question whether the test’s total score is appropriate for guiding diagnostic decision making.  The total test score is usually considered the best estimate of a client’s overall intellectual functioning.  However, there are instances in which, and individuals for whom, the total test score may not be the best representation of overall cognitive functioning.” (p. 106-107).
            The increased emphasis on intelligence test battery composite part scores in SLD and ID diagnosis and classification raises a number of measurement and conceptual issues (Reschly et al., 2002).  For example, what are statistically significant differences?  What is a meaningful difference?  What appropriate cognitive abilities should serve as proxies of general intelligence when the global IQ is questioned?  What should be the magnitude of the total test score? 
Appropriate cognitive abilities will only be the only issue discussed here.  This issue addresses  which component or part scores are more correlated with general intelligence (g)—that is, what component part scores are high g-loaders?  The traditional consensus has been that measures of Gc (crystallized intelligence; comprehension-knowledge) and Gf (fluid intelligence or reasoning) are the highest g-loading measures and constructs and are the most likely candidates for elevated status when diagnosing ID (Reschly et al., 2002).  Although not always stated explicitly, the third method consistency SLD definitions specify that an individual must demonstrate “at least an average level of general cognitive ability or intelligence” (Flanagan et al., 2010, p.745), a statement that implicitly suggests cognitive abilities and component scores with high g-ness.
Table 1 is intended to provide guidance when using component part scores in the diagnosis and classification of SLD and ID (click on images to enlarge and use the browser zoom feature  to view; it is recommended you click here to access a PDF copy of the table..and also zoom in on it).  Table 1 presents a summary of the comprehensive, nationally normed, individually administered intelligence batteries that possess satisfactory psychometric characteristics (i.e., national norm samples, adequate reliability and validity for the composite g-score) for use in the diagnosis of ID and SLD.



The Composite g-score column lists the global general intelligence score provided by each intelligence battery.  This score is the best estimate of a persons general intellectual ability, which currently is most relevant to the diagnosis of ID as per AAIDD.  All composite g-scores listed in Table 1 meet Jensens (1998) psychometric sampling error criteria as valid estimates of general intelligence.  As per Jensens number of tests criterion, all intelligence batteries g-composites are based on a minimum of nine tests that sample at least three primary cognitive ability domains.  As per Jensens variety of tests criterion (i.e., information content, skills and demands for a variety of mental operations), the batteries, when viewed from the perspective of CHC theory, vary in ability domain coveragefour (CAS, SB5), five (KABC-II, WISC-IV, WAIS-IV), six (DAS-II) and seven (WJ III) (Flanagan, Ortiz & Alfonso, 2007; Keith & Reynolds, 2010).   As recommended by Jensen (1998), the particular collection of tests used to estimate g should come as close as possible, with some limited number of tests, to being a representative sample of all types of mental tests, and the various kinds of test should be represented as equally as possible (p. 85).  Users should consult sources such as Flanagan et al. (2007) and Keith and Reynolds, 2010) to determine how each intelligence battery approximates Jensens optimal design criterion, the specific CHC domains measured, and the proportional representation of the CHC domains in each batteries composite g-score.
Also included in Table 1 are the component part scales provided by each battery (e.g., WAIS-IV Verbal Comprehension Index, Perceptual Reasoning Index, Working Memory Index, and Processing Speed Index), followed by their respective within-battery g-loadings.[1]  Examination of the g-ness of composite scores from existing batteries (see last three columns in Table 1) suggests the traditional assumption that measures of Gf and Gc are the best proxies of general intelligence may not hold across all intelligence batteries.[2] 
In the case of the SB5, all five composite part scores are very similar in g-loadings (h2 = .72 to .79).  No single SB5 composite part score appears better than the other SB5 scores for suggesting average general intelligence (when the global IQ score is not used for this purpose).  At the other extreme is the WJ III where the Fluid Reasoning, Comprehension-Knowledge, Long-term Storage and Retrieval cluster scores are the best g-proxies for part-score based interpretation within the WJ III.  The WJ III Visual Processing and Processing Speed clusters are not composite part scores that should be emphasized as indicators of general intelligence.  Across all batteries that include a processing speed component part score (DAS-II, WAIS-IV, WISC-IV, WJ III) the respective processing speed scale is always the weakest proxy for general intelligence and thus, would not be viewed as a good estimate of general intelligence. 
            It is also clear that one cannot assume that composites with similar sounding names of measured abilities should have similar relative g-ness status within different batteries.  For example, the Gv (visual-spatial or visual processing) clusters in the DAS-II (Spatial Ability), SB5 (Visual-Spatial Processing) are relatively strong g-measures within their respective battery, but the same cannot be said for the WJ III Visual Processing cluster.  Even more interesting are the differences in the WAIS-IV and WISC-IV relative g-loadings for similarly sounding index scores. 
For example, the Working Memory Index is the highest g-loading component part score (tied with Perceptual Reasoning Index) in the WAIS-IV but is only third (out of four) in the WISC-IV.   The Working Memory Index is comprised of the Digit Span and Arithmetic subtests in the WAIS-IV and the Digit Span and the Letter-Number Sequencing subtests in the WISC-IV.  The Arithmetic subtest has been reported to be a factorially complex test which may tap fluid intelligence (Gf-RQ—quantitative reasoning), quantitative knowledge (Gq), working memory (Gsm), and possible processing speed (Gs; Keith & Reynolds, 2010; Phelps, McGrew, Knopik & Ford, 2005).   The factorially complex characteristics of the Arithmetic subtest (which, in essence, makes it function like a mini-g proxy) would explain why the WAIS-IV Working Memory Index is a good proxy for g in the WAIS-IV but not in the WISC-IV. The WAIS-IV and WISC-IV Working Memory Index scales, although named the same, are not measuring identical constructs.

A critical caveat is that the g-loadings cannot be compared across different batteries.  g-loadings may change when the mixture of measures included in the analyses change.  Different "flavors" of g can result (Carroll, 1993; Jensen, 1998). The only way to compare the g-ness across batteries is with appropriately designed cross- or joint-battery analysis (e.g., WAIS-IV, SB5 and WJ III analyzed in a common sample).
The above within and across intelligence battery examples illustrates that those who use component part scores as an estimate of a person’s general intelligence must be aware of the composition and psychometric g-ness of the component scores within each intelligence battery.  Not all component part scores in different intelligence batteries are created equal (with regard to g-ness).  Also, not all similarly named factor-based composite scores may measure the same identical construct and may vary in degree of within battery g-ness.  This is not a new problem in the context of naming factors in factor analysis, and by extension, factor-based intelligence test composite scores, Cliff (1983) described this nominalistic fallacy in simple language—“if we name something, this does not mean we understand it” (p. 120). 




[1] As noted in the footnotes in Table 1, all composite score g-loadings were computed by Kevin McGrew by entering the smallest number (and largest age ranges covered) of the published correlation matrices within each intelligence batteries technical manual (note the exception for the WJ III) in order to obtain an average g-loading estimate.  It would have been possible to calculate and report these values for each age-differentiated correlation matrix for each intelligence battery.  However, the purpose of this table is to provide the best possible average value across the entire age-range of each intelligence battery.  Floyd and colleagues have published age-differentiated g-loadings for the DAS-II and WJ III.  Those values were not used as they are based on the use of the principal common factor analysis method, a method that  analyzes the reliable shared variance among tests.  Although principal factor and principal component loadings typically will order measures in the same relative position, the principal factor loadings typically will be lower.  Given that the imperfect manifest composite scale scores are those that are utilized in practice, and to also allow uniformity in the calculation of the g-loadings reported in Table 1, principal component analysis was used in this work. The same rationale was used for not using the latent factor loadings on a higher-order g-factor in SEM/CFA analysis of each test battery.  Loadings from CFA analyses represent the relations between the underlying theoretical ability constructs and g purged of measurement error.  Also, frequently the final CFA solutions reported in a batteries technical manual (or independent journal articles) allow tests to be factorially complex (load on more than one latent factor), a measurement model that does not resemble the real world reality of the manifest/observed composite scores used in practice.  Latent factor loadings on a higher-order g-factor will often differ significantly from principal component loadings based on the manifest measures, both in absolute magnitude and relative size (e.g., see high Ga loading on g in WJ III technical manual which is at variance with the manifest variable based Ga loading reported in Table 1) 
[2] The h2 values are the values that should be used to compare the relative amount of g-variance present in the component part scores within each intelligence battery.

Wednesday, December 14, 2011

Friday, September 09, 2011

The factor structure of the French WISC-IV: Dr. Philippe Golay presentation




Thanks to Dr. Philippe Golay for sharing a copy of a poster presentation ("Revisiting the factor structure of the French WISC-IV: Insights through Bayesian structural equation modeling (BSEM)" he is presenting Monday at the 12th Congress of the Swiss Psychological Society.

Double click on image to enlarge and click here for PDF copy of the poster.


- iPost using BlogPress from Kevin McGrew's iPad

Generated by: Tag Generator


Tuesday, July 19, 2011

More on the problem with the -1 SD [15 SS (3 ss)] IQ subtest discrepancy rule-of-thumb

In a prior post I raised concerns about the use of the 1 SD (15 SS/3 ss) rule-of-thumb for evaluating differences between two IQ subtest scores that are part of the same composite or cluster. My central point was that this simplistic rule-of-thumb fails to incorporate information regarding the cohesiveness or inter-correlation of the tests within a cluster. More importantly, some human ability domains are more cohesive/tight (e.g., Gc) than others (Gv), and the resulting correlation between two compared tests require the use of the SD (diff) formula that incorporates the correlation between the tests within a domain that are to be compared.

I presented estimated SD (diff) values for select subtest comparisons within the WISC-IV and WJ III in different construct domains. The estimates used the SD (diff) formula that includes the correlation between the measures to be compared.

Knowing that some folks don't like formula's and estimates, I decided to make the point more concrete with real data. A picture is worth a thousand words (or equations).

In the prior post I reported an estimated SD (diff) for the comparison of the WJ III Verbal Comprehension and General Information Gc tests of 9.9, based on their average correlation (across all norm subjects) of .78.

Today I went to the WJ III NU norm data and subtracted all General Information SS's from Verbal Information SS. I then calculated summary stats and generated the histogram below. [Click on image to enlarge]



Beautiful...don't you think? A normal distribution centered on zero (Mean = -0.5) and with an actual data-based SD of 9.8 (9.8 is almost identical to the 9.9 value resulting from the equation method).

Study the graph. It clearly shows that if clinicians want to determine if the WJ III Verbal Comprehension and General Information SS's are 1 SD different (1 SD[diff], technically), then a difference of approximately 10 points is what an examiner should look for...not 15! If an examiner uses the inaccurate rule-of-thumb (i.e, difference of 15 points is 1 SD), in reality the examiner, in the case of these two WJ III Gc tests, is actually requiring a difference of -1.5 SD (diff)....or 15 points.

See prior post for lengthier discussion of the logic, equations, and danger in invoking a subtest difference rule-of-thumb of -1 SD=15 (or, -1 SD = 3 for scaled scores).



- iPost using BlogPress from my Kevin McGrew's iPad

Generated by: Tag Generator


Monday, June 20, 2011

IAP 101 Psychometric Brief # 9: The problem with the 1/1.5 SD SS (15/22) subtest comparison "rule-of-thumb"

In regard to my prior "temp" post, I wrote so much in my NASP listserv response that I have decided to take my email response, correct a few typo's, and post it now as blog post. I may return to this later to write a lengthier IAP 101 Research Brief or report.

Psychologists who engage in intelligence testing frequently compare subtest scores to determine if they are statistically and practically different...as part of the clinical interpretation process. Most IQ test publishers provide sound statistical procedures (tables or software for evaluating the statistical difference of two test scores; confidence band comparison rules-of-thumb).

However, traditional and clinical lore has produced a common "rule-of-thumb" that is problematic. The typical scenario is when a clinician subtracts two test SS's (M=100; SD=15) and invokes the rule-of-thumb that the difference needs to be 15 SS points (1 SD) or 22/23 points (1.5 SD). This is not correct.

SS difference scores do NOT have an SD scale of 15! When you subtract two SS's (with mean=100; SD=15) the resultant score distribution has a mean of zero and an SD that is NOT 15 (unless you transform/rescale the distribution to this scale) The size of the difference SD is a function of the correlation between the two measures compared.

The SD(diff) is the statistic that should be used, and there are a number of different forumla for computing this metric. The different SD(diff)'s differ based on the underlying question or assumptions that is the basis for making the comparison.

One way to evaluate score differences is the SEM band overlap approach. This is simple and is based on underlying statistical calculations (averaged across different scenarios to allow for a simple rule of thumb) that incorporates information about the reliability of the difference score. Test publishers also provide tables to evaluate the statistical significance of differences of a certain magnitude for subtests, such as in the various Wechsler manuals and software. These are all psychometrically sound and defensible procedures.......let me say that again...these are all psychometrically sound and defensible procedures. I repeat this phrase as the point I make below was recently misinterpreted at a state SP workshop as me saying there was something wrong with tables in the WISC-IV...which is NOT what I said and is NOT what I am saying here).

However, it is my opinion that in these situations we must do better and there is a more appropriate and better metric for evaluating differences between two different test scores, ESPECIALLY when the underlying assumption is that the two measures should be similar because they form a composite or cluster. This implies "correlation"...and not simple comparison of any two tests.

When one is attempting to evaluate the "unity" of a cluster or composite, an SD(diff) metric should be used that is consistent with the underlying assumption of the question. Namely, one is expecting the scores to be similar because they form a factor. This implies "correlation" between the measures. There is an SD(diff) calculation that incorporates the correlation between the measures being compared. When one uses this approach, the proper SD(diff) can vary from as small as approximately 10 points (for "tight" or highly correlated Gc tests) to as high as approximately 27 pts (for "loose" or weekly correlated tests in a cluster).

The information for this SD(diff) metric comes from a classic 1957 article by Payne and Jones (click here) (thanks to Joel S. for brining it to my attention recently). Also, below are two tables that show the different, and IMHO, more appropriate SD(diff) values that should be used when making some example test comparisons on the WISC-IV and WJ-III. (Click on images to enlarge)






As you see in the tables, the 15 (3 if using scaled scores) and 22 (4.5 if scaled scores) rules-of-thumb will only be correct when the correlation between the two tests being compared is of a moderate magnitude. When the correlation between tests being compared is high (when you have a "tight" ability domain) the appropriate SDdiff metric to evaluate differences can be as low as 9.9 points (for 1 SDdiff) and 14.8 (for 1.5 SDdiff) for the Verbal Comp/Gen Info test from the WJ-III Gc cluster or 2.2 scaled score (1 SDdiff) and 3.3 (1.5 SDdiff) when comparing WISC-IV Sim/Vocab.

In contrast, when the ability domain is very wide or "loose", one would expect more variability since traits/tests are not as correlated. In reviewing the above tables one concludes that the very low test correlations for the tests that comprise the WJ-III Gv and Glr clusters produce a 1 SDdiff that is nearly TWICE the 15 point rule of thumb (27-28 points).

I have argued this point with a number of quants (and some have agreed with me) but believe that the proper SS(diff) to be used is not "one size fits all situations." The confidence band and traditional tables of subtest significant difference approaches are psychometrically sound and work when comparing any two tests. However, when the question becomes one of comparing tests where the fundamental issue revolves around the assumption that the tests scores should be similar because they share a common ability (are correlated), then IMHO, we can do better...there is a better way for these situations. We can improve our practice. We can move forward.

This point is analogous to doing simple t-tests of group means. When one has two independent samples the t-test formula includes a standard error term (in the denominator) that does NOT include any correlation/covariance parameter. However, when one is calculating a dependent sample t-test (which means there is a correlation between the scores), the error term incorporates information about the correlation. It is the same concept.....just applied to group vs individual score comparisons.

I urge people to read the 1957 article, review the tables I have provided above, and chew on the issue. There is a better way. The 15/22 SS rule of thumb is only accurate when a certain moderate level of correlation exists between the two tests being compared and when the comparison implies a common factor or ability. If one uses this simplistic rule of thumb practitioners are likely using a much too stringent rule in the case of highly correlated tests (e.g., Gc) and being overly liberal when evaluating tests from a cluster/composite that are low in correlation (what I call ability domain cohesion--click here for prior post that explains/illustrates this concept). The 15/22 SS rule of thumb is resulting in inaccurate decisions regarding the unusualness of test differences when we fail to incorporate information about the correlation between the compared measures. And, even when such differences are found via this method (or the simple score difference method), this does not necessarily indicate that something is "wrong" and the cluster can't be computed or interpreted. This point was recently made clear in an instructional video by Dr. Joel Schneider on sources of variance in test scores that form composites.

If using the recommended SDdiff metric recommended here is to much work, I would recommend that practitioners steer clear of the 15/22 (1/1.5 SD) rule-of-thumb and instead use the tables provided by the test publishers or use the simple SEM confidence band overlap rule-of-thumb. Sometimes simpler may be better.


- iPost using BlogPress from my Kevin McGrew's iPad