Showing posts sorted by relevance for query Cohesion. Sort by date Show all posts
Showing posts sorted by relevance for query Cohesion. Sort by date Show all posts

Friday, August 10, 2012

AP101 Brief # 15: Beyond CHC: Cognitive-Aptitude-Achievement Trait Complex Analysis: Implications for SLD Assessment and Dx




This is the final post in a series of posts clarifying the nature of cognitive, aptitude, achievement ability constructs.  Readers should consult the preceding post (which contains links to all prior background posts) that defined cognitive abilities, aptitudes, achievement abilities, and CHC cognitive-aptitude-achievement trait complexes (CATTCs).  I apologize for not including the reference list.  These posts are snippets of a manuscript in preparation and I like to post to IQs Corner for feedback that I might incorporate in the final manuscript.  References are the last thing I do.

Beyond CHC:  CHC Cognitive-Aptitude-Achievement Trait Complex Analyses

I have previously argued that alternative non-factor analysis methodological (e.g., multidimensional scaling-MDS) and theoretical lenses need to be applied to validated CHC measures to better understand “both the content and processes underlying performance on diverse cognitive tasks” (McGrew, 2005, p. 172).  When MDS “faceted” methods have been applied to data sets previously analyzed by exploratory or confirmatory factor methods, “new insights into the characteristics of tests and constructs previously obscured by the strong statistical machinery of factor analysis emerge.” (Schneider & McGrew, 2012, p. 110).[1]   

Following the methods similar to that explained and demonstrated by Beauducel, Brocke and Liepmann (2001), Beauducel and Kersting (2002), SÜß and Beauducel (2005), Tucker-Drob and Salthouse (2009; this is an awesome example of MDS analyses side-by-side with factor analsysis of the same set of variables) and Wilhelm (2005), I subjected all WJ-R standardization subjects (McGrew, Werder & Woodcock, 1991) who had complete sets of scores (i.e., listwise deletion of missing data) for the WJ-R Broad Cognitive Ability-Extend (BCA-EXT), Reading Aptitude (RAPT), Math Aptitude (MAPT), Written Language Aptitude (WLAPT), Gf-Gc cognitive factors (Gf, Gc, Glr, Gsm, Gv, Ga, Gs), and Broad Reading (BRDG), Broad Math (BMATH), and Broad Written Language (BWLANG) achievement clusters to a Guttman Radex MDS analysis (n = 4,328 subjects from early school years to late adulthood).[2]  MDS procedures have more relaxed assumptions than linear statistical models and allow for the simultaneous analysis of variables that share common variables or tests—a situation that results in non-convergence problems due to excessive multicolinearity when using linear statistical models.  This feature made it possible to explore the degree of similarity of the WJ-R operationalized measures of the constructs of cognitive abilities, general intelligence (g), scholastic aptitudes, and academic achievement, in a single analysis.  That is, it was possible to explore the relations between and among the core elements of CHC-based cognitive-aptitude-achievement trait complexes (CAATC).  The results are presented in Figure 1. [Click on images to enlarge] 


Figure 1 (Click on image to enlarge)

WJ-R MDS Analysis:  Basic Interpretation

In Guttman Radex models, variables closest to the center of the 2-D plots are the most cognitively complex. Also, the variables are located along two continua or dimensions that often have substantive/theoretical interpretations.  The two dimensions in Figure 1 are labeled A<->B and C<->D.  The following is concluded from a review of Figure 1:

--The WJ-R g-measure (BCA-EXT) is almost directly at the center of the plot and is the most cognitive complex variable.  This makes theoretical sense given that it is a composite comprised of 14 tests from 7 of the CHC Gf-Gc cognitive domains.  Proximity to the center of MDS plots is sometimes considered evidence for g.

--Reading and Writing Aptitude (GRWAPT) and MAPT are also cognitively complex.  Both the GRWAPT[3] and MAPT clusters are comprised of four equally weighted tests of four different Gf-Gc abilities—and thus, the finding that they are also among the most cognitively complex WJ-R measures is not surprising.  The CHC Gf-Gc cognitive measures of Gf and Gc are much more cognitively complex that Gv, Glr, Ga and Gsm.[4]

--The A<->B dimension appears to reflect the ordering of variables as per stimulus content, a common finding in MDS analyses.  The cognitive variables on the left-hand side of the continuum midlines (Gv, Glr, Gf, Gs, MAPT) are comprised of measures with predominant visual-figural or numeric/quantitative characteristics.  The majority of the variables on the right-hand of the continuum midline (GRWAPT, Gc, Ga, Gsm, BRDG, BWLANG) are characterized as more auditory-linguistic, language, or verbal.  This visual-figural/numeric/quantitative-to-auditory-linguistic/language/verbal content dimension is very similar to the verbal, figural, and numeric content facets of the Berlin Model of Intelligence Structure (BIS; SÜß and Beauducel, 2005).[5] 

--The C<->D dimension appears to reflect the ordering of variables as per cognitive operations or processes, another common finding in MDS analyses.  The majority of the cognitive variables above the continua midline (Gv, Glr, Ga, Gc, Gsm, BCAEXT, GRWAPT) are comprised primarily of cognitive ability tasks the involve mental processes or operations.  Conversely, although not as consistent, three of the lowest variables below the continua midline are the achievement ability clusters (BRDG BWLANG; BMATH).  Thus, the C<->D dimension is interpreted as representing a cognitive operations/process-to-acquired knowledge/product dimension.

--In contrast to factor analysis, interpretation of MDS is more is more qualitative and subjective.  Variables that may share a common dimension are typically identified as lying on relatively straight lines or planes, in separate quadrants or partitions, or tight groupings (often represented by circles or ovals or connected as a shape via lines).  Inspecting the four quadrants created by the A<->B C<->D dimensions (see Figure 1) suggests the following.  The AC quadrant is interpreted to represent (excluding BCAEXT which is near the center) cognitive operations with visual-figural content (Gv; Glr).  The CB quadrant is interpreted as representing auditory-linguistic/language/verbal content based cognitive operations.  The BC quadrant only includes the three broad achievement clusters, and is thus an achievement or an acquired knowledge dimension.  Finally, the DA quadrant can be interpreted as cognitive operations that involve quantitative operations or numeric stimuli (e.g., Gf is highly correlated with math achievement; McGrew & Wendling, 2010; one-half of the Gs-P cluster is the Visual Matching test which requires the efficient perceptual processing of numeric stimuli—Glr-N).[6]  The interpretation of these four quadrants is very consistent with the BIS content-faceted content-by-operations model research.

--The theoretical interpretation of the two continua and four quadrants provides potentially important insights into the abilities measured by the WJ-R measures.  More importantly, the conclusions provide potentially important theoretical insights into the nature of human intelligence, insights that typically fail to emerge when using factor analysis methods (see Schneider & McGrew, 2012 and SÜß and Beauducel, 2005). In other MDS analyses I have completed, similar visual-figural/numeric/quantitative-to-auditory-linguistic/language/verbal and cognitive operations/process-to-acquired knowledge/product continua dimensions have emerged (McGrew, 2005; Schneider & McGrew, 2012).  When I have investigated a handful of 3-D MDS[7] models the same two dimensions emerge along with a third automatic-to-deliberate/controlled cognitive processing dimension which is consistent with the prominent dual-process models of cognition and neurocognitive functioning (Evans, 2008, 2011; Barrouillet, 2011; Reyna & Brainerd, 2011; Rico & Overton, 2011; Stanovich, West & Toplak, 2011) that are typically distinguished as Type I/II or System I/II  (see Kahneman’s, 2011, highly acclaimed Thinking, Fast and Slow).[8] 

--These higher-order cognitive processing dimensions, which are not present in the CHC taxonomy, suggest that intermediate strata (or dimensions that cut across broad CHC abilities) might be useful additions to the current three-stratum CHC model.  These higher-order dimensions may be capturing the essence of fundamental neurocognitive processes and argue for moving beyond CHC to integrate neurocognitive research to better understand intellectual performance.


WJ-R MDS Analysis:  Cognitive-Aptitude-Achievement Trait Complex (CAATC) Interpretation

Figure 2 is an extension of the results presented in Figure 1.  Two different CAATCs are suggested.  These were identified by starting first with the BMATH and BRDG/BWLANG achievement variables and next connecting these variables to their respective SAPTs (GRWAPT; MAPT).  Next, the closest cognitive Gf-Gc measures that were in the same general linear path were connected (the goal was to find the math and reading related variables that were closest to lying on a straight line).  Ovals encompassing the entire space comprising the two circle-line-circle traces where superimposed on the figure.  A dotted line that represented the approximate bisection of each of the cognitive-aptitude-achievement trait complex vectors was drawn.  Finally, an approximate correlation (r = .55; see Figure 2) between the two multidimensional CAATC was estimated via measurement of the angle between the CAATC vector dotted lines.[9]

Figure 2 (Click on image to enlarge)

As presented in Figure 2, Math and Reading-Writing CAATCs are suggested as a viable perspective from which to view the relations between cognitive abilities, aptitudes, and achievement abilities.  The primary conclusions, insights, and questions are drawn from Figure 1 and 2 are:

--It appears that the potential exists to empirically identify CAATCs via the use of CHC-grounded theory, the extant CHC COG->ACH relations research, and multidimensional scaling.  It also appears possible to estimate the correlation between different trait complexes (see math/reading-writing trait complex r = .55 in Figure 2).  I suggest these preliminary findings may help the field of cognitive-achievement assessment and research better approximate the multidimensional nature of human cognitive abilities, aptitudes, and achievement abilities.

--Although the WJ-R battery is not as comprehensive a measure of CHC abilities as the WJ III, the cognitive abilities within the respective math and reading/writing CAATCs are very consistent with the extant CHC COG->CHC relations research (McGrew & Wendling, 2010; click here for visual-graphic summary).  The reading-writing trait complex (see Figure 2) includes Ga-PC, Gc-LD/VL, and via the GRWAPT, Gs-P, and Gsm-MS, abilities that are listed as domain-general and domain-specific abilities in Figure 3.  In the case of math, the trait complex includes indicators of Gf-RG, Gv-MV, and via the MAPT, Gs-P (Visual Matching, which might also tap Gs-N) and Gc-LD/VL, abilities that are either domain-general or domain-specific for math in Figure 3.  Working memory (Gsm-WM) is not present (as suggested by Figure 3) as the WJ-R battery did not include a working memory cluster that could enter the analysis.


Figure 3 (Click image to enlarge)

--Also of interest are the three WJ-R cognitive factors (Gsm-MS, Glr-MA, Gs-P) that are excluded from the hyperspace representations of the proposed math and reading-writing CAATCs.  Although highly speculative, it may be possible that their separation from the designated trait complexes suggest, that if known to be related to reading-writing or math achievement, their independence from the narrower trait complexes may be an indication that they represent domain-general abilities.  Glr-MA and Gs-P are both listed as domain-general abilities in Figure 3.  Additional work is needed to determine if the independence (from identified CAATCs) of CHC measures known to be significantly related to achievement indicates domain general abilities.  Alternatively, it is very possible, given the previously demonstrated developmental nuances of CHC COG->ACH relations that the results presented in Figures 1 and 2, which used the entire age range of the WJ-R measures, may mask or distort findings in unknown ways.

--Those knowledgeable of the CHC COG->ACH relations research will obviously note the prior inclusion of certain Gv abilities (Vz, SR, MV) in Figure 3 as well as the inclusion of the WJ-R Gv-MV/CS cluster as part of the proposed math CAATC (Figure 2), despite the lack of consistently reported significant CHC Gv-ACH relations.  McGrew and Wendling (2010) recognized that some Gv abilities have clearly been linked to reading and math achievement (especially the later) in non CHC-organized research.  They speculated that the “Gv Mystery” may be due to certain Gv abilities being threshold abilities or that the cognitive batteries included in their review did not include Gv measures that measured complex Gv related Vz or MV processes.  Given this context, it may be an important finding (via the methods described above) that the WJ-R Gv measure is unexpectedly included in the math CAATC.  This may support the importance of Gv abilities in explaining math and concurrently indicate a problem with the operational Gv measures. 

--The long distance from the WJ-R Gv measure to the center of the diagram (see Figure 2) indicates that the WJ-R Gv measure, which included tests classified as indicators of CS and MV, is not cognitively complex.  This conclusion is consistent with Lohman’s seminal review of Gv abilities (Lohman, 1979) where he specifically mentions CS and MV as representing low level Gv processes and “such tests and their factors consistently fall near the periphery of scaling representations, or at the bottom of a hierarchical model” (Lohman, 1979, 126-127).  I advance the hypothesis that the math CAATC in Figure 2 suggests that Gv is a math-relevant domain, but more complex Gv tests (e.g., 3-D mental “mind’s eye” rotation; complex visual working memory), which would be closer to the center of the MDS hyperspace, need to be developed and included in cognitive batteries.  This suggestion is consistent with Wittmann’s concept of Brunswick Symmetry, which, in turn, is founded on the fundamental concept of symmetry which has been central to success in most all branches of science (Wittmann & SÜß, 1999).  The Brunswick Symmetry model argues that in order to maximize prediction or explanation between predictor and criterion variables, one should match the level of cognitive complexity of the variables in both the predictor and criterion space (Hunt, 2011; Wittmann & SÜß, 1999).  The WJ-R Gv-WJ-R BRMATH relation may represent a low (WJ-R Gv)-to-high (WJ-R BMATH) predictor-criterion complexity mismatch, thus dooming any possible significant relation. 

--Researchers and practitioners in the area of SLD should recognize that when third method POSW “aptitude-achievement” discrepancies are evaluated to determine “consistency”, the combination of domain-general and domain-specific abilities that comprise an aptitude for a specific achievement domain in many ways can be considered a mini-proxy for general intelligence (g).  In Figures 1 and 2 the BCA-EXT and MAPT and GRWAPT variables are in close proximity (which also represents high correlation) and are all near the center of the MDS Radex model.  The manifest correlations between the WJ-R BCA-EXT (in the WJ-R data used to generate the CAATCs in Figure 10) and RAPT, WLAPT, and MAPT clusters are .91, .89 and .91, respectively.  This reflects the reality of the CHC COG->ACH research as in both reading and math achievement, cognitive tests or clusters with high g-loadings (viz., measures of Gc and Gf), as well as shared domain-general abilities, are always in the pool of CHC measures associated with the academic deficit.

--However, the placement of GRWAPT and MAPT in the different content/operations quadrants in Figures 1 and 2 suggests that more differentiated CHC-designed achievement domain SAPT measures might be possible to develop.   The manifest correlations between MAPT and the two GRWAPT measures were .82 to .84, suggesting approximately 69 % shared variance.  GRWAPT and MAPT are strongly related SAPTs, yet there is still unique variance in each.  Furthermore, the WJ-R SAPT measures used in this analysis were equally weighted clusters and not the differentially weighted clusters as in the original WJ.  As presented previously, research suggests that optimal SAPT prediction requires developmentally shifting weights across age.  It is my opinion that the development of developmentally-sensitive CHC-designed SAPTs will result in lower correlations between RAPT and MAPT measures.


Beyond CHC Theory:  Cognitive-Aptitude-Achievement Trait Complexes and SLD Identification Models

The possibility of measuring, mapping and quantifying CAATCs raises intriguing possibilities for re-conceptualizing approaches to the identification of SLD.  Figure 4 presents the generic representation of the prevailing third-method SLD models as well as a formative proposal for a conceptual revision.  As noted previously, the prevailing POSW model (left half of Figure 4), although useful for communication and enhancing understanding of the conceptual approach, is simplistic.   Implementation of the model requires successive calculations of simple (and often multiple) discrepancies which fails to capture the multidimensional and multivariate nature of human cognitive, aptitude, and achievement abilities.  I believe that the CAATC representations in Figure 2, although still clearly imperfect and fallible representations of the non-linear nature of reality, are a better approximation of the complex nature of cognitive-aptitude trait complex relations.  The right-side of Figure 4 is an initial attempt to conceptualize SLD within a CAATC framework.  In this formative model, the bottom two components of the current third-method models (i.e., academic and cognitive weakness) have been combined into a single multidimensional CAATC domain.



Figure 4 (Click on image to enlarge)

CAATCs better operationalize the notion of consistency among the multiple cognitive, aptitude, and achievement elements of an important academic learning domain or domain of SLD.  As noted in the operational definition of a CAATC presented earlier, the emphasis is on a constellation or combination of elements that are related and are combined together in a functional fashion.  These characteristics imply a form of a centrally inward directed force that pulls elements together much like magnetism.  Cohesion appears the most appropriate term for this form of multiple element bonding.  Cohesion is defined, as per the Shorter English Oxford Dictionary (2002), as “the action or condition of sticking together or cohering; a tendency to remain united” (p. 444).  Element bonding and stickiness are also conveyed in the APA Dictionary of Psychology (VandenBos, 2007) definition of cohesion as “the unity or solidarity of a group, as indicated by the strength of the bonds that link group members to the group as a whole” (p. 192).  Thus, in the CAATC-based SLD proposal in Figure 4, the degree of cohesion within a CAATC (as designed by circular icon shape) is considered an integral and critical step to ascertaining if a strong cohesive CAATC, which represents a particular academic domain deficit, is present.  

The stronger the within-CAATC cohesion, the more confidence one could place in the identification of a CAATC as possibly indicative of a SLD.  This focus on quantifying the CAATC cohesion is seen as a necessary, but not sufficient, first step in attempting to identify SLD based on a multivariate POSW.  If the CAATC demonstrates very weak cohesion, the hypothesis of a possible SLD should receive less consideration.  If there is significant (yet to be defined) moderate to strong CAATC cohesion, then the comparison of the CAATC to the cognitive/academic strengths portion of the conceptual model is appropriate for SLD consideration.  To simplify, POSW-based SLD identification would be based first on the identification of a weakness in a cohesive specific CAATC which is then determined to be significantly discrepant from relative strengths in other cognitive and achievement domains.  

Of course, additional variations of this model require further exploration.  For example, should discrepant/discordant comparisons be made between other empirically identified and quantified CAATCs?  Would CAATC-to-CAATC comparisons between high empirical and theoretically correlated CAATCs (e.g., basic reading skills and basic writing skills), when contrasted to less empirically and theoretically correlated CAATC-to-CAATC domains (e.g., basic reading skills and math reasoning), be diagnostically important?  I have more questions than answers at this time.
      
Yes—this proposed framework is speculative and in the formative stages of conceptualization.  It is based on exploratory data analyses, theoretical considerations, and well reasoned logic.  It is not yet ready for applied practice.  Appropriate statistical metrics and methods for operationalizing the degree of domain cohesion are required.  I do not see this as an insurmountable hurdle as methods based on Euclidean distance measures (e.g., Mahalanobis and or Minkowski distance) which can quantify the cohesion between CAATC measures as well as the distance of all the trait complex elements from the centroid of a CAATC exist.  Or, statisticians much smarter than I can might apply centroid-based multivariate statistical measures to quantify and compare CAATC domain cohesion.  I urge those with such skills and interest to pursue the development of these metrics.  Also, the current limited exploratory results with the WJ-R data should be replicated and extended in more contemporary samples with a larger range of both CHC cognitive, aptitude, and achievement tests and clusters.  I would encourage split-sample CAATC model-development and cross-validation in the WJ III norm data.

The proposed CAATC framework, and integration into SLD models is, at this time, simply that—a proposal.  It is not ready for prime-time, in-the-field implementation.  It is presented here as a formative idea that will hopefully encourage others to explore.  Additional research and development, some of which I suggested above, will either prove this to be a promising methodology or an idea with limited validity or one with too many practical constraints that render it hard to implement.  Nevertheless, the results presented here suggest promise.  The results suggest possible incremental progress toward better defining SLD and learning complexes that are more consistent with nature—with the identification of CAATC taxon’s[10] that better approximate “nature carved at the joints” (Meehl, 1973, as quoted and explained by Greenspan, 2006, in the context of MR/ID diagnosis).  Such a development would be consistent with Reynolds and Lakin’s (1987) plea, 25 years ago, for disability identification methods that better represent dispositional taxon’s rather than classes or categories based on specific cutting scores which are grounded in “administrative conveniences with boundaries created out of political and economic considerations” (p. 342). 






[1] See SÜß and Beauducel (2005) and Tucker-Drob and Salthouse (2009) for excellent descriptions of these methods and illustrative results.

[2] The WJ-R battery was analyzed since it was the last version of the WJ series to include scholastic aptitude clusters.

[3] As noted in Figure 1, the Reading and Written Language Aptitude clusters, which were separate variables in the analysis, shared 3 of 4 common tests and nearly overlapped in the MDS plot.  Thus, for simplicity they were combined into the single GRWAPT variable in Figure 1.  This is also consistent the factor analysis of reading and writing achievement variables that typically produce a single Grw factor and not separate reading and writing factors.

[4] The primary narrow abilities measured by each of the cognitive Gf-Gc cluster are included in the label for each cluster.  Contrary to the WJ III, the Gf-Gc clusters were not all operationally constructed as broad Gf-Gc abilities (see McGrew, 1997; McGrew & Woodcock, 2001).  Only the WJ-R Gf and Gc clusters can be interpreted as measuring broad domains as per the requirement that broad measures must include indicators of different narrow abilities (e.g., Concept Formation-I and Analysis-Synthesis-RG).  The other five WJ-R Gf-Gc clusters are now understood to be valid indicators of narrow CHC abilities (Gsm-MS; Ga-PC; Glr-MA; Gv-MV/CS; Gs-P).

[5]  The BIS model is a heuristic framework, derived from both factor analysis and MDS facet analysis, for the classification of performance on different tasks and is not to be considered a trait-like structural model of intelligence as exemplified by the factor-based CHC theory.  Nevertheless, Guttman Radex MDS models often show strong parallels to hierarchical factor based models based on the same set of variables (Kyllonen, 1996; SÜß & Beauducel, 2005; Tucker-Drob & Salthouse, 2009).

[6] The MAPT cluster also includes the two Gf tests and Visual Matching.

[7] WJ III 3-D MDS model for norms subjects aged 9-13 is available at http://www.iqscorner.com/2008/10/wj-iii-guttman-radex-mds-analysis.html

[8] A similar dimension emerged as a plausible higher-order cognitive processing dimension in the previously mentioned Carroll type analysis of 50 WJ III test variables.

[9] Using trigonometry, the cosine of the intersection of the two trait complex vectors was converted to a correlation.  I thank Dr. Joel Schneider for helping fill the gap in my long-lost expertise in basic trigonometry via an excel spreadsheet that converted the measured angle to a correlation.

[10] The Shorter Oxford English Dictionary defines a taxon as “a taxonomic group of any ran, as species, family, class, etc; an organism contained in such a group” (p. 3193) and taxonomy as “classification, esp. in relation to its general laws or principles; the branch of science, or of a particular science or subject, that deals with classification; esp. the systematic classification of living organisms” (p. 3193; italics in original)

Monday, March 28, 2011

Cognitive ability domain cohesion-why composite scores comprised of significantly different subtest scores are still valid

Some excellent discussion has been occurring on the NASP and CHC listservs in response to the "Just say no to averaging IQ subtest scores" blog post and report.

An issue/question that has surfaced (not for the first time) is why markedly discrepant subtest scores that form a composite can still be considered valid indicators of the construct domain. Often clinicians believe that if there is a significant and large discrepancy between tests within a composite, the total score should be considered invalid.

The issue is complex and was touched on briefly in our report and in the NASP and CHC threads by Joel Schneider. Here I mention just ONE concept for consideration.

Below is a 2-D MDS analysis of the WJ III Cog/Ach tests for subjects aged 6-18 in the norm sample. MDS also finds structure as does factor analysis. This 2D model is based on the analysis of the tests correlation matrix. What I think is a major value of MDS, and other spatial statistics, is that one can "see" the numerical relations between tests. Although the metrics are not identical, the visual-spatial map of the WJ III tests does, more-or-less, mirror the intercorrelations between tests. [Double click on image to enlarge]




So....take a look at the Gc, Grw, or Gq tests in this MDS map. All of these tests cluster closely together. Inspection of their intercorrelations finds high correlations among all measures. Conversely, look at the large amount of spatial territory covered by the WJ III Gv tests. Also look at the Ga tests (note that a red line is not connecting Auditory Attention, AA, down in the right-hand quadrant with the other Ga tests). Furthermore, even though most of the Gsm tests are relatively cohesive or tight, Memory for Sentences is further away from the other Gsm tests.

IMHO, these visual-spatial maps, which mirror intercorrelations, tell us than in humans, not all cognitive/ach domains include narrow abilities that are highly interrcorrrelated. I call it "ability domain cohesion." Clearly the different Gv abilities measured by the WJ III Gv tests indicate that the Gv domain is less cohesive (less tight) than the Gc or Grw domain. This does not suggest the tests are flawed..instead it tells us about the varying degrees of cohesiveness present in different ability domains.

Thus, for ability domains that are very very broad (in terms of domain cohesion--e.g., Gv and Ga in this MDS figure), wildly different test scores (e.g., between WJ III Spatial Relations, SR, and Picture Recognition, PR) may be valid and simply reflect that inherent lower cohesiveness (tightness) of these ability domains in human intelligence. Thus, if a person is significantly different in his/her respective Gv SR or PR scores, and these scores are providing valid indications of their relative standing on these measured abilities, then combining them together is appropriate and reflects a valid estimate of the Gv domain....which by nature is broad...and people will often display significant within-domain variability.

Bottom line. Composite scores produced by subtests that are markedly different are likely valid estimates of domains...it is just the nature of human intelligence that some of these domains are more tight or cohesive than others.
- iPost using BlogPress from my Kevin McGrew's iPad

Sunday, August 13, 2017

Fine-Tuning Cross-Battery Assessment Procedures: After Follow-Up Testing, Use All Valid Scores, Cohesive or Not via BrowZine




Another brilliant piece of work by Joel Schneider.  I have been talking about ability domain cohesion for the past decade (http://www.iqscorner.com/search?q=Cohesion)....now Joel has outlined how to deal with the concept psychometrically.  Well done.

Fine-Tuning Cross-Battery Assessment Procedures: After Follow-Up Testing, Use All Valid Scores, Cohesive or Not
Schneider, W. Joel; Roman, Zachary
Journal of Psychoeducational Assessment: Articles in press

University of Minnesota Users:
http://login.ezproxy.lib.umn.edu/login?url=http://journals.sagepub.com/doi/10.1177/0734282917722861

Non-University of Minnesota Users: (Full text may not be available)
http://journals.sagepub.com/doi/10.1177/0734282917722861

Accessed with BrowZine, supported by University of Minnesota.



Monday, October 20, 2008

WJ III: 2-D MDS analysis ages 6-18

As promised, this is a follow-up to my posting of a 3-D Guttman Radex MDS model of WJ III tests. I'm now presenting a 2-D Radex model based on the analysis of all WJ III norm subjects from ages 6-18 (using the WJ III NU norms). You can view/download the pdf file by clicking here.

I could write an entire chapter on implications, hypotheses, etc. Instead, I'm going to make just a few comments and post a few questions in hopes that this approach to examining the characteristics of tests generates some interest. IMHO, MDS is an excellent analytic tool that provides a unique lens by which to augment our factor-analytic based understanding of cognitive ability tests. I wish more of us would complete these analyses with all major intelligence batteries.

A few thoughts/comments/questions:
  • Note that Concept Formation is near the middle of the circle. This whole round of MDS analyses I've been posting is based on a concern (see J. Schneider comments) whether the CF test was a good measure of Gf....and if it was strongly related to g. As per MDS interpretation, the proximity of CF to the middle cross-hairs suggests it is one of the more "cognitively complex" tests in the entire WJ III battery. This would support its interpretation as a strong indicator of Gf and g.
  • Notice that Sound Awareness (Ga/Gsm), Understanding Directions (Gsm/Gc), Applied Problems (Gq/Gf), Quantiative Concepts (Gq/Gf), and Verbal Comprehension (Gc) are also close to the middle - suggesting that they are all cognitively demanding measures in terms of the concept of cogntive complexity. And...interestingly they come from different CHC broad factors. I'm convinced that the reason Sound Awareness and Understanding Directions are cognitively complex is the major working memory load placed on subjects during these tasks. This should serve to remind us that cognitive complexity does NOT necessarily need to be associated with abstract "thinking" (Gf-ish) type tasks. Further notice that Auditory Working Memory is not that far away either. Do these findings support the research that suggest a strong relation between working memory (Gsm-MW) and Gf or g?
  • Notice how "tight" or "cohesive" some of the respective CHC factor tests are. Clearly the Grw, Gq, Gc, Gf, and Gsm (except for MS) tests all tend to hang in the same proximity. In contrast notice the wide degree of distance between the WJ III Gv and Ga tests. Does this suggest that some broad CHC domains are more tight/cohesive while other domains are much broader (lower domain cohesion). What does this mean for test interpretation? What does this mean for understanding the theoretical nature of the different CHC factor domains?
  • Notice the cool cognitive efficency (CE) quandrant. Isn't it sweet how most all the Gs and Gsm tests can be circumscribed in one area. Yet...there is distance between the CE tsts which probably is very informative in understanding differences in the characteristic process/content demands placed on subjects. Isn't this exciting?
  • Is the fact that most Gv tests are far from the cognitive complexity center (as were most of the Wechsler Gv tests in the enclosed slide in the file) helping us understand why traditonal Gv tests are repeatedly found to be unrelated (statistically) to school achievement (esp. mathematics), when we know that considerable research indicates that Gv is important for mathematics. Does this tell us that we have yet, in the world of applied test development, failed to develop sufficiency complex and cognitively demanding Gv tests that would relate more to school achievement (e.g., visual-spatial working memory tests). Curious minds want to know.
  • Like Gv, notice the distances between the Ga tests (which do form a nice psychometric factor when using traditonal factor methods). Incomplete Words is quite far away from Sound Blending, which in turn is closer to the acquired knowledge tests. Does this suggest that IW may be the more "pure" phonetic coding measure while SB is potentially influenced by training and education? Further note the location of Auditory Attention --- I've included it among the cognitive efficiency area. Is this telling us that the sound discrimination (Ga component) of the AA test is minimal while the selective attention (under distraction--ability to resist distractions) component is greater?
  • I'm not comfortable with the interpretation of quandrant 4 in the model. Can others suggest ideas? I think part of the problem is that a 3-D model (like the one I posted the other day) may required to better account for the dimensionality of the complete set of WJ III tests.
I could stare at this forever and generate more thoughts, hypotheses, questions, etc. I'd like to leave that to others. Please feel free to start a thread discussing the potential benefits of examing cognitive and achievement tests via the lens of MDS analysis. It clearly is an under-utilized methodology that can help us better understand our measures. A problem is that most quantoids (myself include) have become seduced by the more sexy contemporary SEM (CFA) methods. Maybe it is time we go "back to the future."


Technorati Tags: , , , , , , , , , , , , , , , , , , , , , , , ,

Thursday, February 18, 2016

How to evaluate the unusualness (base rate) of WJ IV cluster or test score differences: It is a pleasure to use the correct measure - A SlideShare presentation

The WJ IV provides two primary methods for comparing tests or cluster scores.  One is based on a predictive model (the variation and comparison procedures) and the other allows comparisons of SEM confidence bands, which takes into account each measures reliability.  A third method for comparing scores, one that takes into account the correlation between compared measures (ability cohesion model) is not provided, but is frequently used by assessment professionals.  The three types of score comparison methods are described and new information, via a "rule of thumb" summary slide and nomograph, are provided to allow WJ IV users to evaluate scores via all three methods.

A PDF copy of the key WJ IV base rate rule-of-thumb slide can be found here.

Monday, June 20, 2011

IAP 101 Psychometric Brief # 9: The problem with the 1/1.5 SD SS (15/22) subtest comparison "rule-of-thumb"

In regard to my prior "temp" post, I wrote so much in my NASP listserv response that I have decided to take my email response, correct a few typo's, and post it now as blog post. I may return to this later to write a lengthier IAP 101 Research Brief or report.

Psychologists who engage in intelligence testing frequently compare subtest scores to determine if they are statistically and practically different...as part of the clinical interpretation process. Most IQ test publishers provide sound statistical procedures (tables or software for evaluating the statistical difference of two test scores; confidence band comparison rules-of-thumb).

However, traditional and clinical lore has produced a common "rule-of-thumb" that is problematic. The typical scenario is when a clinician subtracts two test SS's (M=100; SD=15) and invokes the rule-of-thumb that the difference needs to be 15 SS points (1 SD) or 22/23 points (1.5 SD). This is not correct.

SS difference scores do NOT have an SD scale of 15! When you subtract two SS's (with mean=100; SD=15) the resultant score distribution has a mean of zero and an SD that is NOT 15 (unless you transform/rescale the distribution to this scale) The size of the difference SD is a function of the correlation between the two measures compared.

The SD(diff) is the statistic that should be used, and there are a number of different forumla for computing this metric. The different SD(diff)'s differ based on the underlying question or assumptions that is the basis for making the comparison.

One way to evaluate score differences is the SEM band overlap approach. This is simple and is based on underlying statistical calculations (averaged across different scenarios to allow for a simple rule of thumb) that incorporates information about the reliability of the difference score. Test publishers also provide tables to evaluate the statistical significance of differences of a certain magnitude for subtests, such as in the various Wechsler manuals and software. These are all psychometrically sound and defensible procedures.......let me say that again...these are all psychometrically sound and defensible procedures. I repeat this phrase as the point I make below was recently misinterpreted at a state SP workshop as me saying there was something wrong with tables in the WISC-IV...which is NOT what I said and is NOT what I am saying here).

However, it is my opinion that in these situations we must do better and there is a more appropriate and better metric for evaluating differences between two different test scores, ESPECIALLY when the underlying assumption is that the two measures should be similar because they form a composite or cluster. This implies "correlation"...and not simple comparison of any two tests.

When one is attempting to evaluate the "unity" of a cluster or composite, an SD(diff) metric should be used that is consistent with the underlying assumption of the question. Namely, one is expecting the scores to be similar because they form a factor. This implies "correlation" between the measures. There is an SD(diff) calculation that incorporates the correlation between the measures being compared. When one uses this approach, the proper SD(diff) can vary from as small as approximately 10 points (for "tight" or highly correlated Gc tests) to as high as approximately 27 pts (for "loose" or weekly correlated tests in a cluster).

The information for this SD(diff) metric comes from a classic 1957 article by Payne and Jones (click here) (thanks to Joel S. for brining it to my attention recently). Also, below are two tables that show the different, and IMHO, more appropriate SD(diff) values that should be used when making some example test comparisons on the WISC-IV and WJ-III. (Click on images to enlarge)






As you see in the tables, the 15 (3 if using scaled scores) and 22 (4.5 if scaled scores) rules-of-thumb will only be correct when the correlation between the two tests being compared is of a moderate magnitude. When the correlation between tests being compared is high (when you have a "tight" ability domain) the appropriate SDdiff metric to evaluate differences can be as low as 9.9 points (for 1 SDdiff) and 14.8 (for 1.5 SDdiff) for the Verbal Comp/Gen Info test from the WJ-III Gc cluster or 2.2 scaled score (1 SDdiff) and 3.3 (1.5 SDdiff) when comparing WISC-IV Sim/Vocab.

In contrast, when the ability domain is very wide or "loose", one would expect more variability since traits/tests are not as correlated. In reviewing the above tables one concludes that the very low test correlations for the tests that comprise the WJ-III Gv and Glr clusters produce a 1 SDdiff that is nearly TWICE the 15 point rule of thumb (27-28 points).

I have argued this point with a number of quants (and some have agreed with me) but believe that the proper SS(diff) to be used is not "one size fits all situations." The confidence band and traditional tables of subtest significant difference approaches are psychometrically sound and work when comparing any two tests. However, when the question becomes one of comparing tests where the fundamental issue revolves around the assumption that the tests scores should be similar because they share a common ability (are correlated), then IMHO, we can do better...there is a better way for these situations. We can improve our practice. We can move forward.

This point is analogous to doing simple t-tests of group means. When one has two independent samples the t-test formula includes a standard error term (in the denominator) that does NOT include any correlation/covariance parameter. However, when one is calculating a dependent sample t-test (which means there is a correlation between the scores), the error term incorporates information about the correlation. It is the same concept.....just applied to group vs individual score comparisons.

I urge people to read the 1957 article, review the tables I have provided above, and chew on the issue. There is a better way. The 15/22 SS rule of thumb is only accurate when a certain moderate level of correlation exists between the two tests being compared and when the comparison implies a common factor or ability. If one uses this simplistic rule of thumb practitioners are likely using a much too stringent rule in the case of highly correlated tests (e.g., Gc) and being overly liberal when evaluating tests from a cluster/composite that are low in correlation (what I call ability domain cohesion--click here for prior post that explains/illustrates this concept). The 15/22 SS rule of thumb is resulting in inaccurate decisions regarding the unusualness of test differences when we fail to incorporate information about the correlation between the compared measures. And, even when such differences are found via this method (or the simple score difference method), this does not necessarily indicate that something is "wrong" and the cluster can't be computed or interpreted. This point was recently made clear in an instructional video by Dr. Joel Schneider on sources of variance in test scores that form composites.

If using the recommended SDdiff metric recommended here is to much work, I would recommend that practitioners steer clear of the 15/22 (1/1.5 SD) rule-of-thumb and instead use the tables provided by the test publishers or use the simple SEM confidence band overlap rule-of-thumb. Sometimes simpler may be better.


- iPost using BlogPress from my Kevin McGrew's iPad