Showing posts with label cognitive complexity. Show all posts
Showing posts with label cognitive complexity. Show all posts

Saturday, November 30, 2024

On making individual tests in #CHC #intelligence test batteries more #cogntivelycomplex: Two approaches



The following information is from a section of the WJ IV techncial manual (McGrew, LaForte & Schrank, 2014) and will again be included in the WJ V technical manual (LaForte, Dailey, McGrew, Q1, 2025).  It was first discussed in McGrew (2012)

On making individual tests in intelligence test batteries more cogntively complex

In the applied intelligence test literature, their are typically two different approaches typically used to increase the cognitive complexity of individual tests (McGrew et al., 2014). The first approach is to deliberately design factorially complex CHC tests, or tests that deliberately include the influence of two or more narrow CHC abilities. This approach is exemplified by Kaufman and Kaufman (2004a) in the development of the Kaufman Assessment Battery for Children–Second Edition (KABC-II), where:

the authors did not strive to develop “pure” tasks for measuring the five CHC broad abilities. In theory, Gv tasks should exclude Gf or Gs, for example, and tests of other broad abilities, like Gc or Glr, should only measure that ability and no other abilities. In practice, however, the goal of comprehensive tests of cognitive abilities like the KABC-II is to measure problem solving in different contexts and under different conditions, with complexity being necessary to assess high-level functioning. (p. 16)

In this approach to test development, construct-irrelevant variance (Benson, 1998; Messick, 1995) is not deliberately minimized or eliminated. Although tests that measure more than one narrow CHC ability typically have lower validity as indicators of CHC abilities, they tend to lend support to other types of validity evidence (e.g., higher predictive validity). The WJ V has several new cognitive tests that use this approach to cognitive complexity. 

The second approach to enhancing the cognitive complexity of tests is to maintain the CHC factor purity of tests or clusters (as much as possible) while concurrently and deliberately increasing the complexity of information processing demands of the tests within the specific broad or narrow CHC domain (McGrew, 2012). As described by Lohman and Lakin (2011), the cognitive complexity of the abilities measured by tests can be increased by (a) increasing the number of cognitive component processes, (b) including differences in speed of component processing, (c) increasing the number of more important component processes (e.g., inference), (d) increasing the demands of attentional control and working memory, or (e) increasing the demands on adaptive functions (assembly, control, and monitoring). This second form of cognitive complexity, not to be confused with factorial complexity, is the inclusion of test tasks that place greater demands on cognitive information processing (i.e., cognitive load), that require greater allocation of key cognitive resources (viz., working memory or attentional control), and that invoke the involvement of more cognitive control or executive functions. Per this second form of cognitive complexity, the objective is to design a test that is more cognitively complex within a CHC domain, not to deliberately make it a mixed measure of two or more CHC abilities.

A large number of prior IQs Corner’s posts regarding the topic of cognitive complexity in intelligence testing can be found here.

Benson, J. (1998). Developing a strong program of construct validation: A test anxiety example. Educational Measurement: Issues and Practice, 17(1), 10–22.

Lohman, D. F., & Lakin, J. (2011). Reasoning and intelligence. In R. J. Sternberg & S. B. Kaufman (Eds.), The Cambridge handbook of intelligence (2nd ed., pp. 419–441). New York, NY: Cambridge University Press.

McGrew, K. S. (2012, September). Implications of 20 years of CHC cognitive-achievement research: Back-to-the-future and beyond CHC. Paper presented at the Richard Woodcock Institute, Tufts University, Medford, MA. (click here to access)

Messick, S. (1995). Validity of psychological assessment: Validation of inferences from persons' responses in performances as scientific inquiry into score meaning. American Psychologist, 50, 741–749.


Friday, November 15, 2024

#WJIV Geometric-Quantoid (#geoquant) #intelligence art: A geoquant interpretation of #cognitive tests is worth a 1000 words—some similar “art parts” will be in #WJV technical manual


(You will need to click on image to enlarge figure to read)

I frequently complete data analyses that never see the light-of-day in a journal article. The results are all I need (at the time) to answer intriguing questions for me, and I then move on…or tantalize psychologists during a workshop or conference presentation.  Thus, this is non-peer reviewed information.  Below is one of my geoquant figures from a series of 2016 analyses (later updated in 2020) I completed on a portion of the WJ IV norm data.  To interpret you should have knowledge of the WJ IV tests—so you can understand the test variable abbbreviation names.  This MDS figure includes numerous interesting cognitive psychology constructs and theoretical principles based on multiple methodological lenses and supporting theory/research.  This was completed before I was introduced to psychometric network analysis methods as yet another visual means to understand intelligence test data.  You can play “where’s Waldo” and look for the following

  • CHC broad cognitive factors
  • Cognitive complexity information re WJ IV tests
  • Kahneman’s two systems of cognition (System I/II thinking)
  • Berlin BIS ability x content facet framework
  • Two of Ackerman’s intelligence dimensions as per PPIK theory (intelligence-as-process; intelligence-as-knowledge)
  • Cattell’s general fluid (gf) and general crystallized (gc) abilities, the two major domains in his five domain triadic theory of intelligence.…..lower case gf/gc notation is deliberate and indicates more “general” capacities (akin, in breadth, to Spearman’s g, who was Cattell’s mentor) and not the Horn and Carroll-like broad Gf and Gc
  • Newlands process and product dominant distinction of cognitive abilities.
Enjoy.  MDS analyses and figures will also be in the forthcoming (Q1 2025)  WJ V technical manual (LaForte, Dailey, & McGrew, 2025, in preparation) but not in the form of these mutiple method/theory synthesis grand figures….stay tunned.  I may create such beautiful geoquant WJ V masterpieces once the WJ V is launched in Q1 2025.  We shall see.  I find these grand synthesis figures particularly useful when interpreting test rests…all critical information in one single figure…would you?

Friday, May 24, 2013

A useful taxonomy for classifying Gf tests: Oliver Wilhelm chapter

This is a post made early in the history of this blog.  Still relevant and important.

In a prior post I summarized a taxonomic lens for analyzing performance on figural/spatial matrix measures of fluid intelligence (Gf). Since then I have had the opportunity to read “Measuring Reasoning Ability” by Oliver Wilhelm (see early blog post on recommended books to read – this chapter is part of the Handbook of Understanding and Measuring Intelligence by Wilhelm and Engle). Below are a few select highlights.

The need for a more systematic framework for understanding Gf measures

As noted by Wilhelm, “there is certainly no lack of reasoning measures” (p. 379). Furthermore, as I learned when classifying tests as per CHC theory with Dr. Dawn Flanagan, the classificaiton of Gf tests as measures of general sequential (deductive) reasoning (RG) inductive reasoning (I), and quantitative reasoning (QR) is very difficult. Kyllonen and Christal’s 1990 statement (presented in the Wilhelm chapter) that the “development of good tests of reasoning ability has been almost an art form, owing more to empirical trial-and-error than to systematic delineation of the requirements which such tests must satisfy” (p.446 in Kyllonen and Christal; p. 379 in Wilhelm). It thus follows that the logical classification of Gf tests is often difficult…or, as we used to say when I was in high school..”no sh____ batman!!!!”

As a result, “scientists and practitioners are left with little advice from test authors as to why a specific test has the form it has. It is easy to find two reasoning tests that are said to measure the same ability but that are vastly different in terms of their features, attributes, and requirements” (p. 379).

Wilhelm’s system for formally classifying reasoning measures

Wilhelm articulates four aspects to consider in the classification of reasoning measures. These are:
  • Formal operation task requirements – this is what most CHC assessment professionals have been encouraged to examine via the CHC lens. Is a test a measure of RG, I, RQ, or a mixture of more than one narrow ability?
  • Content of tasks – this is where Wilhelm’s research group has made one of its many significant contributations during the past decade. Wilhelm et al. have reminded us that just because the Rubik’s cube model of intelligence (Guilford’s SOI model) was found seriously wanting, the analyses of intelligence tests by operation (see above) and content facets is theoretically and empirically sound. I fear that many psychologists, having been burned by the unfulfilled promise of the SOI interpretative framework, have often thrown out the content facet with the SOI bath water. There is clear evidence (see my prior post that presents evidence for content facets based on the analysis of 50 CHC designed measures via a Carroll analyses of the data) that most psychometric tests can be meaningfully classified as per stimulus content – figural, verbal, and quantitative.
  • The instantiation of the reasoning tasks/problems – what is the formal underlying structure of the reasoning tasks? Space does not allow a detailed treatment here, but Wilhelm provides a flavor of this feature when he suggests that one must go through a “decision tree” to ascertain if the problems are concrete vs. abstract. Following the abstract branch, further differentiation might occur vis-à-vis the distinction of “nonsense” vs. “variable” instantiation. Following the concrete branch decision tree, reasoning problem instantiation can be differentiated as to whether they require prior knowledge or not. And so on.
    • As noted by Wilhelm, “it is well established that the form of the instantiation has substantial effects on the difficulty of structurally identical reasoning tasks” (p. 380).
  • Vulnerability of task to reasoning ‘strategies” – all good clinicians know, and have seen, that certain examinees often change the underlying nature of a psychometric task via the deployment of unique metacognitive/learning strategies. I often call this the “expansion of a tests specificity by the examinee.” According to Wilhelm, “if a subgroup of participants chooses a different approach to work on a given test, the consequence is that the test is measuring different abilities for different subgroups…depending on which strategy is chosen, different items are easy and hard, respectively” (p, 381). Unfortunately, research-based protocols for ascertaining which strategies are used during reasoning task performance are more-or-less non-existent.

Ok…that’s enough for this blog post. Readers are encouraged to chew on this taxonomic framework. I do plan (but don’t hold me to the promise…it is a benefit of being the benevolent blog dictator) to summarize additional information from this excellent chapter. Whilhelm’s taxonomy has obvious implications for those who engage in test development. Wilhelm’s framework suggests a structure from which to systematically design/specify Gf tests as per the four dimensions.

On the flip side (applied practice), Whilhelm’s work suggests that our understanding of the abilities measured by existing Gf tests might be facilitated via the classification of different Gf tests as per these dimensions. Work on the “operation” characteristic has been going strong since the mid 1990’s as per the CHC narrow ability classification of tests.

Might not a better understanding of Gf measures emerge if those leading the pack on how to best interpret intelligence tests add (to the CHC operation classifications of Gf tests) the analysis of tests as per the content and instantiation dimensions, as well as identifying the different types of cognitive strategies that might be elicited by different Gf tests by different individuals?

I smell a number of nicely focused and potentially important doctoral dissertations based on the administration of a large collection of available practical Gf measures (e.g., Gf tests from WJ III, KAIT, Wechslers, DAS, CAS, SB5, Ravens, and other prominent “nonverbal” Gf measures) to a decent sample, followed by exploratory and/or confirmatory factor analyses and multidimensional scaling (MDS). Heck….doesn’t someone out there have access to that ubiquitous pool of psychology experiment subjects --- viz., undergraduates in introductory psychology classes? This would be a good place to start.


Tuesday, December 25, 2012

What we've learned from 20 years of CHC COG-ACH relations research: Back to the future and Beyond CHC

A draft of the paper I presented at the 1st Richard Woodcock Institute on Advances in Cognitive Assessment (this past spring at Tufts) can now be read by clicking here. Three of the 12 figures are included below......as a tease :). The final paper will be published by WMF Press.

 

Sunday, November 25, 2012

Implications of 20 Years of CHC Cognitive-Achievement Research: Back-to-the-Future and Beyond CHC

[Click image to enlarge]
 
The key slides from my presentation at the first Richard Woodcock Institute on Cognitive Assessment are now posted at SlideShare.  I thought I had posted these before, but I can't seem to find them.  So here they are for the first (or second) time.  Below is the abstract for the paper that I also submitted--to be published eventually by the WMF Press.


Much has been learned about CHC CHC COG-->ACH relations during the past 20 years (McGrew & Wendling’s, 2010).  This paper built on this extant research by first clarifying the definitions of abilities, cognitive abilities, achievement abilities, and aptitudes.  Differences between domain-general and domain-specific CHC predictors of school achievement were defined.   The promise of Kafuman’s “intelligent” intelligence testing approach was illustrated with two approaches to CHC-based selective referral-focused assessment (SRFA).  Next, a number of new intelligent test design (ITD) principles were described and demonstrated via a series of exploratory data analyses that employed a variety of data analytic tools (multiple regression, SEM causal modeling, multidimensional scaling).  The ITD principles and analyses resulted in the proposal to construct developmentally-sensitive CHC-consistent scholastic aptitude clusters, measures that can play an important role in contemporary third method (pattern of strength and weakness) approaches to SLD identification. 
The need to move beyond simplistic conceptualizations of COG COG-->ACH relations and SLD identification models was argued and demonstrated via the presentation and discussion of CHC COG-->ACH causal SEM models.  Another example was the proposal to identify and quantify cognitive-aptitude-achievement trait complexes (CAATCs).  A revision in current PSW third-method SLD models was proposed that would integrate CAATCs.  Finally, the need to incorporate the degree of cognitive complexity of tests and composite scores within CHC domains in the design and organization of intelligence test batteries (to improve the prediction of school achievement) was proposed.  The various proposals presented in this paper represented a mixture of (a) a call to return to old ideas with new methods (Back-to-the-Future) or (b) the embracing of new ideas, concepts and methods that require psychologists to move beyond the confines of the dominant CHC taxonomy of human cognitive abilities (i.e., Beyond CHC).




Monday, September 10, 2012

AP101 Brief # 16: Beyond CHC: Within-CHC Domain Complexity Optimized Measures

[Note:  This is a working draft of a larger paper (Implications of 20 years of CHC Cognitive-Achievement Research:  Back-to-the-future and Beyond CHC) that will be presented at the first Inaugural Session of the Richard Woodcock Institute for Advancement of Contemporary Cognitive Assessment at Tufts University (Sept, 29, 2012):  The Evolution of CHC Theory and Cognitive Assessment).]   Working knowledge of the WJ III test batery will make this brief easier to understand, but is not necessary.

Beyond CHC:  ITD—Within-CHC Domain Complexity Optimized Measures
            Optimizing Cognitive Complexity of CHC measures
I have recently begun to recognize the contribution that The Brunswick Symmetry derived Berlin Intelligence Structure (BIS) model can make in applied intelligence research, especially for increasing predictor-criteria relations by maximizing these relations via matching the predictor-criteria space on the dimension of cognitive complexity.  What is cognitive complexity?  Why is it important?  More important, what role should it play in designing intelligence batteries to optimize CHC COG-ACH relations?
Cognitive complexity is often operationalized by inspecting individual test loadings on the first principal component from principal component analysis (Jensen, 1998).  The high g-test rationale is that performance on tests that are more cognitively complex “invoke a wider range of elementary cognitive processes (Jensen, 1998; Stankov, 2000, 2005)” (McGrew, 2010b, p. 452).  High g-loading tests are often at the center of MDS (multidimensional scaling) radex models (click here for AP101 Brief Report #15:  Cognitive-Aptitude-Achievement Trait Complexes example)—but this isomorphism does not always hold.   David Lohman, a student of Richard Snow’s, has made extensive use of MDS methods to study intelligence and has one of the best grasps of what cognitive complexity, as represented in the hyperspace of MDS figures, contributes to understanding intelligence and intelligence tests.  According to Lohman (2011), those tests closer to the center are more cognitively complex due five possible factors—larger number of cognitive component processes; accumulation of speed component differences: more important component processes (e.g., inference); increased demands of attentional control and working memory; and/or or more demands on adaptive functions (assembly, control, and monitoring).  Schneider’s (in press) level of abstraction description of broad CHC factors is similar to cognitive complexity.  He uses the simple example of 100 meter hurdle performance.  According to Schneider (in press), one could independently measure 100 meter sprinting speed and then standing still and jumping over a hurdle (both examples of narrow abilities).  However, running a 100 meter race is not the mere sum of the two narrow abilities and as is more of a non-additive combination and integration of narrow abilities.  This analogy captures the essence of cognitively complexity—which, in the realm of cognitive measures, are tasks that have more of the five factors listed by Lohman involvedduring successful task performance.
Of critical importance is the recognition that factor or ability domain breadth (i.e., broad or narrow) is not synonymous with cognitively complexity.  More important, cognitive complexity has not always been a test design concept (as defined by the Brunswick Symmetry and BIS model) explicitly incorporated into "intelligent" intelligence test design (ITD).  A number of tests have incorporated the notion of cognitive complexity in their design plans, but I believe this type of cognitive complexity is different than the within-CHC domain cognitive complexity discussed here.
For example, according to Kaufman and Kaufman (2004), “in developing the KABC-II, the authors did not strive to develop ‘pure’ tasks for measuring the five CHC broad abilities.  In theory, Gv tasks should exclude Gf or Gs, for example, and tests of other broad abilities, like Gc or Glr, should only measure that ability and none other.  In practice, however, the goal of comprehensive tests of cognitive ability like the KABC-II is to measure problem solving in different contexts and under different conditions, with complexity being necessary to assess high-level functioning” (p. 16; italics emphasis added).  Although the Kaufman’s address the importance of cognitively complex measures in intelligence test batteries, their CHC-grounded description defines complex measures as those that are factorially complex or mixed measures of abilities from more than one broad CHC domain.  The Kaufman’s also address cognitive complexity from the non-CHC neurocognitve three-block functional Luria neurocognitive model when they indicate that it is important to provide measurement that evaluates the “dynamic integration of the three blocks” (Kaufman & Kaufman, 2004, p.13).   This emphasis on neurocognitive integration (and thus, complexity) is also an explicit design goal of the latest Wechsler batteries.  As stated in the WAIS-IV manual (Wechsler, 2008), “although there are distinct advantages to the assessment and division of more narrow domains of cognitive functioning, several issues deserve note.  First, cognitive functions are interrelated, functionally and neurologically, making it difficult to measure a pure domain of cognitive functioning” (p. 2).  Furthermore, “measuring psychometrically pure factors of discrete domains may be useful for research, but it does not necessarily result in information that is clinically rich or practical in real world applications (Zachary, 1900)” (Wechsler, p. 3).   Finally, Elliott (2007) similarly argues for the importance of recognizing neurocognitive-based “complex information processing” (p. 15; italics emphasis added) in the design of the DAS-III, which results in tests or composites measuring across CHC-described domains, as important in test design.
The ITD principle explicated and proposed here is that of striving to develop cognitively complex measures within broad CHC domains—that is, not attaining complexity via the blending of abilities across CHC broad domains and not attempting to directly link to neurocognitive network integration.[1]   The Brunswick Symmetry based BIS model provides a framework for attaining this goal via the development and analysis of test complexity by paying attention to cognitive content and operations facets. 
Figure 12 presents the results of a 2-D MDS Radex model of most all key WJ III broad and narrow CHC cognitive and achievement clusters (for all norm subjects from approximately 6 years of age thru late adulthood). [2]   The current focus of the interpretation of the results in Figure 12 is only on the degree of cognitive complexity (proximity to the center of the figure) of the broad and narrow WJ III clusters within the same domain (interpretations of the content and operations facets are not a focus of this current material).  Within a domain the broadest three-test parent clusters are designated by black circles.[3]  Two-test broad clusters are designed by gray circles.  Two test narrow offspring clusters within broad domains are designated by white circles.  All clusters within a domain are connected to the broadest parent broad cluster by lines.  The critically important information is the within-domain cognitive complexity of the respective parent and sibling clusters as represented by their relative distances from the center of the figure.  A number of interesting conclusions are apparent. [Click on image to enlarge]

First, as expected, the WJ III GIA-Ext cluster is almost perfectly centered in the figure—it is clearly the most cognitively complex WJ III cluster.   In comparison, the three WJ III Gv clusters are much weaker in cognitive complexity than all other cognitive clusters with no particular Gv cluster demonstrating a clear cognitive complexity advantage.    As expected, the measured reading and math achievement clusters are primarily cognitively complex measures.  However, those achievement clusters that deal more with basic skills (Math Calculation—MTHCAL; Basic Reading Skills—RDGBS) are less complex that the application clusters (Reading Comprehension-RDGCMP; Math Reasoning-MTHREA). 
The most intriguing findings in Figure 12 are the differential cognitive complexity patterns within CHC domains (with at least one parent and at least one offspring cluster).  For example, the narrow Perceptual Speed (Gs-P) offspring cluster is more cognitively complex than the broad parent Gs cluster.  The broad Gs cluster is comprised of the Visual Matching (Gs-P) and Decision Speed (Gs-R9; Glr-NA) tests, tests that measure different narrow abilities.  In contrast the Perceptual Speed cluster (Gs-P) is comprised of two tests that are classified as both measuring the same narrow ability (perceptual speed).  This finding appears, on first blush, counterintuitive as one would expect a cluster comprised of tests that measure different content and operations (Gs cluster) would be more complex (as per the above definition and discussion) than one comprised of two measures of the same narrow ability (Gs-P).  However, one must task analyze the two Perceptual Speed tests to realize that although both are classified as measuring the same narrow ability (perceptual speed), they differ in both stimulus content and cognitive operations.  Visual Matching requires processing of numeric stimuli.  Cross Out requires the processing of visual-figural stimuli.  These are two different content facets in the BIS model.  The Cross Out visual-figural stimuli are much more spatially challenging than the simple numerals in Visual Matching.  Furthermore, the Visual Matching test requires the examinee to quickly seek out and discover and mark two digit pairs that are identical.  In contrast, in the Cross Out test the subject is provided a target visual-figural shape and the subject must then quickly scan a row of complex visual images and mark two that are identical to the target.  Interesting, in other unpublished  analyses I have completed, the Visual Matching test often loads on or groups with quantitative achievement tests while Cross Out has frequently show to load on a Gv factor.  Thus, task analysis of the content and cognitive operations of the WJ III Perceptual Speed tests suggests that although both are classified as narrow indicators of Gs-P, they differ markedly in task requirements.  More important, the Perceptual Speed cluster tests, when combined, appear to require more cognitively complex processing than the broad Gs cluster.  This finding is consistent with Ackerman, Beier and Boyle’s (2002) research that suggests that perceptual speed has another level of factor breadth via the identification of four subtypes of perceptual speed (i.e., pattern recognition, scanning, memory and complexity; see McGrew 2005 and Schneider & McGrew, 2012 for discussion of a hierarchically organized model of speed abilities).  Based on Bruinswick Symmetry/BIS cognitive complexity principles, one would predict that a Gs-P cluster comprised of two parallel forms of the same task (e.g., two Visual Matching or two Cross Out tests) would be less cognitively complex than broad Gs.  A hint of the possible correctness of this hypothesis is present in the inspection of the Gsm-MS-MW domain results.
The WJ III Gsm cluster is the combination of the Numbers Reversed (MW) and Memory for Words (MS) tests.  In contrast, the WJ III Auditory Memory Span cluster (AUDMS; Gsm-MS) cluster is much less cognitively complex when compared to Gsm (see Figure 12).  Like the Perceptual Speed (Gs-P) cluster described in the context of the processing speed family of clusters, the Auditory Memory Span cluster is comprised of two tests with the same memory span (MS) narrow ability classification (Memory for Words; Memory for Sentences).  Why is this narrow cluster less complex than its broad parent Gsm cluster while the opposite held true for Gs-P and Gs?  Task analysis suggests that the two memory span tests are more alike than the two perceptual speed tests.  The Memory for Words and Memory Sentences tests require the same cognitive operation—simply repeating back, in order, words or sentences spoken to the subject.  This differs from the WJ III Perceptual Speed cluster as the similarly classified narrow Gs-P tests most likely invoke both common and different cognitive component operations.  Also, the Memory Span cluster tests are comprised of stimuli from the same BIS content facet (i.e., words and sentences; auditory-linguistic/verbal).  In contrast, the Gs-P Visual Matching and Cross Out tests involve two different content facets (numeric and visual-figural).
In contrast, the WJ III Working Memory cluster (Gsm-MW) is more cognitively complex than the parent Gsm cluster.  This finding is consistent with the prior WJ III Gs/Perceptual Speed and WJ III Gsm/Auditory Memory Span discussion.  The WJ III Working Memory cluster is comprised of the Numbers Reversed and Auditory Working Memory tests.  Numbers Reversed requires the processing of stimuli from one BIS content facet—numeric stimuli.  In contrast, Auditory Working Memory requires the processing of stimuli from two BIS content factors—numeric and auditory-linguistic/verbal; numbers and words).  The cognitive operations of the two tests also differ.  Both require the holding of the presented stimuli in active working memory space.  Numbers Reversed then requires the simple reproduction of the numbers in reverse order.  In contrast, the Auditory Working Memory test requires the storage of the numbers and words in separate chunks, and then the production of the forward sequence of each respective chunk (numbers or words), one chunk before the other.  Greater reliance on divided attention is most likely occurring during the Auditory Working Memory test. 
In summary, the results presented in Figure 12 suggest that it is possible to develop cluster scores that vary by degree of cognitively complexity within the same broad CHC domain.  More important is the finding that the classification of clusters as broad or narrow does not provide information on the measures cognitive complexity.  Cognitively complexity, as defined in the classification of clusters as broad or narrow does not provide information on the measures cognitive complexity.  Cognitive complexity, as in the Lohman sense, can be achieved within CHC domains without resorting to mixing abilities across CHC domains.  Finally, narrow clusters can be more cognitively complex, and thus likely better predictors of complex school achievement, than broad clusters or other narrow clusters. 

Implications for Test Battery Design and Assessment Strategies
The recognition of cognitive complexity as an important ITD principle suggests that the push to feature broad CHC clusters in contemporary test batteries, or in the construction of cross-battery assessments, fails to recognize the importance of cognitive complexity.  I plead guilty to contributing to this focus via my role in the design of the WJ III which focused extensively on broad CHC domain construct representation—most WJ III narrow CHC clusters require the use of the third WJ III cognitive book (the Diagnostic Supplement; Woodcock, McGrew, Mather & Schrank, 2003).  Similarly, guilty as charged in the dominance of broad CHC factor representation in the development of the original cross-battery assessment principles (Flanagan & McGrew, 1997; McGrew & Flanagan, 1998). 
It is also my conclusion that the narrow is better conclusion of McGrew and Wendling (2010) may need modification.   Revisiting the McGrew and Wendling (2010) results suggest that the narrow CHC clusters that were more predictive of academic achievement likely may have been so not necessarily because they are narrow, but because they are more cognitively complex.  I offer the hypothesis that a more correct principle is that cognitively complex measures are better.   I welcome new research focused on testing this principle.
In retrospect, given the universe of WJ III clusters, a broad+narrow hybrid approach to intelligence battery configuration (or cross-battery assessment) may be more appropriate.  Based exclusively on the results presented in Figure 12, the following clusters would appear those that might better be featured in the “front end” of the WJ III or a selective testing constructed assessment—those clusters that examiners should consider first within each CHC broad domain:  Fluid Reasoning (Gf)[4], Comprehension-Knowledge (Gc), Long-term Retrieval (Glr), Working Memory (Gsm-MW), Phonemic Awareness 3 (Ga-PC), and Perceptual Speed (Gs-P).  No clear winner is apparent for Gv, although the narrow Visualization cluster is slightly more cognitively complex than the Gv and Gv3 clusters.  The above suggests that if broad clusters are desired for the domains of Gs, Gsm and Gv, then additional testing beyond the “front end” or featured tests and clusters would require administration of the necessary Gs (Decision Speed), Gsm (Memory for Words) and Gv (Picture Recognition) tests.

Utilization of the ITD test design principle of optimizing within-CHC cognitively complexity of clusters suggests that a different emphasis and configuration of WJ III tests might be more appropriate.  It is proposed that the above WJ III cluster complexity priority or feature model would likely allow practitioners to administer the best predictors of school achievement.  I further hypothesize that this cognitive complexity based broad+narrow test design principle most likely applies to other intelligence test batteries that have adhered to the primary focus on featuring tests that are the purest indicators of two or more narrow abilities within the provided broad CHC interpretation scheme.  Of course, this is an empirical question that begs research with other batteries.  More useful with be similar MDS Radex cognitive complexity analysis of cross-battery intelligence data sets.[5]

References (not included in this post.  The complete paper will be announced and made available for reading and download in the near future)



[1] This does not mean that cognitive complexity may not be related to the integrity of the human connectome or different brain networks. I am excited about contemporary brain network research (Bressler & Menon, 2010; Cole, Yarkoni, Repovs, Anticevic & Braver, 2012; Toga, Clark, Thompson, Shattuck, & Van Horn, 2012; van den Heuvel & Sporns, 2011), particularly that which has demonstrated links between neural network efficiency and working memory, controlled attention and clinical disorders such as ADHD (Brewer, Worunsky, Gray, Tang, Weber & Kober, 2011; Lutz, Slagter, Dunne, & Davidson, 2008; McVay & Kane, 2012). The Parietal-Frontal Integration (P-FIT) theory of intelligence is particularly intriguing as it has been linked to CHC psychometric measures (Colom, Haier, Head, Álvarez-Linera, Quiroga, Shih, & Jung, 2009; Deary, Penke, & Johnson, 2010; Haier, 2009; Jung & Haier, 2007) and could be linked to CHC cognitively-optimized psychometric measures.
[2] Only reading and math clusters were included to simplify the presentation of the results and the fact, as reported previously, that reading and writing measures typically do not differentiate well in multivariate analysis—and thus the Grw domain in CHC theory.
[3] GIA-Ext is also represented by a black circle.
[4] Although the WJ III Fluid Reasoning 3 cluster (Gf3) is slightly closer to the center of the figure, the difference from Fluid Reasoning (Gf) is not large and time efficiency would argue for the two-test Gf cluster.
[5] It is important to note that the cognitive complexity analysis and interpretation discussed here is specific to within the WJ III battery only. The degree of cognitive complexity in the WJ III cognitive clusters in comparison to composite scores from other intelligence batteries can only be ascertained by cross-battery MDS complexity analysis.

Wednesday, November 11, 2009

MDS analysis of the WJ III: Implications for CHC theory refinement and extension




IAP AP101 # 3 report is now available (click here for all AP101 reports and briefs).  "IAP AP101 Report #3: MDS Analysis of the CHC-based WJ III Battery: Implications for possible refinements and extensions of the CHC model of human intelligence" can be viewed  or downloaded by clicking here.

The PPT files are also viewable and downloadable via SlideShare.

Abstract
The WJ III Battery is comprised of both cognitive (intelligence) and achievement components.  As reported in the technical manual, the Cattell-Horn-Carroll (CHC) theory of cognitive abilities organizational structure of the WJ III has been validated.  The current investigation analyzed the cognitive and achievement tests for all WJ III norm subjects from ages 6-18 years of age.  Multidimensional scaling (MDS—Guttman Radex model) of the 50 WJ III tests suggested new facets from which to interpret the WJ III.  The results suggested three to four higher-order intermediate CHC model stratum abilities that varied along the dimensions of (a) controlled vs automatic cognitive processing and (b) product- vs process-dominant abilities. The results, together with recent similar analysis of the WAIS-IV, support Woodcock’s Cognitive Performance Model (CPM).  Implications for possible minor changes in the CPM model are suggested.  More importantly, the WJ III and WAIS-IV results collectively suggest hypothesized refinements and extensions of the CHC intelligence framework.   Research focused on exploring the compatibility of a combined CHC and Berlin Model of Intelligence Structure (BIS) theory is recommended.
Technorati Tags: , , , , , , , , , , , , , , , , , , ,


Thursday, October 23, 2008

WISC-III/WJ III cross-battery Guttman 2-D Radex analysis

One more WISC-III/WJ III cross-battery analysis--this time a 2-D Guttman Radex MDS model (click here).  As readers have noted, I've been on a bit of a data analysis binge this past week (in preparation for writing a manuscript---and after being refreshed by an actual 2+ week vacation) and have reported:  (a) WISC-III/WJ III cross-battery g+specific cog-ach relations SEM, (b) WJ III 2-D Guttman Radex MDS of WJ III norm sample ages 6-8, and (c) WJ III 3-D Guttman Radex MDS of ages 9-13 of norm sample.  It is hoped these analyses provide useful information in understanding the characteristics of the tests in the WJ III and Wechsler intelligence batteries.

Unfortunately this analysis is based on the WISC-III and not the more recent WISC-IV.  Nevertheless, the results still provide useful information on the WISC-III tests that are still present in the WISC-IV.

Given all I've written regarding the various MDS models, I'm going to only make a few comments and hope others take the presentation of these data to engage in additional discussion, interpretion, etc.----have some fun.

A few observations/comments:
  • Gv tests (both WISC-III and WJ III) continue to surface on the more outer rings of the MDS models---suggesting that they are more lower-level perceptual/processing measures and do not capture complex Gv cognitive processing.  See my Gv comments on this the other day.  The same can be said for Ga tests.
  • WJ III Understanding Directions is consistently one of the more cognitive complex tests.  And, it is largely a language-based measure of working memory (Gsm-MW).  Remember that as per the Radex model, cognitive complexity deals with the amount of elements/components that are processed.....and is not the same as abstract thinking (Gf-ish stuff).  WJ III Numbers Reversed also shows up close to the center, with WISC-III Digit Span not far behind.  Does this support the popular working memory=Gf/g research hypothesis?
  • The Gc tests from both batteries appear similar in placement.
  • As would be expected, the WJ III Gf tests (Concept Formation, Numerical Reasoning [which is a combo of Number Series and Number Matrices], and Analysis-Synthesis are within the center "cognitive complexity" circle.
I'm sure there is much more that can be gleaned, but I'll leave that to the readers to discover, debate, and discuss.  I actually think a 3-D MDS model is necessary to capture the characteristics of the measures...but I've run out of time and steam on these analysis.  Maybe at a later date.

A couple caveats I provided the other day are also relevant here--(a) I'm a coauthor of the WJ III (conflict of interest disclosure) and (b) these results have not been peer-reviewed



Monday, October 20, 2008

WJ III: 2-D MDS analysis ages 6-18

As promised, this is a follow-up to my posting of a 3-D Guttman Radex MDS model of WJ III tests. I'm now presenting a 2-D Radex model based on the analysis of all WJ III norm subjects from ages 6-18 (using the WJ III NU norms). You can view/download the pdf file by clicking here.

I could write an entire chapter on implications, hypotheses, etc. Instead, I'm going to make just a few comments and post a few questions in hopes that this approach to examining the characteristics of tests generates some interest. IMHO, MDS is an excellent analytic tool that provides a unique lens by which to augment our factor-analytic based understanding of cognitive ability tests. I wish more of us would complete these analyses with all major intelligence batteries.

A few thoughts/comments/questions:
  • Note that Concept Formation is near the middle of the circle. This whole round of MDS analyses I've been posting is based on a concern (see J. Schneider comments) whether the CF test was a good measure of Gf....and if it was strongly related to g. As per MDS interpretation, the proximity of CF to the middle cross-hairs suggests it is one of the more "cognitively complex" tests in the entire WJ III battery. This would support its interpretation as a strong indicator of Gf and g.
  • Notice that Sound Awareness (Ga/Gsm), Understanding Directions (Gsm/Gc), Applied Problems (Gq/Gf), Quantiative Concepts (Gq/Gf), and Verbal Comprehension (Gc) are also close to the middle - suggesting that they are all cognitively demanding measures in terms of the concept of cogntive complexity. And...interestingly they come from different CHC broad factors. I'm convinced that the reason Sound Awareness and Understanding Directions are cognitively complex is the major working memory load placed on subjects during these tasks. This should serve to remind us that cognitive complexity does NOT necessarily need to be associated with abstract "thinking" (Gf-ish) type tasks. Further notice that Auditory Working Memory is not that far away either. Do these findings support the research that suggest a strong relation between working memory (Gsm-MW) and Gf or g?
  • Notice how "tight" or "cohesive" some of the respective CHC factor tests are. Clearly the Grw, Gq, Gc, Gf, and Gsm (except for MS) tests all tend to hang in the same proximity. In contrast notice the wide degree of distance between the WJ III Gv and Ga tests. Does this suggest that some broad CHC domains are more tight/cohesive while other domains are much broader (lower domain cohesion). What does this mean for test interpretation? What does this mean for understanding the theoretical nature of the different CHC factor domains?
  • Notice the cool cognitive efficency (CE) quandrant. Isn't it sweet how most all the Gs and Gsm tests can be circumscribed in one area. Yet...there is distance between the CE tsts which probably is very informative in understanding differences in the characteristic process/content demands placed on subjects. Isn't this exciting?
  • Is the fact that most Gv tests are far from the cognitive complexity center (as were most of the Wechsler Gv tests in the enclosed slide in the file) helping us understand why traditonal Gv tests are repeatedly found to be unrelated (statistically) to school achievement (esp. mathematics), when we know that considerable research indicates that Gv is important for mathematics. Does this tell us that we have yet, in the world of applied test development, failed to develop sufficiency complex and cognitively demanding Gv tests that would relate more to school achievement (e.g., visual-spatial working memory tests). Curious minds want to know.
  • Like Gv, notice the distances between the Ga tests (which do form a nice psychometric factor when using traditonal factor methods). Incomplete Words is quite far away from Sound Blending, which in turn is closer to the acquired knowledge tests. Does this suggest that IW may be the more "pure" phonetic coding measure while SB is potentially influenced by training and education? Further note the location of Auditory Attention --- I've included it among the cognitive efficiency area. Is this telling us that the sound discrimination (Ga component) of the AA test is minimal while the selective attention (under distraction--ability to resist distractions) component is greater?
  • I'm not comfortable with the interpretation of quandrant 4 in the model. Can others suggest ideas? I think part of the problem is that a 3-D model (like the one I posted the other day) may required to better account for the dimensionality of the complete set of WJ III tests.
I could stare at this forever and generate more thoughts, hypotheses, questions, etc. I'd like to leave that to others. Please feel free to start a thread discussing the potential benefits of examing cognitive and achievement tests via the lens of MDS analysis. It clearly is an under-utilized methodology that can help us better understand our measures. A problem is that most quantoids (myself include) have become seduced by the more sexy contemporary SEM (CFA) methods. Maybe it is time we go "back to the future."


Technorati Tags: , , , , , , , , , , , , , , , , , , , , , , , ,

Friday, October 17, 2008

WJ III: Guttman Radex MDS analysis

More on "Gf=g revisited" thread (click here for original post) that produced some excellent discussion (click here) on the NASP listserv.

In response to the request for the application of Guttman's Radex MDS model to the Woodcock Johnson III (in age 9-13 norm sample), I looked through my old files and found a 3D MDS WJ III model that I completed a number of years ago. The slides have been posted in a pdf file for viewing. It would take a manuscript to explain and interpret everything....I hope the broad stroke hypotheses (esp. regarding the nature of three dimensions) stimulate some thought and discussion.

Yesterday I completed a new 2D model across all school-age subjects (6-18 years). I hope to post those findings within the week. Stay tunned.

[Conflict of interest disclosure - I'm a coauthor of the WJ III]


Technorati Tags: , , , , , , , , , , , , , , , , , , , , , , , ,


Thursday, October 16, 2008

Gf=g revisited: Schneider, Lopez and Fiorello comment

Joel Schneider provided some excellent thoughts and questions related to my recent "Gf=g: Maybe not" post over on the NASP listserv. His comments were then augmented by Ruben Lopez and Cathy Fiorello. I thought the quality of the comments were so good that they should be preserved for others to read. They are produced below - "as is" from the NASP list.

Joel Schneider comment:

Kevin's recent blog post about the Gf=g hypothesis is interesting and worth reading.

For most hypotheses about the structure of cognitive abilities, I can think of no better dataset on which to test them than on the WJ-III standardization sample. However, in this particular case, I've always had my doubts about WJ-III Gf tests. I am confident that both the primary WJ-III Gf tests are excellent markers of Gf. However, I've always thought that they contained a hint of common variance that was non-Gf related. What that is, I can't quite put my finger on it but it has something to do with executive control of attention. Both involve a need to generate hypotheses and test them in working memory in ways that seem more involved than the traditional matrix Gf tests. Both of them also seem to require math-like thought processes, especially in the more difficult items.

Suppose that the Gf=g hypothesis were true. Let's say that Concept Formation and Analysis-Synthesis both consist of the following sources of variance:

CF = Gf + Something Extra + error
AS = Gf + Something Extra + error

The latent variable that would be constructed to represent Gf in a CFA would thus be: WJ-III Gf = Gf + Something Extra

The chi square test to see if constraining the Gf to g path to 1.0 would be significant, not because the Gf=g hypothesis is wrong, but because the 2 WJ-III Gf subtests were not pure enough markers of Gf. It would only take a little something extra for the chi square test to be significant.

I would think that adding one Raven-like matrix in the Gf mix would reduce the problem (if there actually is a problem). These tests seem less-mathy and more visual-spatialish and thus might reduce the non-Gf common variance.

The tables Kevin links to include a Gf latent variable that consists of:

Concept Formation
Analysis-Synthesis
Numerical Reasoning (Number Matrices + Number Series?)
Applied Problems
Quantitative Concepts

If I am right about CF and AS being mathy and if mathiness is not exactly the same as Gf, then this WJ-III Gf is likely to be WJ-III Gf = Gf + Mathiness

I was very surprised to see how strong an indicator of Gf Quantitative Concepts is, given its Gc-like question format. Perhaps it is glomming onto Gf not because it has a lot of Gf in it but because it is attracted to the math-like elements of the other indicators. Even so, I am very much at a loss to understand why Quantitative Concepts has a higher loading on Gf than does Applied Problems.


Ruben Lopez responds:

Hi Joel,

Maybe the messiness of Gf's measurability even in an exceptional measure like the WJ-III--may have more to do with abstraction and its relationship to g than with a separate Gf.

Consider Dr. David Lohman's discussion of Gf's relationship to Gq in "The Woodcock-Johnson III and the Cognitive Abilities Test (Form 6): A concurrent validity study" (March 2003):

"Recent discussions of the nature of general ability have emphasized the importance of physiological processes (Jensen, 1998), the role of working memory (Kyllonen, 1996), or the congruence between aprimary Inductive Reasoning factor, the stratum II Fluid Ability factor (Gf), and g (Gustafsson, 2002). However, the present study supports Keith and Witta's (1997) hypothesis that quantitative reasoning may be an even better indicator of g. Quantitative reasoning has always been represented in some form in achievement test batteries, and in aptitude tests (such as the SAT) designed to predict academic success. But a broad quantitative knowledge factor (Gq) was not added to Gf-Gc theory until the late 1980s (Horn, 1989). Carroll's (1993) three-stratum theory, on the other hand, considers quantitative reasoning to be part of a broad fluid reasoning (Gf) factor. Confirmatory factor analyses of different ability test batteries mirror this ambivalence. Some studies find g and Gq indistinguishable [as in Keith & Bickley's (1992) factor analysis of the Stanford-Binet IV or Lohman & Hagen's (2002) factor analyses of the CogAT Primary Battery], other studies find Gq to be the best indicator of g (as in Keith & Witta's (1997) factor analyses of the WISC-III or Lohman & Hagen's (2002) factor analyses of the CogAT Multilevel Battery], and yet other studies find distinguishable g and Gq factors [as in Bickley, Keith, & Wolfe's (1995) factor analysis of the Woodcock-Johnson Psychoeducational Battery-Revised].

Paradoxically quantitative reasoning has not been much studied because it is difficult to separate from g unless combined with tests of more specific mathematical knowledge and skill (as in the Gq factor). But it is this overlap with g that makes quantitative reasoning particularly interesting as a vehicle for understanding the nature of g. Perhaps the most salient characteristic of quantitative concepts is abstraction. Even elementary operations like counting require abstraction: two cats are in some way the same as two dogs or two anything. The number line itself is an abstraction, especially when it includes negative numbers. Abstraction is most obvious in understanding concepts such as variable or, later, imaginary number.

Several early definitions of g emphasized abstract thinking or reasoning abilities. And the transition from concrete to abstract thinking figured prominently in Piaget's theory of intelligence. Modern definitions of g emphasize the importance of working memory resources or even of reasoning, but do not have much to say about the role of abstract thinking. These analyses suggest a closer study of quantitative reasoning might be a good place to begin in exploring this possibility.
" (p. 16).

And don't forget Keith and colleagues recommendation that the Arithmetic subtest should be added to the Perceptual Reasoning scale to assess Gf.

Cathy Fiorello chimes in:

Folks may be interested in looking at Guttman's model of intelligence in this context. Some colleagues and I had an article in Intelligence a couple of years ago (Cohen, Fiorello, & Farley, maybe 2006?) with a Smallest Space Analysis of the WISC-IV. Guttman's model was supported, which considers level of abstraction as one dimension of a three-dimensional model.