Showing posts with label measurement. Show all posts
Showing posts with label measurement. Show all posts

Thursday, January 02, 2025

Quote2note: E. L. Thorndike on importance of psychological #measurement

 

Whatever exists at all exists in some amount.  To know it thoroughly involves knowing its quantity as well as its quality


E. L. Thorndike

Thursday, December 12, 2024

#quote2note: #Galton on #measurement

Whenever you can, count.     

              — Sir Francis Galton


Wednesday, December 11, 2024

Applied #psychometrics 101: Strong programs of #constructvalidity—the #theory - #measurement framework with emphasis on #substantive & #structural validity - #WJIV #WJV #shoolpsychology #psychology

The validity of psychological tests “is an overall judgment of the degree to which empirical evidence and theoretical rationales support the adequacy and appropriateness of inferences and actions based on test scores or other modes of assessment” (Messick, 1995, p. 741).

The ability to draw valid inferences regarding theoretical constructs from observable or manifest measures (e.g., test or composite scores) is a function of the extent to which the underlying program of validity research attends to both the theoretical and measurement domains of the focal constructs (Bensen, 1998; Bensen & Hagtvet, 1996; Cronbach, 1971; Cronbach & Meehl, 1955; Loevinger, 1957; Messick, 1995; Nunnally, 1978). 


The theoretical—measurement domain framework that has driven the revisions of the WJ test batteries, particularly from the WJ-R to the forthcoming WJ V cognitive and achievement test batteries (Q1, 2025; COI disclosure: I am a coauthor of the current WJ IV and forthcoming WJ V), is represented in the figures below.  


The goal of this post is to provide visual-graphic (Gv) images that hopefully, if properly studied by the reader (and if I did a decent job), provide the basic concepts of what constitutes the substantive component (and to a more limited extent the structural component) of a strong program of construct validity—in particular, the theoretical-measurement domain mapping framework used in the WJ-R to the forthcoming WJ V. The external stage of construct validity is not highlighted in this current post.  The goal is for conceptual understanding…thus the absence of empirical data, etc.


For those who want written background information, the most succinct conceptual overview of a “strong program of construct validation” is Bensen (1998; click to download and read).  


Otherwise…sit back and enjoy the Gv presenation…where five images equal at least one or more chapter in a technical manual :). 


Be sure to click on each image to enlarge (and make readable)


This figure below was first published in a book on CHC theoretical (then known as Gf-Gc) interpretation of the Wechsler intelligence test batteries (Flanagan, McGrew, & Ortiz, 2000).









Yea…I know.  The following figure uses Gv as the sample cognitive ability construct domain and not Gf as in the prior figure.  I first crafted the figure below in 2005 and don’t have the time (nor attentional control focus bandwidth) to make a new version.  Consider the switch from the Gf domain (above) to Gv (below) as a test of your understanding of the material…and ability to generalize what you have learned.  And yes, I do see there is a spelling error (“on” for “one”)…but it is an old image file and I don’t have time to “clean it up” as noted above.  The primary new feature is the addition of the concept  of developing developmentally (difficulty) ordered sets of test items for the underling ability trait scales for manifest indicator tests C and D under the CHC theoretical narrow ability domain of spatial scanning, under the broad ability domain of Gv. This is where IRT (Rasch model) item scaling is involved.




The following figure is drawn from the WJ IV technical manual (McGrew, LaForte, Schrank, 2014) and illustrates the three-stage structural validity process used in the WJ IV.  The same process, with slightly different age groups and the addition of exploratory hierarchical psychometric network analysis (see exciting and ground-breaking work of Dr. Hudson Golino and colleagues) during stage 2A, will be presented in the WJ V technical manual (LaForte, Dailey & McGrew, Q1-2025).




Thursday, July 12, 2018

Great psychometric resource: The Wiley Handbook of Psychometric Testing.

I just received my two volume set of this excellent resource on psychometric testing.  There are not many good books that cover such a broad array of psychometric measurement issues.  This is not what I would call "easy reading."  This is more like a "must have" resource book to have "at the ready" when seeking to understand contemporary psychometric test development issues.

Tuesday, July 07, 2009

Applied Psych Test Development Series: Parts F--Psychometric/technical statistical analysis: Internal

The sixth in the series Art and Science of Applied Test Development is now available.

The sixth module (Part F--Psychometric/technical statistical analysis:  Internal) is now available.

In addition, I've made some edits and additions (esp. summary "Tools, Tips, and Troubles" and "Advanced Topics" slides) to prior presentations (Part A-E).

This is the sixth in a series of PPT modules explicating the development of psychological tests in the domain of cognitive ability using contemporary methods (e.g., theory-driven test specification; IRT-Rasch scaling; etc.). The presentations are intended to be conceptual and not statistical in nature. Feedback is appreciated.

This project can be tracked on the left-side pane of the blog under the heading of Applied Test Development Test Development Series.

The first module (Part A: Planning, development frameworks & domain/test specification blueprints) was posted previously and is accessible via SlideShare.

The second module (Part B: Test and item development) was posted previously and is accessible via SlideShare.

The third module (Part C--Use of Rasch scaling technology) was posted previously and is accessible via Slideshare.

The fourth module (Part D--Develop norm [standardization] plan) was posted previously and is accessible via Slideshare.

The fifth module (Part E--Calcuate norms and derived scores) was posted previously and is accessible via Slideshare.

You are STRONGLY encouraged to view them in order as concepts, graphic representation of concepts and ideas, etc., build on each other from start to finish.

Enjoy...more to come.

Technorati Tags: , , , , , , , , , , , , , , , , ,



Monday, June 29, 2009

Applied Psych Test Development Series: Part C--Use of Rasch scaling technology

The third in the series Art and Science of Applied Test Development is now available. The third module (Part C: Test and Item Development--Use of Rasch Scaling Technology) is now available.

This is the third in a series of PPT modules explicating the development of psychological tests in the domain of cognitive ability using contemporary methods (e.g., theory-driven test specification; IRT-Rasch scaling; etc.). The presentations are intended to be conceptual and not statistical in nature. Feedback is appreciated.

This project can be tracked on the left-side pane of the blog under the heading of Applied Test Development Test Development Series.

The first module (Part A: Planning, development frameworks & domain/test specification blueprints) was posted previously and is accessible via SlideShare.

The second module (Part B: Test and item development) was posted previously and is accessible via SlideShare.

You are STRONGLY encouraged to view them in order as concepts, graphic representation of concepts and ideas, build on each other from start to finish.

Enjoy...more to come.

Technorati Tags: , , , , , , , , , , , , , , , ,

Friday, June 26, 2009

Applied Psych Test Development Series: Part B-Test and Item Development

The second in the series Art and Science of Applied Test Development is now available. The second module (Part B: Test and Item Development) is now available.

This is the second in a series of PPT modules explicating the development of psychological tests in the domain of cognitive ability using contemporary methods (e.g., theory-driven test specification; IRT-Rasch scaling; etc.). The presentations are intended to be conceptual and not statistical in nature. Feedback is appreciated.

This project can be tracked on the left-side pane of the blog under the heading of Applied Test Development Test Development Series.

The first module (Part A: Planning, development frameworks & domain/test specification blueprints) was posted previously and is accessible via SlideShare.

Enjoy...more to come.


Applied Psych Test Development Series: Part A-Planning, development frameworks & domain/test specification blueprints

Announcement--the Art and Science of Applied Test Development. Let the games begin.

This is the first in a series of PPT modules explicating the development of psychological tests in the domain of cognitive ability using contemporary methods (e.g., theory-driven test specification; IRT-Rasch scaling; etc.). The presentations are intended to be conceptual and not statistical in nature. Feedback is appreciated.

This project can be tracked on the left-side pane of the blog under the heading of Applied Test Development Test Development Series.

The first module (Part A: Planning, development frameworks & domain/test specification blueprints) is now available for viewing via SlideShare.

Stay tuned.


Monday, April 27, 2009

New IRT (item response theory) book from Guilford Press


New IRT (item response theory) book available from Guilford. This is an FYI post. I've not read this book nor have I read any reviews. If anyone reads it and has comments, please feel free to add a comment at this blog.

Technorati Tags: , , , , , , ,

Friday, April 10, 2009

The attack of the psychometricians: Psychological measurement

I'm just in the processing of reading Borsboom's (2006; Psychometrika) provocative article "The attack of the psychometricians."  The article abstract is below.  As I'm reading, I'm loving a number of statements meant to get the attention of psychologists.  Here is the most recent favorite. 

"psychologists have a tendency to endow obsolete techniques with obscure interpretations"

Abstract:  This paper analyzes the theoretical, pragmatic, and substantive factors that have hampered the integration between psychology and psychometrics. Theoretical factors include the operationalist mode of thinking which is common throughout psychology, the dominance of classical test theory, and the use of “construct validity” as a catch-all category for a range of challenging psychometric problems. Pragmatic factors include the lack of interest in mathematically precise thinking in psychology, inadequate representation of psychometric modeling in major statistics programs, and insufficient mathematical training in the psychological curriculum. Substantive factors relate to the absence of psychological theories that are sufficiently strong to motivate the structure of psychometric models. Following the identification of these problems, a number of promising recent developments are discussed, and suggestions are made to further the integration of psychology and psychometrics.

Technorati Tags: , , , , , ,

Tuesday, December 30, 2008

ITEMS - Instructional Topics in Educational Measurement Series

Regardless whether you are a user of educational measurement technology or teach courses in educational and psychological measurement, if you want to read relatively brief overview modules on select measurement topics, you should check out the free on-line NCME ITEMS modules.  The goal of ITEMS is to improve the understanding of educational measurement principles by providing brief instructional units on timely topics in the field, modules developed for use by college faculty and students as well as by workshop leaders and participants.  ITEMS are a product provided by the National Council on Measurement in Education (NCME)

Below is information I lifted from the NCME ITEMS web page:

Instructional modules are designed to be learner-oriented and consist of an abstract, tutorial content, exercises, and annotated references. The teaching aids accompanying most modules are designed to support the use of the instructional modules in teaching and workshop settings by providing supplemental student exercises, references, test items, and figures or masters for transparencies.

The ITEMS modules can be downloaded as PDF files below (you can use Adobe Reader to view them).

Get Adobe Reader

Technorati Tags: , , , , , , , , ,

Saturday, June 23, 2007

More on reading comprehension (Grw)

Regular readers of this blog will notice that recently I've been particularly focused on reading articles dealing with the development and assessment of reading comprehension (Grw-RC; click here and here).

Today I stumbled across a special 2006 issue of the journal Scientific Studies of Reading dealing with the topic of reading comprehension assessment. A copy of the articles and abstracts can be found by clicking here. Dr. Jack Fletcher provides a nice summary of the content of the entire special issue.

Check it out. A good issue to read.


Technorati Tags: , , , , , , , ,

Powered by ScribeFire.

Thursday, May 24, 2007

Nonword (Ga/Gsm) repetition tasks - literature to track

Sorry for my very inconsistent posting over the past few months. This summer has been crazy as I work with my lovely fiance to plan a wedding, sell two houses, and build a new house :)

The purpose of this post is to alert readers to a trend I've detected (I may be late in this detection...but...at least I've now noticed it..better late than never)---an increasing body of empirical literature that implicates the abilities measured by non-word repetition tasks in the identification of children with specific language impairments (SLI). Today I ran across a meta-analysis by Estes et al. (2007; click here to view) that continues to highlight the importance of these abilities and measurement tasks. The abstract is reproduced below.

Something important seems to be measured by non-word repetition tasks, although what these abilities are is a matter of debate. As noted by Estes et al.:
  • "There has been considerable debate surrounding the nature of the skills tapped in nonword repetition, whether it recruits phonological working memory (Bishop et al., 1996; Botting & Conti-Ramsden, 2001; Montgomery, 1995b; Van der Lely & Howard, 1993), phonological encoding (Kamhi & Catts, 1986), phonological awareness or sensitivity (e.g., Metsala, 1999), or a general phonological processing ability (e.g., Bowey, 1996, 2001). Many authors have also acknowledged that the act of repeating nonwords involves multiple processes (e.g., Briscoe, Bishop, & Norbury, 2001; Edwards & Lahey, 1998; Gathercole, Willis, Baddeley, & Emslie, 1994; Snowling, Chiat, & Hulme, 1991). A child's ability to repeat a novel word may be affected by any of the component skills involved in the process of hearing, encoding, and producing a word form: the ability to perceive speech distinctions; the preciseness, robustness, or organization of phonological and morphological representations; the ability to store the word form; and motor planning and articulation skills. The impairments of children with SLI may affect performance at any point or at many points in this process."
I concur. Task analysis suggests that, from a CHC factor analysis perspective, non-word repetition tasks may garner their diagnostic sensitivity from their factorial complexity (i.e., they measure multiple important abilities/constructs). These may include such Ga (auditory processing) narrow abilities as phonetic coding (PC), speech sound discrimination (US), memory for sound patterns (UM), and temporal tracking (UK). In addition, clearly the Gsm narrow ability of working memory (MW; what is often called the phonological working memory or articulatory loop) is implicated. Other CHC candidate abilities included efficacy of accessing a person's lexicon (aka; speed of lexical access or naming facility-Glr: NA). For users of the WJ-III battery [conflict of interest disclosure - I'm a coauthor], we have a test called Sound Awareness that has been found to be very predictive of academic achievement and diagnostic classification (normal vs some kind of disorder)...primarily, I believe, because it is a CHC ability-complex measure of multiple narrow abilities (at a minimum, PC and MW). Measures that are not factorially "pure" can still be important and useful for other assessment purposes - diagnosis and prediction.

I would encourage readers to continue to track the emerging non-word repetition practical and theoretical literature. Another important article to read is by Gathercole (2006). Also, I've previously blogged about a non-word repetition article in the journal Dyslexia that, IMHO, suffered from serious methodological flaws and should not be taken seriously. Finally, as my awareness of this literature has grown I recently ran a search of the IAP Reference Database for other articles that may be related (as you will see..there is no shortage of literature to read in this area).

Estes et al. (2007) Abstract
  • Purpose: This study presents a meta-analysis of the difference in nonword repetition performance between children with and without specific language impairment (SLI). The authors investigated variability in the effect sizes (i.e., the magnitude of the difference between children with and without SLI) across studies and its relation to several factors: type of nonword repetition task, age of SLI sample, and nonword length. Method: The authors searched computerized databases and reference sections and requested unpublished data to find reports of nonword repetition tasks comparing children with and without SLI. Results: Children with SLI exhibited very large impairments in nonword repetition, performing an average (across 23 studies) of 1.27 standard deviations below children without SLI. A moderator analysis revealed that different versions of the nonword repetition task yielded significantly different effect sizes, indicating that the measures are not interchangeable. The second moderator analysis found no association between effect size and the age of children with SLI. Finally, an exploratory meta-analysis found that children with SLI displayed difficulty repeating even short nonwords, with greater difficulty for long nonwords. Conclusions: These findings have potential to affect how nonword repetition tasks are used and interpreted, and suggest several directions for future research.
Technorati Tags: , , , , , , , , , , , , , , , , , , , , ,

Powered by ScribeFire.

Wednesday, April 11, 2007

Differential Ability Scales-2nd edition: One more in the CHC column

Kudos to Colin Elliott for the recent publication of the second edition of the Differential Ability Scales (DAS-2). As I wrote about the DAS in my 1997 CHC broad-narrow analysis of all major intelligence batteries (in Flanagan et al.'s, 1997 CIA book), I considered it to be the second most comprehensive battery of CHC abilities...the first, of course, being the deliberately CHC-designed WJ-R and WJ -III (obligatory conflict of interest - I'm a coauthor of the WJ III). I'm not surprised to see that it has now joined the growing crowd (see my CHC bandwagon post) of deliberately CHC-designed intelligence batteries, as it was, IMHO, the next-best instrument (from a CHC perspective) at the time (1997).

As usual, check out the Willis and Dumont web site for additional DAS-2 related information.

I'd love to see a copy. Hint...hint......is anyone from Psych. Corp. listening? Don't you like the free publicity I just gave the DAS-2? I'd sure love a complimentary copy to examine.





Technorati Tags: , , , , , , , , , , ,



Powered by ScribeFire.

Friday, January 26, 2007

Quantoids corner-bifactor and second-order FA comparisons-Guest post by Matthew Reynolds

The following is a guest blog post by Matthew Reynolds, one of Tim Keith's Doctoral Student in Educational Psychology (School Psychology & Quantitative Methods) at the University of Texas at Austin, Department of Educational Psychology.

This is an excellent post by a future quantoid to be reckoned with in the field of school/educational psychology research. Kudos to Dr. Tim Keith for suggesting that one of his doctoral students make a quest blog post. This is the first such doctoral student virtual scholar post. If there are other professors who would like to entertain the idea of doctoral students being assigned articles to review and prepare for guest posts on IQ's Corner, then drop me an email.... iap@earthlink.net
  • Chen, F. F., West, S. G., & Sousa, K. H. (2006). A comparison of bifactor and second-order models of quality of life. Multivariate Behavioral Research, 41, 189-225. (click to view)

Although not directly related to intelligence, this article compares two confirmatory factor analytic (CFA) models frequently used in psychometric research of intelligence: bifactor and second-order models. Chen et al. (2006) describe the bifactor model as having a general factor that accounts for the communality in all items and domain specific factors that account for influences above and beyond the general factor. The second-order model is described as having interrelated first-order factors with a general factor that accounts for those relations.

Study 1 compared the two models by applying the factor structure to a quality of life measurement from the AIDS Time-Oriented Health Outcome Study. Study 2 was a Monte Carlo study investigating whether there was enough power to detect differences in the bifactor and second-order model. Previous research had suggested that it was empirically impossible to distinguish between the two in typical samples used in social science research (i.e. Mulaik & Quartetti, 1997).

Results from Study 1:

  • Bifactor and second-order factor models were imposed on a 17 item health-care related quality of life survey. The models had a general overall quality of life factor and four domain specific factors. The four domain-specific factors included cognition, vitality, mental health, and disease worry.
  • The results from the bifactor model suggested that the mental health factor did not provide unique information above and beyond the general factor. Therefore, the model was re-specified without a mental health factor.
  • The second-order factor model was specified with four first-order factors and a general quality of life factor that accounted for the relations among the first-order factors. The residual variance for the mental health factor, however, was statistically significant suggesting that there was some unique contribution of this factor (although the general factor accounted for 91.4% of the variance in that factor). Note this finding was different from the bifactor model. In the bifactor model the mental health factor did not provide unique information. Therefore, to be consistent with the bifactor model the authors also re-specified the second-order model so that only three factors, and the subtests related to mental health factor loaded directly on the second-order factor.
  • The results comparing the two different models showed that both the bifactor and second-order factor models provided adequate fit. Because the second-order model is a more constrained version of the bi-factor model, the likelihood ratio test (i.e., chi-squared difference test) was used to compare the fit of the models. The second-order model fit worse than did the bifactor suggesting that the constraints applied to the bifactor model to get to the second-order model were too restrictive. Also, a power analysis suggested that there was adequate power to detect the difference.
  • Next, the authors used these models to predict social functioning. Both models resulted in almost identical standardized estimates. This finding was rather reassuring in regards to the interpretability of the ability factors.

Study 2:

  • The findings suggested that even with a sample size of 200 there appears to be enough power to detect differences between the bifactor and second-order models.


Discussion:

  • The authors concluded that the bifactor model offers several advantages over the second-order model. One advantage was that it identified three factors instead of four. I am not quite convinced that this is necessarily an advantage. Two, they noted that researchers may miss potential non-significant first-order factor variances when looking at their results. I thought this was a good point by the authors; however, I also have had the same concern about using bifactor models. For example, a not-so-careful researcher may not consider the non-significant domain specific factor loadings as well as a non-significant domain specific factor variance.
  • The second advantage was that the bifactor model fit better. That is, the relations between the general factor and the items could not be fully mediated by the first-order factors.
  • Third, they stated that the bi-actor model is easier to interpret when predicting external criteria because the domain factors are represented as common factors in bifactor models whereas they are residualized factors in the higher-order model. Although true, I think the point is rather minor.
  • Last, and perhaps most importantly, they conclude that BOTH models are useful in research. I agree completely with this point as CFA models should be consistent with theoretical models.
  • In general, the article provides great information for those interested in hierarchical factor analysis, and it is provided in a straightforward manner. I think that the advantages of the bifactor model were a bit overstated. I do agree that it is useful to examine both models in research, especially since the second-order model can be derived from the bifactor model.
  • In my own research, one weakness of the bifactor model has been related to empirical under-identification. I believe that perhaps it runs into some of the same difficulties as the multi-method multi-trait models in that they are over-parameterized. A recent study that used the bifactor model to test for method effects also found that the bifactor model may fit well even when it is an incorrect model (Maydeu-Olivares & Coffman, 2006).
  • In terms of research in psychometric intelligence, the interpretation of the two models is slightly different as well. For example, in a bifactor model all of the effects of the general factor are direct. In intelligence research it seems to me that the contemporary theories are more consistent with the higher order model in which the general factor explains the interrelations of the broad abilities and its relation to test performance is mediated through the broad abilities.
  • To make this more germane to intelligence researchers I have included some output of analyses that I performed using the Holzinger & Swineford correlation matrix reported in their 1937 study. Shown are the specifications, the models with standardized loadings, and the unstandardized loadings, variances, and the total effects shown separately. Just as a warning, the models are not in publication form, but suffice for a demonstration. I hope these models help to clarify how the second-order model is in fact a more constrained version of the bifactor model. See Yung, Thissen, and McLeod (1999) for a more technical account.
  • Last, as an aside, I thought I would share the last two sentences from the Holzinger & Swineford 1937 article in Psychometrika. In this article the authors introduced the bifactor model:
  • The Bi-factor analysis illustrated above is not only very simple, but the calculation is relatively easy as compared with other methods. The total time for computation, done by one person, was less than ten hours for the present example.”
  • I just ran a bifactor model in Amos 5, and other than setting the model up, the actual computational time took 0.29 seconds. You have to appreciate all of the time and patience that researchers have put in over the years to get us where we are today!
Technorati Tags: , , , , , , , , ,

powered by performancing firefox