Thursday, October 01, 2026

Another AI evaluation framework informed by the human CHC cognitive abilities taxonomy



According to my tracking of the literature (of course, most likely not complete), here is a third article (open access from the Journal of Intelligence) that proposes using the CHC cognitive ability taxonomy (or parts of it) to evaluate an aspect of AI performance—in this case AI-assisted organizational decision-making.  The other two proposed AI evaluation framework articles using the CHC taxonomy (alone or as part of a much grander framework) can be found here and here. 

In a related development, Kameron Green—the developer behind an ambitious and grand AI evaluation framework (in second to last link above)—today posted a link to an AI generated video summary overview of his HCQM taxonomic model—which would help most readers understand his lengthy HCQM paper.


Cognitive Performance Under AI Advice: Development and Initial Validation of a CHC-Informed Assessment for Organizational Decision-Making

Abstract

Artificial intelligence (AI) is increasingly embedded in organizational decision-making, requiring employees not only to use AI-generated recommendations but also to evaluate their quality and determine when reliance is appropriate. Although established research examines behavioral reliance on algorithmic and AI advice, fewer studies have approached performance under AI advice as an individual-differences assessment problem integrating psychometric structure, cognitive correlates, process indicators, and criterion-related evidence. This study developed and initially validated a CHC-informed, performance-based assessment of cognitive performance under AI advice using 24 organizational decision scenarios. The assessment was designed around three closely related content/performance dimensions—AI error detection, evidence integration, and cognitive control and adaptive reliance—while also capturing confidence, response time, and reliance behavior. The validation sample comprised 780 employed adults in Türkiye. Psychometric analyses included confirmatory factor analysis, multidimensional item response theory, response-time analyses, scenario-level logistic regression, measurement invariance, differential item functioning, and internal cross-validation. Results indicated a dominant general cognitive-performance component together with additional structure corresponding to the three theoretically specified dimensions. Assessment performance was positively associated with established cognitive measures, including ICAR-16 reasoning performance, working memory, processing speed, and attentional control, whereas associations with AI-related self-reports were generally weaker. Dimension-aligned analyses supported the expected associations of ICAR-16 with AI error detection and working memory with evidence integration. Performance was also moderately associated with concurrently assessed organizational decision quality (r = 0.408) and explained additional variance in this criterion beyond demographic and work characteristics, AI experience, conventional cognitive-performance measures, and AI-related self-reports (ΔR2 = 0.103, p < .001). In contrast, several hypothesized scenario-specific associations involving AI confidence, time pressure, interruptions, and resistance to confidently inaccurate advice were not supported. Overall, the findings provide initial evidence for a performance-based approach to assessing how employees evaluate and respond to AI advice, while indicating that the proposed scenario-specific mechanisms and group-comparability findings require further replication before consequential applications are considered.