Publications
Preprints (under review)
Trautwein, T., & Schroeders, U. (2026, September 29). A comparison of mixed-integer linear programming and metaheuristics in parallel short-form construction. PsyArXiv.
Constructing parallel test forms requires balancing different construction goals, including content coverage, measurement precision, and equivalence across forms. We evaluated the accuracy and efficiency of two different psychometric approaches in parallel test constructions using two simulations: mixed-integer linear programming (MILP) uses fixed full-pool item-parameter estimates, whereas ant colony optimization (ACO) re-estimates item parameters for candidate subsets. In Simulation 1, we compared their parameter recovery across 1,000 randomly sampled item subsets per condition to quantify item-parameter uncertainty, varying (1) effective sample size per item (100, 200, 500, 1,000), (2) model complexity (1PL, 2PL, 3PL), (3) latent trait distribution (normal, left-skewed), and (4) item-set proportion (10%, 20%, 30%). Full-pool calibration generally yielded more accurate and less variable parameter recovery, particularly for discrimination parameters in more complex models and smaller item sets. In Simulation 2, we assembled three parallel 40-item forms across 30 independently generated datasets in two exemplary conditions. MILP and ACO performed similarly in the 2PL condition, whereas in the 3PL condition ACO showed a slight advantage in parallelism and MILP achieved higher minimum test information. Across both conditions, MILP with full-pool calibration showed more consistent item-parameter and test-information recovery. Evaluating the same final ACO-selected items with full-pool rather than subset-specific re-estimates improved recovery on nearly all metrics. These findings indicate that repeated subset-specific re-estimation may not improve the psychometric properties of parallel short tests, but increase the risk of overfitting by allowing the search to capitalize on sample-specific estimation error.
Schroeders, U., Achaa-Amankwaa, P., & Wilhelm, O. (2026, September 10). The long shadow of the wall: Socialization signatures in public-events knowledge across East and West Germany. PsyArXiv.
Germany’s division and subsequent reunification provide a rare historical possibility to examine how regionally segregated information environments are reflected in what people know and how long such differences persist. In a national sample of 1,012 German adults aged 30–80 years, participants completed a 120-item multiple-choice test of public-events knowledge spanning 1960–2019, including items specifically targeting East- and West-German events from the pre-1990 period. While overall knowledge levels were similar among East- and West-socialized respondents, the item-level differences were pronounced, especially for East-specific items in pop culture, arts and literature. Nested cross-validated elastic-net models yielded near-perfect discrimination between East and West socialization in older cohorts (ages 51–80; AUCs > .96), slightly lower discrimination among ages 41–50 (AUC = .93), and a sharp drop in the youngest cohort (ages 30–40; AUC = .67). East-specific items contributed most strongly to classification, which is consistent with the concept of cultural hegemony and an asymmetric information flow. Despite post-1990 convergence, persistent geographic patterns in knowledge profiles remain. Public-events knowledge thus carries a cohort-dependent socialization signature, leaving detectable traces of historical division even after societal convergence.
Schroeders, U., Einsiedler, J., & Gnambs, T. (2026, September 2). Factor structure and symptom profiles of the Montgomery–Åsberg Depression Rating Scale (MADRS): A meta-analytic mean and covariance structure analysis. PsyArXiv.
The Montgomery–Åsberg Depression Rating Scale (MADRS) is one of the most widely used clinician-rated instruments for assessing depressive symptom severity, particularly in pharmacological intervention studies. Despite its widespread use, evidence on the dimensional structure of the MADRS remains inconsistent. In this Registered Report, we propose a meta-analytic mean and covariance structure analysis of the MADRS to examine (a) which latent structure best represents MADRS responses and how reliably its factors are measured, (b) whether items or symptom dimensions function differently across clinical (diagnostic group, pharmacological treatment, assessment time point within treatment) and measurement-related (language version) moderators, and (c) whether diagnostic groups, pharmacological treatment regimens, and treatment stages are associated with distinct latent symptom profiles.
Jankowsky, K., Fuhrmann, M. F., Schroeders, U., & Zimmermann, J. (2026, August 27). A matter of time? The impact of time frames on the self-report assessment of personality pathology. PsyArXiv.
Diagnosing a personality disorder requires that impairments in personality functioning have persisted over an extended period—at least 2 years according to the ICD-11, with an onset traceable to adolescence or early adulthood according to the Alternative DSM-5 Model for Personality Disorders. Common self-report measures of personality pathology, however, specify no time frame in their instructions, leaving unclear which reference periods respondents adopt and whether instructed time frames alter their responses. In a preregistered experiment, N = 1,447 U.S. adults currently receiving or awaiting treatment for a mental health condition completed the Level of Personality Functioning Scale-Brief Form 2.0 (LPFS-BF 2.0), the modified Personality Inventory for DSM-5-Brief Form (PID5BF+M), and the Personality Disorder Severity ICD-11 scale (PDS-ICD-11) after random assignment to one of three instruction conditions: the original instructions (no specific time frame), a short time frame (“in the last 2 weeks”), or a time frame matching diagnostic criteria (“in the last 2 years”/“since early adulthood”). Multigroup confirmatory factor analyses supported strict measurement invariance across conditions for all three instruments. Correlations with depressive symptoms (Patient Health Questionnaire-9) and factor means did not differ systematically between conditions, a pattern that replicated in two sensitivity analyses (one of them including only participants correctly recalling the time frame). Notably, only 45% of participants recalled the instructed time frame correctly. Time frame instructions thus appear to have little impact on self-reported personality pathology, possibly because respondents disregard them or because trait-like item content elicits semanticized self-knowledge irrespective of the instructed reference period.
Trummer, B. T. T., Wendt, L. P., Schroeders, U., & Zimmermann, J. (2026, August 21). Measurement invariance and sociodemographic differences at the lowest level of the psychopathology hierarchy: An analysis of the German HiTOP-SR. PsyArXiv.
Comparisons of mental health problems between sociodemographic groups are interpretable only if measures function equivalently across groups, yet measurement invariance has rarely been tested at the lowest level of the psychopathology hierarchy. Using the validation sample of the German Hierarchical Taxonomy of Psychopathology Self-Report (HiTOP-SR; N = 1,049 general-population adults), we tested measurement invariance across sex, age, and socioeconomic status (SES) for all 87 facets in ordinal multigroup confirmatory factor analyses and, where scalar invariance held, estimated standardized latent mean differences. Of the 261 cases (87 facets × 3 grouping variables), 74.3% reached scalar invariance, 14.6% showed non-invariance (usually small to medium in magnitude) and 11.1% were not estimable owing to low base rates. On invariant facets, women scored higher on internalizing and lower on externalizing facets, symptom levels declined with age, and higher SES was associated with lower symptom levels, with facet-level heterogeneity revealing where these differences were most pronounced. The results identify which facets permit meaningful sociodemographic comparisons and inform future revisions of the instrument.
Schroeders, U., Breuer, J., Einsiedler, J., Frank, M., Gnambs, T., Haim, M., Hommel, B., Jankowsky, K., Knöpfle, P., & Schönbrodt, F. (2026, August 8). Seven guiding questions for reproducible AI-assisted research in the social and behavioral sciences. PsyArXiv.
Artificial Intelligence (AI) is increasingly used in social and behavioral research for tasks including literature search, stimulus generation, response annotation, code generation, and writing. These uses challenge conventional understandings of reproducibility because AI-assisted procedures may shape not only the final analysis but also the empirical materials, data-processing steps, and analytical inputs on which reported findings depend. In this opinion article, we propose seven guiding questions for deciding when, what, and how to document AI-assisted research to support reproducibility. We distinguish between two poles of a continuum: AI as a research partner, where transparent disclosure may often be sufficient, and AI as a methodological instrument, where AI outputs enter the evidential chain and therefore require more detailed documentation. The questions are structured across three stages: (a) clarifying the level of reproducibility required for AI use, (b) determining what needs to be reproducible, and (c) deciding how AI assistance should be documented in manuscripts and supplementary materials. Examples from AI-assisted research demonstrate how these questions guide documentation decisions across different reproducibility requirements. Overall, our proposal shifts the focus from generic AI disclosure to role-based documentation that makes AI-assisted workflows understandable, auditable, and, where necessary, reproducible. It also underscores that reproducibility should not be reduced to reporting prompts and AI model settings alone, but requires attention to output variability, changes introduced through human oversight, and the infrastructures through which AI systems are accessed, versioned, and preserved.
Schroeders, U., Achaa-Amankwaa, P., Walter, J., Endlich, D., Hasselhorn, M., Golle, J., & Goecke, B. (2025, June 10). PINGUIN – Assessing elementary students’ initial competencies. PsyArXiv.
Assessing initial competencies at school entry is essential for guiding targeted support. PINGUIN is a modular screening instrument including three modules that assess early competencies in (1) auditory language processing, (2) reading and writing, and (3) mathematical literacy. Limited literacy and attention spans, heterogeneous abilities, and shifting indicators make assessment challenging. We therefore developed age-appropriate tablet-based tasks with audio instructions and touchscreen responses. These tasks were administered to N = 604 students at two measurement occasions across the first school year. After horizontal and vertical linking via item response theory, we compared module-wise unidimensional and multidimensional scaling. Correlated-factors models best represented the structure of the competence domains, but only unidimensional scaling yielded reliabilities sufficient for individual-level feedback. Module scores showed the expected convergent and discriminant relations with established measures of reading, mathematics, and reasoning, validating PINGUIN as a reliable, age-appropriate screening tool.
Schroeders, U., Gnambs, T., & Einsiedler, J. (2026, June 8). MASEMiner – AI-assisted data extraction for meta-analytic structural equation modeling. PsyArXiv.
Meta-analytic structural equation modeling (MASEM) synthesizes multivariate evidence across studies, allowing researchers to examine theory-driven path models, including mediated and predictive effects, and to evaluate the dimensional structure of psychological and behavioral measures. A major bottleneck is the labor-intensive extraction of statistical information from primary studies. Large language models and vision-language models offer new opportunities for extracting such heterogeneous statistics from textual and visual document content. Building on these capabilities, we introduce MASEMiner, an open-source application for AI-assisted extraction and verification of data required for MASEM that can be run locally or accessed through a hosted web interface. MASEMiner supports the extraction of directly reported correlations and related effect sizes as well as indirect information (i.e., factor loadings and factor correlations) used to reconstruct correlation matrices for MASEM. We evaluated MASEMiner’s extraction accuracy across four use cases covering directly and indirectly reported statistics. Outputs were compared with human consensus codings from published meta-analytic databases, using separate study sets for prompt development and evaluation. Precision was near-perfect for indirectly reported information (0.95–1.00) but lower for some directly reported quantities such as effect sizes (~0.65), where the tool over-extracted candidate values. Recovery was high, with sensitivity above 0.90 for almost all extracted variables. MASEMiner supports a human-in-the-loop workflow by extracting candidate statistics that can be efficiently checked against the source documents, thereby reducing manual coding effort while preserving transparency and reproducibility in research synthesis.
Goecke, B., Golle, J., Zettler, I., & Schroeders, U. (2026, May 10). Boys and things, girls and people? Gender-related interest patterns in elementary school. PsyArXiv.
Gender-related interest patterns are well documented in adolescence and adulthood, but little is known about whether comparable patterns can already be identified in elementary school, whether they extend beyond subject-specific academic interests, and whether they are organized along the same broad dimensions as in older samples. In this study, we developed an age-appropriate self-report measure of RIASEC-based activity preferences for elementary school children. In a sample of more than 1,000 third graders, interests were assessed at the beginning and end of one school term. Items were constructed to represent the six RIASEC domains and selected for the final instrument using Ant Colony Optimization. We examined the measure’s psychometric properties, including structural validity, as well as gender differences, short-term temporal stability, and associations with academic interests and school grades. Results revealed a RIASEC-like circumplex structure in this young age group for both girls and boys. At the beginning of Grade 3, children showed distinct gender-related interest profiles consistent with patterns reported in older populations, with boys showing stronger things-oriented profiles and girls showing stronger people- and ideas-oriented profiles. These gender differences remained essentially stable across the term, and vocational-like interests showed small-to-moderate, theoretically plausible associations with subject-specific academic interests and school grades. Overall, the findings suggest that gender-typed organization of broad activity preferences is already in place in middle childhood, well before children make consequential educational choices, and that interest assessment beyond curricular labels can capture motivational structures not visible in subject-specific measures.
Trautwein, T., Steger, D., Wilhelm, O., & Schroeders, U. (2026, April 21). Predicting item-level characteristics of knowledge tests using large-language models. PsyArXiv.
Efficient test construction depends on accurate estimates of item difficulty and discrimination, yet obtaining such estimates through piloting with human participants is costly and time-consuming. Whereas item parameters of decontextualized reasoning tasks can often be approximated from formal construction rules, comparable predictor sets for contextualized knowledge items remain scarce. We investigated whether or not large language models (LLMs) can predict item difficulty and item-total correlations for declarative knowledge items from item text. We compared zero-shot predictions from a generic GPT-4o model with predictions from a fine-tuned version trained on empirically calibrated items. We reanalyzed data from a smartphone-based assessment comprising 2,534 German-language knowledge items spanning 32 knowledge domains, administered to 6,039 participants. Predictive accuracy was evaluated on an independent holdout sample of 534 items. A fine-tuned GPT-4o model predicted item difficulty with high accuracy (R² = .63) and discrimination with moderate accuracy (R² = .32), substantially outperforming zero-shot prompting. Including response options in addition to the question further improved prediction (for difficulty: R² = .38-.63; for discrimination: R² = .23-.32), suggesting that distractor characteristics provide additional cues for parameter estimation. Predictions generalized across knowledge domains, and even smaller training subsets yielded reasonably accurate difficulty predictions. Fine-tuned LLMs can support preliminary item screening and pre-sorting in knowledge test development, thereby reducing the need for empirical pilot testing and cheapening and accelerating test development.
Achaa-Amankwaa, P., Hommel, B., Robitzsch, A., Schipolowski, S., & Schroeders, U. (2026, March 20). Predicting item difficulties in C-tests using linguistic features and transformer-based language models. PsyArXiv.
C-Tests, a specific form of cloze tests requiring the completion of truncated word endings in short text passages, are widely used in educational measurement as indicators of reading comprehension, general language proficiency, and crystallized intelligence. In this paper, we examine the feasibility of a more rational construction of C-Tests by estimating item difficulty solely from textual features, without relying on prior empirical testing. For that, we use item-level (e.g., word length, word category, and word frequency), sentence-level (e.g., existence of sentence negation), and text-level (e.g., readability) surface features to predict item difficulty and compare these predictions to empirical difficulty estimates. Furthermore, we evaluate the added predictive performance derived from two transformer-based language models: BERT and GPT-4o. We reanalyze data from a large-scale educational study in which 1,197 German secondary school students worked on 16 C-Tests. Using elastic net regression with nested resampling, we found that surface features explained approximately 20% of the variance in out-of-sample predictions. While BERT estimates did not yield any incremental predictive performance over linguistic features, GPT-based predictions added an increment of 6%, resulting in a total explained variance of R² = 26% in out-of-sample predictions. We discuss the strengths and limitations of using large language models in language assessment in specific and test construction in general, as well as the challenges that must be addressed to fully realize their potential.
Breit, M., Brunner, M., Preuß, J., Daseking, M., Esser, G., Grob, A., Freudenstein, J.-P., Kieschke, U., Kreuzpointner, L., Pauls, F., Perleth, C., Ricken, G., Schipolowski, S., Schroeders, U., Walter, F., Wilhelm, O., Wyschkon, A., & Preckel, F. (2025, November 25). General intelligence contributions to cognitive performance across ability levels and age: An IPD meta-analysis of child and adolescent norming data. PsyArXiv.
Cognitive differentiation effects describe differences in the strength of the correlations among cognitive performances depending on general ability level and age. In this context, variations in the contribution of general intelligence (g) to specific cognitive performances have attracted significant interest, yet have not been investigated meta-analytically. The present study provides an examination of moderated g-factor loadings through a two-stage individual participant data (IPD) meta-analysis of 29 samples from 20 German test standardizations in children and adolescents (N = 42,207). Standardization samples ensured population representativeness, sufficient sample size, and representation of different cognitive tests. Using nonlinear factor analysis, we examined the moderation of g-loadings by age, ability, and their interaction. Subsequent multi-level meta-regressions revealed decreasing g-loadings with increasing ability levels (meta-analytic λ2 = –.046 [95% CI -.064, -.028]). Age moderation effects were inconsistent and nonsignificant (meta-analytic λ3 = .003 [95% CI -.024, .030]). A negative age × ability interaction was present in some samples but not significant meta-analytically (meta-analytic λ4 = –.005 [95% CI -.012, .002]). Moderator analyses indicated that effects differed between test families but rarely between CHC abilities. The results highlight ability level as a key source of differences in the contribution of general intelligence to specific cognitive performances, underscoring the need for conditional reliability estimates in cognitive assessment and for cognitive theories on the nature and development of general intelligence to account for this effect. Age differences in g-factor loadings were less consequential, suggesting limited relevance.
Achaa-Amankwaa, P., Walter, J., Hübner, V., Hübner, P., & Schroeders, U. (2025, November 5). Modeling Item Difficulty in Large-Scale Game-Based Cognitive Ability Assessment: An Illustration Using Wordle. PsyArXiv.
Building on the long and successful tradition of incorporating games into cognitive ability assessment, we investigated the popular word puzzle game Wordle, which integrates verbal abilities (lexical knowledge, orthographic processing), strategic word retrieval, and reasoning. Specifically, we examined the extent to which item difficulty in the German adaptation, GridWords, can be predicted using linguistic features, using a large online database that comprised 872 items collected from 76,095 players over two years. Item difficulty was operationalized via the number of attempts needed to solve an item, with scores adjusted for possible within-person practice effects across repeated gameplay. Moving beyond previous approaches that were limited to narrow sets of predictors, we integrated a broad set of features from psycholinguistic research on word recognition, such as word and letter frequencies, orthographic neighborhood size, and letter repetitions. Using regularized machine-learning methods in a nested cross-validation approach, we found out-of-sample predictive performances of R² = 52% (elastic net regression) and R² = 58% (gradient boosting machines) in item difficulty prediction. The cumulative consonant and unigram position frequencies, number of letter repetitions, and number of orthographic neighbors of the target words ranked among the most important features. We discuss feature importances as indicators of the underlying cognitive processes during gameplay, particularly with regard to the relevance of implicit knowledge of lexical regularities and the use of deliberate strategies. We also consider the strengths and limitations of using online games for cognitive ability assessment.
Schroeders, U. (2025, September 6). Crystallized intelligence: Exploring the dark matter of intelligence. PsyArXiv.
Do you know the capital of Mozambique? What does the word Надежда mean? Which author voluntarily declined the Nobel Prize in Literature? Although these questions may appear unrelated, they could all serve as indicators of crystallized intelligence (gc). At the same time, they illustrate some of the challenges involved in defining, assessing, and modeling gc. Several questions arise: How can such diverse content be traced back to a single underlying construct? What role do language skills play? And does the knowledge really reflect widely shared cultural content or more specialized expertise?It also becomes evident that the likelihood of answering the above questions correctly depends on factors such as birthplace, native language, and year of birth. Someone born in Maputo, a Russian speaker, or a person who witnessed the media controversy in 1964, when Sartre famously declined the Nobel Prize on the day it was announced, will likely know the correct answers. A fair assessment of gc across national borders, language groups, and historical eras is therefore challenging. If the construct of gc is so multifaceted, difficult to measure fairly, and often vague in its operationalization, why should we study gc at all? In this chapter, we aim to address this question by first outlining prominent theoretical perspectives on gc (Section 1), followed by a discussion about its structure and organization (Section 2). The subsequent sections examine gc from a developmental perspective (Section 3) and review empirical findings on its determinants and consequences (Section 4). We then turn to the main focus of the chapter—the measurement and psychometric modeling of gc (Sections 5). Finally, we discuss how large language models can inspire future work on how to understand, measure, and model gc (Section 6). Taken together, these perspectives demonstrate that gc remains not only a central pillar of intelligence research but also a construct whose mysteries can be re-examined in light of advances in artificial intelligence and large language models.
Peer-reviewed articles
Advance online publications
Achaa-Amankwaa, P., Walter, J., & Schroeders, U. (2026). Wordle — A game-based assessment of verbal ability? Collabra: Psychology. https://doi.org/10.1525/collabra.167610
There is a long and successful tradition of incorporating games into cognitive ability assessment, with recent research exploring the potential of online games as engaging and valid measures of cognitive ability. Building on this work, we examined the popular word puzzle game Wordle as a possible measure of verbal ability in a sample of N = 370 German-speaking adults in this Registered Report. First, we investigated the psychometric properties of the game (e.g., dimensionality, reliability) within an item response theory (IRT) framework. Second, we located Wordle proficiency within a nomological network of cognitive abilities, including vocabulary, declarative knowledge, verbal fluency, and figural reasoning. Third, we examined the correlations between Wordle proficiency and individual differences in age, gender, education, daily reading time, game experience, self-reported strategy use, and an entropy-based measure of strategy efficiency. The findings supported a unidimensional model of Wordle proficiency with moderate reliability (rxx = .68). In a multidimensional IRT model, Wordle proficiency showed its strongest correlations with figural reasoning (r = .50) and verbal fluency (r = .31), a smaller correlation with vocabulary (r = .22), and no meaningful correlation with declarative knowledge (r = .06). Additionally, Wordle proficiency was positively correlated with education, prior experience with the game, and both strategy use and strategy efficiency, and it was negatively correlated with age. This correlational pattern and a response time analysis at the attempt level suggest that Wordle can be understood as a multifaceted problem-solving task that integrates heuristic word retrieval and lexical access with reasoning processes.
Jankowsky, K., Speck, K., Cummins, J., & Schroeders, U. (2026). Registered reports reduce bias and improve trust in machine learning. Nature Human Behaviour. https://doi.org/10.1038/s41562-026-02586-2
Applying Registered Reports to machine learning (ML) may seem counterintuitive, but many challenges faced in ML research (including information leakage, unstable performance estimates and selective reporting) mirror longstanding concerns in behavioural science. We argue that preregistering modelling workflows combined with an early peer review can reduce bias, improve robustness and make predictions more reproducible.
Speck, K.-L., Jankowsky, K., Scharf, F., & Schroeders, U. (2026). Beyond the hype: A simulation study evaluating the predictive performance of machine learning models in psychology. Psychological Methods. Advance online publication. https://doi.org/10.1037/met0000832
Citations per year
FWCI — citations received divided by those expected for works of the same type, subfield and year (2026); 1.0 is the world average. The figure stays provisional until three years after publication.
Percentile — this work's rank within that same comparison group.
Although machine learning (ML) methods are gaining popularity in psychological research, the debate about their usefulness ranges from hype to disillusionment. The discrepancy between the hopes placed in ML methods and the empirical reality is often attributed to the quality of psychological data sets, which tend to be small and subject to imprecise measurement. In this simulation study, we examined the data requirements necessary for ML methods to perform well. We compared the performance of elastic net regressions with and without prespecified interactions, random forests, and gradient boosting machines for different data-generating processes (including interaction, stepwise, or piecewise linear effects) and under various conditions: (a) sample size, (b) number of irrelevant predictors, (c) predictor reliability, (d) effect size, and (e) nature of the data-generating process (i.e., linear vs. nonlinear effects). We investigated whether the models achieved the highest level of predictive performance attainable under the given simulated conditions. There were two main takeaways from our results: First, the maximum possible predictive performance was only achieved under optimal simulation conditions (N = 1,000, perfectly reliable predictors, predominantly linear effects, and an exceptionally large effect size of R² = .80), which are arguably rarely met in psychological research. Second, each ML model outperformed the others under certain conditions, but none was consistently superior or entirely robust to suboptimal data characteristics. We stress that data quality fundamentally limits predictive performance and discuss the interpretation of comparisons between flexible ML models and simpler (regularized linear) baselines in psychological research.
Schroeders, U., & Achaa-Amankwaa, P. (2026). Developing NOVA: A Next-Generation Open Vocabulary Assessment. European Journal of Psychological Assessment. Advance online publication. https://doi.org/10.1027/1015-5759/a000937
In psychological assessment, vocabulary tests are commonly used as reliable and efficient indicators of crystallized intelligence, as retrospective proxies for premorbid intelligence, and as measures of language proficiency. However, many of the widely used German vocabulary tests are outdated, proprietary, and lack a clear rationale for item selection. To address these limitations, we developed a new, openly available vocabulary test: the Next-Generation Open Vocabulary Assessment (NOVA). We constructed 110 multiple-choice vocabulary items with support from ChatGPT and administered them to 1,052 German-speaking adults using a multiple-matrix design, along with a declarative knowledge test for validation purposes. Using Ant Colony Optimization, we assembled two parallel 30-item short forms by optimizing reliability as well as item difficulty and discrimination parameters within a three-parameter logistic item response model. The resulting test forms provided unidimensional and reliable measurement, covered a broad ability range, showed no gender differences, and correlated strongly with declarative knowledge. A Shiny app is available to calculate norm-referenced scores based on individual test results. Additional analyses revealed that 61.2% of the variance in item difficulty was explained by word frequency and word length, underscoring their potential utility in guiding future vocabulary test development.
Achaa-Amankwaa, P., Trautwein, T., Lenhard, W., & Schroeders, U. (2025). Balancing between categorical and dimensional assessment in short-scale construction using Ant Colony Optimization. European Journal of Psychological Assessment. Advance online publication. https://doi.org/10.1027/1015-5759/a000892
Citations per year
FWCI — citations received divided by those expected for works of the same type, subfield and year (2025); 1.0 is the world average. The figure stays provisional until three years after publication.
Percentile — this work's rank within that same comparison group.
Language proficiency assessment poses particular challenges for test developers in selecting items that allow for a clear assignment of individuals to language proficiency levels (categorical assessment), while at the same time providing a reliable and comprehensive dimensional assessment of language proficiency. We show how Ant Colony Optimization (ACO) can be used to achieve a balance between these measurement goals, using a German entry-level language assessment as a working example. We tailored competing ACO algorithms to develop short scales of different lengths that met several pre-specified criteria, including model fit, composite reliability, and criterion validity. In optimizing the short scales, we favored either accurate dimensional assessment (model fit and composite reliability), between-category classification accuracy (a high polychoric correlation between model-predicted and independently assessed proficiency levels), or a balance of both. We argue that scale optimization strategies such as ACO are essential for balancing conflicting measurement goals such as optimizing between categorical and dimensional assessment.
Schroeders, U., Loos, A., Wiedemann, S., & Jankowsky, K. (2024). Is it just a game? Development and validation of a deductive version of mastermind as measure of reasoning ability. European Journal of Psychological Assessment. Advance online publication. https://doi.org/10.1027/1015-5759/a000855
Citations per year
FWCI — citations received divided by those expected for works of the same type, subfield and year (2024); 1.0 is the world average. The figure stays provisional until three years after publication.
Percentile — this work's rank within that same comparison group.
Building on the long common history of board games and intelligence research, we developed a new deductive reasoning test based on the popular game Mastermind. The research questions of this registered report were: (a) Is a psychometrically sound measurement of the ability to solve Mastermind items possible (i.e., a reliable, uni-dimensional measurement with a good coverage of difficulty)? (b) Is the ability to solve Mastermind items substantially related to other measures of cognitive ability (i.e., matrix test, knowledge test) and need for cognition? (c) Can item difficulty be predicted by the number of colors, positions, premises, and a newly proposed entropy-based index? Based on the results of a pilot study, we developed 30 items and administered them to 351 participants in the preregistered main study using a multiple matrix sampling design. The deductive Mastermind test proved to be (a) a reliable and efficient measure of reasoning across a wide ability range, and (b) showed expectation-consistent patterns for the convergent and divergent measures. (c) The entropy-based index allowed for the prediction of item difficulty to a considerable degree (R² = .47). We discuss the ideas of information theory, including entropy, as constructing principles for the rational test development of reasoning tests.
2026
Trautwein, T., Scharf, F., & Schroeders, U. (2026). Model specification search in correlated factor models using Bee Swarm Optimization. Structural Equation Modeling: A Multidisciplinary Journal, 33(4), 515–530. https://doi.org/10.1080/10705511.2025.2612170
We introduce a Bee Swarm Optimization (BSO), a metaheuristic algorithm for correlated factor models that optimizes test construction by simultaneously determining the dimensionality and the optimal item set of a measure. In an extensive simulation study, we systematically varied the number of factors and indicators, the structural complexity, the magnitude of factor correlations, cross-loadings, and the extent of noise items. We benchmarked the BSO performance against four EFA–CFA pipelines. Performance was assessed via multiple criteria, including global and local model fit, maximal factor correlation, and item retention. BSO consistently recovered optimal models, achieving higher factor-retention accuracy while removing problematic cross-loading and noise items. In 97.2% of datasets and conditions, BSO outperformed the EFA–CFA pipelines. Based on a second simulation, we provide recommendations for hyperparameters. Overall, BSO offers a robust, data-driven alternative to prevalent approaches for developing measures with a correlated factor structure that are interpretable and parsimonious.
Schroeders, U.*, & Walter, J.* (2026). Developing BOLT – A matrix test using Boolean operations to assess logical thinking. Intelligence, 117, Article 102026. https://doi.org/10.1016/j.intell.2026.102026 [* Shared 1st authorship]
Matrix reasoning tasks are widely used measures of fluid reasoning. Despite variation in construction principles, four difficulty components recur across tests: the number of elements, the number of rules, rule type, and perceptual organization. Among these, rule type is particularly central and can be formally classified into unary, binary, and ternary operations. Accordingly, we introduce BOLT (Boolean Operations to assess Logical Thinking), a matrix test grounded in Boolean algebra as the formal framework for item generation. Across two online pilot studies (Study 1: N = 473, 45 items; Study 2: N = 430, 42 items) and one operational administration (Study 3: N = 7150, 39 items), we estimated Rasch item difficulties and examined feature-based explanations of these estimates. Study 1 confirmed that the number of binary rules was the primary driver of difficulty, whereas unary and ternary rules showed no consistent effects. For Studies 2–3, we combined cross-validated LASSO feature selection with linear logistic test models (LLTMs) including random item effects to quantify contributions of structural features (number of elements, number of rules, rule types) and perceptual-organizational indices extracted from stimulus files. Item difficulty could be predicted with sufficient accuracy (LLTM R² = . 74 in Study 2; R² = . 55 in Study 3). Binary operations and perceptual organization emerged as converging determinants of item difficulty in BOLT items. Future work should replicate the finding that image-derived computer-vision indices make an incremental contribution beyond rule-based predictors and connect them to underlying processes of visual organization and rule induction
Altgassen, E.*, Hartung, J.*, Steger, D., Schroeders, U., & Wilhelm, O. (2026). From school lessons to life lessons: School knowledge, life knowledge and their relation to biographical experiences. Intelligence, 116, Article 102007. https://doi.org/10.1016/j.intell.2026.102007 [* Shared 1st authorship]
Crystallized intelligence (gc) is typically defined as the breadth and depth of a person’s knowledge and skills within a culture. Contemporary assessments of gc in adults often focus on school-based knowledge acquired through formal education. To broaden this scope, we developed a gc measure that captures life knowledge acquired through biographical experiences outside formal schooling. A sample of 348 adults completed items assessing school knowledge and newly developed items targeting life knowledge, tailored to specific biographical experiences, covering five areas: humanities, natural sciences, life sciences, arts, and social sciences. Besides the knowledge items, participants also answered questions about the associated biographical experiences. Latent factors for both knowledge item sets correlated perfectly with each other. Nonetheless, correlations between biographical experiences and corresponding life knowledge items were slightly stronger than correlations between biographical items and non-corresponding life or school knowledge items. We discuss the often-neglected relevance of the source of knowledge acquisition and biographical learning opportunities, and consider their implications for the psychometric modeling of gc.
Müller, S., Schroeders, U., Bachrach, N., Benecke, C., Cuevas, L., Doering, S., Elklit, A., Gutiérrez, F., Hengartner, M. P., Hogue, T. E., Hopwood, C. J., Mihura, J. L., Oltmanns, T. F., Paap, M. C. S., Pedersen, G., Renn, D., Ringwald, W. R., Rossi, G., Samuels, J., … Zimmermann, J. (2026). Revisiting the structure of DSM-5 section II personality disorder criteria using individual participant data meta-analysis. Personality Disorders: Theory, Research, and Treatment, 17(1), 1–16. https://doi.org/10.1037/per0000736
Citations per year
FWCI — citations received divided by those expected for works of the same type, subfield and year (2026); 1.0 is the world average. The figure stays provisional until three years after publication.
Percentile — this work's rank within that same comparison group.
The differentiation–dedifferentiation hypothesis of general cognitive ability has been widely studied, but comparable research on crystallized intelligence is scarce. To close this gap, we conducted an empirical test of the age differentiation hypothesis of declarative knowledge as proposed in Cattell’s investment theory, which predicts that knowledge differentiates into diverse forms after compulsory education ends. Thereto, a cross-sectional sample of 1,629 participants aged 18–70 years (M = 45.3) completed a comprehensive knowledge test comprising 120 broadly sampled questions from 12 knowledge domains, as well as a measure of openness/intellect. To investigate age-related differences in the level and structure of knowledge, we performed invariance tests in local structural equation models. The results did not provide any evidence for age-related differentiation of declarative knowledge but indicated age-related differences in the mean structure. Higher levels in openness/intellect were associated with higher levels in knowledge but not with more differentiated structure of knowledge. Contrary to predictions of the investment theory, our results suggest that declarative knowledge is age invariant across major parts of the adult lifespan.
Schroeders, U., Mariss, A., Sauter, J., & Jankowsky, K. (2026). Predicting juvenile delinquency and criminal behavior in adulthood using machine learning. International Journal of Behavioral Development, 50(1), 126–139. https://doi.org/10.1177/01650254251339392
Citations per year
FWCI — citations received divided by those expected for works of the same type, subfield and year (2026); 1.0 is the world average. The figure stays provisional until three years after publication.
Percentile — this work's rank within that same comparison group.
By violating social norms, deviant behavior is an important issue that affects society as a whole and has serious consequences for its individuals. Different scientific disciplines have proposed theories of deviant behavior that often fall short of predicting actual behavior. In this registered report, we used data from the longitudinal National Study of Adolescent to Adult Health (Add Health) to examine the predictability of juvenile delinquency (Wave I) and adult criminal behavior (Wave V), distinguishing between drug, property, and violent offenses. Comparing the predictive accuracy of traditional regression models with different machine learning algorithms (elastic net regression and gradient boosting machines), we found the elastic net regressions with item-level data performed best. The prediction of juvenile delinquency was relatively accurate for drug offenses (R² = .57), violent offenses (R² = .44), and property offenses (R² = .39), while the performance declined significantly for adult delinquency, with R² values ranging from .16 to .13. Key predictors of juvenile delinquency versus adult criminal behavior were clearly different from each other. Early risk factors for adult criminal behavior included prior juvenile delinquency, particularly drug-related offenses, sex, and school-related issues such as suspension or expulsion. We discuss the findings in the context of relevant theories on the causes and development of criminal behavior and explore potential approaches for prevention and early intervention, particularly within the framework of the “Central Eight.”
2025
Busse, F., Zimny, L., Schroeders, U., & Wilhelm, O. (2025). Cloze test performance and cognitive abilities: A comprehensive meta-analysis. Intelligence, 113, Article 101962. https://doi.org/10.1016/j.intell.2025.101962
Cloze tests have a long history and have been used to measure various abilities, including intelligence, reading comprehension, and language proficiency. To locate cloze tests within a nomological network of cognitive abilities, we conducted a multilevel random effects meta-analysis covering 110 years of research. Studies were eligible if they provided a measure of association between a cognitive fill-in-the-blank test and any cognitive ability test. We synthesized manifest correlations from 89 studies (N = 37,912, k = 634) and found an average correlation of r = .54 (95% CI [.49, .59], k = 485) with crystallized intelligence, r = .48 (95% CI [.42, .54], k = 69) with fluid intelligence, and r =.61 (95% CI [.46, .77], k = 32) with general intelligence. While today’s application of the typical cloze is to measure reading comprehension, our results revealed a similarly strong association with a broad range of crystallized abilities. Of the key moderators we investigated—text base, administration mode, deletion pattern, and response type—only the response type showed a significant effect. Sensitivity analyses supported the robustness of our findings. We conclude by revisiting the origin of the cloze test and highlighting the need for systematic studies on how different cloze test designs affect construct validity. Whereas the meta-analytic database predominantly originates from language research, where cloze tests are entrenched as markers of language proficiency, we propose reframing cloze tests as a versatile intelligence test format—just like multiple-choice tests constitute a testing method—that can be tailored to assess various specific cognitive abilities.
Jankowsky, K., Zimmermann, J., Jaeger, U., Mestel, R., & Schroeders, U. (2025). First impressions count: Therapists’ impression on patients’ motivation and helping alliance predicts psychotherapy dropout. Psychotherapy Research, 35(8), 1326–1338. https://doi.org/10.1080/10503307.2024.2411985
Citations per year
FWCI — citations received divided by those expected for works of the same type, subfield and year (2025); 1.0 is the world average. The figure stays provisional until three years after publication.
Percentile — this work's rank within that same comparison group.
Objective. With meta-analytically estimated rates of about 25%, dropout in psychotherapies is a major concern for individuals, clinicians, and the healthcare system at large. To be able to counteract dropout in psychotherapy, accurate insights about its predictors are needed. Method. We compared logistic regression models with two machine learning algorithms (elastic net regressions and gradient boosting machines) in the prediction of therapy dropout in two large inpatient samples (N = 1,691 and N = 12,473) using baseline and initial process variables reported by patients and therapists. Results. Predictive accuracies of the two machine learning algorithms were similar and higher than for logistic regressions: Therapy dropout could be predicted with an AUC of .73 and .83 for Sample 1 and 2, respectively. The initial evaluation of patients’ motivation and the therapeutic alliance rated by the respective therapist were the most important predictors of dropout.Conclusions. Therapy dropout in naturalistic inpatient settings can be predicted to a considerable degree by using baseline indicators and therapists’ first impressions. Feature selection via regularization leads to higher predictive performances whereas non-linear or interaction effects are dispensable. The most promising point of intervention to reduce therapy dropouts seems to be patients’ motivation and the therapeutic alliance.
Schroeders, U., & Gnambs, T. (2025). Sample size planning in item response theory: A tutorial. Advances in Methods and Practices in Psychological Science, 8(1). Article 25152459251314798. https://doi.org/10.1177/25152459251314798
Citations per year
FWCI — citations received divided by those expected for works of the same type, subfield and year (2025); 1.0 is the world average. The figure stays provisional until three years after publication.
Percentile — this work's rank within that same comparison group.
Although item-response-theory (IRT) models offer well-established psychometric advantages over traditional scoring methods, they remain underused in practice. Following a brief introduction to the IRT framework, we emphasize its major advantages and explore potential applications in various research areas. The main part of this tutorial provides a comprehensive, step-by-step guide to Monte Carlo simulation-based sample-size estimation in IRT, which is essential for obtaining precise estimates of item and person parameters, structural effects, and model fit. Accurate a priori sample-size estimation is also crucial for effective study planning, especially in preregistration and registered reports. We highlight 10 key decisions, organized into four areas: (a) determining the data-generation model, (b) defining the test design and the process of missing values, (c) selecting the IRT model and parameters of interest, and (d) setting up and running the Monte Carlo simulation. The procedure is illustrated with examples from educational, personality, and clinical psychology. An extensively annotated and easily customizable syntax is available in an online repository.
2024
Schroeders, U., Morgenstern, M., Jankowsky, K., & Gnambs, T. (2024). Short-scale construction using meta-analytic Ant Colony Optimization: A demonstration with the Need for Cognition Scale. European Journal of Psychological Assessment, 40(5), 376–395. https://doi.org/10.1027/1015-5759/a000818
Citations per year
FWCI — citations received divided by those expected for works of the same type, subfield and year (2024); 1.0 is the world average. The figure stays provisional until three years after publication.
Percentile — this work's rank within that same comparison group.
The Need for Cognition Scale (NCS) is a self-report scale measuring individual differences in the tendency to engage in and enjoy thinking. The shortened version with 18 items (NCS-18; Cacioppo et al., 1984 ) has widely been administered in research on persuasion, critical thinking, and educational achievement. Whereas most studies advocated for essential uni-dimensionality, the question remains which psychometric model yields the best representation of the NCS-18. In the present study, we compared six competing measurement models for the NCS-18 with meta-analytic structural equation models using summary data of 87 samples (N = 90,215). Results demonstrated that the negatively worded items introduced considerable measurement bias that was best accounted for with an acquiescence model. In a further analytical step, we showcased how the pooled correlation matrix can be used to compile short versions of the NCS-18 via Ant Colony Optimization. We examined model fit and reliability of short scales with varying item numbers (between 4 and 15) and a balanced ratio of positively and negatively worded items. We discuss the potentials and limits of the newly proposed method.
Gnambs, T., & Schroeders, U. (2024). Reliability and factorial validity of the Core Self-Evaluations Scale: A meta-analytic investigation of wording effects. European Journal of Psychological Assessment, 40(5), 343–359. https://doi.org/10.1027/1015-5759/a000783
Citations per year
FWCI — citations received divided by those expected for works of the same type, subfield and year (2024); 1.0 is the world average. The figure stays provisional until three years after publication.
Percentile — this work's rank within that same comparison group.
The Core Self-Evaluations Scale (CSES) measures a broad personality trait reflecting individuals’ self-appraisals of their worth, capabilities, and control of their lives. Although the CSES was designed to capture a single trait, factor analytic studies often found more complex measurement structures. These either referred to different content facets or methodological artifacts due to the item wording. The present random-effects meta-analysis summarized correlation matrices from 53 samples including 31,843 respondents. After accounting for acquiescent responding, meta-analytic confirmatory factor analyses revealed a single common factor for all items. The factor was highly reliable (ω = .87) and demonstrated partial metric measurement invariance across English, German, and Spanish language versions as well as cultural tendencies of individualism and flexibility. However, Chinese and Romanian translations exhibited substantially lower factor loadings. These results corroborate the use of the CSES as a unidimensional measure, albeit systematic investigations of measurement invariance are recommended before its use in cross-cultural research.
Zimny, L., Schroeders, U., & Wilhelm, O. (2024). Ant Colony Optimization for parallel test assembly. Behavior Research Methods, 56, 5834–5848. https://doi.org/10.3758/s13428-023-02319-7
Citations per year
FWCI — citations received divided by those expected for works of the same type, subfield and year (2024); 1.0 is the world average. The figure stays provisional until three years after publication.
Percentile — this work's rank within that same comparison group.
Ant colony optimization (ACO) algorithms have previously been used to compile single short scales of psychological constructs. In the present article, we showcase the versatility of the ACO to construct multiple parallel short scales that adhere to several competing and interacting criteria simultaneously. Based on an initial pool of 120 knowledge items, we assembled three 12-item tests that (a) adequately cover the construct at the domain level, (b) follow a unidimensional measurement model, (c) allow reliable and (d) precise measurement of factual knowledge, and (e) are gender-fair. Moreover, we aligned the test characteristic and test information functions of the three tests to establish the equivalence of the tests. We cross-validated the assembled short scales and investigated their association with the full scale and covariates that were not included in the optimization procedure. Finally, we discuss potential extensions to metaheuristic test assembly and the equivalence of parallel knowledge tests in general.
Achaa-Amankwaa, P., Steger, D., Wilhelm, O., & Schroeders, U. (2024). Public events knowledge in an age-heterogeneous sample: Reminiscence bump or bummer? Psychology and Aging, 39(1), 72–87. https://doi.org/10.1037/pag0000786
The reminiscence bump describes an increased recollection of autobiographic experiences made in adolescence and early adulthood. It is unclear if this phenomenon can also be found in declarative knowledge of past public events. To answer this question, we assessed public events knowledge (PEK) about the past 6 decades with a 120-item knowledge test across six domains in a sample of 1,012 Germans that were sampled uniformly across the ages of 30-80 years. General and domain-specific PEK scores were analyzed as a function of age-at-event. Scores were lower for public events preceding participants’ birth and stayed stable from the age-at-event of 5-10 years onward. There was no significant peak in PEK in adolescence or early adulthood, arguing against an extension of the reminiscence effect to factual knowledge. We examined associations between PEK and relevant variables such as crystallized intelligence (Gc), news consumption, and openness to experience with structural equation models. Strong associations between PEK and Gc were established, whereas the associations of PEK with news consumption and openness were mainly driven by their link to declarative knowledge.
Jankowsky, K., Steger, D., & Schroeders, U. (2024). Predicting lifetime suicide attempts in a community sample of adolescents using machine learning algorithms. Assessment, 31(3), 557–573. https://doi.org/10.1177/10731911231167490
Citations per year
FWCI — citations received divided by those expected for works of the same type, subfield and year (2024); 1.0 is the world average. The figure stays provisional until three years after publication.
Percentile — this work's rank within that same comparison group.
Suicide is a major global health concern and a prominent cause of death in adolescents. Previous research on suicide prediction has mainly focused on clinical or adult samples. To prevent suicides at an early stage, however, it is important to screen for risk factors in a community sample of adolescents. We compared the accuracy of logistic regressions, elastic net regressions, and gradient boosting machines in predicting suicide attempts by 17-year-olds in the Millennium Cohort Study (N = 7,347), combining a large set of self- and other-reported variables from different categories. Both machine learning algorithms outperformed logistic regressions and achieved similar balanced accuracies (.76 when using data 3 years before the self-reported lifetime suicide attempts and .85 when using data from the same measurement wave). We identified essential variables that should be considered when screening for suicidal behavior. Finally, we discuss the usefulness of complex machine learning models in suicide prediction.
Schroeders, U., Scharf, F., & Olaru, G. (2024). Model specification searches in structural equation modeling using Bee Swarm Optimization. Educational and Psychological Measurement, 84(1), 40–61. https://doi.org/10.1177/00131644231160552
Citations per year
FWCI — citations received divided by those expected for works of the same type, subfield and year (2024); 1.0 is the world average. The figure stays provisional until three years after publication.
Percentile — this work's rank within that same comparison group.
Metaheuristics are optimization algorithms that efficiently solve a variety of complex combinatorial problems. In psychological research, metaheuristics have been applied in short-scale construction and model specification search. In the present study, we propose a bee swarm optimization (BSO) algorithm to explore the structure underlying a psychological measurement instrument. The algorithm assigns items to an unknown number of nested factors in a confirmatory bifactor model, while simultaneously selecting items for the final scale. To achieve this, the algorithm follows the biological template of bees’ foraging behavior: Scout bees explore new food sources, whereas onlooker bees search in the vicinity of previously explored, promising food sources. Analogously, scout bees in BSO introduce major changes to a model specification (e.g., adding or removing a specific factor), whereas onlooker bees only make minor changes (e.g., adding an item to a factor or swapping items between specific factors). Through this division of labor in an artificial bee colony, the algorithm aims to strike a balance between two opposing strategies diversification (or exploration) versus intensification (or exploitation). We demonstrate the usefulness of the algorithm to find the underlying structure in two empirical data sets (Holzinger-Swineford and short dark triad questionnaire, SDQ3). Furthermore, we illustrate the influence of relevant hyperparameters such as the number of bees in the hive, the percentage of scouts to onlookers, and the number of top solutions to be followed. Finally, useful applications of the new algorithm are discussed, as well as limitations and possible future research opportunities.
Gnambs, T., & Schroeders, U. (2024). Accuracy and precision of fixed and random effects in meta-analyses of randomized control trials for continuous outcomes. Research Synthesis Methods, 15(1), 86–106. https://doi.org/10.1002/jrsm.1673
Citations per year
FWCI — citations received divided by those expected for works of the same type, subfield and year (2024); 1.0 is the world average. The figure stays provisional until three years after publication.
Percentile — this work's rank within that same comparison group.
Meta-analyses of treatment effects in randomized control trials are often faced with the problem of missing information required to calculate effect sizes and their sampling variances. Particularly, correlations between pre- and posttest scores are frequently not available. As an ad-hoc solution, researchers impute a constant value for the missing correlation. As an alternative, we propose adopting a multivariate meta-regression approach that models independent group effect sizes and accounts for the dependency structure using robust variance estimation or three-level modeling. A comprehensive simulation study mimicking realistic conditions of meta-analyses in clinical and educational psychology suggested that imputing a fixed correlation 0.8 or adopting a multivariate meta-regression with robust variance estimation work well for estimating the pooled effect but lead to slightly distorted between-study heterogeneity estimates. In contrast, three-level meta-regressions resulted in largely unbiased fixed effects but more inconsistent prediction intervals. Based on these results recommendations for meta-analytic practice and future meta-analytic developments are provided.
Jankowsky, K., Krakau, L., Schroeders, U., Zwerenz, R., & Beutel, M. E. (2024). Predicting treatment response using machine learning: A registered report. British Journal of Clinical Psychology, 63(2), 137–155. https://doi.org/10.1111/bjc.12452
Citations per year
FWCI — citations received divided by those expected for works of the same type, subfield and year (2024); 1.0 is the world average. The figure stays provisional until three years after publication.
Percentile — this work's rank within that same comparison group.
Objective: Previous research on psychotherapy treatment response has mainly focused on outpatients or clinical trial data which may have low ecological validity regarding naturalistic inpatient samples. To reduce treatment failures by proactively screening for patients at risk of low treatment response, gain more knowledge about risk factors and to evaluate treatments, accurate insights about predictors of treatment response in naturalistic inpatient samples are needed. Methods: We compared the performance of different machine learning algorithms in predicting treatment response, operationalized as a substantial reduction in symptom severity as expressed in the Patient Health Questionnaire Anxiety and Depression Scale. To achieve this goal, we used different sets of variables-(a) demographics, (b) physical indicators, (c) psychological indicators and (d) treatment-related variables-in a naturalistic inpatient sample (N = 723) to specify their joint and unique contribution to treatment success. Results: = .12. Treatment-related variables were the most predictive, followed psychological indicators. Physical indicators and demographics were negligible. Conculsions: Treatment response in naturalistic inpatient settings can be predicted to a considerable degree by using baseline indicators. Regularization via machine learning algorithms leads to higher predictive performances as opposed to including nonlinear and interaction effects. Heterogenous aspects of mental health have incremental predictive value and should be considered as prognostic markers when modelling treatment processes.
2023
Watrin, L., Schroeders, U., & Wilhelm, O. (2023). Gc at its boundaries: A cross-national investigation of declarative knowledge. Learning and Individual Differences, 102, Article 102267. https://doi.org/10.1016/j.lindif.2023.102267
Citations per year
FWCI — citations received divided by those expected for works of the same type, subfield and year (2023); 1.0 is the world average. The figure stays provisional until three years after publication.
Percentile — this work's rank within that same comparison group.
Crystallized intelligence (gc) is considered culture-specific, but this notion is rarely substantiated empirically. We empirically investigated to what extent the measurement of declarative knowledge depends on the national specificity of its indicators and individuals’ affinity for certain countries, respectively. Therefore, we administered a knowledge test with 75 items to participants from Germany, France, and the USA (Ntotal = 906). Each of the 15 domains was measured with both country-specific and global items. We found a strong national specificity of knowledge in the social sciences and humanities but no systematic differences in the natural sciences. Country-specific knowledge shared common variance beyond a general knowledge factor and could be predicted by the respective country of residence. No association of affinity for and knowledge about other countries was observed. Regional variations in the composition of knowledge tests pose a substantial threat to cross-national comparison but might foster our understanding of gc.
Steger, D., Jankowsky, K., Schroeders, U., & Wilhelm, O. (2023). The road to hell is paved with good intentions: How common practices in scale construction hurt validity. Assessment, 30(6), 1811–1824. https://doi.org/10.1177/10731911221124846
Citations per year
FWCI — citations received divided by those expected for works of the same type, subfield and year (2023); 1.0 is the world average. The figure stays provisional until three years after publication.
Percentile — this work's rank within that same comparison group.
Sound scale construction is pivotal to the measurement of psychological constructs. Common item sampling procedures emphasize aspects of reliability to the disadvantage of aspects of validity, which are less tangible. We use a health knowledge test as an example to demonstrate how item sampling strategies that focus on either factor saturation or construct coverage influence scale composition and demonstrate how to find a trade-off between these two opposing needs. More specifically, we compile three 75-item health knowledge scales using Ant Colony Optimization, a metaheuristic algorithm that is inspired by the foraging behavior of ants, to optimize factor saturation, construct coverage, or a compromise of both. We demonstrate that our approach is well suited to balance out construct coverage and factor saturation when constructing a health knowledge test. Finally, we discuss conceptual problems with the modeling of declarative knowledge and provide recommendations for the assessment of health knowledge.
Wendt, L. P., Jankowsky, K., Schroeders, U., Nolte, T., Fonagy, P., Montague, P. R., Zimmermann, J., & Olaru, G. (2023). Mapping established psychopathology scales onto the Hierarchical Taxonomy of Psychopathology (HiTOP). Personality and Mental Health, 17(2), 117–134. https://doi.org/10.1002/pmh.1566
Citations per year
FWCI — citations received divided by those expected for works of the same type, subfield and year (2023); 1.0 is the world average. The figure stays provisional until three years after publication.
Percentile — this work's rank within that same comparison group.
The Hierarchical Taxonomy of Psychopathology (HiTOP) organizes phenotypes of mental disorder based on empirical covariation, offering a comprehensive organizational framework from narrow symptoms to broader patterns of psychopathology. We argue that established self-report measures of psychopathology from the pre-HiTOP era should be systematically integrated into HiTOP to foster cumulative research and further the understanding of psychopathology structure. Hence, in this study, we mapped 92 established psychopathology (sub)scales onto the current HiTOP working model using data from an extensive battery of self-report assessments that was completed by community participants and outpatients (N = 909). Content validity ratings of the item pool were used to select indicators for a bifactor-(S-1) model of the p factor and five HiTOP spectra (i.e., internalizing, thought disorder, detachment, disinhibited externalizing, and antagonistic externalizing). The content-based HiTOP scales were validated against personality disorder diagnoses as assessed by standardized interviews. We then located established scales within the taxonomy by estimating the extent to which scales reflected higher-level HiTOP dimensions. The analyses shed light on the location of established psychopathology scales in HiTOP, identifying pure markers and blends of HiTOP spectra, as well as pure markers of the p factor (i.e., scales assessing mentalizing impairment and suspiciousness/epistemic mistrust).
Goecke, B., Schroeders, U., Zettler, I., Schipolowski, S., Golle, J., & Wilhelm, O. (2023). The nomological net of knowledge, self-reported knowledge, and overclaiming in children. Journal of Personality Assessment, 105(5), 702–713. https://doi.org/10.1080/00223891.2022.2144332
Citations per year
FWCI — citations received divided by those expected for works of the same type, subfield and year (2023); 1.0 is the world average. The figure stays provisional until three years after publication.
Percentile — this work's rank within that same comparison group.
Research on self-reported knowledge and overclaiming in children is sparse. With the current study, we aim to close this gap by developing an overclaiming questionnaire measuring self-reported knowledge and overclaiming that is tailored to children. Moreover, we examine the nomological net of self-reported knowledge and overclaiming in childhood discussing three perspectives: Overclaiming as (a) a result of deliberate self-enhancement tendencies, (b) a proxy for declarative knowledge, and (c) an indicator of creative engagement. We juxtaposed overclaiming, as indicated by claiming familiarity with non-existent terms, and self-reported knowledge with fluid and crystallized intelligence, creativity, and personality traits in a sample of 897 children attending third grade. The results of several latent variable analyses were similar to findings known from adult samples: We found no strong evidence for any of the competing perspectives on overclaiming. Just like in adults, individual differences in self-reported knowledge were strongly inflated by overclaiming, and only weakly related to declarative knowledge.
2022
Hartung, J., Goecke, B., Schroeders, U., Schmitz, F., & Wilhelm, O. (2022). Latin Square Tasks: A multi-study evaluation. Intelligence, 94, Article 101683. https://doi.org/10.1016/j.intell.2022.101683
Citations per year
FWCI — citations received divided by those expected for works of the same type, subfield and year (2022); 1.0 is the world average. The figure stays provisional until three years after publication.
Percentile — this work's rank within that same comparison group.
The Latin Square Task (LST) was proposed as a theoretically well-grounded paradigm for measuring fluid intelligence. In four studies (total N = 3,439) we systematically investigated the psychometric properties of LSTs. Results provided evidence that (a) the construct is unidimensional, (b) the administered stimulus types and the rotation of the item matrix only played a minor role, and (c) the relations with other measures of reasoning ability were in the expected range (about r = .50), confirming the validity of LSTs. For unsupervised automated test construction, item difficulty was only insufficiently accounted for by relational complexity and the number of steps that need to be memorized for solving an item. We discuss possible remedies for improving the paradigm including its generalizability. Furthermore, we propose that the binding-hypothesis of working memory is theoretically suited to account for item difficulties in LSTs.
Olaru, G., Robitzsch, A., Hildebrandt, A., & Schroeders, U. (2022). Examining moderators of vocabulary acquisition from kindergarten throughout elementary school using Local Structural Equation Modeling. Learning and Individual Differences, 95, Article 102136. https://doi.org/10.1016/j.lindif.2022.102136
Citations per year
FWCI — citations received divided by those expected for works of the same type, subfield and year (2022); 1.0 is the world average. The figure stays provisional until three years after publication.
Percentile — this work's rank within that same comparison group.
Parental socio-economic status (SES) is often found to be associated with children’s language competence in the first decade of life. To examine the effect of SES on children’s vocabulary development, as well as potential compensatory effects of schooling and learning-related activities, we examined the joint and unique effects of parental education, occupational status, and learning environment at home on children’s receptive vocabulary competence and growth in early childhood. We used latent growth curve models to assess pre-school receptive vocabulary and growth across primary school. Analyses were based on data from the German National Educational Panel Study (NEPS), a large-scale longitudinal study assessing vocabulary competence and family background from Kindergarten to the 3rd grade of elementary school. To examine the moderating effects of parental education, occupational status, and learning environment at home, we used local structural equation modeling. Results revealed a moderate to strong positive association between parental education and children’s receptive vocabulary competence, which fully explained the effect of occupational status on this language skill. With the exception of the activity of reading aloud, we found no effect of learning environment at home. Initially lower performing children showed steeper growth trajectories across school, but rank-orders were relatively stable across time. In summary, the results suggest large initial differences in receptive vocabulary between children from different educational backgrounds, which are reduced, but not fully overcome across elementary school.
Schroeders, U., Zimmermann, J., Wicke, T., Schaumburg, M., Lang, E., Trenkwalder, C., & Mollenhauer, B. (2022). Dynamic interplay of cognitive functioning and depressive symptoms in patients with Parkinson’s disease. Neuropsychology, 36(4), 266–278. https://doi.org/10.1037/neu0000795
Citations per year
FWCI — citations received divided by those expected for works of the same type, subfield and year (2022); 1.0 is the world average. The figure stays provisional until three years after publication.
Percentile — this work's rank within that same comparison group.
Objective: We examine the trajectories of and the dynamic interplay between cognitive functioning and depressive symptoms in patients with Parkinson’s disease (PD) in comparison to healthy controls (HC) from an intraindividual perspective. Method: The DeNoPa study is a single-center, observational, longitudinal study with biennial follow-ups over 8 years. The present analyses are based on 123 PD (79 male) and 107 HC (64 male) with a mean age of 64.1 years (SD = 8.3). PD and HC completed a battery of neuropsychological tests and scales assessing depressive symptoms. We used a random-intercept crosslagged panel model (RI-CLPM) to study their trajectories and the dynamic interplay. Results: Cognitive abilities of PD were on average d = −0.67 worse at baseline and d = −1.22 at 8-years follow-up in comparison to HC. Depressive symptoms in PD showed large variability and followed a U-shaped trajectory. From an intraindividual perspective, greater impairments in cognitive abilities were subsequently associated with increased depressive symptoms (b = −0.60, p = .03), whereas the effect in the opposite direction was not significant. Conclusions: We found indication that a decline on a global composite scale of cognition can be seen as a precursor of depressive symptoms in patients with PD. To counter cognitive losses and the subsequent mood deterioration, patient education and early cognitive (and behavioral) enrichment seem promising candidates for treatment.
Schroeders, U., & Jansen, M. (2022). Science self-concept – More than the sum of its parts? The Journal of Experimental Education, 90(2), 435–451. https://doi.org/10.1080/00220973.2020.1740967
Citations per year
FWCI — citations received divided by those expected for works of the same type, subfield and year (2022); 1.0 is the world average. The figure stays provisional until three years after publication.
Percentile — this work's rank within that same comparison group.
Academic self-concept is understood as a multidimensional, hierarchical construct. Multidimensionality refers to the subject-specific differentiation of academic self-concepts, whereas hierarchy refers to the aggregation of more specific facets of self-concepts into more general ones. Previous research demonstrated that students distinguish between their self-concepts in biology, chemistry, and physics, if taught as separate school subjects, as is done in Germany. However, large-scale international educational studies, such as PISA, often use a monolithic science self-concept measure. It is yet unclear whether an aggregate of subject-specific self-concepts is equivalent to a directly measured science self-concept. We assessed the subject-specific and a global science self-concept of 1,229 German grade 10 students. A higher-order factor model and a bifactor model demonstrated a very high correlation between the “inferred” and the explicitly assessed global science self-concept. Despite the high empirical overlap, we argue for a more nuanced view of the science self-concept, because statistical unity is not to be confused with causal unity. Moreover, from a methodological perspective, we used multi-group confirmatory factor analysis to examine the mean structure and local structural equation models to study measurement invariance across science ability. Implications for the theoretical status of self-concept as a hierarchical construct are discussed.
Schroeders, U., Schmidt, C., & Gnambs, T. (2022). Detecting careless responding in survey data using stochastic gradient boosting. Educational and Psychological Measurement, 82(1), 29–56. https://doi.org/10.1177/00131644211004708
Citations per year
FWCI — citations received divided by those expected for works of the same type, subfield and year (2022); 1.0 is the world average. The figure stays provisional until three years after publication.
Percentile — this work's rank within that same comparison group.
Careless responding is a bias in survey responses that disregards the actual item content, constituting a threat to the factor structure, reliability, and validity of psychological measurements. Different approaches have been proposed to detect aberrant responses such as probing questions that directly assess test-taking behavior (e.g., bogus items), auxiliary or paradata (e.g., response times), or data-driven statistical techniques (e.g., Mahalanobis distance). In the present study, gradient boosted trees, a state-of-the-art machine learning technique, are introduced to identify careless respondents. The performance of the approach was compared with established techniques previously described in the literature (e.g., statistical outlier methods, consistency analyses, and response pattern functions) using simulated data and empirical data from a web-based study, in which diligent versus careless response behavior was experimentally induced. In the simulation study, gradient boosting machines outperformed traditional detection mechanisms in flagging aberrant responses. However, this advantage did not transfer to the empirical study. In terms of precision, the results of both traditional and the novel detection mechanisms were unsatisfactory, although the latter incorporated response times as additional information. The comparison between the results of the simulation and the online study showed that responses in real-world settings seem to be much more erratic than can be expected from the simulation studies. We critically discuss the generalizability of currently available detection methods and provide an outlook on future research on the detection of aberrant response patterns in survey research.
Schroeders, U., Kubera, F., & Gnambs, T. (2022). The structure of the Toronto Alexithymia Scale (TAS-20): A meta-analytic confirmatory factor analysis. Assessment, 29(8), 1806–1823. https://doi.org/10.1177/10731911211033894
Citations per year
FWCI — citations received divided by those expected for works of the same type, subfield and year (2022); 1.0 is the world average. The figure stays provisional until three years after publication.
Percentile — this work's rank within that same comparison group.
Alexithymia is defined as the inability of persons to describe their emotional states, to identify the feelings of others, and a utilitarian type of thinking. The most popular instrument to assess alexithymia is the Toronto Alexithymia Scale (TAS-20). Despite its widespread use, an ongoing controversy pertains to its internal structure. The TAS-20 was originally constructed to capture three different factors, but several studies suggested different factor solutions, including bifactor models and models with a method factor for the reversely keyed items. The present study examined the dimensionality of the TAS-20 using summary data of 88 samples from 62 studies (total N = 69,722) with meta-analytic structural equation modeling. We found support for the originally proposed three-dimensional solution, whereas more complex models produced inconsistent factor loadings. Because a major source of misfit stems from translated versions, the results are discussed with respect to generalizations across languages and cultural contexts.
Jankowsky, K., & Schroeders, U. (2022). Validation and generalizability of machine learning prediction models on attrition in longitudinal studies. International Journal of Behavioral Development, 46(2), 169–176. https://doi.org/10.1177/01650254221075034
Citations per year
FWCI — citations received divided by those expected for works of the same type, subfield and year (2022); 1.0 is the world average. The figure stays provisional until three years after publication.
Percentile — this work's rank within that same comparison group.
Attrition in longitudinal studies is a major threat to the representativeness of the data and the generalizability of the findings. Typical approaches to address systematic nonresponse are either expensive and unsatisfactory (e.g., oversampling) or rely on the unrealistic assumption of data missing at random (e.g., multiple imputation). Thus, models that effectively predict who most likely drops out in subsequent occasions might offer the opportunity to take countermeasures (e.g., incentives). With the current study, we introduce a longitudinal model validation approach and examine whether attrition in two nationally representative longitudinal panel studies can be predicted accurately. We compare the performance of a basic logistic regression model with a more flexible, data-driven machine learning algorithm—gradient boosting machines. Our results show almost no difference in accuracies for both modeling approaches, which contradicts claims of similar studies on survey attrition. Prediction models could not be generalized across surveys and were less accurate when tested at a later survey wave. We discuss the implications of these findings for survey retention, the use of complex machine learning algorithms, and give some recommendations to deal with study attrition.
Watrin, L., Schroeders, U., & Wilhelm, O. (2022). Structural invariance of declarative knowledge across the adult lifespan. Psychology and Aging, 37(3), 283–297. https://doi.org/10.1037/pag0000660
Citations per year
FWCI — citations received divided by those expected for works of the same type, subfield and year (2022); 1.0 is the world average. The figure stays provisional until three years after publication.
Percentile — this work's rank within that same comparison group.
The differentiation–dedifferentiation hypothesis of general cognitive ability has been widely studied, but comparable research on crystallized intelligence is scarce. To close this gap, we conducted an empirical test of the age differentiation hypothesis of declarative knowledge as proposed in Cattell’s investment theory, which predicts that knowledge differentiates into diverse forms after compulsory education ends. Thereto, a cross-sectional sample of 1,629 participants aged 18–70 years (M = 45.3) completed a comprehensive knowledge test comprising 120 broadly sampled questions from 12 knowledge domains, as well as a measure of openness/intellect. To investigate age-related differences in the level and structure of knowledge, we performed invariance tests in local structural equation models. The results did not provide any evidence for age-related differentiation of declarative knowledge but indicated age-related differences in the mean structure. Higher levels in openness/intellect were associated with higher levels in knowledge but not with more differentiated structure of knowledge. Contrary to predictions of the investment theory, our results suggest that declarative knowledge is age invariant across major parts of the adult lifespan.
2021
Schroeders, U., Watrin, L., & Wilhelm, O. (2021). Age-related nuances in knowledge assessment. Intelligence, 85, Article 101526. https://doi.org/10.1016/j.intell.2021.101526
Citations per year
FWCI — citations received divided by those expected for works of the same type, subfield and year (2021); 1.0 is the world average. The figure stays provisional until three years after publication.
Percentile — this work's rank within that same comparison group.
Although crystallized intelligence (gc) is a prominent factor in contemporary theories of individual differences in intelligence, its structure and optimal measurement are elusive. Analogously to the personality trait hierarchy, we propose the following hierarchy of declarative fact knowledge as a key component of gc: a general fact knowledge factor at the apex, followed by broad knowledge areas (e.g., natural sciences, social sciences, humanities), knowledge domains (e.g., chemistry, law, art), and nuances. In most scientific contexts we are predominantly concerned with aggregate levels, but we argue that the sampling of knowledge items strongly affects distinctions at higher levels of the hierarchy. We illustrate the magnitude of item-level heterogeneity by predicting chronological age differences through knowledge differences at different levels of the hierarchy. Analyses were based on an online sample of 1629 participants between age 18 and 70 who completed 120 broadly sampled declarative knowledge items across twelve domains. The results of linear and elastic net regressions, respectively, demonstrated that the majority of the age variance was located at the item level, and the strength of the prediction decreased with increasing aggregation. Knowledge nuances seem to tap important variance that is not covered by aggregate scores (e.g., sum or factor scores) and that is useful in the prediction of age. In turn, these effects extend our understanding how knowledge is acquired and imparted. On a more general stance, to gain new insights into the nature of knowledge, its optimal measurement and psychometric representation, item and person sampling issues should be considered.
Achaa-Amankwaa, P., Olaru, G., & Schroeders, U. (2021). Coffee or tea? Examining cross-cultural differences in personality nuances across former colonies of the British Empire. European Journal of Personality, 35(3), 383–397. https://doi.org/10.1177/0890207020962327
Citations per year
FWCI — citations received divided by those expected for works of the same type, subfield and year (2021); 1.0 is the world average. The figure stays provisional until three years after publication.
Percentile — this work's rank within that same comparison group.
Cross-cultural comparisons often focus on differences in broad personality traits across countries. However, many cross-cultural studies report differential item functioning which suggests that considerable group differences are not accounted for by the overarching personality factors. We argue that this may reflect cross-cultural personality differences at a lower level of personality, namely personality nuances. To investigate the degree of cultural similarities and differences between participants of 10 English speaking countries (of which nine formerly belonged to the British Empire), we scrutinized participants’ personality scores on the domain, facet, and nuance level of the personality hierarchy. More specifically, we used the responses of 9110 participants on the IPIP-NEO 300-item personality inventory in cross-validated and regularized logistic regressions. Based on the trait domain and facet scores, we were able to identify the country of residence for 60% and 73% of the participants, respectively. By using the nuance level of personality, we correctly identified the nationality of 89% of the participants. This pattern of results explains the lack of measurement invariance in cross-cultural studies. We discuss implications for cross-cultural personality research and whether the high degree of cross-cultural item-level differences compromises the universality of the personality structure.
Steger, D., Schroeders, U., & Wilhelm, O. (2021). Caught in the act: Predicting cheating in unproctored knowledge assessment. Assessment, 28(3), 1004–1017. https://doi.org/10.1177/1073191120914970
Citations per year
FWCI — citations received divided by those expected for works of the same type, subfield and year (2021); 1.0 is the world average. The figure stays provisional until three years after publication.
Percentile — this work's rank within that same comparison group.
Cheating is a serious threat in unproctored ability assessment, irrespective of countermeasures taken, anticipated consequences (high vs. low stakes), and test modality (paper-pencil vs. computer-based). In the present study, we examined the power of (a) self-report-based indicators (i.e., Honesty-Humility and Overclaiming scales), (b) test data (i.e., performance with extremely difficult items), and (c) para data (i.e., reaction times, switching between browser tabs) to predict participants’ cheating behavior. To this end, 315 participants worked on a knowledge test in an unproctored online assessment and subsequently in a proctored lab assessment. We used multiple regression analysis and an extended latent change score model to assess the potential of the different indicators to predict cheating. In summary, test data and para data performed best, while traditional self-report-based indicators were not predictive. We discuss the findings with respect to unproctored testing in general and provide practical advice on cheating detection in online ability assessments.
2020
Goecke, B., Weiss, S., Steger, D., Schroeders, U., & Wilhelm, O. (2020). Testing competing claims about overclaiming. Intelligence, 81, Article 101470. https://doi.org/10.1016/j.intell.2020.101470
Citations per year
FWCI — citations received divided by those expected for works of the same type, subfield and year (2020); 1.0 is the world average. The figure stays provisional until three years after publication.
Percentile — this work's rank within that same comparison group.
Overclaiming has been described as people’s tendency to overestimate their cognitive abilities in general and their knowledge in particular. We discuss four different perspectives on the phenomenon of overclaiming that have been proposed in the research literature: Overclaiming as a result of a) self-enhancement tendencies, b) as a cognitive bias (e.g., hindsight bias, memory bias), c) as proxy for cognitive abilities, and d) as sign of creative engagement. Moreover, we discuss two different scoring methods for an OCQ (signal detection theory vs. familiarity ratings). To distinguish between the different viewpoints of what overclaiming is, we juxtaposed overclaiming, as indicated by claiming familiarity with non-existent terms, with fluid and crystallized intelligence, self-reported knowledge, creativity, faking ability, and personality. Overclaiming was measured with a newly comprised overclaiming questionnaire. Results of several latent variable analyses based upon a multivariate study with 298 participants were: First, overclaiming is neither predicted by honesty-humility nor faking ability and therefore reflects something different than mere self-enhancement tendencies. Second, overclaiming is not predicted by crystallized intelligence, but is highly predictive of self-reported knowledge and, thus, not suitable as an index or a proxy for cognitive abilities. Finally, overclaiming is neither related to divergent thinking and originality, and only moderately predicted by self-reported openness creativity from the HEXACO which means that overclaiming does not reflect creative ability. In sum, our results favor an interpretation of overclaiming as a phenomenon that requires more than self-enhancement motivation, in contrast to the claim that was initially proposed in the literature.
Weiss, S., Steger, D., Schroeders, U., & Wilhelm, O. (2020). A reappraisal of the threshold hypothesis of creativity and intelligence. Journal of Intelligence, 8(4), Article 38. https://doi.org/10.3390/jintelligence8040038
Citations per year
FWCI — citations received divided by those expected for works of the same type, subfield and year (2020); 1.0 is the world average. The figure stays provisional until three years after publication.
Percentile — this work's rank within that same comparison group.
Intelligence has been declared as a necessary but not sufficient condition for creativity, which was subsequently (erroneously) translated into the so-called threshold hypothesis. This hypothesis predicts a change in the correlation between creativity and intelligence at around 1.33 standard deviations above the population mean. A closer inspection of previous inconclusive results suggests that the heterogeneity is mostly due to the use of suboptimal data analytical procedures. Herein, we applied and compared three methods that allowed us to handle intelligence as a continuous variable. In more detail, we examined the threshold of the creativity-intelligence relation with (a) scatterplots and heteroscedasticity analysis, (b) segmented regression analysis, and (c) local structural equation models in two multivariate studies (N1 = 456; N2 = 438). We found no evidence for the threshold hypothesis of creativity across different analytical procedures in both studies. Given the problematic history of the threshold hypothesis and its unequivocal rejection with appropriate multivariate methods, we recommend the total abandonment of the threshold.
Weiss, S., Steger, D., Kaur, Y., Hildebrandt, A., Schroeders, U., & Wilhelm, O. (2020). On the trail of creativity: Dimensionality of divergent thinking and its relation with cognitive abilities, personality and insight. European Journal of Personality, 35(3), 291–314. https://doi.org/10.1002/per.2288
Citations per year
FWCI — citations received divided by those expected for works of the same type, subfield and year (2020); 1.0 is the world average. The figure stays provisional until three years after publication.
Percentile — this work's rank within that same comparison group.
Divergent thinking (DT) is an important constituent of creativity that captures aspects of fluency and originality. The literature lacks multivariate studies that report relationships between DT and its aspects with relevant covariates, such as cognitive abilities, personality traits (e.g. openness), and insight. In two multivariate studies (N = 152 and N = 298), we evaluate competing measurement models for a variety of DT tests and examine the relationship between DT and established cognitive abilities, personality traits, and insight. A nested factor model with a general DT and a nested originality factor described the data well. In Study 1, DT was moderately related with working memory, fluid intelligence, crystallized intelligence, and mental speed. In Study 2, we replicate these results and add insight, openness, extraversion, and honesty–humility as covariates. DT was associated with insight, extraversion, and honesty–humility, whereas crystallized intelligence mediated the relationship between openness and DT. In contrast, the nested originality factor (i.e. the specificity of originality tasks beyond other DT tasks) had low variance and was not meaningfully related with any other constructs in the nomological net. We highlight avenues for future research by discussing issues of measurement and scoring.
Jankowsky, K., Olaru, G., & Schroeders, U. (2020). Compiling measurement invariant short scales in cross-cultural personality assessment using Ant Colony Optimization. European Journal of Personality, 34(3), 470–485. https://doi.org/10.1002/per.2260
Citations per year
FWCI — citations received divided by those expected for works of the same type, subfield and year (2020); 1.0 is the world average. The figure stays provisional until three years after publication.
Percentile — this work's rank within that same comparison group.
Examining the influence of culture on personality and its unbiased assessment is the main subject of cross–cultural personality research. Recent large–scale studies exploring personality differences across cultures share substantial methodological and psychometric shortcomings that render it difficult to differentiate between method and trait variance. One prominent example is the implicit assumption of cross–cultural measurement invariance in personality questionnaires. In the rare instances where measurement invariance across cultures was tested, scalar measurement invariance—which is required for unbiased mean–level comparisons of personality traits—did not hold. In this article, we present an item sampling procedure, ant colony optimization, which can be used to select item sets that satisfy multiple psychometric requirements including model fit, reliability, and measurement invariance. We constructed short scales of the IPIP–NEO–300 for a group of countries that are culturally similar (USA, Australia, Canada, and UK) as well as a group of countries with distinct cultures (USA, India, Singapore, and Sweden). In addition to examining factor mean differences across countries, we provide recommendations for cross–cultural research in general. From a methodological perspective, we demonstrate ant colony optimization’s versatility and flexibility as an item sampling procedure to derive measurement invariant scales for cross–cultural research.
Schroeders, U., & Gnambs, T. (2020). Degrees of freedom in multigroup confirmatory factor analyses: Are models of measurement invariance testing correctly specified? European Journal of Psychological Assessment, 36(1), 105–113. https://doi.org/10.1027/1015-5759/a000500
Citations per year
FWCI — citations received divided by those expected for works of the same type, subfield and year (2020); 1.0 is the world average. The figure stays provisional until three years after publication.
Percentile — this work's rank within that same comparison group.
Measurement invaraiance is a key concept in psychological assessment and a fundamental prerequisite for meaningful comparisons across groups. In the prevalent approach, multigroup confirmatory factor analysis (MGCFA), specific measurement parameters are constrained to equality across groups. The degrees of freedom ( df) for these models readily follow from the hypothesized measurement model and the invariance constraints. In light of research questioning the soundness of statistical reporting in psychology, we examined how often reported df match with the df recalcualted based on information given in the publications. More specifically, we reviewed 128 studies from six leading peer-reviewed journals focusing on psychological assessment and recalculated the df for 302 measurement invariance testing procedures. Overall, about a quarter of all articles included at least one discrepancy with metric and scalar invariance being more frequently affected. We discuss moderators of these discrepancies and identify typical pitfalls in measurement invariance testing. Moreover, we provide example syntax for different methods of scaling latent variables and introduce a tool that allows for the recalculation of df in common MGCFA models to improve the statistical soundness of invariance testing in psychological research.
Steger, D., Schroeders, U., & Gnambs, T. (2020). A meta-analysis of test scores in proctored and unproctored ability assessment. European Journal of Psychological Assessment, 36(1), 174–184. https://doi.org/10.1027/1015-5759/a000494
Citations per year
FWCI — citations received divided by those expected for works of the same type, subfield and year (2020); 1.0 is the world average. The figure stays provisional until three years after publication.
Percentile — this work's rank within that same comparison group.
Unproctored, web-based assessments are frequently compromised by a lack of control over the participants’ test-taking behavior. It is likely that participants cheat if personal consequences are high. This meta-analysis summarizes findings on context effects in unproctored and proctored ability assessments and examines mean score differences and correlations between both assessment contexts. As potential moderators, we consider (a) the perceived consequences of the assessment, (b) countermeasures against cheating, (c) the susceptibility to cheating of the measure itself, and (d) the use of different test media. For standardized mean differences, a three-level random-effects meta-analysis based on 109 effect sizes from 49 studies (total N = 100,434) identified a pooled effect of Δ = 0.20, 95% CI [0.10, 0.31], indicating higher scores in unproctored assessments. Moderator analyses revealed significantly smaller effects for measures that are difficult to research on the Internet. These results demonstrate that unproctored ability assessments are biased by cheating. Unproctored assessments may be most suitable for tasks that are difficult to search on the Internet.
2019
Sewasew, D., & Schroeders, U. (2019). The developmental interplay of academic self-concept and achievement within and across domains among primary school students. Contemporary Educational Psychology, 58, 204–212. https://doi.org/10.1016/j.cedpsych.2019.03.009
Citations per year
FWCI — citations received divided by those expected for works of the same type, subfield and year (2019); 1.0 is the world average. The figure stays provisional until three years after publication.
Percentile — this work's rank within that same comparison group.
The reciprocal internal/external frame of reference model (RI/EM) extends the internal/external frame of reference model (I/EM) over time and the reciprocal effects model (REM) across domains. The RI/EM postulates positive developmental relations between academic achievement and self-concept within a domain and negative relations across two non-matching domains (e.g., math and English). However, until now, empirical investigations of the RI/EM had only focused on secondary school students from specific countries. In the present study, we test whether the RI/EM also applies to primary school students and to students in the United States, by using a representative longitudinal data set: the Early Childhood Longitudinal Study-Kindergarten (ECLS-K: 1998-1999). We found positive reciprocal relations between academic self-concept and standardized test scores within a domain, whereas the effect of prior achievement on self-concept was much stronger (skill-development part) than the effect of self-concept on achievement (self-enhancement). Furthermore, we found negative effects of achievement on subsequent self-concepts across domains (I/E frame of references). Overall, the findings of the study strongly support the RI/EM for primary school students. Our results are compared to previous findings in the literature for secondary school students and are discussed with regard to self-concept formation in primary school.
Olaru, G., Schroeders, U., Hartung, J., & Wilhelm, O. (2019). Ant Colony Optimization and local weighted structural equation modeling. A tutorial on novel item and person sampling procedures for personality research. European Journal of Personality, 33(3), 400–419. https://doi.org/10.1002/per.2195
Citations per year
FWCI — citations received divided by those expected for works of the same type, subfield and year (2019); 1.0 is the world average. The figure stays provisional until three years after publication.
Percentile — this work's rank within that same comparison group.
Measurement in personality development faces many psychometric problems. First, theory–based measurement models do not fit the empirical data in terms of traditional confirmatory factor analysis. Second, measurement invariance across age, which is necessary for a meaningful interpretation of age–associated personality differences, is rarely accomplished. Finally, continuous moderator variables, such as age, are often artificially categorized. This categorization leads to bias when interpreting differences in personality across age. In this tutorial, we introduce methods to remedy these problems. We illustrate how Ant Colony Optimization can be used to sample indicators that meet prespecified demands such as model fit. Further, we use Local Structural Equation Modeling to resample and weight subjects to study differences in the measurement model across age as a continuous moderator variable. We also provide a detailed illustration for both tools with the Neuroticism scale of the openly available International Personality Item Pool – NEO inventory using data from the UK sample (N = 15 827). Combined, both tools can remedy persistent problems in research on personality and its development. In addition to a step–by–step illustration, we provide commented syntax for both tools.
Steger, D., Schroeders, U., & Wilhelm, O. (2019). On the dimensionality of crystallized intelligence: A smartphone-based assessment. Intelligence, 72, 76–85. https://doi.org/10.1016/j.intell.2018.12.002
Citations per year
FWCI — citations received divided by those expected for works of the same type, subfield and year (2019); 1.0 is the world average. The figure stays provisional until three years after publication.
Percentile — this work's rank within that same comparison group.
Crystallized intelligence (gc) is a prominent factor in consensual theories on the structure of intelligence. Although declarative knowledge is arguably a core aspect of gc, little is known about the dimensionality of knowledge in adults; the proposed dimensional models vary broadly from unidimensionality, to three-dimensional models (science, humanities, and civics), to a six-dimensional model with an overarching g-factor. While previous studies were mostly based on narrow item samples once administered to a specific sample within a restricted time frame, we used a smartphone-based approach to investigate the dimensionality of knowledge based on a large set of items administered to a heterogeneous sample. More specifically, questions were randomly drawn from a pool of 4050 items from 34 subject domains such as chemistry, arts, and politics and administered to an age- and ability-heterogeneous sample of 1117 participants. We calculated Weighted Likelihood Estimates separately for each domain and then estimated a series of principal component analyses with increasing number of factors. The component solution at different levels match models reported in previous studies on the dimensionality of knowledge. We conclude that the dimensionality of declarative knowledge highly depends on the item and person sample. Finally, we discuss different approaches to model gc, give advice on the measurement of gc in general and discuss weaknesses and strengths of mobile assessments.
Jansen, M., Schroeders, U., Lüdtke, O., & Marsh, H. W. (2019). The dimensional structure of students’ self-concept and interest in science depends on course composition. Learning and Instruction, 60, 20–28. https://doi.org/10.1016/j.learninstruc.2018.11.001
Citations per year
FWCI — citations received divided by those expected for works of the same type, subfield and year (2019); 1.0 is the world average. The figure stays provisional until three years after publication.
Percentile — this work's rank within that same comparison group.
Both academic self-concept and interest are considered domain-specific constructs. Previous research has not yet explored how the composition of the courses affects the domain-specificity of these constructs. Using data from a large-scale study in Germany, we compared ninth-grade students who were taught science as an integrated subject with students who were taught biology, chemistry, and physics separately with regard to the dimensional structure of their self-concepts and interests. Whereas the structure of the constructs was six-dimensional in both groups (self-concept and interest factors for biology, chemistry, and physics), the correlations between the domain-specific factors were higher in the integrated group. Furthermore, the pattern of gender differences differed across groups. Whereas male students generally showed higher self-concept and interest in physics and chemistry, a small advantage for male students in biology was only present in integrated science teaching group. We conclude that aspects of the learning environment such as course composition may affect the dimensional structure of motivational constructs.
Olaru, G., Schroeders, U., Ostendorf, F., & Wilhelm, O. (2019). “Grandpa, do you like roller coasters?”: Identifying age-appropriate personality indicators. European Journal of Personality, 33(3), 264–278. https://doi.org/10.1002/per.2185
Citations per year
FWCI — citations received divided by those expected for works of the same type, subfield and year (2019); 1.0 is the world average. The figure stays provisional until three years after publication.
Percentile — this work's rank within that same comparison group.
Personality development research heavily relies on the comparison of scale means across age. This approach implicitly assumes that the scales are strictly measurement invariant across age. We questioned this assumption by examining whether appropriate personality indicators change over the lifespan. Moreover, we identified which types of items (e.g. dispositions, behaviours, and interests) are particularly prone to age effects. We reanalyzed the German Revised NEO Personality Inventory normative sample (N = 11,724) and applied a genetic algorithm to select short scales that yield acceptable model fit and reliability across locally weighted samples ranging from 16 to 66 years of age. We then examined how the item selection changes across age points and item types. Emotion–type items seemed to be interchangeable and generally applicable to people of all ages. Specific interests, attitudes, and social effect items—most prevalent within the domains of Extraversion, Agreeableness, and Openness—seemed to be more prone to measurement variations over age. A large proportion of items were systematically discarded by the item–selection procedure, indicating that, independent of age, many items are problematic measures of the underlying traits. The implications for personality assessment and personality development research are discussed.
2018
Olaru, G., Schroeders, U., Wilhelm, O., & Ostendorf, F. (2018). A confirmatory examination of age-associated personality differences: Deriving age-related measurement invariant solutions using Ant Colony Optimization. Journal of Personality, 86(6), 1037–1049. https://doi.org/10.1111/jopy.12373
Citations per year
FWCI — citations received divided by those expected for works of the same type, subfield and year (2018); 1.0 is the world average. The figure stays provisional until three years after publication.
Percentile — this work's rank within that same comparison group.
Objective: The goal of this study was to examine age-associated personality differences using a measurement-invariant representation of the higher-order structure of the Five-Factor Model. Method: We reanalyzed the German NEO-PI-R norm sample (N = 11,724) and applied ant colony optimization in a multigroup confirmatory factor analysis setting in order to select three items per first-order factor that would optimize model fit and measurement invariance across 18 age groups ranging from 16 to 65 years of age. Results: Ant colony optimization substantially improved absolute and relative model fit under measurement invariance constraints. However, the results showed that even when selecting items, measurement invariance across a large age span could not be guaranteed. Strong measurement invariance for Extraversion and Agreeableness could not be established. The age-associated mean-level differences of the first-order factors of Neuroticism and Conscientiousness supported the maturity hypothesis. The mean levels of the first-order factors of Openness varied substantially from each other across age. Conclusions: Findings on age differences in personality can be particularly distorted in older age groups. Testing for and ensuring measurement invariance with item selection procedures can help solve this problem. The higher-order structure of personality should be accounted for when personality development is examined.
Sewasew, D., Schroeders, U., Schiefer, I. M., Weirich, S., & Artelt, C. (2018). Development of sex differences in math achievement, self-concept, and interest from grade 5 to 7. Contemporary Educational Psychology, 54, 55–65. https://doi.org/10.1016/j.cedpsych.2018.05.003
Citations per year
FWCI — citations received divided by those expected for works of the same type, subfield and year (2018); 1.0 is the world average. The figure stays provisional until three years after publication.
Percentile — this work's rank within that same comparison group.
Sex differences in mathematics achievement have been a controversial topic in educational psychology for a long time. This study sheds light on the developmental aspects of sex differences in math achievement and domain-specific motivational variables such as self-concept and interest. Using a Reciprocal Effects Model (REM), we analyzed 2,342 German fifth to seventh grade students who participated in a large-scale longitudinal study. Math self-concept was validated as a consistent predictor of subsequent achievement and interest for both sexes, supporting the self-enhancement part of the REM using test-scores and teacher-assigned grades. However, math achievement affects subsequent self-concept inconsistently (i.e., the skill development part). Although the bivariate relationships between the constructs were homogeneous across sex and over time, there were large sex differences in the motivational constructs, but not in the achievement measure regardless of achievement measures. The present findings underline the importance of considering both the mean and the covariance structure when describing sex differences in academic achievement. In addition, they also stress the impact of motivational constructs on educational achievement, which also have implications for sex-specific intervention programs in general.
Hartung, J., Doebler, P., Schroeders, U., & Wilhelm, O. (2018). Dedifferentiation and differentiation of intelligence in adults across age and years of education. Intelligence, 69, 37–49. https://doi.org/10.1016/j.intell.2018.04.003
Citations per year
FWCI — citations received divided by those expected for works of the same type, subfield and year (2018); 1.0 is the world average. The figure stays provisional until three years after publication.
Percentile — this work's rank within that same comparison group.
The extent that the structure of cognitive abilities changes across the lifespan or across ability levels is an ongoing debate in intelligence research. The differentiation-dedifferentiation theory states that cognitive abilities differentiate until the beginning of maturity, after which relations increase or dedifferentiate until late adulthood. Spearman’s law of diminishing returns proposes that cognitive abilities are more differentiated at higher ability levels. However, the evidence for ability differentiation and age dedifferentiation, in particular, is mixed. A prerequisite for the evaluation of dedifferentiation processes, expressed as changes in ability factor correlations, is the invariance of intelligence across age. However, a strong interpretation of the dedifferentiation hypothesis states that changes in model parameters, such as factor loadings and residual variances, can also be indicative of dedifferentiation. Traditional statistical tools for testing measurement invariance are not feasible for studying parameter changes over a continuous context variable, such as age. However, a recently developed non-parametric method, Local Structural Equation Modeling (LSEM), closes this gap. LSEM is a powerful and versatile method for studying structural changes; it is an improvement over competing methods, because it avoids artificial categorization of a moderator that is continuous in nature and also renounces the determination of a priori parameter functions. Using cross-sectional data from the standardization sample of the Woodcock-Johnson IV Intelligence Test Battery, we present an application and extension of the LSEM approach to accommodate two moderator variables. Specifically, we studied measurement invariance across two context variables, age and years of education, in models testing the unique moderating effect of each, and in a model examining their combined effects. We found no significant moderating effects of age and no effects of years of education on the relation between fluid and crystallized intelligence. Moreover, an interaction of age and years of education with respect to model parameter change cannot be supported.
Moehring, A., Schroeders, U., & Wilhelm, O. (2018). Knowledge is power for medical assistants: Crystallized and fluid intelligence as predictors of vocational knowledge. Frontiers in Psychology, 9. Article 28. https://doi.org/10.3389/fpsyg.2018.00028
Citations per year
FWCI — citations received divided by those expected for works of the same type, subfield and year (2018); 1.0 is the world average. The figure stays provisional until three years after publication.
Percentile — this work's rank within that same comparison group.
Medical education research has focused almost entirely on the education of future physicians. In comparison, findings on other health-related occupations, such as medical assistants, are scarce. With the current study, we wanted to examine the knowledge-is-power hypothesis in a real life educational setting and add to the sparse literature on medical assistants. Acquisition of vocational knowledge in vocational education and training (VET) was examined for medical assistant students (n = 448). Differences in domain-specific vocational knowledge were predicted by crystallized and fluid intelligence in the course of VET. A multiple matrix design with three year-specific booklets was used for the vocational knowledge tests of the medical assistants. The unique and joint contributions of the predictors were investigated with structural equation modeling. Crystallized intelligence emerged as the strongest predictor of vocational knowledge at every stage of VET, while fluid intelligence only showed weak effects. The present results support the knowledge-is-power hypothesis, even in a broad and more naturalistic setting. This emphasizes the relevance of general knowledge for occupations, such as medical assistants, which are more focused on learning hands-on skills than the acquisition of academic knowledge.
Gnambs, T., Scharl, A., & Schroeders, U. (2018). The structure of the Rosenberg Self-Esteem Scale: A cross-cultural meta-analysis. Zeitschrift für Psychologie, 226(1), 14–29. https://doi.org/10.1027/2151-2604/a000317
Citations per year
FWCI — citations received divided by those expected for works of the same type, subfield and year (2018); 1.0 is the world average. The figure stays provisional until three years after publication.
Percentile — this work's rank within that same comparison group.
The Rosenberg Self-Esteem Scale (RSES; Rosenberg, 1965 ) intends to measure a single dominant factor representing global self-esteem. However, several studies have identified some form of multidimensionality for the RSES. Therefore, we examined the factor structure of the RSES with a fixed-effects meta-analytic structural equation modeling approach including 113 independent samples (N = 140,671). A confirmatory bifactor model with specific factors for positively and negatively worded items and a general self-esteem factor fitted best. However, the general factor captured most of the explained common variance in the RSES, whereas the specific factors accounted for less than 15%. The general factor loadings were invariant across samples from the United States and other highly individualistic countries, but lower for less individualistic countries. Thus, although the RSES essentially represents a unidimensional scale, cross-cultural comparisons might not be justified because the cultural background of the respondents affects the interpretation of the items.
Edossa, A., Schroeders, U., Weinert, S., & Artelt, C. (2018). The development of emotional and behavioral self-regulation and their effects on academic achievement in childhood. International Journal of Behavioral Development, 42(2), 192–202. https://doi.org/10.1177/0165025416687412
Citations per year
FWCI — citations received divided by those expected for works of the same type, subfield and year (2018); 1.0 is the world average. The figure stays provisional until three years after publication.
Percentile — this work's rank within that same comparison group.
Self-regulation is an essential ability of children to cope with various developmental challenges. This study examines the developmental interplay between emotional and behavioral self-regulation during childhood and the relationship with academic achievement using data from the longitudinal Millennium Cohort Study (UK). Using cross-lagged panel analyses, we found that emotional and behavioral self-regulation were separate and stable constructs. In addition, both emotional and behavioral self-regulation had positive cross-lagged effects from ages 3 to 7. At an early developmental stage (ages 3 to 5), emotional regulation affected behavioral regulation more strongly than later developmental stages. However, the difference between the reciprocal effects was small from ages 5 to 7. Moreover, behavioral regulation during the third year of primary education (age 7) had a substantial and positive effect on teachers’ evaluations of educational achievement during the last year of primary school (age 11). In contrast, emotional self-regulation only had a small indirect and positive effect via behavioral self-regulation. The current study suggests the structure of self-regulation was multidimensional and its facets are mutually dependent in the child’s development. In order to gain a complete picture of the development of self-regulation and its effect on educational achievement, the facets emotional and behavioral regulation should both be studied in concert.
2017
Gnambs, T., & Schroeders, U. (2017). Cognitive abilities explain wording effects in the Rosenberg Self-Esteem Scale. Assessment, 27(2), 404–418. https://doi.org/10.1177/1073191117746503
Citations per year
FWCI — citations received divided by those expected for works of the same type, subfield and year (2017); 1.0 is the world average. The figure stays provisional until three years after publication.
Percentile — this work's rank within that same comparison group.
There is consensus that the 10 items of the Rosenberg Self-Esteem Scale (RSES) reflect wording effects resulting from positively and negatively keyed items. The present study examined the effects of cognitive abilities on the factor structure of the RSES with a novel, nonparametric latent variable technique called local structural equation models. In a nationally representative German large-scale assessment including 12,437 students competing measurement models for the RSES were compared: a bifactor model with a common factor and a specific factor for all negatively worded items had an optimal fit. Local structural equation models showed that the unidimensionality of the scale increased with higher levels of reading competence and reasoning, while the proportion of variance attributed to the negatively keyed items declined. Wording effects on the factor structure of the RSES seem to represent a response style artifact associated with cognitive abilities.
Lenhard, W., Schroeders, U., & Lenhard, A. (2017). Equivalence of screen versus print reading comprehension depends on task complexity and proficiency. Discourse Processes, 54(5-6), 427–445. https://doi.org/10.1080/0163853X.2017.1319653
Citations per year
FWCI — citations received divided by those expected for works of the same type, subfield and year (2017); 1.0 is the world average. The figure stays provisional until three years after publication.
Percentile — this work's rank within that same comparison group.
As reading and reading assessment become increasingly implemented on electronic devices, the question arises whether reading on screen is comparable with reading on paper. To examine potential differences, we studied reading processes on different proficiency and complexity levels. Specifically, we used data from the standardization sample of the German reading comprehension test ELFE II (n = 2,807), which assesses reading at word, sentence, and text level with separate speeded subtests. Children from grades 1 to 6 completed either a test version on paper or via computer under time constraints. In general, children in the screen condition worked faster but at the expense of accuracy. This difference was more pronounced for younger children and at the word level. Based on our results, we suggest that remedial education and interventions for younger children using computer-based approaches should likewise foster speed and accuracy in a balanced way.
2016
Schroeders, U., Wilhelm, O., & Olaru, G. (2016). Meta-heuristics in short scale construction: Ant Colony Optimization and Genetic Algorithm. PLOS ONE, 11(11), Article e0167110. https://doi.org/10.1371/journal.pone.0167110
Citations per year
FWCI — citations received divided by those expected for works of the same type, subfield and year (2016); 1.0 is the world average. The figure stays provisional until three years after publication.
Percentile — this work's rank within that same comparison group.
The advent of large-scale assessment, but also the more frequent use of longitudinal and multivariate approaches to measurement in psychological, educational, and sociological research, caused an increased demand for psychometrically sound short scales. Shortening scales economizes on valuable administration time, but might result in inadequate measures because reducing an item set could: a) change the internal structure of the measure, b) result in poorer reliability and measurement precision, c) deliver measures that cannot effectively discriminate between persons on the intended ability spectrum, and d) reduce test-criterion relations. Different approaches to abbreviate measures fare differently with respect to the above-mentioned problems. Therefore, we compare the quality and efficiency of three item selection strategies to derive short scales from an existing long version: a Stepwise COnfirmatory Factor Analytical approach (SCOFA) that maximizes factor loadings and two metaheuristics, specifically an Ant Colony Optimization (ACO) with a tailored user-defined optimization function and a Genetic Algorithm (GA) with an unspecific cost-reduction function. SCOFA compiled short versions were highly reliable, but had poor validity. In contrast, both metaheuristics outperformed SCOFA and produced efficient and psychometrically sound short versions (unidimensional, reliable, sensitive, and valid). We discuss under which circumstances ACO and GA produce equivalent results and provide recommendations for conditions in which it is advisable to use a metaheuristic with an unspecific out-of-the-box optimization function.
Moehring, A., Schroeders, U., Leichtmann, B., & Wilhelm, O. (2016). Ecological momentary assessment of digital literacy: Influence of fluid and crystallized intelligence, domain-specific knowledge, and computer usage. Intelligence, 59, 170–180. https://doi.org/10.1016/j.intell.2016.10.003
Citations per year
FWCI — citations received divided by those expected for works of the same type, subfield and year (2016); 1.0 is the world average. The figure stays provisional until three years after publication.
Percentile — this work's rank within that same comparison group.
The ability to comprehend new information is closely related to the successful acquisition of new knowledge. With the ubiquitous availability of the Internet, the procurement of infor-mation online constitutes a key aspect in education, work, and our leisure time. In order to investigate individual differences in digital literacy, testtakers were presented with health-related comprehension problems with task-specific time restrictions. Instead of reading a given text, they were instructed to search the Internet for the information required to answer the questions. We investigated the relationship between this newly developed test and fluid and crystallized intelligence, while controlling for computer usage, in two studies with adults (n1 = 120) and vocational high school students (n2 = 171). Structural equation modeling was used to investigate the amount of unique variance explained by each predictor. In both studies, about 85% of the variance in the digital literacy factor could be explained by reasoning and knowledge while computer usage did not add to the variance explained. In Study 2, prior health-related knowledge was included as a predictor instead of general knowledge. While the influence of fluid intelligence remained significant, prior knowledge strongly influenced digital literacy (β = .81). Together both predictor variables explained digital literacy exhaustively. These findings are in line with the view that knowledge is a major determinant of higher-level cognition. Further implications about the influence of the restrictiveness of the testing environment are discussed.
Schroeders, U., Schipolowski, S., Zettler, I., Golle, J., & Wilhelm, O. (2016). Do the smart get smarter? Development of fluid and crystallized intelligence in 3rd grade. Intelligence, 59, 84–95. https://doi.org/10.1016/j.intell.2016.08.003
Citations per year
FWCI — citations received divided by those expected for works of the same type, subfield and year (2016); 1.0 is the world average. The figure stays provisional until three years after publication.
Percentile — this work's rank within that same comparison group.
There are conflicting theoretical assumptions about the development of general cognitive abilities in childhood: On the one hand, a higher initial level of abilities has been suggested to facilitate ability improvement, for example, prior knowledge fosters the acquisition of new knowledge (Matthew effect). On the other hand, it has been argued that school education with its special focus on promoting less able students results in a compensation effect. A third hypothesis is that the development of cognitive abilities is—as an outcome of the opposing effects—overall independent of the initial state. In this study, 1,102 elementary students in 3rd Grade worked on two versions of the Berlin Test of Fluid and Crystallized Intelligence at two time points with an interval of five months. Beside the question of how initial state and growth are related (Matthew vs. compensation effect), we considered performance gains in fluid intelligence (gf) and crystallized intelligence (gc) as well as cross-lagged effects in a bivariate latent change score model. Both for gf and gc there was a strong compensation effect. Mean change was more pronounced in gf than in gc. We considered student characteristics (interest and self-concept), family background (socio-economic status, parental education) and classroom characteristics (teaching styles) in a series of prediction models to explain these changes in gf and gc. Although several predictors were included, only few had a significant contribution. Several methodological and content-related reasons are discussed to account for the unexpectedly negligible effects found for most of the covariates.
Schroeders, U., Wilhelm, O., & Olaru, G. (2016). The influence of item sampling on sex differences in knowledge tests. Intelligence, 58, 22–32. https://doi.org/10.1016/j.intell.2016.06.003
Citations per year
FWCI — citations received divided by those expected for works of the same type, subfield and year (2016); 1.0 is the world average. The figure stays provisional until three years after publication.
Percentile — this work's rank within that same comparison group.
Few topics in psychology have generated as much controversy as sex differences in intelligence. For fluid intelligence, researchers emphasize the high overlap between the ability distributions of males and females, whereas research on sex differences in declarative knowledge often uncovers a male advantage. However, on the level of knowledge domains, a more nuanced picture emerged: while females perform better in health-related topics (e.g., aging, medicine), males outperform females in domains of natural sciences (e.g., engineering, physics). In this paper we show that sex differences vary substantially depending on item sampling. Analyses were based on a sample of n = 3306 German high-school students (Grades 9 and 10) who worked on the 64 declarative knowledge items of the Berlin Test of Fluid and Crystallized Intelligence (BEFKI) assessing knowledge within three broad content domains (science, humanities, social studies). Using two strategies of item sampling - stepwise confirmatory factor analysis and ant colony optimization algorithm - we deliberately manipulate sex differences in multi-group structural equation models. Results show that sex differences considerably vary depending on the indicators drawn from the item pool. Furthermore, ant colony optimization outperforms the simple stepwise selection strategy since it can optimize several criteria simultaneously (model fit, reliability, and preset sex differences). Taken together, studies reporting sex differences in declarative knowledge fail to acknowledge item sampling issues. On a more general stance, handling item sampling hinges on profound considerations of the content of measures.
Jansen, M., Lüdtke, O., & Schroeders, U. (2016). Evidence for a positive relation between interest and achievement: Examining between-person and within-person variation in five domains. Contemporary Educational Psychology, 46, 116–127. https://doi.org/10.1016/j.cedpsych.2016.05.004
Citations per year
FWCI — citations received divided by those expected for works of the same type, subfield and year (2016); 1.0 is the world average. The figure stays provisional until three years after publication.
Percentile — this work's rank within that same comparison group.
We examined the incremental effect of academic interest on achievement beyond general cognitive ability and students’ background characteristics in five domains (math, German, biology, chemistry, and physics). We analyzed a nationally representative German dataset of 39,192 ninth-grade students and found a unique effect of interest over and above the other predictors across the five domains, both for class grades and standardized test scores. The effect was present between persons (in a given domain, students with higher interest showed higher achievement) and within persons (the same student showed a higher achievement in domains she/he was more interested in). The effects were stronger for grades than test scores and stronger in math than in other domains. The results emphasize the positive relation between interest and academic achievement in different domains. Furthermore, they expand the literature by emphasizing the role of the achievement measure and the domain as moderators of the interest–achievement relation and by showing that interest can predict both inter- and intraindividual variation in achievement.
Wolf, K., Schroeders, U., & Kriegbaum, K. (2016). Metaanalyse zur Wirksamkeit einer Förderung der phonologischen Bewusstheit in der deutschen Sprache. Zeitschrift für Pädagogische Psychologie, 30, 9–33. https://doi.org/10.1024/1010-0652/a000165
Citations per year
FWCI — citations received divided by those expected for works of the same type, subfield and year (2016); 1.0 is the world average. The figure stays provisional until three years after publication.
Percentile — this work's rank within that same comparison group.
Zusammenfassung. Internationale Metaanalysen belegen bedeutsame Effekte einer Förderung der phonologischen Bewusstheit auf den Schriftspracherwerb. In der vorliegenden Metaanalyse mit 27 Primärstudien wurden die Effekte phonologischer Fördermaßnahmen auf die frühen schriftsprachlichen Kompetenzen im Deutschen quantifiziert. Die Analysen belegten generell höhere Effekte für Fördermaßnahmen, die vor der Einschulung stattfanden. Eine vorschulische Förderung besaß zwar keine signifikanten Effekte auf die Dekodierfähigkeit; jedoch konnten geringe Effekte auf die Rechtschreibkompetenz nachgewiesen werden, was auf die unterschiedliche Konsistenz der Graphem-Phonem- und Phonem-Graphem-Zuordnungen zurückzuführen sein könnte. Analysen zum Einfluss von Moderatoren ergaben, dass Kinder mit schwachen und guten Ausgangskompetenzen gleichermaßen von der Förderung profitieren, und dass die Kombination mit einem Buchstaben-Laut-Training keine inkrementellen Effekte im Vergleich zu rein phonologischen Fördermaßnahmen hat. Insgesamt fielen die metaanalytischen Trainingseffekte der deutschsprachigen Förderprogramme deutlich niedriger aus als in den internationalen Metaanalysen. Für diesen Befund werden sprachliche und methodische Erklärungen gegeben.
Akukwe, B., & Schroeders, U. (2016). Socio-economic, cultural, social, and cognitive aspects of family background and the biology competency of ninth-graders in Germany. Learning and Individual Differences, 45, 185–192. https://doi.org/10.1016/j.lindif.2015.12.009
Citations per year
FWCI — citations received divided by those expected for works of the same type, subfield and year (2016); 1.0 is the world average. The figure stays provisional until three years after publication.
Percentile — this work's rank within that same comparison group.
Students’ academic achievement is related to different family background factors such as socio-economic, cultural, social, and cognitive factors. Research on family background has mainly focused on socio-economic factors, often neglecting the significance of providing a cognitively activating home environment. As a supplement to a largescale study assessing the competency of ninth-graders in Biology in Germany, 543 parents provided information on their socio-economic, cultural, and social background and worked on a domain-specific competency test. By means of hierarchical regression analyses, we established the separate and combined effects of the different background variables on students’ performance. Including all predictors simultaneously in a prediction model, only two—the number of books in home (β = .11) and the biology competency of parents (β = .26)—significantly predicted differences in their children’s competency in biology. Based on the results, we advocate a more comprehensive assessment of family background.
2015
Schroeders, U., Schipolowski, S., & Böhme, K. (2015). Typical intellectual engagement and achievement in math and the sciences in secondary education. Learning and Individual Differences, 43, 31–38. https://doi.org/10.1016/j.lindif.2015.08.030
Citations per year
FWCI — citations received divided by those expected for works of the same type, subfield and year (2015); 1.0 is the world average. The figure stays provisional until three years after publication.
Percentile — this work's rank within that same comparison group.
Typical Intellectual Engagement (TIE) is considered a key trait in explaining individual differences in educational achievement in advanced academic or professional settings. Research in secondary education, however, has focused on cognitive and conative factors rather than personality. In the present large-scale study, we investigated the relation between TIE and achievement tests in math and science in Grade 9. A three-dimensional model (reading, contemplation, intellectual curiosity) provided high theoretical plausibility and satisfactory model fit. We quantified the predictive power of TIE with hierarchical regression models. After controlling for gender, migration background, and socioeconomic status, TIE contributed substantially to the explanation of math and science achievement. However, this effect almost disappeared after fluid intelligence and interest were added into the model. Thus, we found only limited support for the significance of TIE on educational achievement, at least for subjects more strongly relying on fluid abilities such as math and science.
Jansen, M., Schroeders, U., Lüdtke, O., & Marsh, H. (2015). Contrast and assimilation effects of dimensional comparisons in five subjects: An extension of the I/E model. Journal of Educational Psychology, 107(4), 1086–1101. https://doi.org/10.1037/edu0000021
Citations per year
FWCI — citations received divided by those expected for works of the same type, subfield and year (2015); 1.0 is the world average. The figure stays provisional until three years after publication.
Percentile — this work's rank within that same comparison group.
Students evaluate their achievement in a specific domain in relation to their achievement in other domains and form their self-concepts accordingly. These comparison processes have been termed dimensional comparisons and shown to be an important source of academic self-concepts in addition to social and temporal comparisons. Research on the internal/external frame of reference model (I/E model) has frequently found negative effects of students’ achievement on their academic self-concept between different scholastic domains (mathematics and the language of instruction) that are interpreted as contrast effects of dimensional comparisons. There is mixed evidence with regard to whether negative contrast effects or positive assimilation effects occur when students compare their achievement in domains that are more similar. In this study, we extended the original I/E model with 3 science domains (biology, chemistry, and physics). Using structural equation modeling, we analyzed the domain-specific selfconcepts, grades, and test scores of a representative sample of 9th-grade students in Germany (N 20,050) across 5 domains. Mathematics, physics, and chemistry showed contrast effects to German, whereas small assimilation effects were found between mathematics, physics, and chemistry. This effect pattern was present for both grades and test scores. Achievement in mathematics and the language of instruction affected self-concepts in the sciences, whereas achievement in the sciences had no effect on self-concepts in other subjects. The results support the hypotheses derived from dimensional comparison theory that both contrast and assimilation effects can result from dimensional comparisons and that the 3 science subjects are affected differentially by these comparisons.
Schroeders, U., Schipolowski, S., & Wilhelm, O. (2015). Age-related changes in the mean and covariance structure of fluid and crystallized intelligence in childhood and adolescence. Intelligence, 48, 15–29. https://doi.org/10.1016/j.intell.2014.10.006
Citations per year
FWCI — citations received divided by those expected for works of the same type, subfield and year (2015); 1.0 is the world average. The figure stays provisional until three years after publication.
Percentile — this work's rank within that same comparison group.
Evidence on age-related differentiation in the structure of cognitive abilities in childhood and adolescence is still inconclusive. Previous studies often focused on the interrelations or the g-saturation of broad ability constructs, neglecting abilities on lower strata. In contrast, we investigated differentiation in the internal structure of fluid intelligence/gf (with verbal, numeric, and figural reasoning) and crystallized intelligence/gc (with knowledge in the natural sciences, humanities, and social studies). To better understand the development of reasoning and knowledge during secondary education, we analyzed data from 11,756 students attending Grades 5 to 12. Changes in both the mean structure and the covariance structure were estimated with locally-weighted structural equation models that allow handling age as a continuous context variable. To substantiate a potential influence of school tracking (i.e., different learning environments), analyses were additionally conducted separated by school track (academic vs. nonacademic). Mean changes in gf and gc were approximately linear in the total sample, with a steeper slope for the latter. There was little indication of age-related differentiation for the different reasoning facets and knowledge domains. The results suggest that the relatively homogeneous scholastic learning environment in secondary education prevents the development of more pronounced ability or knowledge profiles.
Jansen, M., Scherer, R., & Schroeders, U. (2015). Students’ self-concept and self-efficacy in the sciences: Differential relations to antecedents and educational outcomes. Contemporary Educational Psychology, 41, 13–24. https://doi.org/10.1016/j.cedpsych.2014.11.002
Citations per year
FWCI — citations received divided by those expected for works of the same type, subfield and year (2015); 1.0 is the world average. The figure stays provisional until three years after publication.
Percentile — this work's rank within that same comparison group.
Self-concept and self-efficacy are two of the most important motivational predictors of educational outcomes. As most research has studied these constructs separately, little is known about their differential relations to peer ability, opportunities-to-learn in classrooms, and educational outcomes. We investigated these relations by applying (multilevel) structural equation modeling to the German PISA 2006 data set. We found a correlation of ρ = .57 between self-concept and self-efficacy in science, advocating distinguishable constructs. Furthermore, science self-concept was better predicted by the average peer achievement (Big-Fish-Little-Pond Effect), whereas science self-efficacy was more strongly affected by inquirybased learning opportunities. There were also differences in the predictive potential for educational outcomes: Self-concept was a better predictor of future-oriented motivation to aspire a career in the sciences, whereas self-efficacy was a better predictor of current ability. The study at hand provides strong evidence for the related but distinct nature of the two constructs and extends existing research on students’ competence beliefs toward social comparisons and opportunities-to-learn. Further implications for the relevance of inquiry-based classroom activities and for the assessment of competence beliefs are discussed.
2014
Schroeders, U., Robitzsch, A., & Schipolowski, S. (2014). A comparison of different psychometric approaches to modeling testlet structures: An example with c-tests. Journal of Educational Measurement, 51(4), 400–418. https://doi.org/10.1111/jedm.12054
Citations per year
FWCI — citations received divided by those expected for works of the same type, subfield and year (2014); 1.0 is the world average. The figure stays provisional until three years after publication.
Percentile — this work's rank within that same comparison group.
C‐tests are a specific variant of cloze tests that are considered time‐efficient, valid indicators of general language proficiency. They are commonly analyzed with models of item response theory assuming local item independence. In this article we estimated local interdependencies for 12 C‐tests and compared the changes in item difficulties, reliability estimates, and person parameter estimates for different modeling approaches: (a) Rasch, (b) testlet, (c) partial credit, and (d) copula models. The results are complemented with findings of a simulation study in which sample size, number of testlets, and strength of residual correlations between items were systematically manipulated. Results are discussed with regard to the pivotal question whether residual dependencies between items are an artifact or part of the construct.
Schipolowski, S., Schroeders, U., & Wilhelm, O. (2014). Pitfalls and challenges in constructing short forms of cognitive ability measures. Journal of Individual Differences, 35(4), 190–200. https://doi.org/10.1027/1614-0001/a000134
Citations per year
FWCI — citations received divided by those expected for works of the same type, subfield and year (2014); 1.0 is the world average. The figure stays provisional until three years after publication.
Percentile — this work's rank within that same comparison group.
Especially in survey research and large-scale assessment there is a growing interest in short scales for the cost-efficient measurement of psychological constructs. However, only relatively few standardized short forms are available for the measurement of cognitive abilities. In this article we point out pitfalls and challenges typically encountered in the construction of cognitive short forms. First we discuss item selection strategies, the analysis of binary response data, the problem of floor and ceiling effects, and issues related to measurement precision and validity. We subsequently illustrate these challenges and how to deal with them based on an empirical example, the development of short forms for the measurement of crystallized intelligence. Scale shortening had only small effects on associations with covariates. Even for an ultra-short six-item scale, a unidimensional measurement model showed excellent fit and yielded acceptable reliability. However, measurement precision on the individual level was very low and the short forms were more likely to produce skewed score distributions in ability-restricted subpopulations. We conclude that short scales may serve as proxies for cognitive abilities in typical research settings, but their use for decisions on the individual level should be discouraged in most cases.
Schipolowski, S., Wilhelm, O., & Schroeders, U. (2014). On the nature of crystallized intelligence: The relationship between verbal ability and factual knowledge. Intelligence, 46, 156–168. https://doi.org/10.1016/j.intell.2014.05.014
Citations per year
FWCI — citations received divided by those expected for works of the same type, subfield and year (2014); 1.0 is the world average. The figure stays provisional until three years after publication.
Percentile — this work's rank within that same comparison group.
While crystallized intelligence (gc) is recognized in many contemporary intelligence frameworks, there is no consensus as to the nature and contents of the construct. Originally conceptualized as capturing acquired skills and declarative knowledge in different content domains, more recent definitions and typical indicators focus on verbal ability. We investigated the relationship between verbal ability and declarative knowledge under consideration of individual differences in fluid intelligence in a large-scale assessment study with 6,701 adolescents. Structural equation modeling was used to examine the factorial distinctness of verbal ability and declarative knowledge with three analytical strategies: (i) Estimating correlations between latent variables, (ii) estimating the amount of unique variance in each factor after accounting for differences in the other ability constructs, and (iii) investigating associations with covariates including school achievement, students’ characteristics, and psychological traits. The correlation between latent variables representing verbal ability, measured with items from six language domains, and knowledge in 16 content domains was very high (ρ = .91), but significantly different from unity. About 17% of the variance in the knowledge factor was independent of individual differences in verbal ability and fluid intelligence. Associations with covariates revealed unique correlational patterns for each ability construct. The findings suggest that verbal ability and knowledge are closely related, but empirically distinguishable facets of crystallized intelligence. The discussion focuses on the construct validity of verbal tests for the measurement of gc and the interpretation of the common factor of a broad knowledge assessment as a causal variable.
Jansen, M., Schroeders, U., Lüdtke, O., & Pant, H. A. (2014). Interdisziplinäre Beschulung und die Struktur des akademischen Selbstkonzepts in den naturwissenschaftlichen Fächern. Zeitschrift für Pädagogische Psychologie, 28, 43–49. https://doi.org/10.1024/1010-0652/a000120
Citations per year
FWCI — citations received divided by those expected for works of the same type, subfield and year (2014); 1.0 is the world average. The figure stays provisional until three years after publication.
Percentile — this work's rank within that same comparison group.
Das akademische Selbstkonzept wird als multidimensionales, fachspezifisches Konstrukt aufgefasst. Welchen Einfluss die schulische Fächerstruktur auf die Struktur des Selbstkonzepts hat, ist jedoch noch unklar. In diesem Beitrag wird untersucht, ob bei Schülern, die interdisziplinären Naturwissenschaftsunterricht erhalten, ein weniger stark ausdifferenziertes naturwissenschaftliches Selbstkonzept vorliegt als bei Schülern, die Unterricht in Biologie, Chemie und Physik getrennt erhalten. Dazu wurden 326 Schüler, die in der gesamten Sekundarstufe im Fächerverbund beschult wurden, mit einer Vergleichsgruppe von 4361 fächergetrennt beschulten Schülern verglichen. Mit konfirmatorischen (Mehrgruppen-) Faktorenanalysen wird gezeigt, dass Schüler beider Gruppen in ihrem Selbstkonzept zwischen den drei Fächern differenzieren, in der interdisziplinär beschulten Gruppe die Zusammenhänge zwischen den Selbstkonzeptfaktoren aber deutlich höher sind. Interdisziplinärer Unterricht scheint also mit einer Vereinheitlichung der Selbstkonzeptstruktur auf Schülerseite einherzugehen. Implikationen für die Selbstkonzepttheorie werden diskutiert.
Jansen, M., Schroeders, U., & Lüdtke, O. (2014). Academic self-concept in science: Multidimensionality, relations to achievement measures, and gender differences. Learning and Individual Differences, 30, 11–21. https://doi.org/10.1016/j.lindif.2013.12.003
Citations per year
FWCI — citations received divided by those expected for works of the same type, subfield and year (2014); 1.0 is the world average. The figure stays provisional until three years after publication.
Percentile — this work's rank within that same comparison group.
Students’ academic self-concept is a good predictor of academic achievement and a desirable educational outcome per se. In this study, we take a closer look at the nature of the academic self-concept in the natural sciences by examining its dimensional structure, its relation to achievement, and gender differences. We analyzed data from self-concept measures, grades and standardized achievement tests of 6036 German 10th graders across three science subjects – biology, chemistry, and physics – using structural equation modeling. Results indicate that (a) a 3-dimensional, subject-specific measurement model of the self-concept in science is preferable to a 1-dimensional model, (b) the relations between the self-concept and achievement are substantial and subjectspecific when grades are used as achievement indicators, and (c) female students possess a lower self-concept in chemistry and physics even after controlling for achievement measures. Therefore, we recommend conceptualizing the self-concept in science as a multidimensional, subject-specific construct both in educational research and in science classes.
< 2013
Schroeders, U., Bucholtz, N., Formazin, M., & Wilhelm, O. (2013). Modality specificity of comprehension abilities in the sciences. European Journal of Psychological Assessment, 29(1), 3–11. https://doi.org/10.1027/1015-5759/a000114
Citations per year
FWCI — citations received divided by those expected for works of the same type, subfield and year (2013); 1.0 is the world average. The figure stays provisional until three years after publication.
Percentile — this work's rank within that same comparison group.
The measurement of science achievement is often unnecessarily restricted to the presentation of reading comprehension items that are sometimes enriched with graphs, tables, and figures. In a newly developed viewing comprehension task, participants watched short videos covering different science topics and were subsequently asked several multiple-choice comprehension questions. Research questions were whether viewing comprehension (1) can be measured adequately, (2) is perfectly collinear with reading comprehension, and (3) can be regarded as a linear function of reasoning and acquired knowledge. High-school students (N = 216) worked on a paper-based reading comprehension task, a viewing comprehension task delivered on handheld devices, a sciences knowledge test, and three fluid intelligence measures. The data show that, first, the new viewing comprehension test worked psychometrically fine; second, performance in both comprehension tasks was essentially perfectly collinear; third, fluid intelligence and domain-specific knowledge fully accounted for the ability to comprehend texts and videos. We conclude that neither test medium (paper-pencil versus handheld device) nor test modality (reading versus viewing) are decisive for comprehension ability in the natural sciences. Fluid intelligence and, even more strongly, domain-specific knowledge turned out to be exhaustive predictors of comprehension performance.
Schipolowski, S., Wilhelm, O., Schroeders, U., Kovaleva, A., Kemper, C. J., & Rammstedt, B. (2013). BEFKI GC-K. Eine Kurzskala zur Messung kristalliner Intelligenz [BEFKI GC-K. A Short Scale for the Measurement of Crystallized Intelligence]. Methoden, Daten, Analysen, 7, 153–181. https://doi.org/10.12758/mda.2013.010
Citations per year
FWCI — citations received divided by those expected for works of the same type, subfield and year (2013); 1.0 is the world average. The figure stays provisional until three years after publication.
Percentile — this work's rank within that same comparison group.
Crystallized intelligence (gc) is a well-established cognitive ability factor that has been conceptualized as reflecting influences of learning, education, and acculturation. In this article, we describe the development of a short knowledge scale for the measurement of gc in five minutes administration time using declarative knowledge items from the sciences, the humanities, and civics. Based on a large item pool we compiled a 32-item knowledge test that was subsequently presented to a nationally representative sample of 1,134 German adults. In the next step, this data were used to derive a short 12-item knowledge scale. A unidimensional measurement model had satisfactory model fit and showed high reliability of the latent factor. There were no substantial floor or ceiling effects in the adult German population. Similar to the full scale, the short scale correlated highly positively with education (ISCED-97) and socio-economic status (ISEI) and was meaningfully related to self-reported knowledge and the Big Five personality traits. Therefore, the short knowledge scale allows for an efficient and valid measurement of crystallized intelligence in survey research.
Ehlert, A., Schroeders, U., & Fritz-Stratmann, A. (2012). Kritik am Diskrepanzkriterium in der Diagnostik von Legasthenie und Dyskalkulie. Lernen und Lernstörungen, 1, 169–184. https://doi.org/10.1024/2235-0977/a000018
Citations per year
FWCI — citations received divided by those expected for works of the same type, subfield and year (2012); 1.0 is the world average. The figure stays provisional until three years after publication.
Percentile — this work's rank within that same comparison group.
Zusammenfassung: Laut den beiden großen internationalen Klassifikationssystemen psychischer Störungen, ICD-10 und DSM-IV-TR, muss zur Diagnose einer Dyskalkulie oder Legasthenie u. a. eine Diskrepanz zwischen der auf Grund allgemeiner kognitiver Leistung zu erwartenden und der tatsächlichen Rechenleistung vorliegen. Im ersten Teil dieses Beitrags soll die national und international geäußerte Kritik zu inhaltlichen und methodischen Schwächen aufgegriffen und gebündelt diskutiert werden. Die Annahme einer Diskrepanz impliziert auch, dass sich rechenschwache Kinder, deren Rechenleistung zusätzlich das Diskrepanzkriterium erfüllt, von anderen rechenschwachen Kindern abgrenzen lassen. Um diese Vorstellung zu entkräften, soll in einem zweiten Schritt empirisch überprüft werden, inwiefern eine auf der Diskrepanz zwischen erwarteter und gemessener Rechenleistung basierende Gruppenzuordnung künstlich oder gerechtfertigt ist. Dazu wurden die rechenschwachen Kinder einer geschichteten Stichprobe bestehend aus 458 Kindern zwei Gruppen zugewiesen: Die Kinder der ersten Gruppe zeigen schwache Rechenleistungen in einem curricular orientierten Rechentest und schneiden auch schlechter ab als auf Grund eines Intelligenztests zu erwarten wäre. Die rechenschwachen Kinder der zweiten Gruppe erfüllen das Diskrepanzkriterium hingegen nicht. Die beiden Gruppen werden hinsichtlich ihres Verständnisses mathematischer Konzepte mit Hilfe eines kriterienorientierten Tests, der auf dem mathematischen Kompetenzstufenmodell von Fritz und Ricken (2008*) beruht ( Fritz, Ricken, Balzer, Leutner & Willmes, 2012 ), miteinander verglichen. Es zeigt sich, dass die Kinder der beiden rechenschwachen Gruppen über dieselben mathematischen Konzepte verfügen. Deswegen legen sowohl die wissenschaftlich-theoretische Diskussion als auch die Ergebnisse der empirischen Studie den Schluss nahe, dass die Verwendung eines Diskrepanzkriteriums in der Diagnostik einer Teilleistungsschwäche fraglich ist und durch eine kriterienorientierte Diagnostik, die auf basale numerische Grundfähigkeiten fokussiert, ersetzt werden sollte.
Formazin, M., Schroeders, U., Köller, O., Wilhelm, O., & Westmeyer, H. (2011). Studierendenauswahl im Fach Psychologie. Testentwicklung und Validitätsbefunde. Psychologische Rundschau, 62, 221–236. https://doi.org/10.1026/0033-3042/a000093
Citations per year
FWCI — citations received divided by those expected for works of the same type, subfield and year (2011); 1.0 is the world average. The figure stays provisional until three years after publication.
Percentile — this work's rank within that same comparison group.
Zusammenfassung. Die internationale Forschung im Bereich der Hochschulzulassung zeigt eindrücklich, dass Leistungstests gute Prädiktoren der späteren Studienleistung sind und inkrementelle Validität über Schulnoten hinaus aufweisen. An deutschen Hochschulen ist der Einsatz standardisierter Leistungstests jedoch nach wie vor die Ausnahme. In der vorliegenden Arbeit schildern wir die Entwicklung und Validierung einer Testbatterie für die Zulassung von Psychologiestudierenden an deutschen Hochschulen. Im Rahmen der Testung von 1187 Bewerberinnen und Bewerbern für die Vergabe von 60 Studienplätzen prüfen wir mit Strukturgleichungsmodellen und Regressionsanalysen die prädiktive und inkrementelle Validität der neuen Testbatterie. Neben einem allgemeinen Faktor für das schlussfolgernde Denken kann auf der Prädiktorseite ein zweiter, geschachtelter Faktor für relevantes Vorwissen etabliert werden. Beide latenten Faktoren tragen nennenswert zur Vorhersage der Studienleistungen bei. Die Ergebnisse unterstützen nachdrücklich die Forderung, bei der Zulassung zu Studiengängen mit hohem Bewerberandrang Leistungstests einzusetzen. Neben schlussfolgerndem Denken verdient das relevante Vorwissen besondere Beachtung.
Schroeders, U., & Wilhelm, O. (2011). Equivalence of reading and listening comprehension across test media. Educational and Psychological Measurement, 71(5), 849–869. https://doi.org/10.1177/0013164410391468
Citations per year
FWCI — citations received divided by those expected for works of the same type, subfield and year (2011); 1.0 is the world average. The figure stays provisional until three years after publication.
Percentile — this work's rank within that same comparison group.
Whether an ability test delivered on either paper or computer provides the same information is an important question in applied psychometrics. Besides the validity, it is also the fairness of a measure that is at stake if the test medium affects performance. This study provides a comprehensive review of existing equivalence research in the field of reading and listening comprehension in English as a foreign language and specifies factors that are likely to have an impact on equivalence. Taking into account these factors, comprehension measures were developed and tested with N = 442 high school students. Using multigroup confirmatory factor analysis, it is shown that reading and listening comprehension both were measurement invariant across test media. Nevertheless, it is argued that equivalence of data gathered on paper and computer depends on the specific measure or construct, the participants or the recruitment mechanisms, and the software and hardware realizations. Therefore, equivalence research is required for specific instantiations unless generalizable knowledge about factors affecting equivalence is available. Multigroup confirmatory factor analysis is an appropriate and effective tool for the assessment of the comparability of test scores across test media.
Schroeders, U., & Wilhelm, O. (2011). Computer usage questionnaire: Structure, correlates, and gender differences. Computers in Human Behavior, 27(2), 899–904. https://doi.org/10.1016/j.chb.2010.11.015
Citations per year
FWCI — citations received divided by those expected for works of the same type, subfield and year (2011); 1.0 is the world average. The figure stays provisional until three years after publication.
Percentile — this work's rank within that same comparison group.
Computer usage, computer experience, computer familiarity, and computer anxiety are often discussed as constructs potentially compromising computer-based ability assessment. After presenting and discussing these constructs and associated measures we introduce a brief new questionnaire assessing computer usage. The self-report measure consists of 18 questions asking for the frequency of different computer activities and software usage. Participants were N = 976 high school students who completed the questionnaire and several covariates. Based on theoretical considerations and data driven adjustments a model with a general computer usage factor and three nested content factors (Office, Internet, and Games) is established for a subsample (n = 379) and cross-validated with the remaining sample (n = 597). Weak measurement invariance across gender groups could be established using multi-group confirmatory factor analysis. Differential relations between the questionnaire factors and self-report scales of computer usage, self-concept, and evaluation are reported separately for females and males. It is concluded that computer usage is distinct from other behavior oriented measurement approaches and that it shows a diverging, gender-specific pattern of relations with fluid and crystallized intelligence.
Schroeders, U., & Wilhelm, O. (2010). Testing reasoning ability with handheld computers, notebooks, and paper and pencil. European Journal of Psychological Assessment, 26(4), 284–292. https://doi.org/10.1027/1015-5759/a000038
Citations per year
FWCI — citations received divided by those expected for works of the same type, subfield and year (2010); 1.0 is the world average. The figure stays provisional until three years after publication.
Percentile — this work's rank within that same comparison group.
Electronic devices can be used to enhance or improve cognitive ability testing. We compared three reasoning-ability measures delivered on handheld computers, notebooks, and paper-and-pencil to test whether or not the same underlying abilities were measured irrespective of the test medium. Rational item-generative principles were used to generate parallel item samples for a verbal, a numerical, and a figural reasoning test, respectively. All participants, 157 high school students, completed the three measures on each test medium. Competing measurement models were tested with confirmatory factor analyses. Results show that 2 test-medium factors for tests administrated via notebooks and handheld computers, respectively, had small to negligible loadings, and that the correlation between these factors was not substantial. Overall, test medium was not a critical source of individual differences. Perceptual and motor skills are discussed as potential causes for test-medium factors.
Schroeders, U., Wilhelm, O., & Bucholtz, N. (2010). Reading, listening, and viewing comprehension in English as a foreign language: One or more constructs? Intelligence, 38(6), 562–573. https://doi.org/10.1016/j.intell.2010.09.003
Citations per year
FWCI — citations received divided by those expected for works of the same type, subfield and year (2010); 1.0 is the world average. The figure stays provisional until three years after publication.
Percentile — this work's rank within that same comparison group.
Receptive foreign language proficiency is usually measured with reading and listening comprehension tasks. A novel approach to assess such proficiencies – viewing comprehension is based on the presentation of short instructional videos followed by one or more comprehension questions concerning the preceding video stimulus. In order to evaluate a newly developed viewing comprehension test 485 German high school students completed reading, listening, and viewing comprehension tests, all measuring the receptive proficiency in English as a foreign language. Fluid and crystallized intelligence were measured as predictors of performance. Relative to traditional comprehension tasks, the viewing comprehension task has similar psychometric qualities. The three comprehension tests are very highly but not perfectly correlated with each other. Relations with fluid and crystallized intelligence show systematic differences between the three comprehension tasks. The high overlap between foreign language comprehension measures and between crystallized intelligence and language comprehension ability can be taken as support for a uni-dimensional interpretation. Implications for the assessment of language proficiency are discussed.
Kunina, O., Wilhelm, O., Formazin, M., Jonkmann, K., & Schroeders, U. (2007). Extended criteria and predictors in college admission: Exploring the structure of study success and investigating the validity of domain knowledge. Psychology Science, 49, 88–114.
The utility of aptitude tests and intelligence measures in the prediction of the success in college is one of the empirically best supported results in ability research. However, the structure of the criterion “study success” has not been appropriately investigated so far. Moreover, it remains unclear which aspect of intelligence fluid intelligence or crystallized intelligence – has the major impact on the prediction. In three studies we have investigated the dimensionality of the criterion achievements as well as the relative contributions of competing ability predictors. In the first study, the dimensionality of college grades was explored in a sample of 629 alumni. A measurement model with two correlated latent factors distinguishing undergraduate college grades on the one hand from graduate college grades on the other hand had the best fit to the data. In the second study, a group of 179 graduate students completed a Psychology knowledge test and provided available college grades in undergraduate studies. A model separating a general latent factor for Psychology knowledge from a nested method factor for college grades, and a second nested factor for “experimental orientation” had the best fit to the data. In the third study the predictive power of domain specific knowledge tests in Mathematics, English, and Biology was investigated. A sample of 387 undergraduate students in this prospective study additionally completed a compilation of fluid intelligence tests. The results of this study indicate as expected that: a) ability measures are incrementally predictive over school grades in predicting exam grades; and b) that knowledge tests from relevant domains were incrementally predictive over fluid intelligence. The results of these studies suggest that criteria for college admission tests deserve and warrant more attention, and that domain specific ability indicators can contribute to the predictive validity of established admission tests.
Monographs
Schroeders, U., Schipolowski, S., & Wilhelm, O. (2020). BEFKI 5-7. Berliner Test zur Erfassung fluider und kristalliner Intelligenz (Testform 5-7). Hogrefe.
Schipolowski, S., Wilhelm, O., & Schroeders, U. (2020). BEFKI 11-12+. Berliner Test zur Erfassung fluider und kristalliner Intelligenz (Testform 11-12+). Hogrefe.
Wilhelm, O., Schroeders, U., & Schipolowski, S. (2014). BEFKI 8-10. Berliner Test zur Erfassung fluider und kristalliner Intelligenz (Testform 8-10). Hogrefe.
Pant, H. A., Stanat, P., Schroeders, U., Roppelt, A., Siegle, T., & Pöhlmann, C. (Hrsg.) (2013). IQB-Ländervergleich 2012. Mathematische und naturwissenschaftliche Kompetenzen am Ende der Sekundarstufe I. Waxmann.
Schroeders, U. (2010). Measurement of cognitive abilities using modern technologies: artifacts, equivalence, and new constructs. Dissertationsschrift. Humboldt-Universität zu Berlin.
Schroeders, U., & Schneider, W. (2008). TeDDy-PC. Test zur Diagnose von Dyskalkulie. Hogrefe.
Schroeders, U. (2004). Entwicklung von TEDDY – Ein computergestützter Test zur Diagnose von Dyskalkulie in der 1. Klasse. Unveröffentlichte Diplomarbeit, Julius-Maximilians-Universität Würzburg.
Contributions to edited books
Olaru, G., Robitzsch, A., Hildebrandt, A., & Schroeders, U. (2023). An illustration of Local Structural Equation Modeling for longitudinal data: Examining differences in competence development in secondary schools. In S. Weinert, G. J. Blossfeld, & H.-P. Blossfeld (Eds.), Education, Competence Development and Career Trajectories (pp. 153–176). Springer International Publishing. https://doi.org/10.1007/978-3-031-27007-9_7
Citations per year
FWCI — citations received divided by those expected for works of the same type, subfield and year (2023); 1.0 is the world average. The figure stays provisional until three years after publication.
Percentile — this work's rank within that same comparison group.
In this chapter, we discuss how a combination of longitudinal modeling and local structural equation modeling (LSEM) can be used to study how students’ context influence their growth in educational achievement. LSEM is a nonparametric approach that allows for the moderation of a structural equation model over a continuous variable (e.g., socio-economic status; cultural identity; age). Thus, it does not require the categorization of continuous moderators as applied in multi-group approaches. In contrast to regression-based approaches, it does not impose a particular functional form (e.g., linear) on the mean-level differences and can spot differences in the variance-covariance structure. LSEM can be used to detect nonlinear moderation effects, to examine sources of measurement invariance violations, and to study moderation effects on all parameters in the model. We showcase how LSEM can be implemented with longitudinal of the National Educational Panel Study (NEPS) using the R-package sirt. In more detail, we examine the effect of parental education on math and reading competence in secondary school across three measurement occasions, comparing LSEM to regression based approaches and multi-group confirmatory factor analysis. Results provide further evidence of the strong influence of the educational background of the family. This chapter offers a new approach to study inter-individual differences in educational development.
Wilhelm, O., & Schroeders, U. (2019). Intelligence. In R. J. Sternberg & J. Funke (Eds.), The Psychology of Human Thought (pp. 255–275), Heidelberg University Publishing. https://doi.org/10.17885/heiUP.470
Citations per year
FWCI — citations received divided by those expected for works of the same type, subfield and year (2019); 1.0 is the world average. The figure stays provisional until three years after publication.
Percentile — this work's rank within that same comparison group.
The “Psychology of Human Thought” is an “open access” collection of peer-reviewed chapters from all areas of higher cognitive processes. The book is intended to be used as a textbook in courses on higher process, complex cognition, human thought, and related courses. Chapters include concept acquisition, knowledge representation, inductive and deductive reasoning, problem solving, metacognition, language, expertise, intelligence, creativity, wisdom, development of thought, affect and thought, and sections about history and about methods. The chapters are written by distinguished scholarly experts in their respective fields, coming from such diverse regions as North America, Great Britain, France, Germany, Norway, Israel, and Australia. The level of the chapters is addressed to advanced undergraduates and beginning graduate students.
Schroeders, U. (2018). Ability. In M. H. Bornstein (Ed.), The SAGE Encyclopedia of Lifespan Human Development (pp. 1–5). SAGE Publications, Inc. https://doi.org/10.4135/9781506307633.n8
Citations per year
FWCI — citations received divided by those expected for works of the same type, subfield and year (2018); 1.0 is the world average. The figure stays provisional until three years after publication.
Percentile — this work's rank within that same comparison group.
Schipolowski, S., Wilhelm, O., & Schroeders, U. (2016). Sprachliche Fähigkeiten und Intelligenz. In J. Kilian, B. Brouër, & D. Lüttenberg (Hrsg.), Handbuch Sprache in der Bildung (S. 523–543). Walter de Gruyter. https://doi.org/10.1515/9783110296358-027
Citations per year
FWCI — citations received divided by those expected for works of the same type, subfield and year (2016); 1.0 is the world average. The figure stays provisional until three years after publication.
Percentile — this work's rank within that same comparison group.
Im vorliegenden Beitrag beleuchten wir zum einen die Rolle der Sprache in Intelligenztheorien und -tests, zum anderen sollen die Zusammenhange zwischen Konstrukten aus der Intelligenzforschung und der empirischen Bildungsforschung diskutiert werden. Im ersten Teil des Beitrags gehen wir auf die Rolle sprachlicher Fahigkeiten in etablierten Intelligenztheorien ein. Dabei wird herausgearbeitet, dass alle Strukturtheorien sprachliche Fahigkeiten berucksichtigen, jedoch mit unterschiedlicher Gewichtung. Im zweiten Teil fassen wir empirische Befunde zu den Beziehungen zwischen etablierten Faktoren der Intelligenzforschung und sprachgebundenen Denkleistungen, wie sie in grosen Schulleistungsstudien im Bildungsbereich untersucht werden, zusammen. Im dritten Abschnitt thematisieren wir schlies- lich die Rolle der Sprache bei der Entwicklung, Anwendung und Auswertung von Intelligenztests. Aus dieser Zusammenschau wird deutlich, dass kein psychologisches Testverfahren vollends sprach- und kulturunabhangig ist. Abschliesend gehen wir auf die Messung der kognitiven Leistungsfahigkeit in sprachlich und kulturell heterogenen Populationen ein.
Schipolowski, S., Wilhelm, O., Schroeders, U., Kovaleva, A., Kemper, C. J., & Rammstedt, B. (2014). Eine kurze Skala zur Messung kristalliner Intelligenz. Die Kurzskala gc des Berliner Tests zur Erfassung Fluider und Kristalliner Intelligenz (BEFKI GC-K). GESIS Working Papers 2014 (29). GESIS.
Schipolowski, S., Wilhelm, O., & Schroeders, U. (2013). BEFKI GC-K. Berliner Test zur Erfassung fluider und kristalliner Intelligenz – GC-Kurzskala. In C. J. Kemper, E. Brähler & M. Zenger (Hrsg.), Psychologische und sozialwissenschaftliche Kurzskalen – Standardisierte Erhebungsinstrumente für Wissenschaft und Praxis (S. 30–34). Medizinisch Wissenschaftliche Verlagsgesellschaft.
Siegle, T., Schroeders, U., & Roppelt, A. (2013). Anlage und Durchführung des Ländervergleichs. In H. A. Pant, P. Stanat, U. Schroeders, A. Roppelt, T. Siegle, & C. Pöhlmann (Hrsg.), IQB-Ländervergleich 2012. Mathematische und naturwissenschaftliche Kompetenzen am Ende der Sekundarstufe I (S. 101–121). Waxmann.
Schroeders, U., Hecht, M., Heitmann, P., Jansen, M., Kampa, N., Klebba, N., Lenski, A. E., & Siegle, T. (2013). Der Ländervergleich in den naturwissenschaftlichen Fächern. In H. A. Pant, P. Stanat, U. Schroeders, A. Roppelt, T. Siegle, & C. Pöhlmann (Hrsg.), IQB-Ländervergleich 2012. Mathematische und naturwissenschaftliche Kompetenzen am Ende der Sekundarstufe I (S. 141–158). Waxmann.
Pant, H. A., Stanat, P., Pöhlmann, C., Hecht, M., Jansen, M., Kampa, N., … Ziemke, A. (2013). Der Blick in die Länder. In H. A. Pant, P. Stanat, U. Schroeders, A. Roppelt, T. Siegle, & C. Pöhlmann (Hrsg.), IQB-Ländervergleich 2012. Mathematische und naturwissenschaftliche Kompetenzen am Ende der Sekundarstufe I (S. 159–247). Waxmann.
Schroeders, U., Penk, C., Jansen, M., & Pant, H. A. (2013). Geschlechtsbezogene Disparitäten. In H. A. Pant, P. Stanat, U. Schroeders, A. Roppelt, T. Siegle, & C. Pöhlmann (Hrsg.), IQB-Ländervergleich 2012. Mathematische und naturwissenschaftliche Kompetenzen am Ende der Sekundarstufe I (S. 249–274). Waxmann.
Schroeders, U., Siegle, T., Weirich, S., & Pant, H. A. (2013). Der Einfluss von Kontext- und Schülermerkmalen auf die naturwissenschaftlichen Kompetenzen. In H. A. Pant, P. Stanat, U. Schroeders, A. Roppelt, T. Siegle, & C. Pöhlmann (Hrsg.), IQB-Ländervergleich 2012. Mathematische und naturwissenschaftliche Kompetenzen am Ende der Sekundarstufe I (S. 331–346). Waxmann.
Jansen, M., Schroeders, U., & Stanat, P. (2013). Motivationale Schülermerkmale in Mathematik und den Naturwissenschaften. In H. A. Pant, P. Stanat, U. Schroeders, A. Roppelt, T. Siegle, & C. Pöhlmann (Hrsg.), IQB-Ländervergleich 2012. Mathematische und naturwissenschaftliche Kompetenzen am Ende der Sekundarstufe I (S. 348–365). Waxmann.
Pant, H. A., Stanat, P., Pöhlmann, C., Roppelt, A., Schroeders, U., & Siegle, T. (2013). Der IQB-Ländervergleich 2012: Zusammenfassung und Einordnung der Befunde. In H. A. Pant, P. Stanat, U. Schroeders, A. Roppelt, T. Siegle, & C. Pöhlmann (Hrsg.), IQB-Ländervergleich 2012. Mathematische und naturwissenschaftliche Kompetenzen am Ende der Sekundarstufe I (S. 403–414). Waxmann.
Schroeders, U., Wilhelm, O., & Schipolowski, S. (2010). Internet-based ability testing. In S. D. Gosling, & J. A. Johnson (Eds.), Advanced methods for behavioral research on the Internet (pp. 131–148). American Psychological Association. https://doi.org/10.1037/12076-009
Citations per year
FWCI — citations received divided by those expected for works of the same type, subfield and year (2010); 1.0 is the world average. The figure stays provisional until three years after publication.
Percentile — this work's rank within that same comparison group.
Wilhelm, O., & Schroeders, U. (2010). Bildung in Zeiten neuer Medien aus kognitionspsychologischer Sicht. In J.-D. Gauer, & J. Kraus (Hrsg.), Bildung und Unterricht in Zeiten von Google und Wikipedia (S. 27–45). Konrad-Adenauer-Stiftung.
Schroeders, U. (2009). Testing for equivalence of test data across media. In F. Scheuermann & J. Björnsson (Eds.), The transition to computer-based assessment. Lessons learned from the PISA 2006 computer-based assessment of science (CBAS) and implications for large scale testing (pp. 164–170). JRC Scientific and Technical Report EUR 23679 EN.
Schroeders, U. (2009). Computergestützte Diagnostik und Förderung rechenschwacher Kinder am Beispiel von TeDDy-PC. In Verband Dyslexie Schweiz (Hrsg.), Dyskalkulie – Ansätze zu Diagnostik und Förderung in einer integrativen Schule (pp. 53–60). Verband Dyslexie Schweiz.
Wilhelm, O., & Schroeders, U. (2008). Computerized ability measurement: Some substantive dos and don’ts. In F. Scheuermann & A. G. Pereira (Eds.), Towards a research agenda in computer-based assessment. Challenges and needs for European Educational Measurement (pp. 76–84). JRC Scientific and Technical Report 23306 EN.
Formazin, M., Wilhelm, O., Schroeders, U., Kunina, O., Hildebrandt, A., & Köller, O. (2008). Validitäts- und Nützlichkeitsüberlegungen zur Studierendenauswahl im Allgemeinen mit Präzisierungen für das Fach Psychologie im Besonderen. In H. Schuler & B. Hell (Hrsg.), Studierendenauswahl und Studienentscheidung (S. 204–214). Hogrefe.
Others
Gewehr, E., Molz, G., & Schroeders, U. (2022). TBS-DTK-Rezension: Wechsler Adult Intelligence Scale - Fourth Edition (WAIS-IV). Psychologische Rundschau, 73(1), 92–94. https://doi.org/10.1026/0033-3042/a000581
Schroeders, U.* & Wilhelm*, O. (2020). Es gibt drei Arten von Lügen: Lügen, verdammte Lügen und Statistiken - Kommentar zu Klauk (2019). Wirtschaftspsychologie, 2, 45–48. [* Shared 1st authorship]
Molz, G., Schulze, R., Schroeders, U., & Wilhelm, O. (2010). TBS-TK Rezension: Wechsler Intelligenztest für Erwachsene WIE. Deutschsprachige Bearbeitung und Adaptation des WAIS-III von David Wechsler. Psychologische Rundschau, 61(4), 229–230. https://doi.org/10.1026/0033-3042/a000042
Citations per year
FWCI — citations received divided by those expected for works of the same type, subfield and year (2010); 1.0 is the world average. The figure stays provisional until three years after publication.
Percentile — this work's rank within that same comparison group.