1 Validity in Psychological Tests: Sources of Evidence
What does it mean to say that a test has evidence of validity? Where do the concepts of validity in psychology come from? How can we seek evidence of validity for our instruments? In this chapter, we will answer these questions.
1.1 Constructs, Behavior and Traits
In physics, we usually have an instrument that physically exists and measures physical properties. For example, an instrument that measures length uses this property (i.e., length) to measure the length of another object. Therefore, there is no need to prove that this property is congruent with the same property of the object being measured.
However, there are some cases where this is not so clear. For example, if we are measuring speed using the Doppler effect (Doppler Effect is a physical wave phenomenon that occurs when there is relative approach or distance between a source of waves and an observer), where the approach/distance of the spectral lines of the galaxy’s lights is the instrument. In this case, we have the problem of the validity of an instrument, as we need to know whether or not it is true that the distance between the spectral lines is related to speed. To do this, we have to prove it empirically. Validity is common in areas of knowledge that use indirect or derived measures. The same thing that happens with the Doppler effect is very common in behavioral and psychological sciences (for example, psychology and education), especially if we are using the concept of construct (for example, happiness, anxiety or attraction).
As argued by Maraun & Peters (2005), in scientific research, concepts are essential tools for identifying and organizing phenomena. Scientific outputs—such as observations, hypotheses, and theories—are expressed through language, which in turn can affect the reproducibility of scientific results (Schmalz et al., 2024). To study these concepts, the researchers in psychology have been pursuing the development of the quantification of the human mind and its related measures (Michell, 2014). Measurement, as a fundamental scientific method, provides a means to determine, with varying degrees of accuracy and precision, the level of an attribute present in the object (or objects) being studied. In psychology and the social sciences, we also investigate measures’ validity (the extent to which they reflect a concept) and reliability (the degree to which they produce consistent results), making explicit the strong reliance of these measurement perspectives on previous theories (Michell, 2001).
This strong reliance on previous theories makes measurement in psychology mostly based on the definition of constructs that explain or describe individual differences (Trendler, 2013). A construct is defined as “a conceptual system that refers to a set of entities—the construct referents—that are regarded as meaningfully related in some ways or for some purpose although they never occur all at once and that are therefore considered only on more abstract levels as a joint entity” (Uher, 2022). This definition can be summarized in three characteristics of a construct (Cronbach & Meehl, 1955): (i) it is not defined by a single observable referent; (ii) it cannot be directly observed; and (iii) its observable referents are not all-inclusive. Because psychological measurement and theory, mostly understood in terms of psychometric practices, are usually formalized concerning unobservables, they depend on a series of qualitative assumptions that are not directly testable (Franco et al., 2022). And because these assumptions are not testable to some degree, it has been argued that actual scientific measurement in psychology is an empirical impossibility (Trendler, 2013).
Of course, we have many ways to measure constructs, a common way is through questionnaires, where people respond each item on a scale of 1 (strongly agree) to 5 (strongly disagree), for example. Let’s say we’re going to measure self-efficacy in the workplace. We developed the items based on the definition of self-efficacy and then what? How can we know what our test results mean? Is self-efficacy a single phenomenon or can it be divided into different aspects? This is the role of seeking validity.
The need for valid measures seems obvious enough, given that to test theories that relate theoretical constructs (e.g., construct A influences construct B for individuals drawn from population P under conditions C), it is necessary to have valid measures of these constructs. Thus, even successful and replicable tests of a theory may be false if the measures lack construct validity; that is, they do not measure what researchers assume they are measuring (Schimmack, 2021).
1.2 Psychometrics at the Crossroads: Confronting Psychology’s Measurement Crisis
The replicability of scientific claims is a core principle of scientific progress (Lakatos & Musgrave, 1970). In recent years, the field of psychology has become increasingly aware of a significant challenge—what is now recognized as a replication crisis (Schimmack, 2021). Since 2012, it has been evident that many published psychological findings fail to replicate when subjected to rigorous testing (Schimmack, 2021). Due to concerns like these, along with other factors, several scientists have suggested that questionable research practices (such as p-hacking and HARKing) contribute significantly to the rise in “false positive” findings, thereby diminishing replication rates (Nosek et al., 2012). In response, advocates for science reform have promoted new “open science” practices designed to expose and mitigate these questionable practices (Munafò et al., 2017).
However, the replication crisis is not the sole issue confronting psychological science. A result can simultaneously be reproducible, robust, replicable, and yet still invalid since they do not ensure that the interventions were effective as intended, that the measures accurately captured the desired outcomes, or that the interpretations align with the evidence generated (Nosek et al., 2022). Thus, robust measurement, alongside adequate study design and statistical modeling proficiency, has a central role in the scientific process. If the measures used in a study lack validity or reliability, the conclusions drawn from it may be questionable (Flake et al., 2017). Schimmack (2021) argues that psychology is facing a validity crisis, particularly concerning the validity of psychological instruments.
Cronbach & Meehl (1955) expressed early skepticism regarding the construct validity of many psychological measures. They emphasized the absence of adequate criteria for validating tests designed to measure psychological constructs, noting that many such tests remain unvalidated or are supported only by a network of rationalizations that do not constitute proper validation. Despite advancements in psychological measurement, the concerns raised by Schimmack (2021) suggest that many of Cronbach and Meehl’s criticisms remain unresolved. This issue is also emphasized, for example, by Flake et al. (2017), who reviewed contemporary practices and found that reliability is often the sole criterion used to claim construct validity. However, reliability is a necessary but not a sufficient condition for validity.
Due to these limitations, it has been argued that many studies fail to provide robust evidence for construct validity (Flake & Fried, 2020). Even when validity is claimed, the degree to which a measure is valid remains unclear. This ongoing issue is further highlighted by the continued use of measures developed decades ago, without quantitative validation to confirm their enduring validity. For instance, Rosenberg’s self-esteem scale (Rosenberg, 1965) remains the most widely used measure of self-esteem (Bosson et al., 2000; Schimmack, 2021), yet its construct validity has never been fully quantified, leaving it uncertain whether it outperforms other self-esteem measures (Schimmack, 2021).
While there is broad recognition of the limitations inherent in traditional psychometric practices (Kane, 2017; Maul, 2017; Schmittmann et al., 2013), there is no consensus on the best approach to address the validation crisis. Some authors question whether conventional psychometric practices can support the strong quantitative measurement claims that psychologists often make, emphasizing conceptual and causal limitations in standard approaches (Maul, 2017; Trendler, 2009). Others focus on changing the models used to represent psychological phenomena—for example, network models treat indicators as potentially interacting components of a system rather than assuming that their associations necessarily arise from a single common latent cause (Schmittmann et al., 2013).
Despite the possible paths to follow, it is reasonable to infer that to advance psychological science, a concerted effort is needed to develop a strong research agenda focused on measurement and its caveats. Such an agenda should prioritize the continuous evaluation of the validity of psychological instruments, thereby ensuring that the measures used in research genuinely reflect the constructs they are intended to assess. In addition, careful attention must be given to the underlying assumptions that inform these instruments. Without rigorous scrutiny of these psychometric foundations, the validity of psychological assessments remains questionable, and any conclusions drawn from such measures could be fundamentally flawed. Therefore, advancing the field not only requires robust validation practices but also a deep and ongoing concern with the psychometric assumptions that underpin these practices.
1.3 A Brief Note on the History of Validity
A. 1900–1950: The hegemony of content validity
At that time, personality theories were on the rise. Most theories (such as psychoanalytic, gestalt, and phenomenology) generally had little empirical reasoning. In this context, personality trait tests were considered valid to the extent that the content of the test corresponded to the content of the theoretically defined traits.
B. 1950–1970: Prevalence of criterion validity
Behaviorism was very influential for Psychology and, of course, for Psychometrics. The tests were composed with a sample of behaviors that were expected to predict other behaviors or future behaviors. These tests were valid if they accurately predicted behavior in the future (or in another time), becoming the new path of validity (called criterion validity). It didn’t matter why the test predicted the behavior, as long as they predicted it, and that was enough for its validity. As we can imagine, there was a shift from theoretical thinking to a focus on statistics. Rather than constructing a test to measure a construct, items were selected from a pool of items that appeared to refer to what they wanted to measure, essentially using statistical analysis to solve their problems.
C. 1970-Today: The rise of construct validity
After an article by Cronbach and Meehl in 1955 on a trinitarian model of validity (content, criterion, and construct), there was a change in the way of thinking about validity. The theory was back in play due to factors such as:
The need to develop a theory of personality and intelligence on an empirical basis, using factor analysis.
Studies of cognitive processes.
Studies of information processes.
Dissatisfaction with the results of using the test in education and work situations.
The impact of Item Response Theory.
Cronbach & Meehl (1955) note that construct validation is necessary
whenever a test is to be interpreted as a measure of some attribute or quality which is not “operationally defined (p. 282).
This definition makes clear that there are other types of validity (e.g., criterion validity) and that not all measures require construct validity. However, studies of psychological theories that relate constructs require valid measures of these constructs to test psychological theories. Thus, construct validity is the relationship between variation in observed scores on a measure (e.g., scores on a Likert scale) and a latent variable that reflects corresponding variation in a theoretical construct (e.g., Extraversion; i.e., people who feel more energized by social interactions).
However, the problem of construct validity can be illustrated with the development of IQ tests (Schimmack, 2021). IQ scores can have predictive validity (e.g., graduate school performance) without making any claims about the construct being measured (IQ tests measure whatever they measure, and what they measure predicts important outcomes). However, IQ tests are often treated as measures of intelligence. For IQ tests to be valid measures of intelligence, it is necessary to define the construct of intelligence and demonstrate that observed IQ scores are related to unobserved variation in intelligence. Thus, construct validation requires clear definitions of constructs that are independent of the measures being validated. Without a clear definition of constructs, the meaning of a measure essentially reverts to “whatever the measure is measuring”, as in the old adage “Intelligence is whatever IQ tests are measuring” (Schimmack, 2021).
1.4 What Is Validity?
A useful starting point is to stop asking whether a test itself is valid. Contemporary validity theory asks whether a particular interpretation and use of test scores is sufficiently supported by theory and empirical evidence (American Educational Research Association et al., 2014; Baptista & Villemor-Amaral, 2019). A test can therefore have strong support for one interpretation, population, or purpose and much weaker support for another. Validity is not a permanent certification attached to an instrument; it is an argument that must be supported, challenged, and updated as evidence accumulates.
This perspective also changes how we talk about the familiar “types of validity.” Content, response processes, internal structure, relations with other variables, and consequences are better understood as sources of validity evidence, not as independent kinds of validity that a test either possesses or lacks. The sources complement one another. A convincing factor structure, for example, cannot compensate for items that respondents systematically misunderstand, just as clearly written items cannot by themselves demonstrate that scores relate to external variables in theoretically expected ways.
The goal of validation is not to collect one result from each category and declare a test “valid.” The goal is to build a coherent argument showing that the proposed interpretation and use of scores remain plausible when examined from several different angles.
1.5 Sources of Validity Evidence
1.5.1 Evidence Based on Test Content
Evidence based on test content asks whether the items adequately represent the construct domain that the score is intended to reflect. This requires more than checking whether items merely sound related to the construct. Researchers should define the construct and its boundaries, identify the relevant facets or behaviors, and examine whether the item set is sufficiently representative without being contaminated by irrelevant content.
Common sources of content evidence include theoretical and empirical literature, test blueprints or specification tables, judgments from subject-matter experts, and feedback from members of the target population. Quantitative summaries such as agreement percentages, content-validity indices, or kappa coefficients can help organize judgments, but they do not replace the conceptual reasoning that links each item to the intended construct.
Bastos et al. (2022) provides a useful example. In developing the Self-Perception of Prejudice and Discrimination Scale, the authors first defined self-perceived prejudice and discrimination, reviewed existing measures, developed items for different social groups, and submitted the items to expert evaluation. The final selection of items was therefore informed by an explicit construct definition and judgments about item relevance, rather than by statistical performance alone.
1.5.2 Evidence Based on Response Processes
Evidence based on response processes concerns whether the psychological and behavioral processes used to answer an item are compatible with the processes assumed by the intended score interpretation. Two people can choose the same response option for very different reasons; conversely, respondents may understand an item in a way that was never intended by its authors. Closed-ended response data alone usually cannot reveal why a particular option was selected.
Response-process studies therefore investigate how respondents comprehend the item, retrieve relevant information, form a judgment, and map that judgment onto the available response options. Methods can include cognitive interviewing, think-aloud protocols, verbal or written probing, observation, and other process data. For example, Noble et al. (2014) examined fifth-grade English-language learners responding to science test items. Interview data showed that linguistic features sometimes led students to interpretations different from those intended by the test developers, producing incorrect responses even when students demonstrated relevant science knowledge. This is precisely the kind of problem that a factor analysis of item scores would be unlikely to diagnose.
1.5.2.1 The Response-Process-Evaluation Method
Wolf et al. (2025) offers a particularly useful framework for collecting this kind of evidence. Wolf and colleagues argue that response-process evidence is often neglected in favor of quantitative analyses, despite the fact that favorable reliability coefficients, factor models, or correlations cannot establish that respondents interpreted items as intended. Their response-process-evaluation (RPE) method turns item pretesting into an iterative and documented validation procedure (Wolf et al., 2025).
The RPE workflow can be summarized as follows:
- Specify the intended meaning. Before evaluating an item, researchers document its intended interpretation, intended use, and target population.
- Probe the response process. Participants answer the item and then respond to targeted probes. Depending on the item, these may examine comprehension, paraphrasing, category selection or rejection, specific interpretations, experiential relevance, and additional feedback.
- Evaluate responses holistically. Multiple trained coders examine the complete set of probes rather than relying on a single answer and classify whether the item was understood as intended.
- Revise and retest. Misinterpretations are treated as information about the item, not as mistakes made by participants. Problematic items or probes are revised and tested with new respondents.
- Evaluate the final version. Wolf and colleagues tentatively recommend evaluating the final version of an item with about 20 respondents and considering at least 80% understood or likely understood as a practical benchmark, while noting that larger samples or stricter thresholds may be appropriate for some uses (Wolf et al., 2025).
- Document the evidence. The final item-validation report records the item wording, intended interpretation and use, target population, proportion understood as intended, examples of interpretations, and common misinterpretations (Wolf et al., 2025).
The value of the iterative approach is visible in the paper’s compassion-item example. The item was revised repeatedly after respondents revealed interpretations that differed from the researchers’ intentions. By the sixth and final iteration, 19 of 20 participants (95%) understood the item as intended; the authors note that an earlier version would have left roughly one quarter of respondents misunderstanding the item (Wolf et al., 2025). This example illustrates an important principle: response-process evidence can change the item itself, rather than merely provide another statistic after data collection is complete.
The RPE method is not a replacement for cognitive interviewing in every context. It trades some of the flexibility of live follow-up questions for greater standardization and easier administration to larger samples. Wolf and colleagues also emphasize that multiple probes are often needed because a single probe may not provide enough information to reconstruct how a respondent understood and answered an item (Wolf et al., 2025).
1.5.3 Evidence Based on Internal Structure
Evidence based on internal structure asks whether relationships among items and components of the test are consistent with the structure implied by the construct definition and the proposed scoring model. If a scale is interpreted as unidimensional, for example, its internal structure should support treating the items as indicators of a common dimension. If a test proposes several distinct but related subscales, the observed structure should be compatible with those distinctions.
Exploratory factor analysis (EFA), confirmatory factor analysis (CFA), and item response theory (IRT) models are common tools for investigating internal structure. Depending on the theory and test design, researchers may also examine local dependence, factor correlations, item discrimination, dimensionality, measurement invariance, or other features. Importantly, good model fit is not equivalent to validity: it answers a structural question about the responses but does not establish that the construct was adequately represented or that respondents interpreted the items as intended.
Selau et al. (2020) illustrates this source of evidence with the Adaptive Functioning Scale for Intellectual Disability (EFA-DI), developed for children aged 7 to 15. The authors investigated whether the observed item responses supported the proposed structure of adaptive functioning using CFA and IRT, alongside evidence involving external variables.
1.5.4 Evidence Based on Relations With Other Variables
Evidence based on relations with other variables asks whether test scores relate to external measures in ways predicted by theory. The important point is not simply whether a correlation is statistically significant, but whether its direction and magnitude are compatible with the nomological network surrounding the construct.
Common patterns include:
- Convergent evidence: scores should relate relatively strongly to credible measures of the same or closely related constructs.
- Discriminant evidence: scores should show weaker relationships with measures of theoretically distinct constructs.
- Criterion-related evidence: scores should relate to relevant behaviors, outcomes, classifications, or criteria. When the criterion is measured at approximately the same time, this is often described as concurrent evidence; when scores are used to anticipate a later outcome, it is often described as predictive evidence.
- Nomological evidence: a broader pattern of associations should correspond to the theoretical network in which the construct is embedded. This can include both positive and negative associations and, importantly, theoretically expected near-zero relationships.
These expectations should be specified from theory rather than decided after inspecting the correlations. For example, Beymer et al. (2022) examined a measure of college students’ perceptions of cost and tested whether its relationships with expectancy and value variables followed theoretically expected patterns.
1.5.5 Evidence Concerning the Consequences of Testing
Evidence concerning the consequences of testing examines what happens when scores are interpreted and used in practice. This includes intended benefits as well as unintended outcomes for individuals, groups, institutions, or decision systems. The central question is not simply whether a test has good or bad social consequences. An undesirable outcome does not automatically invalidate a score interpretation, and a desirable outcome does not establish validity.
Consequences become especially relevant to a validity argument when they reveal problems such as construct underrepresentation, construct-irrelevant variance, systematic barriers for particular groups, or decisions that extend beyond the interpretation and use for which evidence was established (American Educational Research Association et al., 2014). For high-stakes applications, researchers should therefore ask not only whether scores are psychometrically defensible, but also whether the decision rule built on those scores is justified for the intended population and purpose.
This distinction is useful because it separates two questions that are often conflated: Does the score support the intended interpretation? and Is the proposed use of that score appropriate and sufficiently supported? A responsible validation program considers both.
1.6 Concluding Remarks
Validity is not a property that a test acquires once and retains indefinitely. Rather, validation is an ongoing process of building and evaluating an argument about what test scores mean and whether their proposed uses are justified. This distinction is particularly important in psychology, where the attributes of interest are often theoretical constructs that cannot be observed directly. A statistically well-behaved instrument may still provide a poor representation of the construct it was designed to measure.
For this reason, no single analysis can establish validity. Evidence based on test content, response processes, internal structure, relations with other variables, and the consequences of testing addresses different questions about the proposed interpretation of scores. These sources should be considered complementary rather than treated as independent boxes to be checked. Factor analysis, reliability estimates, or correlations with external variables may provide important information, but favorable quantitative results cannot compensate for poorly defined constructs or items that respondents systematically interpret differently from what researchers intended (Flake et al., 2017; Wolf et al., 2025).
Perhaps the most important implication is that measurement should be treated as part of the substantive theory rather than as a preliminary technical step that precedes the “real” research. Psychological theories are tested through observations generated by measurement procedures; therefore, uncertainty about what those observations represent necessarily becomes uncertainty about the theories built from them. The concerns raised by Cronbach & Meehl (1955) more than half a century ago consequently remain relevant: construct validation requires continually confronting our interpretations with empirical evidence rather than relying on convention, statistical convenience, or the historical popularity of an instrument.
Throughout the remainder of this book, this principle will serve as a recurring theme. The psychometric methods introduced in the following chapters—classical test theory, latent-variable models, factor analysis, item response models, measurement invariance, and other approaches—should not be viewed merely as techniques for producing coefficients or achieving acceptable model fit. They are tools for asking increasingly precise questions about measurement. Their value ultimately depends on how effectively they help us understand what our scores represent, why they behave as they do, and whether the conclusions we draw from them are scientifically defensible.