References
American Educational Research Association, American Psychological
Association, & National Council on Measurement in Education. (2014).
Standards for educational and psychological testing. American
Educational Research Association.
Ark, L. A. van der. (2007). Mokken scale analysis in r. Journal of
Statistical Software, 20(11), 1–19. https://doi.org/10.18637/jss.v020.i11
Ark, L. A. van der. (2012). New developments in mokken scale analysis in
r. Journal of Statistical Software, 48(5), 1–27. https://doi.org/10.18637/jss.v048.i05
Baptista, M. N., & Villemor-Amaral, A. E. de. (2019). Compêndio
de avaliação psicológica. Editora Vozes.
Bartholomew, D. J. (2004). Measuring intelligence: Facts and
fallacies. Cambridge University Press.
Bartlett, M. S. (1954). A note on the multiplying factors for various
chi square approximations. Journal of the Royal Statistical Society:
Series B (Methodological), 16(2), 296–298. https://doi.org/10.1111/j.2517-6161.1954.tb00174.x
Bastos, R. V. S., Novaes, F. C., & Natividade, J. C. (2022).
Self-perception of prejudice and discrimination scale: Evidence of
validity and other psychometric properties. Trends in
Psychology, 1–19. https://doi.org/10.1007/s43076-022-00190-7
Baumgartner, H., & Steenkamp, J. (2001). Response styles in
marketing research: A cross national investigation. Journal of
Marketing Research, 38(2), 143–156. https://doi.org/10.1509/jmkr.38.2.143.18840
Beauducel, A., & Herzberg, P. Y. (2006). On the performance of
maximum likelihood versus means and variance adjusted weighted least
squares estimation in CFA. Structural Equation Modeling: A
Multidisciplinary Journal, 13(2), 186–203. https://doi.org/10.1207/s15328007sem1302_2
Bentler, P. M. (1990). Comparative fit indexes in structural models.
Psychological Bulletin, 107(2), 238–246. https://doi.org/10.1037/0033-2909.107.2.238
Beymer, P. N., Ferland, M., & Flake, J. K. (2022). Validity evidence
for a short scale of college students’ perceptions of cost. Current
Psychology, 41(11), 7937–7956. https://doi.org/10.1007/s12144-020-01218-w
Billiet, J. B., & Davidov, E. (2008). Testing the stability of an
acquiescence style factor behind two interrelated substantive variables
in a panel design. Sociological Methods Research,
36(4), 542–562. https://doi.org/10.1177/0049124107313901
Billiet, J. B., & McClendon, M. J. (2000). Modeling acquiescence in
measurement models for two balanced sets of items. Structural
Equation Modeling, 7, 608–628. https://doi.org/10.1207/S15328007SEM0704_5
Bollen, K. A. (1989). Structural equations with latent
variables. Wiley.
Bollen, K. A., & Ting, K. (1993). Confirmatory tetrad analysis.
Sociological Methodology, 23, 147–175. https://doi.org/10.2307/271009
Bonifay, W., & Cai, L. (2017). On the complexity of item response
theory models. Multivariate Behavioral Research,
52(4), 465–484. https://doi.org/10.1080/00273171.2017.1309262
Bork, R. van, Wijsen, L. D., & Rhemtulla, M. (2017). Toward a causal
interpretation of the common factor model. Disputatio,
9(47), 581–601. https://doi.org/10.1515/disp-2017-0019
Borsboom, D. (2005). Measuring the mind: Conceptual issues in
contemporary psychometrics. Cambridge University Press.
Borsboom, D. (2006). The attack of the psychometricians.
Psychometrika, 71(3), 425–440. https://doi.org/10.1007/s11336-006-1447-6
Borsboom, D., & Cramer, A. O. J. (2013). Network analysis: An
integrative approach to the structure of psychopathology. Annual
Review of Clinical Psychology, 9, 91–121. https://doi.org/10.1146/annurev-clinpsy-050212-185608
Borsboom, D., & Mellenbergh, G. J. (2004). Why psychometrics is not
pathological: A comment on michell. Theory & Psychology,
14(1), 105–120. https://doi.org/10.1177/0959354304040200
Borsboom, D., Mellenbergh, G. J., & Heerden, J. van. (2004). The
concept of validity. Psychological Review, 111(4),
1061–1071. https://doi.org/10.1037/0033-295X.111.4.1061
Borsboom, D., & Zand Scholten, A. (2008). The rasch model and
conjoint measurement theory from the perspective of psychometrics.
Theory & Psychology, 18(1), 111–117. https://doi.org/10.1177/0959354307086925
Bosson, J. K., Swann, Jr., William B., & Pennebaker, J. W. (2000).
Stalking the perfect measure of implicit self-esteem: The blind men and
the elephant revisited? Journal of Personality and Social
Psychology, 79(4), 631–643. https://doi.org/10.1037/0022-3514.79.4.631
Brennan, R. L. (2011). Generalizability theory and classical test
theory. Applied Measurement in Education, 24(1), 1–21.
https://doi.org/10.1080/08957347.2011.532417
Brentano, F. (1874). Psychology from an empirical standpoint.
Humanities Press.
Bringmann, L. F., & Eronen, M. I. (2016). Heating up the measurement
debate: What psychologists can learn from the history of physics.
Theory & Psychology, 26(1), 27–43. https://doi.org/10.1177/0959354315617253
Brown, T. A. (2015). Confirmatory factor analysis for applied
research (2nd ed.). The Guilford Press.
Brown, W. (1910). Some experimental results in the correlation of mental
abilities. British Journal of Psychology, 3(3),
296–322. https://doi.org/10.1111/j.2044-8295.1910.tb00207.x
Cambré, B., Welkenhuysen-Gybels, J., & Billiet, J. (2002). Is it
content or style? An evaluation of two competitive measurement models
applied to a balanced set of ethnocentrism items. International
Journal of Comparative Sociology, 43, 1–20. https://doi.org/10.1177/002071520204300101
Cattell, J. M. (1890). Mental tests and measurements. Mind,
15(59), 373–381. https://doi.org/10.1093/mind/os-XV.59.373
Cattell, R. B. (1966). The scree test for the number of factors.
Multivariate Behavioral Research, 1(2), 245–276. https://doi.org/10.1207/s15327906mbr0102_10
Chang, L. (1995). Connotatively consistent and reversed connotatively
inconsistent items are not fully equivalent: Generalizability study.
Educational and Psychological Measurement, 55,
991–997. https://doi.org/10.1177/0013164495055006007
Chen, C., Shin-ying, L., & Stevenson, H. W. (1995). Response style
and cross-cultural comparisons of rating scales among east asian and
north american students. Psychological Science, 6,
170–175. https://doi.org/10.1111/j.1467-9280.1995.tb00327.x
Chen, J., & Chen, Z. (2008). Extended bayesian information criteria
for model selection with large model spaces. Biometrika,
95(3), 759–771. https://doi.org/10.1093/biomet/asn034
Cheung, G. W., & Lau, R. S. (2012). A direct comparison approach for
testing measurement invariance. Organizational Research
Methods, 15(2), 167–198. https://doi.org/10.1177/1094428111421987
Christensen, A. P., Garrido, L. E., & Golino, H. (2023). Unique
variable analysis: A network psychometrics method to detect local
dependence. Multivariate Behavioral Research, 58(6),
1165–1182. https://doi.org/10.1080/00273171.2023.2194606
Christensen, A. P., & Golino, H. (2021). Estimating the stability of
psychological dimensions via bootstrap exploratory graph analysis: A
monte carlo simulation and tutorial. Psych, 3(3),
479–500. https://doi.org/10.3390/psych3030032
Christensen, A. P., Golino, H., & Silvia, P. J. (2020). A
psychometric network perspective on the validity and validation of
personality trait questionnaires. European Journal of
Personality, 34(6), 1095–1108. https://doi.org/10.1002/per.2265
Cliff, N. (1992). Abstract measurement theory and the revolution that
never happened. Psychological Science, 3(3), 186–190.
https://doi.org/10.1111/j.1467-9280.1992.tb00024.x
Connelly, B. S., & Chang, L. (2016). A meta‐analytic multitrait
multirater separation of substance and style in social desirability
scales. Journal of Personality, 84(3), 319–334. https://doi.org/10.1111/jopy.12161
Cortina, J. M. (1993). What is coefficient alpha? An examination of
theory and applications. Journal of Applied Psychology,
78(1), 98–104. https://doi.org/10.1037/0021-9010.78.1.98
Crocker, L., & Algina, J. (1986). Introduction to classical and
modern test theory. Holt, Rinehart; Winston.
Cronbach, L. J. (1942). Studies of acquiescence as a factor in the
true-false test. Journal of Educational Psychology,
33, 401–415. https://doi.org/10.1037/h0054677
Cronbach, L. J. (1951). Coefficient alpha and the internal structure of
tests. Psychometrika, 16(3), 297–334. https://doi.org/10.1007/BF02310555
Cronbach, L. J., & Meehl, P. E. (1955). Construct validity in
psychological tests. Psychological Bulletin, 52(4),
281–302. https://doi.org/10.1037/h0040957
Danner, D., Aichholzer, J., & Rammstedt, B. (2015). Acquiescence in
personality questionnaires: Relevance, domain specificity, and
stability. Journal of Research in Personality, 57,
119–130. https://doi.org/10.1016/j.jrp.2015.05.004
Degobi, E. B., Franco, V. R., & Student, S. R. (2026). A model
fit approach to test the adequacy of interval-level measurement in
psychometric models. https://osf.io/jufy5/
Degobi, E. B., & Valentini, F. (2023). Simulations for two
theoretically sound controls for social desirability: MIMIC and
forced-choice [Master's thesis]. Universidade São Francisco.
DeVellis, R. F., & Thorpe, C. T. (2021). Scale development:
Theory and applications (5th ed.). SAGE.
Domingue, B. W. (2014). Evaluating the equal-interval hypothesis with
test score scales. Psychometrika, 79(1), 1–19. https://doi.org/10.1007/s11336-013-9342-4
Dunn, T. J., Baguley, T., & Brunsden, V. (2014). From alpha to
omega: A practical solution to the pervasive problem of internal
consistency estimation. British Journal of Psychology,
105(3), 399–412. https://doi.org/10.1111/bjop.12046
Edwards, A. L. (1953). The relationship between the judged desirability
of a trait and the probability that the trait will be endorsed.
Journal of Applied Psychology, 37(2), 90. https://doi.org/10.1037/h0058073
Edwards, J. R., & Bagozzi, R. P. (2000). On the nature and direction
of relationships between constructs and measures. Psychological
Methods, 5(2), 155–174. https://doi.org/10.1037/1082-989X.5.2.155
Epskamp, S. (2022). semPlot: Path diagrams and visual analysis of
various SEM packages’ output. CRAN. https://CRAN.R-project.org/package=semPlot
Epskamp, S., Maris, G., Waldorp, L. J., & Borsboom, D. (2018).
Network psychometrics. In The wiley handbook of psychometric
testing: A multidisciplinary reference on survey, scale and test
development (pp. 953–986). Wiley. https://doi.org/10.1002/9781118489772.ch30
Fabrigar, L. R., Wegener, D. T., MacCallum, R. C., & Strahan, E. J.
(1999). Evaluating the use of exploratory factor analysis in
psychological research. Psychological Methods, 4(3),
272–299. https://doi.org/10.1037/1082-989X.4.3.272
Fechner, G. T. (1860). Elemente der psychophysik. Breitkopf und
Härtel.
Ferrando, P. J., Condon, L., & Chico, E. (2004). The convergent
validity of acquiescence: An empirical study relating balanced scales
and separate acquiescence scales. Personality and Individual
Differences, 37(7), 1331–1340. https://doi.org/10.1016/j.paid.2004.01.003
Ferrando, P. J., Lorenzo-Seva, U., & Chico, E. (2003). Unrestricted
factor analytic procedures for assessing acquiescent responding in
balanced, theoretically unidimensional personality scales.
Multivariate Behavioral Research, 38(2), 353–374. https://doi.org/10.1207/S15327906MBR3803_04
Ferrando, P. J., Lorenzo-Seva, U., & Chico, E. (2009). A general
factor-analytic procedure for assessing response bias in questionnaire
measures. Structural Equation Modeling: A Multidisciplinary
Journal, 16(2), 364–381. https://doi.org/10.1080/10705510902751374
Fischer, G. H. (1995). Some neglected problems in IRT.
Psychometrika, 60(4), 459–487. https://doi.org/10.1007/BF02294324
Flake, J. K., & Fried, E. I. (2020). Measurement schmeasurement:
Questionable measurement practices and how to avoid them. Advances
in Methods and Practices in Psychological Science, 3(4),
456–465. https://doi.org/10.1177/2515245920952393
Flake, J. K., Pek, J., & Hehman, E. (2017). Construct validation in
social and personality research: Current practice and recommendations.
Social Psychological and Personality Science, 8(4),
370–378. https://doi.org/10.1177/1948550617693063
Franco, V. R. (2021). É possível identificar o nível de medida de
variáveis latentes? Avaliação Psicológica, 20(2), a–d.
https://doi.org/10.15689/ap.2021.2002.ed
Franco, V. R., Bastos, R. V. S., & Jiménez, M. (2023). Tetrad
fit index for factor analysis models. Paper presented at Virtual
MathPsych/ICCM 2023. https://mathpsych.org/presentation/1297
Franco, V. R., Laros, J. A., & Bastos, R. V. S. (2022). Theoretical
and practical foundations of mokken scale analysis in psychology.
Paidéia (Ribeirão Preto), 32, e3223. https://doi.org/10.1590/1982-4327e3223
Friborg, O., Martinussen, M., & Rosenvinge, J. H. (2006).
Likert-based vs. Semantic differential-based scorings of positive
psychological constructs: A psychometric comparison of two versions of a
scale measuring resilience. Personality and Individual
Differences, 40(5), 873–884. https://doi.org/10.1016/j.paid.2005.08.015
Friedman, J., Hastie, T., & Tibshirani, R. (2008). Sparse inverse
covariance estimation with the graphical lasso. Biostatistics,
9(3), 432–441. https://doi.org/10.1093/biostatistics/kxm045
Galilei, G. (1864). Il saggiatore. G. Barbèra.
Galton, F. (1879). Psychometric experiments. Brain,
2(2), 149–162. https://doi.org/10.1093/brain/2.2.149
Gibson, W. A. (1959). Three multivariate models: Factor analysis, latent
structure analysis, and latent profile analysis. Psychometrika,
24(3), 229–252. https://doi.org/10.1007/BF02289845
Glymour, C., & Scheines, R. (1986). Causal modeling with the TETRAD
program. Synthese, 68, 37–63.
Glymour, C., Scheines, R., & Spirtes, P. (2014). Discovering
causal structure: Artificial intelligence, philosophy of science, and
statistical modeling. Academic Press.
Golino, H. F., & Epskamp, S. (2017). Exploratory graph analysis: A
new approach for estimating the number of dimensions in psychological
research. PLOS ONE, 12(6), e0174035. https://doi.org/10.1371/journal.pone.0174035
Golino, H., & Christensen, A. (2026). EGAnet: Exploratory graph
analysis—a framework for estimating the number of dimensions in
multivariate data using network psychometrics (Version 2.5.0). https://r-ega.net
Golino, H., & Christensen, A. P. (2023). EGAnet: Exploratory
graph analysis – a framework for estimating the number of dimensions in
multivariate data using network psychometrics. R package.
Golino, H., Moulder, R., Shi, D., Christensen, A. P., Garrido, L. E.,
Nieto, M. D., Sadana, R., Thiyagarajan, J. A., Martínez-Molina, A.,
& Boker, S. M. (2021). Entropy fit indices: New fit measures for
assessing the structure and dimensionality of multiple latent variables.
Multivariate Behavioral Research, 56(6), 874–902. https://doi.org/10.1080/00273171.2020.1779642
Golino, H., Shi, D., Christensen, A. P., Garrido, L. E., Nieto, M. D.,
Sadana, R., Thiyagarajan, J. A., & Martínez-Molina, A. (2020).
Investigating the performance of exploratory graph analysis and
traditional techniques to identify the number of latent factors: A
simulation and tutorial. Psychological Methods, 25(3),
292–320. https://doi.org/10.1037/met0000255
Graham, J. M. (2006). Congeneric and (essentially) tau-equivalent
estimates of score reliability: What they are and how to use them.
Educational and Psychological Measurement, 66(6),
930–944. https://doi.org/10.1177/0013164406288165
Graziano, W. G., & Tobin, R. M. (2002). Agreeableness: Dimension of
personality or social desirability artifact?. Journal of
Personality, 70(5), 695–728. https://doi.org/10.1111/1467-6494.05021
Green, S. B., Redell, N., Thompson, M. S., & Levy, R. (2016).
Accuracy of revised and traditional parallel analyses for assessing
dimensionality with binary data. Educational and Psychological
Measurement, 76(1), 5–21. https://doi.org/10.1177/0013164415581898
Hair, J. F., Black, W. C., Babin, B. J., Anderson, R. E., & Tatham,
R. L. (2006). Multivariate data analysis (6th ed.). Pearson
Prentice Hall.
Hambleton, R. K., Swaminathan, H., & Rogers, H. J. (1991).
Fundamentals of item response theory. SAGE.
Hancock, G. R. (1997). Structural equation modeling methods of
hypothoesis testing of latent variable means. Measurement &
Evaluation in Counseling & Development (American Counseling
Association), 30(2), 91–105. https://doi.org/10.1080/07481756.1997.12068926
Harvill, L. M. (1991). Standard error of measurement. Educational
Measurement: Issues and Practice, 10(2), 33–41. https://doi.org/10.1111/j.1745-3992.1991.tb00195.x
Haslam, N., Holland, E., & Kuppens, P. (2012). Categories versus
dimensions in personality and psychopathology: A quantitative review of
taxometric research. Psychological Medicine, 42(5),
903–920. https://doi.org/10.1017/S0033291711001966
Hastings, W. K. (1970). Monte carlo sampling methods using markov chains
and their applications. Biometrika, 57(1), 97–109. https://doi.org/10.1093/biomet/57.1.97
Hebert, J. R., Ma, Y., Clemow, L., Ockene, I. S., Saperia, G., Stanek,
E. J., Merriam, P. A., & Ockene, J. K. (1997). Gender differences in
social desirability and social approval bias in dietary self-report.
American Journal of Epidemiology, 146(12), 1046–1055.
https://doi.org/10.1093/oxfordjournals.aje.a009233
Hinz, A., Michalski, D., Schwarz, R., & Herzberg, P. Y. (2007). The
acquiescence effect in responding to a questionnaire. GMS
Psycho-Social Medicine, 4.
Holzinger, K. J., & Swineford, F. (1939). A study in factor
analysis: The stability of a bi-factor solution. University of
Chicago Press.
Horn, J. L. (1965). A rationale and test for the number of factors in
factor analysis. Psychometrika, 30(2), 179–185. https://doi.org/10.1007/BF02289447
Hu, L., & Bentler, P. M. (1999). Cutoff criteria for fit indexes in
covariance structure analysis: Conventional criteria versus new
alternatives. Structural Equation Modeling: A Multidisciplinary
Journal, 6(1), 1–55. https://doi.org/10.1080/10705519909540118
Hughes, G. D. (2009). The impact of incorrect responses to reverse-coded
survey items. Research in the Schools.
Jonas, K. G., & Markon, K. E. (2016). A descriptivist approach to
trait conceptualization and inference. Psychological Review,
123(1), 90–96. https://doi.org/10.1037/a0039542
Jorgensen, T. D., Pornprasertmanit, S., Schoemann, A. M., & Rosseel,
Y. (2022). semTools: Useful tools for structural equation
modeling. CRAN. https://CRAN.R-project.org/package=semTools
Kaiser, H. F. (1960). The application of electronic computers to factor
analysis. Educational and Psychological Measurement,
20(1), 141–151. https://doi.org/10.1177/001316446002000116
Kaiser, H. F. (1970). A second generation little jiffy.
Psychometrika, 35(4), 401–415. https://doi.org/10.1007/BF02291817
Kaiser, H. F., & Rice, J. (1974). Little jiffy, mark IV.
Educational and Psychological Measurement, 34(1),
111–117. https://doi.org/10.1177/001316447403400115
Kam, C., Schermer, J. A., Harris, J., & Vernon, P. A. (2013).
Heritability of acquiescence bias and item keying response style
associated with the HEXACO personality scale. Twin Research and
Human Genetics, 16(4), 790–798.
Kam, C., Zhou, X., Zhang, X., & Ho, M. Y. (2012). Examining the
dimensionality of self-construals and individualistic–collectivistic
values with random intercept item factor analysis. Personality and
Individual Differences, 53(6), 727–733. https://doi.org/10.1016/j.paid.2012.05.023
Kan, K.-J., Jonge, H. de, Maas, H. L. J. van der, Levine, S. Z., &
Epskamp, S. (2020). How to compare psychometric factor and network
models. Journal of Intelligence, 8(4), 35. https://doi.org/10.3390/jintelligence8040035
Kane, M. T. (2001). Current concerns in validity theory. Journal of
Educational Measurement, 38(4), 319–342. https://doi.org/10.1111/j.1745-3984.2001.tb01130.x
Kane, M. T. (2006). Validation. In R. L. Brennan (Ed.), Educational
measurement (4th ed., pp. 17–64). American Council on Education;
Praeger.
Kane, M. T. (2013). Validating the interpretations and uses of test
scores. Journal of Educational Measurement, 50(1),
1–73. https://doi.org/10.1111/jedm.12000
Kane, M. T. (2017). Causal interpretations of psychological attributes.
Measurement: Interdisciplinary Research and Perspectives,
15, 79–82. https://doi.org/10.1080/15366367.2017.1369771
Kant, I. (1786). Metaphysical foundations of natural science.
Karabatsos, G. (2001). The rasch model, additive conjoint measurement,
and new models of probabilistic measurement theory. Journal of
Applied Measurement, 2(4), 389–423.
Kline, P. (1998). The new psychometrics: Science, psychology and
measurement. Psychology Press.
Kline, P. (2000). The handbook of psychological testing (2nd
ed.). Routledge.
Kline, R. B. (2023). Principles and practice of structural equation
modeling (5th ed.). Guilford Press.
Knight, R. G., Chisholm, B. J., Marsh, N. V., & Godfrey, H. P.
(1988). Some normative, reliability, and factor analytic data for the
revised UCLA loneliness scale. Journal of Clinical Psychology,
44(2), 203–206. https://doi.org/10.1002/1097-4679(198803
Krantz, D. H., Luce, R. D., Suppes, P., & Tversky, A. (1971).
Foundations of measurement, volume i: Additive and polynomial
representations. Academic Press.
Kruis, J., & Maris, G. (2016). Three representations of the ising
model. Scientific Reports, 6, 34175. https://doi.org/10.1038/srep34175
Külpe, O. (1895). Outlines of psychology. Sonnenschein.
Kyngdon, A. (2008a). Conjoint measurement, error and the rasch model: A
reply to michell, and borsboom and zand scholten. Theory &
Psychology, 18(1), 125–131. https://doi.org/10.1177/0959354307086927
Kyngdon, A. (2008b). The rasch model from the perspective of the
representational theory of measurement. Theory &
Psychology, 18(1), 89–109. https://doi.org/10.1177/0959354307086924
Lakatos, I., & Musgrave, A. (Eds.). (1970). Criticism and the
growth of knowledge: Proceedings of the international colloquium in the
philosophy of science, london 1965, volume 4. Cambridge University
Press. https://doi.org/10.1017/CBO9781139171434
Lange, F., & Dewitte, S. (2019). Measuring pro-environmental
behavior: Review and recommendations. Journal of Environmental
Psychology, 63, 92–100. https://doi.org/10.1016/j.jenvp.2019.04.009
Lanz, L., Thielmann, I., & Gerpott, F. H. (2022). Are social
desirability scales desirable? A meta‐analytic test of the validity of
social desirability scales in the context of prosocial behavior.
Journal of Personality, 90(2), 203–221. https://doi.org/10.1111/jopy.12662
Leite, W. L., & Cooper, L. A. (2010). Detecting social desirability
bias using factor mixture models. Multivariate Behavioral
Research, 45(2), 271–293. https://doi.org/10.1080/00273171003680245
Li, A., & Bagger, J. (2006). Using the BIDR to distinguish the
effects of impression management and self‐deception on the criterion
validity of personality measures: A meta‐analysis. International
Journal of Selection and Assessment, 14(2), 131–141. https://doi.org/10.1111/j.1468-2389.2006.00339.x
Linden, W. J. van der, & Hambleton, R. K. (Eds.). (1997).
Handbook of modern item response theory. Springer.
Loevinger, J. (1948). The technique of homogenous tests compared with
some aspects of “scale analysis” and factor analysis.
Psychological Bulletin, 45(6), 507–530. https://doi.org/10.1037/h0055827
Loken, E., & Molenaar, P. C. M. (2008). Categories or continua? The
correspondence between mixture models and factor models. In G. R.
Hancock & K. M. Samuelsen (Eds.), Advances in latent variable
mixture models (pp. 277–298). Information Age Publishing.
Lord, F. M., & Novick, M. R. (1968). Statistical theories of
mental test scores. Addison-Wesley.
Lubke, G. H., & Muthén, B. O. (2004). Applying multigroup
confirmatory factor models for continuous outcomes to likert scale data
complicates meaningful group comparisons. Structural Equation
Modeling, 11(4), 514–534. https://doi.org/10.1207/s15328007sem1104_2
Luce, R. D., & Tukey, J. W. (1964). Simultaneous conjoint
measurement: A new type of fundamental measurement. Journal of
Mathematical Psychology, 1, 1–27. https://doi.org/10.1016/0022-2496(64)90015-X
Maraun, M. D., & Gabriel, S. M. (2013). Illegitimate concept
equating in the partial fusion of construct validation theory and latent
variable modeling. New Ideas in Psychology, 31(1),
32–42. https://doi.org/10.1016/j.newideapsych.2011.02.006
Maraun, M. D., & Peters, J. (2005). What does it mean that an issue
is conceptual in nature? Journal of Personality Assessment,
85(2), 128–133. https://doi.org/10.1207/s15327752jpa8502_04
Marsh, H. W. (1996). Positive and negative global self-esteem: A
substantively meaningful distinction or artifactors?. Journal of
Personality and Social Psychology, 70(4), 810–819. https://doi.org/10.1037/0022-3514.70.4.810
Maul, A. (2017). Moving beyond traditional methods of survey validation.
Measurement: Interdisciplinary Research and Perspectives,
15, 103–109. https://doi.org/10.1080/15366367.2017.1369786
McDonald, R. P. (1999). Test theory: A unified treatment.
Lawrence Erlbaum Associates.
McFarland, D. J. (2020). The effects of using partial or uncorrected
correlation matrices when comparing network and latent variable models.
Journal of Intelligence, 8(1), 7. https://doi.org/10.3390/jintelligence8010007
McFarland, L. A., & Ryan, A. M. (2000). Variance in faking across
noncognitive measures. Journal of Applied Psychology,
85(5), 812–821. https://doi.org/10.1037/0021-9010.85.5.812
Meehl, P. E. (1995). Bootstraps taxometrics: Solving the classification
problem in psychopathology. American Psychologist,
50(4), 266–275. https://doi.org/10.1037/0003-066X.50.4.266
Mellenbergh, G. J. (1989). Item bias and item response theory.
International Journal of Educational Research, 13(2),
127–143. https://doi.org/10.1016/0883-0355(89)90002-5
Meredith, W. (1993). Measurement invariance, factor analysis and
factorial invariance. Psychometrika, 58, 525–543. https://doi.org/10.1007/BF02294825
Meredith, W., & Tisak, J. (1990). Latent curve analysis.
Psychometrika, 55, 107–122. https://doi.org/10.1007/BF02294746
Messick, S. (1989). Validity. In R. L. Linn (Ed.), Educational
measurement (3rd ed., pp. 13–103). American Council on Education;
Macmillan.
Metropolis, N., Rosenbluth, A. W., Rosenbluth, M. N., Teller, A. H.,
& Teller, E. (1953). Equation of state calculations by fast
computing machines. The Journal of Chemical Physics,
21(6), 1087–1092. https://doi.org/10.1063/1.1699114
Michell, J. (2000). Normal science, pathological science and
psychometrics. Theory & Psychology, 10(5),
639–667. https://doi.org/10.1177/0959354300105004
Michell, J. (2001). Teaching and misteaching measurement in psychology.
Australian Psychologist, 36(3), 211–218. https://doi.org/10.1080/00050060108259657
Michell, J. (2004). Item response models, pathological science and the
shape of error: Reply to borsboom and mellenbergh. Theory &
Psychology, 14(1), 121–129. https://doi.org/10.1177/0959354304040201
Michell, J. (2008). Is psychometrics pathological science?
Measurement: Interdisciplinary Research and Perspectives,
6(1–2), 7–24. https://doi.org/10.1080/15366360802035489
Michell, J. (2014). An introduction to the logic of psychological
measurement. Psychology Press.
Millsap, R. E., & Everson, H. T. (1993). Methodology review:
Statistical approaches for assessing measurement bias. Applied
Psychological Measurement, 17(4), 297–334. https://doi.org/10.1177/014662169301700401
Mirowsky, J., & Ross, C. E. (1991). Eliminating defense and
agreement bias from measures of the sense of control: A 2 x 2 index.
Social Psychology Quarterly, 54(2), 127. https://doi.org/10.2307/2786931
Mokken, R. J. (1971). A theory and procedure of scale analysis.
De Gruyter. https://doi.org/10.1515/9783110813203
Molenaar, I. W. (1991). A weighted loevinger h-coefficient extending
mokken scaling to multicategory items. Kwantitatieve Methoden,
37, 97–117.
Munafò, M. R., Nosek, B. A., Bishop, D. V. M., Button, K. S., Chambers,
C. D., Percie du Sert, N., Simonsohn, U., Wagenmakers, E.-J., Ware, J.
J., & Ioannidis, J. P. A. (2017). A manifesto for reproducible
science. Nature Human Behaviour, 1(1), 0021. https://doi.org/10.1038/s41562-016-0021
Muthén, B. O., & Muthén, L. K. (2000). Integrating person-centered
and variable-centered analyses: Growth mixture modeling with latent
trajectory classes. Alcoholism: Clinical and Experimental
Research, 24(6), 882–891. https://doi.org/10.1111/j.1530-0277.2000.tb02070.x
Muthén, B., & Kaplan, D. (1985). A comparison of some methodologies
for the factor analysis of non-normal likert variables. British
Journal of Mathematical and Statistical Psychology, 38,
171–189. https://doi.org/10.1111/j.2044-8317.1985.tb00832.x
Narens, L., & Luce, R. D. (1993). Further comments on the
“nonrevolution” arising from axiomatic measurement theory.
Psychological Science, 4(2), 127–130. https://doi.org/10.1111/j.1467-9280.1993.tb00475.x
Nederhof, A. J. (1985). Methods of coping with social desirability bias:
A review. European Journal of Social Psychology, 15,
263–280. https://doi.org/10.1002/ejsp.2420150303
Newman, M. E. J. (2006). Modularity and community structure in networks.
Proceedings of the National Academy of Sciences,
103(23), 8577–8582. https://doi.org/10.1073/pnas.0601602103
Newton, P. E., & Shaw, S. D. (2013). Validity in educational and
psychological assessment. SAGE.
Nickerson, C. A., & McClelland, G. H. (1984). Scaling distortion in
numerical conjoint measurement. Applied Psychological
Measurement, 8(2), 183–198. https://doi.org/10.1177/014662168400800207
Noble, T., Rosebery, A., Suarez, C., Warren, B., & O’Connor, M. C.
(2014). Science assessments and english language learners: Validity
evidence based on response processes. Applied Measurement in
Education, 27(4), 248–260. https://doi.org/10.1080/08957347.2014.944309
Nosek, B. A., Hardwicke, T. E., Moshontz, H., Allard, A., Corker, K. S.,
Dreber, A., Fidler, F., Hilgard, J., Kline Struhl, M., Nuijten, M. B.,
Rohrer, J. M., Romero, F., Scheel, A. M., Scherer, L. D., Schönbrodt, F.
D., & Vazire, S. (2022). Replicability, robustness, and
reproducibility in psychological science. Annual Review of
Psychology, 73(1), 719–748. https://doi.org/10.1146/annurev-psych-020821-114157
Nosek, B. A., Spies, J. R., & Motyl, M. (2012). Scientific utopia:
II. Restructuring incentives and practices to promote truth over
publishability. Perspectives on Psychological Science,
7(6), 615–631. https://doi.org/10.1177/1745691612459058
Nowick, K., Gernat, T., Almaas, E., & Stubbs, L. (2009). Differences
in human and chimpanzee gene expression patterns define an evolving
network of transcription factors in brain. Proceedings of the
National Academy of Sciences, 106(52), 22358–22363. https://doi.org/10.1073/pnas.0911376106
Oberski, D. (2016). Mixture models: Latent profile and latent class
analysis. In J. Robertson & M. Kaptein (Eds.), Modern
statistical methods for HCI (pp. 275–287). Springer. https://doi.org/10.1007/978-3-319-26633-6_12
Ones, D. S., Viswesvaran, C., & Reiss, A. D. (1996). Role of social
desirability in personality testing for personnel selection: The red
herring. Journal of Applied Psychology, 81(6), 660. https://doi.org/10.1037/0021-9010.81.6.660
Osborne, J. W. (2014). Best practices in exploratory factor
analysis. CreateSpace Independent Publishing.
Pasquali, L. (2017). Psicometria: Teoria dos testes na psicologia e
na educação. Editora Vozes Limitada.
Paulhus, D. L. (1984). Two-component models of socially desirable
responding. Journal of Personality and Social Psychology,
46(3), 598–609. https://doi.org/10.1037/0022-3514.46.3.598
Paulhus, D. L. (1991). Measurement and control of response bias. In
Measures of personality and social psychological attitudes (pp.
17–59). Academic Press. https://doi.org/10.1016/B978-0-12-590241-0.50006-X
Paulhus, D. L., & John, O. P. (1998). Egoistic and moralistic biases
in self-perception: The interplay of self-deceptive styles with basic
traits and motives. Journal of Personality, 66(6),
1025–1060. https://doi.org/10.1111/1467-6494.00041
Peabody, D. (1967). Trait inferences: Evaluative and descriptive
aspects. Journal of Personality and Social Psychology,
4, Pt. https://doi.org/10.1037/h0025230
Pearl, J. (2009). Causality: Models, reasoning, and inference
(2nd ed.). Cambridge University Press.
Pearson, K. (1978). The history of statistics in the seventeenth and
eighteenth centuries. Griffin.
Peterson, R. A., & Kerin, R. A. (1981). The quality of self-report
data: Review and synthesis. In Review of marketing (pp. 5–20).
American Marketing Association.
Pettersson, E., Mendle, J., Turkheimer, E., Horn, E. E., Ford, D. C.,
Simms, L. J., & Clark, L. A. (2014). Do maladaptive behaviors exist
at one or both ends of personality traits? Psychological
Assessment, 26(2), 433–446. https://doi.org/10.1037/a0035587
Pettersson, E., Turkheimer, E., Horn, E. E., & Menatti, A. R.
(2012). The general factor of personality and evaluation. European
Journal of Personality, 26(3), 292–302. https://doi.org/10.1002/per.839
Podsakoff, P. M., MacKenzie, S. B., Lee, J. Y., & Podsakoff, N. P.
(2003). Common method biases in behavioral research: A critical review
of the literature and recommended remedies. Journal of Applied
Psychology, 88(5), 879–903. https://doi.org/10.1037/0021-9010.88.5.879
Pons, P., & Latapy, M. (2006). Computing communities in large
networks using random walks. Journal of Graph Algorithms and
Applications, 10(2), 191–218. https://doi.org/10.7155/jgaa.00124
Ramsay, J. O. (1975). Review of foundations of measurement, volume i.
Psychometrika, 40, 257–262.
Ramsay, J. O. (1991). Reviews of foundations of measurement, volumes II
and III. Psychometrika, 56, 355–358.
Rasch, G. (1960). Probabilistic models for some intelligence and
attainment tests. Danish Institute for Educational Research.
Raymark, P. H., & Tafero, T. L. (2009). Individual differences in
the ability to fake on personality measures. Human Performance,
22(1), 86–103. https://doi.org/10.1080/08959280802541039
Revelle, W. (2023a). Psych: Procedures for psychological,
psychometric, and personality research. Northwestern University. https://CRAN.R-project.org/package=psych
Revelle, W. (2023b). psychTools: Tools to accompany the ’psych’
package for psychological research. Northwestern University. https://CRAN.R-project.org/package=psychTools
Riggio, R. E., Salinas, C., & Tucker, J. (1988). Personality and
deception ability. Personality and Individual Differences,
9(1), 189–191. https://doi.org/10.1016/0191-8869(88
Robinson, J. P., Shaver, P. R., & Wrightsman, L. S. and. (1991).
Measures of social psychological attitudes. (Vol. 1. Measures
of personality; social psychological atitudes). Academic Press.
Rosenberg, M. (1965). Society and the adolescent self-image.
Princeton University Press.
Rosseel, Y. (2012). Lavaan: An r package for structural equation
modeling. Journal of Statistical Software, 48(2),
1–36. https://doi.org/10.18637/jss.v048.i02
Ruscio, J., & Kaczetow, W. (2009). Differentiating categories and
dimensions: Evaluating the robustness of taxometric analyses.
Multivariate Behavioral Research, 44(2), 259–280. https://doi.org/10.1080/00273170902794248
Ruscio, J., Ruscio, A. M., & Meron, M. (2007). Applying the
bootstrap to taxometric analysis: Generating empirical sampling
distributions to help interpret results. Multivariate Behavioral
Research, 42(2), 349–386. https://doi.org/10.1080/00273170701360795
Salazar, M. S. (2015). The dilemma of combining positive and negative
items in scales. Psicothema, 27(2), 192–199. https://doi.org/10.7334/psicothema2014.266
Saucier, G., Ostendorf, F., & Peabody, D. (2001). The non-evaluative
circumplex of personality adjectives. Journal of Personality,
69(4), 537–582. https://doi.org/10.1111/1467-6494.694155
Savalei, V., & Falk, C. F. (2014). Recovering substantive factor
loadings in the presence of acquiescence bias: A comparison of three
approaches. Multivariate Behavioral Research, 49(5),
407–424. https://doi.org/10.1080/00273171.2014.931800
Scheiblechner, H. (1995). Isotonic ordinal probabilistic models (ISOP).
Psychometrika, 60(2), 281–304.
Scheiblechner, H. (1998). Corrections of theorems in scheiblechner’s
treatment of ISOP models and comments on junker’s remarks.
Psychometrika, 63(1), 87–91. https://doi.org/10.1007/BF02295439
Scheiblechner, H. (1999). Additive conjoint isotonic probabilistic
models (ADISOP). Psychometrika, 64(3), 295–316. https://doi.org/10.1007/BF02294297
Schimmack, U. (2021). The validation crisis in psychology.
Meta-Psychology, 5.
Schmalz, X., Breuer, J., Haim, M., Hildebrandt, A., Knöpfle, P., Leung,
A. Y., & Roettger, T. B. (2024). Let’s talk about language—and
its role for replicability. https://osf.io/preprints/metaarxiv/w2gb9
Schmittmann, V. D., Cramer, A. O. J., Waldorp, L. J., Epskamp, S.,
Kievit, R. A., & Borsboom, D. (2013). Deconstructing the construct:
A network perspective on psychological phenomena. New Ideas in
Psychology, 31(1), 43–53. https://doi.org/10.1016/j.newideapsych.2011.02.007
Schriesheim, C. A., & Hill, K. D. (1981). Controlling acquiescence
response bias by item reversals: The effect on questionnaire validity.
Educational and Psychological Measurement, 41(4),
1101–1114. https://doi.org/10.1177/001316448104100420
Schwager, K. W. (1991). The representational theory of measurement: An
assessment. Psychological Bulletin, 110(3), 618–626.
https://doi.org/10.1037/0033-2909.110.3.618
Selau, T., Silva, M. A. da, Mendonça Filho, E. J. de, & Bandeira, D.
R. (2020). Evidence of validity and reliability of the adaptive
functioning scale for intellectual disability (EFA-DI).
Psicologia: Reflexão e Crítica, 33, 26. https://doi.org/10.1186/s41155-020-00164-7
Sijtsma, K. (2009). On the use, the misuse, and the very limited
usefulness of cronbach’s alpha. Psychometrika, 74(1),
107–120. https://doi.org/10.1007/s11336-008-9101-0
Sijtsma, K. (2012). Psychological measurement between physics and
statistics. Theory & Psychology, 22(6), 786–809.
https://doi.org/10.1177/0959354312454353
Sijtsma, K., & Molenaar, I. W. (2002). Introduction to
nonparametric item response theory. Sage. https://doi.org/10.4135/9781412984676
Snell, A. F., Sydell, E. J., & Lueke, S. B. (1999). Towards a theory
of applicant faking: Integrating studies of deception. Human
Resource Management Review, 9(2), 219–242. https://doi.org/10.1016/S1053-4822(99
Soto, C. J., John, O. P., Gosling, S. D., & Potter, J. (2008). The
developmental psychometrics of big five self-reports: Acquiescence,
factor structure, coherence, and differentiation from ages 10 to 20.
Journal of Personality and Social Psychology, 94(4),
718–737. https://doi.org/10.1037/0022-3514.94.4.718
Spearman, C. (1904). “General intelligence,” objectively
determined and measured. The American Journal of Psychology,
15(2), 201–292. https://doi.org/10.2307/1412107
Spearman, C. (1910). Correlation calculated from faulty data.
British Journal of Psychology, 3(3), 271–295. https://doi.org/10.1111/j.2044-8295.1910.tb00206.x
Spearman, C. (1937). Psychology down the ages (Vol. 1).
Macmillan.
Steiger, J. H., & Lind, J. C. (1980). Statistically based tests
for the number of common factors.
Stevens, S. S. (1946). On the theory of scales of measurement.
Science, 103(2684), 677–680. https://doi.org/10.1126/science.103.2684.677
Stevens, S. S. (1951). Mathematics, measurement and psychophysics. In S.
S. Stevens (Ed.), Handbook of experimental psychology (pp.
1–49). Wiley.
Stevens, S. S. (1959). Measurement, psychophysics and utility. In C. W.
Churchman & P. Ratoosh (Eds.), Measurement: Definitions and
theories (pp. 18–63). Wiley.
Straat, J. H., Ark, L. A. van der, & Sijtsma, K. (2013). Comparing
optimization algorithms for item selection in mokken scale analysis.
Journal of Classification, 30, 72–99. https://doi.org/10.1007/s00357-013-9122-y
Suárez-Alvarez, J., Pedrosa, I., Lozano Fernández, L. M., García-Cueto,
E., Cuesta, M., & Muñiz, J. (2018). Using reversed items in likert
scales: A questionable practice. Psicothema, 30(2),
149–158. https://doi.org/10.7334/psicothema2018.33
Svetina, D., Rutkowski, L., & Rutkowski, D. (2020). Multiple-group
invariance with categorical outcomes using updated guidelines: An
illustration using m plus and the lavaan/semtools packages.
Structural Equation Modeling: A Multidisciplinary Journal,
27(1), 111–130. https://doi.org/10.1080/10705511.2019.1602776
Tabachnick, B. G., & Fidell, L. S. (2007). Using multivariate
statistics (5th ed.). Allyn & Bacon.
Ten Berge, J. M. F., & Kiers, H. A. L. (1991). A numerical approach
to the approximate and the exact minimum rank of a covariance matrix.
Psychometrika, 56, 309–315. https://doi.org/10.1007/BF02294464
Thomson, W. (1891). Popular lectures and addresses (Vol. 1).
Macmillan.
Titchener, E. B. (1905). Experimental psychology: A manual of
laboratory practice. Macmillan.
Tourangeau, R., & Yan, T. (2007). Sensitive questions in surveys.
Psychological Bulletin, 133(5), 859–883. https://doi.org/10.1037/0033-2909.133.5.859
Trendler, G. (2009). Measurement theory, psychology and the revolution
that cannot happen. Theory & Psychology, 19(5),
579–599. https://doi.org/10.1177/0959354309341926
Trendler, G. (2013). Measurement in psychology: A case of Ignoramus et Ignorabimus? A rejoinder. Theory
& Psychology, 23(5), 591–615. https://doi.org/10.1177/0959354313490451
Tucker, L. R., & Lewis, C. (1973). A reliability coefficient for
maximum likelihood factor analysis. Psychometrika,
38(1), 1–10. https://doi.org/10.1007/BF02291170
Uher, J. (2021). Problematic research practices in psychology:
Misconceptions about data collection entail serious fallacies in data
analysis. Theory & Psychology, 31(3), 411–416. https://doi.org/10.1177/09593543211014963
Uher, J. (2022). Rating scales institutionalise a network of logical
errors and conceptual problems in research practices: A rigorous
analysis showing ways to tackle psychology’s crises. Frontiers in
Psychology, 13, 1009893. https://doi.org/10.3389/fpsyg.2022.1009893
Uher, J. (2023). What’s wrong with rating scales? Psychology’s
replication and confidence crisis cannot be solved without transparency
in data generation. Social and Personality Psychology Compass,
17(5), e12740. https://doi.org/10.1111/spc3.12740
Valentini, F. (2017). Influência e controle da aquiescência na análise
fatorial [influence and control of acquiesncence in factor analysis].
Avaliação Psicológica, 16(2), 6–12. https://doi.org/10.15689/ap.2017.1602.ed
Valentini, F., & Hauck Filho, N. (2020). O impacto da aquiescência
na estimação de coeficientes de validade [acquiescence impact in the
estimation of validity coefficients]. Avaliação Psicológica,
19(1), 1–3. http://dx.doi.org/10.15689/ap.2020.1901.ed
Van Sonderen, E., Sanderman, R., & Coyne, J. C. (2013).
Ineffectiveness of reverse wording of questionnaire items: Let’s learn
from cows in the rain. PloS One, 8(7), e68967. https://doi.org/10.1371/journal.pone.0068967
Vecina, M. L., Chacón, F., & Pérez-Viejo, J. M. (2016). Moral
absolutism, self-deception, and moral self-concept in men who commit
intimate partner violence: A comparative study with an opposite sample.
Violence Against Women, 22(1), 3–16. https://doi.org/10.1177/1077801215597791
Verweij, A. C., Sijtsma, K., & Koops, W. (1996). A mokken scale for
transitive reasoning suited for longitudinal research. International
Journal of Behavioral Development, 19(1), 219–238. https://doi.org/10.1177/016502549601900115
Vries, R. E. de, Zettler, I., & Hilbig, B. E. (2014). Rethinking
trait conceptions of social desirability scales: Impression management
as an expression of honesty-humility. Assessment,
21(3), 286–299. https://doi.org/10.1177/1073191113504619
Weijters, B., & Baumgartner, H. (2012). Misresponse to reversed and
negated items in surveys: A review. Journal of Marketing
Research, 49(5), 737–747. https://doi.org/10.1509/jmr.11.0368
Weijters, B., Cabooter, E., & Schillewaert, N. (2010). The effect of
rating scale format on response styles: The number of response
categories and response category labels. International Journal of
Research in Marketing, 27(3), 236–247. https://doi.org/10.1016/j.ijresmar.2010.02.004
Williams, E. A., Pillai, R., Lowe, K. B., Jung, D., & Herst, D.
(2009). Crisis, charisma, values, and voting behavior in the 2004
presidential election. The Leadership Quarterly,
20(2), 70–86. https://doi.org/10.1016/j.leaqua.2009.01.002
Wolf, M. G., Ihm, E., Maul, A., & Taves, A. (2025). The
response-process-evaluation method: A new approach to survey-item
validation. Advances in Methods and Practices in Psychological
Science, 8(3), 1–20. https://doi.org/10.1177/25152459251353135
Woods, C. M. (2006). Careless responding to reverse-worded items:
Implications for confirmatory factor analysis. Journal of
Psychopathology and Behavioral Assessment, 28(3), 186–191.
https://doi.org/10.1007/s10862-005-9004-7
Wu, H., & Estabrook, R. (2016). Identification of confirmatory
factor analysis models of different levels of invariance for ordered
categorical outcomes. Psychometrika, 81(4), 1014–1045.
https://doi.org/10.1007/s11336-016-9506-0
Xia, Y., & Yang, Y. (2019). RMSEA, CFI, and TLI in structural
equation modeling with ordered categorical data: The story they tell
depends on the estimation methods. Behavior Research Methods,
51, 409–428. https://doi.org/10.3758/s13428-018-1055-2
Zhang, X., & Savalei, V. (2016). Improving the factor structure of
psychological scales: The expanded format as an alternative to the
likert scale format. Educational and Psychological Measurement,
76(3), 357–386. https://doi.org/10.1177/0013164415596421
Zickar, M. J., & Robie, C. (1999). Modeling faking good on
personality items: An item-level analysis. Journal of Applied
Psychology, 84(4), 551. https://doi.org/10.1037/0021-9010.84.4.551
Ziegler, M., & Buehner, M. (2009). Modeling socially desirable
responding and its effects. Educational and Psychological
Measurement, 69(4), 548–565. https://doi.org/10.1177/0013164408324469
Ziegler, M., Maccann, C., & Roberts, R. D. (2011). New
perspectives on faking in personality assessment. Oxford University
Press.
Zwick, W. R., & Velicer, W. F. (1986). Comparison of five rules for
determining the number of components to retain. Psychological
Bulletin, 99(3), 432–442. https://doi.org/10.1037/0033-2909.99.3.432