Referencias

American Educational Research Association, American Psychological Association, & National Council on Measurement in Education. (2014). Standards for educational and psychological testing. American Educational Research Association.
Ark, L. A. van der. (2007). Mokken scale analysis in r. Journal of Statistical Software, 20(11), 1–19. https://doi.org/10.18637/jss.v020.i11
Ark, L. A. van der. (2012). New developments in mokken scale analysis in r. Journal of Statistical Software, 48(5), 1–27. https://doi.org/10.18637/jss.v048.i05
Baptista, M. N., & Villemor-Amaral, A. E. de. (2019). Compêndio de avaliação psicológica. Editora Vozes.
Bartholomew, D. J. (2004). Measuring intelligence: Facts and fallacies. Cambridge University Press.
Bartlett, M. S. (1954). A note on the multiplying factors for various chi square approximations. Journal of the Royal Statistical Society: Series B (Methodological), 16(2), 296–298. https://doi.org/10.1111/j.2517-6161.1954.tb00174.x
Bastos, R. V. S., Novaes, F. C., & Natividade, J. C. (2022). Self-perception of prejudice and discrimination scale: Evidence of validity and other psychometric properties. Trends in Psychology, 1–19. https://doi.org/10.1007/s43076-022-00190-7
Baumgartner, H., & Steenkamp, J. (2001). Response styles in marketing research: A cross national investigation. Journal of Marketing Research, 38(2), 143–156. https://doi.org/10.1509/jmkr.38.2.143.18840
Beauducel, A., & Herzberg, P. Y. (2006). On the performance of maximum likelihood versus means and variance adjusted weighted least squares estimation in CFA. Structural Equation Modeling: A Multidisciplinary Journal, 13(2), 186–203. https://doi.org/10.1207/s15328007sem1302_2
Bentler, P. M. (1990). Comparative fit indexes in structural models. Psychological Bulletin, 107(2), 238–246. https://doi.org/10.1037/0033-2909.107.2.238
Beymer, P. N., Ferland, M., & Flake, J. K. (2022). Validity evidence for a short scale of college students’ perceptions of cost. Current Psychology, 41(11), 7937–7956. https://doi.org/10.1007/s12144-020-01218-w
Billiet, J. B., & Davidov, E. (2008). Testing the stability of an acquiescence style factor behind two interrelated substantive variables in a panel design. Sociological Methods Research, 36(4), 542–562. https://doi.org/10.1177/0049124107313901
Billiet, J. B., & McClendon, M. J. (2000). Modeling acquiescence in measurement models for two balanced sets of items. Structural Equation Modeling, 7, 608–628. https://doi.org/10.1207/S15328007SEM0704_5
Bollen, K. A. (1989). Structural equations with latent variables. Wiley.
Bollen, K. A., & Ting, K. (1993). Confirmatory tetrad analysis. Sociological Methodology, 23, 147–175. https://doi.org/10.2307/271009
Bonifay, W., & Cai, L. (2017). On the complexity of item response theory models. Multivariate Behavioral Research, 52(4), 465–484. https://doi.org/10.1080/00273171.2017.1309262
Bork, R. van, Wijsen, L. D., & Rhemtulla, M. (2017). Toward a causal interpretation of the common factor model. Disputatio, 9(47), 581–601. https://doi.org/10.1515/disp-2017-0019
Borsboom, D. (2005). Measuring the mind: Conceptual issues in contemporary psychometrics. Cambridge University Press.
Borsboom, D. (2006). The attack of the psychometricians. Psychometrika, 71(3), 425–440. https://doi.org/10.1007/s11336-006-1447-6
Borsboom, D., & Cramer, A. O. J. (2013). Network analysis: An integrative approach to the structure of psychopathology. Annual Review of Clinical Psychology, 9, 91–121. https://doi.org/10.1146/annurev-clinpsy-050212-185608
Borsboom, D., & Mellenbergh, G. J. (2004). Why psychometrics is not pathological: A comment on michell. Theory & Psychology, 14(1), 105–120. https://doi.org/10.1177/0959354304040200
Borsboom, D., Mellenbergh, G. J., & Heerden, J. van. (2004). The concept of validity. Psychological Review, 111(4), 1061–1071. https://doi.org/10.1037/0033-295X.111.4.1061
Borsboom, D., & Zand Scholten, A. (2008). The rasch model and conjoint measurement theory from the perspective of psychometrics. Theory & Psychology, 18(1), 111–117. https://doi.org/10.1177/0959354307086925
Bosson, J. K., Swann, Jr., William B., & Pennebaker, J. W. (2000). Stalking the perfect measure of implicit self-esteem: The blind men and the elephant revisited? Journal of Personality and Social Psychology, 79(4), 631–643. https://doi.org/10.1037/0022-3514.79.4.631
Brennan, R. L. (2011). Generalizability theory and classical test theory. Applied Measurement in Education, 24(1), 1–21. https://doi.org/10.1080/08957347.2011.532417
Brentano, F. (1874). Psychology from an empirical standpoint. Humanities Press.
Bringmann, L. F., & Eronen, M. I. (2016). Heating up the measurement debate: What psychologists can learn from the history of physics. Theory & Psychology, 26(1), 27–43. https://doi.org/10.1177/0959354315617253
Brown, T. A. (2015). Confirmatory factor analysis for applied research (2nd ed.). The Guilford Press.
Brown, W. (1910). Some experimental results in the correlation of mental abilities. British Journal of Psychology, 3(3), 296–322. https://doi.org/10.1111/j.2044-8295.1910.tb00207.x
Cambré, B., Welkenhuysen-Gybels, J., & Billiet, J. (2002). Is it content or style? An evaluation of two competitive measurement models applied to a balanced set of ethnocentrism items. International Journal of Comparative Sociology, 43, 1–20. https://doi.org/10.1177/002071520204300101
Cattell, J. M. (1890). Mental tests and measurements. Mind, 15(59), 373–381. https://doi.org/10.1093/mind/os-XV.59.373
Cattell, R. B. (1966). The scree test for the number of factors. Multivariate Behavioral Research, 1(2), 245–276. https://doi.org/10.1207/s15327906mbr0102_10
Chang, L. (1995). Connotatively consistent and reversed connotatively inconsistent items are not fully equivalent: Generalizability study. Educational and Psychological Measurement, 55, 991–997. https://doi.org/10.1177/0013164495055006007
Chen, C., Shin-ying, L., & Stevenson, H. W. (1995). Response style and cross-cultural comparisons of rating scales among east asian and north american students. Psychological Science, 6, 170–175. https://doi.org/10.1111/j.1467-9280.1995.tb00327.x
Chen, J., & Chen, Z. (2008). Extended bayesian information criteria for model selection with large model spaces. Biometrika, 95(3), 759–771. https://doi.org/10.1093/biomet/asn034
Cheung, G. W., & Lau, R. S. (2012). A direct comparison approach for testing measurement invariance. Organizational Research Methods, 15(2), 167–198. https://doi.org/10.1177/1094428111421987
Christensen, A. P., Garrido, L. E., & Golino, H. (2023). Unique variable analysis: A network psychometrics method to detect local dependence. Multivariate Behavioral Research, 58(6), 1165–1182. https://doi.org/10.1080/00273171.2023.2194606
Christensen, A. P., & Golino, H. (2021). Estimating the stability of psychological dimensions via bootstrap exploratory graph analysis: A monte carlo simulation and tutorial. Psych, 3(3), 479–500. https://doi.org/10.3390/psych3030032
Christensen, A. P., Golino, H., & Silvia, P. J. (2020). A psychometric network perspective on the validity and validation of personality trait questionnaires. European Journal of Personality, 34(6), 1095–1108. https://doi.org/10.1002/per.2265
Cliff, N. (1992). Abstract measurement theory and the revolution that never happened. Psychological Science, 3(3), 186–190. https://doi.org/10.1111/j.1467-9280.1992.tb00024.x
Connelly, B. S., & Chang, L. (2016). A meta‐analytic multitrait multirater separation of substance and style in social desirability scales. Journal of Personality, 84(3), 319–334. https://doi.org/10.1111/jopy.12161
Cortina, J. M. (1993). What is coefficient alpha? An examination of theory and applications. Journal of Applied Psychology, 78(1), 98–104. https://doi.org/10.1037/0021-9010.78.1.98
Crocker, L., & Algina, J. (1986). Introduction to classical and modern test theory. Holt, Rinehart; Winston.
Cronbach, L. J. (1942). Studies of acquiescence as a factor in the true-false test. Journal of Educational Psychology, 33, 401–415. https://doi.org/10.1037/h0054677
Cronbach, L. J. (1951). Coefficient alpha and the internal structure of tests. Psychometrika, 16(3), 297–334. https://doi.org/10.1007/BF02310555
Cronbach, L. J., & Meehl, P. E. (1955). Construct validity in psychological tests. Psychological Bulletin, 52(4), 281–302. https://doi.org/10.1037/h0040957
Danner, D., Aichholzer, J., & Rammstedt, B. (2015). Acquiescence in personality questionnaires: Relevance, domain specificity, and stability. Journal of Research in Personality, 57, 119–130. https://doi.org/10.1016/j.jrp.2015.05.004
Degobi, E. B., Franco, V. R., & Student, S. R. (2026). A model fit approach to test the adequacy of interval-level measurement in psychometric models. https://osf.io/jufy5/
Degobi, E. B., & Valentini, F. (2023). Simulations for two theoretically sound controls for social desirability: MIMIC and forced-choice [Master's thesis]. Universidade São Francisco.
DeVellis, R. F., & Thorpe, C. T. (2021). Scale development: Theory and applications (5th ed.). SAGE.
Domingue, B. W. (2014). Evaluating the equal-interval hypothesis with test score scales. Psychometrika, 79(1), 1–19. https://doi.org/10.1007/s11336-013-9342-4
Dunn, T. J., Baguley, T., & Brunsden, V. (2014). From alpha to omega: A practical solution to the pervasive problem of internal consistency estimation. British Journal of Psychology, 105(3), 399–412. https://doi.org/10.1111/bjop.12046
Edwards, A. L. (1953). The relationship between the judged desirability of a trait and the probability that the trait will be endorsed. Journal of Applied Psychology, 37(2), 90. https://doi.org/10.1037/h0058073
Edwards, A. L. (1957). The social desirability variable in personality assessment and research. Dryden Press.
Edwards, J. R., & Bagozzi, R. P. (2000). On the nature and direction of relationships between constructs and measures. Psychological Methods, 5(2), 155–174. https://doi.org/10.1037/1082-989X.5.2.155
Epskamp, S. (2022). semPlot: Path diagrams and visual analysis of various SEM packages’ output. CRAN. https://CRAN.R-project.org/package=semPlot
Epskamp, S., Maris, G., Waldorp, L. J., & Borsboom, D. (2018). Network psychometrics. In The wiley handbook of psychometric testing: A multidisciplinary reference on survey, scale and test development (pp. 953–986). Wiley. https://doi.org/10.1002/9781118489772.ch30
Fabrigar, L. R., Wegener, D. T., MacCallum, R. C., & Strahan, E. J. (1999). Evaluating the use of exploratory factor analysis in psychological research. Psychological Methods, 4(3), 272–299. https://doi.org/10.1037/1082-989X.4.3.272
Fechner, G. T. (1860). Elemente der psychophysik. Breitkopf und Härtel.
Ferrando, P. J., Condon, L., & Chico, E. (2004). The convergent validity of acquiescence: An empirical study relating balanced scales and separate acquiescence scales. Personality and Individual Differences, 37(7), 1331–1340. https://doi.org/10.1016/j.paid.2004.01.003
Ferrando, P. J., Lorenzo-Seva, U., & Chico, E. (2003). Unrestricted factor analytic procedures for assessing acquiescent responding in balanced, theoretically unidimensional personality scales. Multivariate Behavioral Research, 38(2), 353–374. https://doi.org/10.1207/S15327906MBR3803_04
Ferrando, P. J., Lorenzo-Seva, U., & Chico, E. (2009). A general factor-analytic procedure for assessing response bias in questionnaire measures. Structural Equation Modeling: A Multidisciplinary Journal, 16(2), 364–381. https://doi.org/10.1080/10705510902751374
Fischer, G. H. (1995). Some neglected problems in IRT. Psychometrika, 60(4), 459–487. https://doi.org/10.1007/BF02294324
Flake, J. K., & Fried, E. I. (2020). Measurement schmeasurement: Questionable measurement practices and how to avoid them. Advances in Methods and Practices in Psychological Science, 3(4), 456–465. https://doi.org/10.1177/2515245920952393
Flake, J. K., Pek, J., & Hehman, E. (2017). Construct validation in social and personality research: Current practice and recommendations. Social Psychological and Personality Science, 8(4), 370–378. https://doi.org/10.1177/1948550617693063
Franco, V. R. (2021). É possível identificar o nível de medida de variáveis latentes? Avaliação Psicológica, 20(2), a–d. https://doi.org/10.15689/ap.2021.2002.ed
Franco, V. R., Bastos, R. V. S., & Jiménez, M. (2023). Tetrad fit index for factor analysis models. Paper presented at Virtual MathPsych/ICCM 2023. https://mathpsych.org/presentation/1297
Franco, V. R., Laros, J. A., & Bastos, R. V. S. (2022). Theoretical and practical foundations of mokken scale analysis in psychology. Paidéia (Ribeirão Preto), 32, e3223. https://doi.org/10.1590/1982-4327e3223
Friborg, O., Martinussen, M., & Rosenvinge, J. H. (2006). Likert-based vs. Semantic differential-based scorings of positive psychological constructs: A psychometric comparison of two versions of a scale measuring resilience. Personality and Individual Differences, 40(5), 873–884. https://doi.org/10.1016/j.paid.2005.08.015
Friedman, J., Hastie, T., & Tibshirani, R. (2008). Sparse inverse covariance estimation with the graphical lasso. Biostatistics, 9(3), 432–441. https://doi.org/10.1093/biostatistics/kxm045
Galilei, G. (1864). Il saggiatore. G. Barbèra.
Galton, F. (1879). Psychometric experiments. Brain, 2(2), 149–162. https://doi.org/10.1093/brain/2.2.149
Gibson, W. A. (1959). Three multivariate models: Factor analysis, latent structure analysis, and latent profile analysis. Psychometrika, 24(3), 229–252. https://doi.org/10.1007/BF02289845
Glymour, C., & Scheines, R. (1986). Causal modeling with the TETRAD program. Synthese, 68, 37–63.
Glymour, C., Scheines, R., & Spirtes, P. (2014). Discovering causal structure: Artificial intelligence, philosophy of science, and statistical modeling. Academic Press.
Golino, H. F., & Epskamp, S. (2017). Exploratory graph analysis: A new approach for estimating the number of dimensions in psychological research. PLOS ONE, 12(6), e0174035. https://doi.org/10.1371/journal.pone.0174035
Golino, H., & Christensen, A. (2026). EGAnet: Exploratory graph analysis—a framework for estimating the number of dimensions in multivariate data using network psychometrics (Version 2.5.0). https://r-ega.net
Golino, H., & Christensen, A. P. (2023). EGAnet: Exploratory graph analysis – a framework for estimating the number of dimensions in multivariate data using network psychometrics. R package.
Golino, H., Moulder, R., Shi, D., Christensen, A. P., Garrido, L. E., Nieto, M. D., Sadana, R., Thiyagarajan, J. A., Martínez-Molina, A., & Boker, S. M. (2021). Entropy fit indices: New fit measures for assessing the structure and dimensionality of multiple latent variables. Multivariate Behavioral Research, 56(6), 874–902. https://doi.org/10.1080/00273171.2020.1779642
Golino, H., Shi, D., Christensen, A. P., Garrido, L. E., Nieto, M. D., Sadana, R., Thiyagarajan, J. A., & Martínez-Molina, A. (2020). Investigating the performance of exploratory graph analysis and traditional techniques to identify the number of latent factors: A simulation and tutorial. Psychological Methods, 25(3), 292–320. https://doi.org/10.1037/met0000255
Graham, J. M. (2006). Congeneric and (essentially) tau-equivalent estimates of score reliability: What they are and how to use them. Educational and Psychological Measurement, 66(6), 930–944. https://doi.org/10.1177/0013164406288165
Graziano, W. G., & Tobin, R. M. (2002). Agreeableness: Dimension of personality or social desirability artifact?. Journal of Personality, 70(5), 695–728. https://doi.org/10.1111/1467-6494.05021
Green, S. B., Redell, N., Thompson, M. S., & Levy, R. (2016). Accuracy of revised and traditional parallel analyses for assessing dimensionality with binary data. Educational and Psychological Measurement, 76(1), 5–21. https://doi.org/10.1177/0013164415581898
Hair, J. F., Black, W. C., Babin, B. J., Anderson, R. E., & Tatham, R. L. (2006). Multivariate data analysis (6th ed.). Pearson Prentice Hall.
Hambleton, R. K., Swaminathan, H., & Rogers, H. J. (1991). Fundamentals of item response theory. SAGE.
Hancock, G. R. (1997). Structural equation modeling methods of hypothoesis testing of latent variable means. Measurement & Evaluation in Counseling & Development (American Counseling Association), 30(2), 91–105. https://doi.org/10.1080/07481756.1997.12068926
Harvill, L. M. (1991). Standard error of measurement. Educational Measurement: Issues and Practice, 10(2), 33–41. https://doi.org/10.1111/j.1745-3992.1991.tb00195.x
Haslam, N., Holland, E., & Kuppens, P. (2012). Categories versus dimensions in personality and psychopathology: A quantitative review of taxometric research. Psychological Medicine, 42(5), 903–920. https://doi.org/10.1017/S0033291711001966
Hastings, W. K. (1970). Monte carlo sampling methods using markov chains and their applications. Biometrika, 57(1), 97–109. https://doi.org/10.1093/biomet/57.1.97
Hebert, J. R., Ma, Y., Clemow, L., Ockene, I. S., Saperia, G., Stanek, E. J., Merriam, P. A., & Ockene, J. K. (1997). Gender differences in social desirability and social approval bias in dietary self-report. American Journal of Epidemiology, 146(12), 1046–1055. https://doi.org/10.1093/oxfordjournals.aje.a009233
Hinz, A., Michalski, D., Schwarz, R., & Herzberg, P. Y. (2007). The acquiescence effect in responding to a questionnaire. GMS Psycho-Social Medicine, 4.
Holzinger, K. J., & Swineford, F. (1939). A study in factor analysis: The stability of a bi-factor solution. University of Chicago Press.
Horn, J. L. (1965). A rationale and test for the number of factors in factor analysis. Psychometrika, 30(2), 179–185. https://doi.org/10.1007/BF02289447
Hu, L., & Bentler, P. M. (1999). Cutoff criteria for fit indexes in covariance structure analysis: Conventional criteria versus new alternatives. Structural Equation Modeling: A Multidisciplinary Journal, 6(1), 1–55. https://doi.org/10.1080/10705519909540118
Hughes, G. D. (2009). The impact of incorrect responses to reverse-coded survey items. Research in the Schools.
Jonas, K. G., & Markon, K. E. (2016). A descriptivist approach to trait conceptualization and inference. Psychological Review, 123(1), 90–96. https://doi.org/10.1037/a0039542
Jorgensen, T. D., Pornprasertmanit, S., Schoemann, A. M., & Rosseel, Y. (2022). semTools: Useful tools for structural equation modeling. CRAN. https://CRAN.R-project.org/package=semTools
Kaiser, H. F. (1960). The application of electronic computers to factor analysis. Educational and Psychological Measurement, 20(1), 141–151. https://doi.org/10.1177/001316446002000116
Kaiser, H. F. (1970). A second generation little jiffy. Psychometrika, 35(4), 401–415. https://doi.org/10.1007/BF02291817
Kaiser, H. F., & Rice, J. (1974). Little jiffy, mark IV. Educational and Psychological Measurement, 34(1), 111–117. https://doi.org/10.1177/001316447403400115
Kam, C., Schermer, J. A., Harris, J., & Vernon, P. A. (2013). Heritability of acquiescence bias and item keying response style associated with the HEXACO personality scale. Twin Research and Human Genetics, 16(4), 790–798.
Kam, C., Zhou, X., Zhang, X., & Ho, M. Y. (2012). Examining the dimensionality of self-construals and individualistic–collectivistic values with random intercept item factor analysis. Personality and Individual Differences, 53(6), 727–733. https://doi.org/10.1016/j.paid.2012.05.023
Kan, K.-J., Jonge, H. de, Maas, H. L. J. van der, Levine, S. Z., & Epskamp, S. (2020). How to compare psychometric factor and network models. Journal of Intelligence, 8(4), 35. https://doi.org/10.3390/jintelligence8040035
Kane, M. T. (2001). Current concerns in validity theory. Journal of Educational Measurement, 38(4), 319–342. https://doi.org/10.1111/j.1745-3984.2001.tb01130.x
Kane, M. T. (2006). Validation. In R. L. Brennan (Ed.), Educational measurement (4th ed., pp. 17–64). American Council on Education; Praeger.
Kane, M. T. (2013). Validating the interpretations and uses of test scores. Journal of Educational Measurement, 50(1), 1–73. https://doi.org/10.1111/jedm.12000
Kane, M. T. (2017). Causal interpretations of psychological attributes. Measurement: Interdisciplinary Research and Perspectives, 15, 79–82. https://doi.org/10.1080/15366367.2017.1369771
Kant, I. (1786). Metaphysical foundations of natural science.
Karabatsos, G. (2001). The rasch model, additive conjoint measurement, and new models of probabilistic measurement theory. Journal of Applied Measurement, 2(4), 389–423.
King, M. F., & Bruner, G. C. (2000). Social desirability bias: A neglected aspect of validity testing. Psychology & Marketing, 17(2), 79–103. https://doi.org/10.1002/(SICI
Kline, P. (1998). The new psychometrics: Science, psychology and measurement. Psychology Press.
Kline, P. (2000). The handbook of psychological testing (2nd ed.). Routledge.
Kline, R. B. (2023). Principles and practice of structural equation modeling (5th ed.). Guilford Press.
Knight, R. G., Chisholm, B. J., Marsh, N. V., & Godfrey, H. P. (1988). Some normative, reliability, and factor analytic data for the revised UCLA loneliness scale. Journal of Clinical Psychology, 44(2), 203–206. https://doi.org/10.1002/1097-4679(198803
Krantz, D. H., Luce, R. D., Suppes, P., & Tversky, A. (1971). Foundations of measurement, volume i: Additive and polynomial representations. Academic Press.
Kruis, J., & Maris, G. (2016). Three representations of the ising model. Scientific Reports, 6, 34175. https://doi.org/10.1038/srep34175
Külpe, O. (1895). Outlines of psychology. Sonnenschein.
Kyngdon, A. (2008a). Conjoint measurement, error and the rasch model: A reply to michell, and borsboom and zand scholten. Theory & Psychology, 18(1), 125–131. https://doi.org/10.1177/0959354307086927
Kyngdon, A. (2008b). The rasch model from the perspective of the representational theory of measurement. Theory & Psychology, 18(1), 89–109. https://doi.org/10.1177/0959354307086924
Lakatos, I., & Musgrave, A. (Eds.). (1970). Criticism and the growth of knowledge: Proceedings of the international colloquium in the philosophy of science, london 1965, volume 4. Cambridge University Press. https://doi.org/10.1017/CBO9781139171434
Lange, F., & Dewitte, S. (2019). Measuring pro-environmental behavior: Review and recommendations. Journal of Environmental Psychology, 63, 92–100. https://doi.org/10.1016/j.jenvp.2019.04.009
Lanz, L., Thielmann, I., & Gerpott, F. H. (2022). Are social desirability scales desirable? A meta‐analytic test of the validity of social desirability scales in the context of prosocial behavior. Journal of Personality, 90(2), 203–221. https://doi.org/10.1111/jopy.12662
Leite, W. L., & Cooper, L. A. (2010). Detecting social desirability bias using factor mixture models. Multivariate Behavioral Research, 45(2), 271–293. https://doi.org/10.1080/00273171003680245
Li, A., & Bagger, J. (2006). Using the BIDR to distinguish the effects of impression management and self‐deception on the criterion validity of personality measures: A meta‐analysis. International Journal of Selection and Assessment, 14(2), 131–141. https://doi.org/10.1111/j.1468-2389.2006.00339.x
Linden, W. J. van der, & Hambleton, R. K. (Eds.). (1997). Handbook of modern item response theory. Springer.
Loevinger, J. (1948). The technique of homogenous tests compared with some aspects of “scale analysis” and factor analysis. Psychological Bulletin, 45(6), 507–530. https://doi.org/10.1037/h0055827
Loken, E., & Molenaar, P. C. M. (2008). Categories or continua? The correspondence between mixture models and factor models. In G. R. Hancock & K. M. Samuelsen (Eds.), Advances in latent variable mixture models (pp. 277–298). Information Age Publishing.
Lord, F. M., & Novick, M. R. (1968). Statistical theories of mental test scores. Addison-Wesley.
Lubke, G. H., & Muthén, B. O. (2004). Applying multigroup confirmatory factor models for continuous outcomes to likert scale data complicates meaningful group comparisons. Structural Equation Modeling, 11(4), 514–534. https://doi.org/10.1207/s15328007sem1104_2
Luce, R. D., & Tukey, J. W. (1964). Simultaneous conjoint measurement: A new type of fundamental measurement. Journal of Mathematical Psychology, 1, 1–27. https://doi.org/10.1016/0022-2496(64)90015-X
Maraun, M. D., & Gabriel, S. M. (2013). Illegitimate concept equating in the partial fusion of construct validation theory and latent variable modeling. New Ideas in Psychology, 31(1), 32–42. https://doi.org/10.1016/j.newideapsych.2011.02.006
Maraun, M. D., & Peters, J. (2005). What does it mean that an issue is conceptual in nature? Journal of Personality Assessment, 85(2), 128–133. https://doi.org/10.1207/s15327752jpa8502_04
Marsh, H. W. (1996). Positive and negative global self-esteem: A substantively meaningful distinction or artifactors?. Journal of Personality and Social Psychology, 70(4), 810–819. https://doi.org/10.1037/0022-3514.70.4.810
Maul, A. (2017). Moving beyond traditional methods of survey validation. Measurement: Interdisciplinary Research and Perspectives, 15, 103–109. https://doi.org/10.1080/15366367.2017.1369786
McDonald, R. P. (1999). Test theory: A unified treatment. Lawrence Erlbaum Associates.
McFarland, D. J. (2020). The effects of using partial or uncorrected correlation matrices when comparing network and latent variable models. Journal of Intelligence, 8(1), 7. https://doi.org/10.3390/jintelligence8010007
McFarland, L. A., & Ryan, A. M. (2000). Variance in faking across noncognitive measures. Journal of Applied Psychology, 85(5), 812–821. https://doi.org/10.1037/0021-9010.85.5.812
Meehl, P. E. (1995). Bootstraps taxometrics: Solving the classification problem in psychopathology. American Psychologist, 50(4), 266–275. https://doi.org/10.1037/0003-066X.50.4.266
Mellenbergh, G. J. (1989). Item bias and item response theory. International Journal of Educational Research, 13(2), 127–143. https://doi.org/10.1016/0883-0355(89)90002-5
Meredith, W. (1993). Measurement invariance, factor analysis and factorial invariance. Psychometrika, 58, 525–543. https://doi.org/10.1007/BF02294825
Meredith, W., & Tisak, J. (1990). Latent curve analysis. Psychometrika, 55, 107–122. https://doi.org/10.1007/BF02294746
Messick, S. (1989). Validity. In R. L. Linn (Ed.), Educational measurement (3rd ed., pp. 13–103). American Council on Education; Macmillan.
Metropolis, N., Rosenbluth, A. W., Rosenbluth, M. N., Teller, A. H., & Teller, E. (1953). Equation of state calculations by fast computing machines. The Journal of Chemical Physics, 21(6), 1087–1092. https://doi.org/10.1063/1.1699114
Michell, J. (2000). Normal science, pathological science and psychometrics. Theory & Psychology, 10(5), 639–667. https://doi.org/10.1177/0959354300105004
Michell, J. (2001). Teaching and misteaching measurement in psychology. Australian Psychologist, 36(3), 211–218. https://doi.org/10.1080/00050060108259657
Michell, J. (2004). Item response models, pathological science and the shape of error: Reply to borsboom and mellenbergh. Theory & Psychology, 14(1), 121–129. https://doi.org/10.1177/0959354304040201
Michell, J. (2008). Is psychometrics pathological science? Measurement: Interdisciplinary Research and Perspectives, 6(1–2), 7–24. https://doi.org/10.1080/15366360802035489
Michell, J. (2014). An introduction to the logic of psychological measurement. Psychology Press.
Millsap, R. E., & Everson, H. T. (1993). Methodology review: Statistical approaches for assessing measurement bias. Applied Psychological Measurement, 17(4), 297–334. https://doi.org/10.1177/014662169301700401
Mirowsky, J., & Ross, C. E. (1991). Eliminating defense and agreement bias from measures of the sense of control: A 2 x 2 index. Social Psychology Quarterly, 54(2), 127. https://doi.org/10.2307/2786931
Mokken, R. J. (1971). A theory and procedure of scale analysis. De Gruyter. https://doi.org/10.1515/9783110813203
Molenaar, I. W. (1991). A weighted loevinger h-coefficient extending mokken scaling to multicategory items. Kwantitatieve Methoden, 37, 97–117.
Munafò, M. R., Nosek, B. A., Bishop, D. V. M., Button, K. S., Chambers, C. D., Percie du Sert, N., Simonsohn, U., Wagenmakers, E.-J., Ware, J. J., & Ioannidis, J. P. A. (2017). A manifesto for reproducible science. Nature Human Behaviour, 1(1), 0021. https://doi.org/10.1038/s41562-016-0021
Muthén, B. O., & Muthén, L. K. (2000). Integrating person-centered and variable-centered analyses: Growth mixture modeling with latent trajectory classes. Alcoholism: Clinical and Experimental Research, 24(6), 882–891. https://doi.org/10.1111/j.1530-0277.2000.tb02070.x
Muthén, B., & Kaplan, D. (1985). A comparison of some methodologies for the factor analysis of non-normal likert variables. British Journal of Mathematical and Statistical Psychology, 38, 171–189. https://doi.org/10.1111/j.2044-8317.1985.tb00832.x
Narens, L., & Luce, R. D. (1993). Further comments on the “nonrevolution” arising from axiomatic measurement theory. Psychological Science, 4(2), 127–130. https://doi.org/10.1111/j.1467-9280.1993.tb00475.x
Navarro-Gonzalez, D., Vigil-Colet, A., Ferrando, P. J., Lorenzo-Seva, U., & Tendeiro, J. N. (2021). Vampyr: Factor analysis controlling the effects of response bias. CRAN. https://CRAN.R-project.org/package=vampyr
Nederhof, A. J. (1985). Methods of coping with social desirability bias: A review. European Journal of Social Psychology, 15, 263–280. https://doi.org/10.1002/ejsp.2420150303
Newman, M. E. J. (2006). Modularity and community structure in networks. Proceedings of the National Academy of Sciences, 103(23), 8577–8582. https://doi.org/10.1073/pnas.0601602103
Newton, P. E., & Shaw, S. D. (2013). Validity in educational and psychological assessment. SAGE.
Nickerson, C. A., & McClelland, G. H. (1984). Scaling distortion in numerical conjoint measurement. Applied Psychological Measurement, 8(2), 183–198. https://doi.org/10.1177/014662168400800207
Noble, T., Rosebery, A., Suarez, C., Warren, B., & O’Connor, M. C. (2014). Science assessments and english language learners: Validity evidence based on response processes. Applied Measurement in Education, 27(4), 248–260. https://doi.org/10.1080/08957347.2014.944309
Nosek, B. A., Hardwicke, T. E., Moshontz, H., Allard, A., Corker, K. S., Dreber, A., Fidler, F., Hilgard, J., Kline Struhl, M., Nuijten, M. B., Rohrer, J. M., Romero, F., Scheel, A. M., Scherer, L. D., Schönbrodt, F. D., & Vazire, S. (2022). Replicability, robustness, and reproducibility in psychological science. Annual Review of Psychology, 73(1), 719–748. https://doi.org/10.1146/annurev-psych-020821-114157
Nosek, B. A., Spies, J. R., & Motyl, M. (2012). Scientific utopia: II. Restructuring incentives and practices to promote truth over publishability. Perspectives on Psychological Science, 7(6), 615–631. https://doi.org/10.1177/1745691612459058
Nowick, K., Gernat, T., Almaas, E., & Stubbs, L. (2009). Differences in human and chimpanzee gene expression patterns define an evolving network of transcription factors in brain. Proceedings of the National Academy of Sciences, 106(52), 22358–22363. https://doi.org/10.1073/pnas.0911376106
Oberski, D. (2016). Mixture models: Latent profile and latent class analysis. In J. Robertson & M. Kaptein (Eds.), Modern statistical methods for HCI (pp. 275–287). Springer. https://doi.org/10.1007/978-3-319-26633-6_12
Ones, D. S., Viswesvaran, C., & Reiss, A. D. (1996). Role of social desirability in personality testing for personnel selection: The red herring. Journal of Applied Psychology, 81(6), 660. https://doi.org/10.1037/0021-9010.81.6.660
Osborne, J. W. (2014). Best practices in exploratory factor analysis. CreateSpace Independent Publishing.
Pasquali, L. (2017). Psicometria: Teoria dos testes na psicologia e na educação. Editora Vozes Limitada.
Paulhus, D. L. (1984). Two-component models of socially desirable responding. Journal of Personality and Social Psychology, 46(3), 598–609. https://doi.org/10.1037/0022-3514.46.3.598
Paulhus, D. L. (1991). Measurement and control of response bias. In Measures of personality and social psychological attitudes (pp. 17–59). Academic Press. https://doi.org/10.1016/B978-0-12-590241-0.50006-X
Paulhus, D. L., & John, O. P. (1998). Egoistic and moralistic biases in self-perception: The interplay of self-deceptive styles with basic traits and motives. Journal of Personality, 66(6), 1025–1060. https://doi.org/10.1111/1467-6494.00041
Peabody, D. (1967). Trait inferences: Evaluative and descriptive aspects. Journal of Personality and Social Psychology, 4, Pt. https://doi.org/10.1037/h0025230
Pearl, J. (2009). Causality: Models, reasoning, and inference (2nd ed.). Cambridge University Press.
Pearson, K. (1978). The history of statistics in the seventeenth and eighteenth centuries. Griffin.
Peterson, R. A., & Kerin, R. A. (1981). The quality of self-report data: Review and synthesis. In Review of marketing (pp. 5–20). American Marketing Association.
Pettersson, E., Mendle, J., Turkheimer, E., Horn, E. E., Ford, D. C., Simms, L. J., & Clark, L. A. (2014). Do maladaptive behaviors exist at one or both ends of personality traits? Psychological Assessment, 26(2), 433–446. https://doi.org/10.1037/a0035587
Pettersson, E., Turkheimer, E., Horn, E. E., & Menatti, A. R. (2012). The general factor of personality and evaluation. European Journal of Personality, 26(3), 292–302. https://doi.org/10.1002/per.839
Podsakoff, P. M., MacKenzie, S. B., Lee, J. Y., & Podsakoff, N. P. (2003). Common method biases in behavioral research: A critical review of the literature and recommended remedies. Journal of Applied Psychology, 88(5), 879–903. https://doi.org/10.1037/0021-9010.88.5.879
Pons, P., & Latapy, M. (2006). Computing communities in large networks using random walks. Journal of Graph Algorithms and Applications, 10(2), 191–218. https://doi.org/10.7155/jgaa.00124
Ramsay, J. O. (1975). Review of foundations of measurement, volume i. Psychometrika, 40, 257–262.
Ramsay, J. O. (1991). Reviews of foundations of measurement, volumes II and III. Psychometrika, 56, 355–358.
Rasch, G. (1960). Probabilistic models for some intelligence and attainment tests. Danish Institute for Educational Research.
Raymark, P. H., & Tafero, T. L. (2009). Individual differences in the ability to fake on personality measures. Human Performance, 22(1), 86–103. https://doi.org/10.1080/08959280802541039
Revelle, W. (2023a). Psych: Procedures for psychological, psychometric, and personality research. Northwestern University. https://CRAN.R-project.org/package=psych
Revelle, W. (2023b). psychTools: Tools to accompany the ’psych’ package for psychological research. Northwestern University. https://CRAN.R-project.org/package=psychTools
Riggio, R. E., Salinas, C., & Tucker, J. (1988). Personality and deception ability. Personality and Individual Differences, 9(1), 189–191. https://doi.org/10.1016/0191-8869(88
Robinson, J. P., Shaver, P. R., & Wrightsman, L. S. and. (1991). Measures of social psychological attitudes. (Vol. 1. Measures of personality; social psychological atitudes). Academic Press.
Rosenberg, M. (1965). Society and the adolescent self-image. Princeton University Press.
Rosseel, Y. (2012). Lavaan: An r package for structural equation modeling. Journal of Statistical Software, 48(2), 1–36. https://doi.org/10.18637/jss.v048.i02
Ruscio, J., & Kaczetow, W. (2009). Differentiating categories and dimensions: Evaluating the robustness of taxometric analyses. Multivariate Behavioral Research, 44(2), 259–280. https://doi.org/10.1080/00273170902794248
Ruscio, J., Ruscio, A. M., & Meron, M. (2007). Applying the bootstrap to taxometric analysis: Generating empirical sampling distributions to help interpret results. Multivariate Behavioral Research, 42(2), 349–386. https://doi.org/10.1080/00273170701360795
Salazar, M. S. (2015). The dilemma of combining positive and negative items in scales. Psicothema, 27(2), 192–199. https://doi.org/10.7334/psicothema2014.266
Saucier, G., Ostendorf, F., & Peabody, D. (2001). The non-evaluative circumplex of personality adjectives. Journal of Personality, 69(4), 537–582. https://doi.org/10.1111/1467-6494.694155
Savalei, V., & Falk, C. F. (2014). Recovering substantive factor loadings in the presence of acquiescence bias: A comparison of three approaches. Multivariate Behavioral Research, 49(5), 407–424. https://doi.org/10.1080/00273171.2014.931800
Scheiblechner, H. (1995). Isotonic ordinal probabilistic models (ISOP). Psychometrika, 60(2), 281–304.
Scheiblechner, H. (1998). Corrections of theorems in scheiblechner’s treatment of ISOP models and comments on junker’s remarks. Psychometrika, 63(1), 87–91. https://doi.org/10.1007/BF02295439
Scheiblechner, H. (1999). Additive conjoint isotonic probabilistic models (ADISOP). Psychometrika, 64(3), 295–316. https://doi.org/10.1007/BF02294297
Schimmack, U. (2021). The validation crisis in psychology. Meta-Psychology, 5.
Schmalz, X., Breuer, J., Haim, M., Hildebrandt, A., Knöpfle, P., Leung, A. Y., & Roettger, T. B. (2024). Let’s talk about language—and its role for replicability. https://osf.io/preprints/metaarxiv/w2gb9
Schmittmann, V. D., Cramer, A. O. J., Waldorp, L. J., Epskamp, S., Kievit, R. A., & Borsboom, D. (2013). Deconstructing the construct: A network perspective on psychological phenomena. New Ideas in Psychology, 31(1), 43–53. https://doi.org/10.1016/j.newideapsych.2011.02.007
Schriesheim, C. A., & Hill, K. D. (1981). Controlling acquiescence response bias by item reversals: The effect on questionnaire validity. Educational and Psychological Measurement, 41(4), 1101–1114. https://doi.org/10.1177/001316448104100420
Schwager, K. W. (1991). The representational theory of measurement: An assessment. Psychological Bulletin, 110(3), 618–626. https://doi.org/10.1037/0033-2909.110.3.618
Selau, T., Silva, M. A. da, Mendonça Filho, E. J. de, & Bandeira, D. R. (2020). Evidence of validity and reliability of the adaptive functioning scale for intellectual disability (EFA-DI). Psicologia: Reflexão e Crítica, 33, 26. https://doi.org/10.1186/s41155-020-00164-7
Sijtsma, K. (2009). On the use, the misuse, and the very limited usefulness of cronbach’s alpha. Psychometrika, 74(1), 107–120. https://doi.org/10.1007/s11336-008-9101-0
Sijtsma, K. (2012). Psychological measurement between physics and statistics. Theory & Psychology, 22(6), 786–809. https://doi.org/10.1177/0959354312454353
Sijtsma, K., & Molenaar, I. W. (2002). Introduction to nonparametric item response theory. Sage. https://doi.org/10.4135/9781412984676
Snell, A. F., Sydell, E. J., & Lueke, S. B. (1999). Towards a theory of applicant faking: Integrating studies of deception. Human Resource Management Review, 9(2), 219–242. https://doi.org/10.1016/S1053-4822(99
Soto, C. J., John, O. P., Gosling, S. D., & Potter, J. (2008). The developmental psychometrics of big five self-reports: Acquiescence, factor structure, coherence, and differentiation from ages 10 to 20. Journal of Personality and Social Psychology, 94(4), 718–737. https://doi.org/10.1037/0022-3514.94.4.718
Spearman, C. (1904). “General intelligence,” objectively determined and measured. The American Journal of Psychology, 15(2), 201–292. https://doi.org/10.2307/1412107
Spearman, C. (1910). Correlation calculated from faulty data. British Journal of Psychology, 3(3), 271–295. https://doi.org/10.1111/j.2044-8295.1910.tb00206.x
Spearman, C. (1937). Psychology down the ages (Vol. 1). Macmillan.
Steiger, J. H., & Lind, J. C. (1980). Statistically based tests for the number of common factors.
Stevens, S. S. (1946). On the theory of scales of measurement. Science, 103(2684), 677–680. https://doi.org/10.1126/science.103.2684.677
Stevens, S. S. (1951). Mathematics, measurement and psychophysics. In S. S. Stevens (Ed.), Handbook of experimental psychology (pp. 1–49). Wiley.
Stevens, S. S. (1959). Measurement, psychophysics and utility. In C. W. Churchman & P. Ratoosh (Eds.), Measurement: Definitions and theories (pp. 18–63). Wiley.
Straat, J. H., Ark, L. A. van der, & Sijtsma, K. (2013). Comparing optimization algorithms for item selection in mokken scale analysis. Journal of Classification, 30, 72–99. https://doi.org/10.1007/s00357-013-9122-y
Suárez-Alvarez, J., Pedrosa, I., Lozano Fernández, L. M., García-Cueto, E., Cuesta, M., & Muñiz, J. (2018). Using reversed items in likert scales: A questionable practice. Psicothema, 30(2), 149–158. https://doi.org/10.7334/psicothema2018.33
Svetina, D., Rutkowski, L., & Rutkowski, D. (2020). Multiple-group invariance with categorical outcomes using updated guidelines: An illustration using m plus and the lavaan/semtools packages. Structural Equation Modeling: A Multidisciplinary Journal, 27(1), 111–130. https://doi.org/10.1080/10705511.2019.1602776
Tabachnick, B. G., & Fidell, L. S. (2007). Using multivariate statistics (5th ed.). Allyn & Bacon.
Ten Berge, J. M. F., & Kiers, H. A. L. (1991). A numerical approach to the approximate and the exact minimum rank of a covariance matrix. Psychometrika, 56, 309–315. https://doi.org/10.1007/BF02294464
Thomson, W. (1891). Popular lectures and addresses (Vol. 1). Macmillan.
Titchener, E. B. (1905). Experimental psychology: A manual of laboratory practice. Macmillan.
Tourangeau, R., & Yan, T. (2007). Sensitive questions in surveys. Psychological Bulletin, 133(5), 859–883. https://doi.org/10.1037/0033-2909.133.5.859
Trendler, G. (2009). Measurement theory, psychology and the revolution that cannot happen. Theory & Psychology, 19(5), 579–599. https://doi.org/10.1177/0959354309341926
Trendler, G. (2013). Measurement in psychology: A case of Ignoramus et Ignorabimus? A rejoinder. Theory & Psychology, 23(5), 591–615. https://doi.org/10.1177/0959354313490451
Tucker, L. R., & Lewis, C. (1973). A reliability coefficient for maximum likelihood factor analysis. Psychometrika, 38(1), 1–10. https://doi.org/10.1007/BF02291170
Uher, J. (2021). Problematic research practices in psychology: Misconceptions about data collection entail serious fallacies in data analysis. Theory & Psychology, 31(3), 411–416. https://doi.org/10.1177/09593543211014963
Uher, J. (2022). Rating scales institutionalise a network of logical errors and conceptual problems in research practices: A rigorous analysis showing ways to tackle psychology’s crises. Frontiers in Psychology, 13, 1009893. https://doi.org/10.3389/fpsyg.2022.1009893
Uher, J. (2023). What’s wrong with rating scales? Psychology’s replication and confidence crisis cannot be solved without transparency in data generation. Social and Personality Psychology Compass, 17(5), e12740. https://doi.org/10.1111/spc3.12740
Valentini, F. (2017). Influência e controle da aquiescência na análise fatorial [influence and control of acquiesncence in factor analysis]. Avaliação Psicológica, 16(2), 6–12. https://doi.org/10.15689/ap.2017.1602.ed
Valentini, F., & Hauck Filho, N. (2020). O impacto da aquiescência na estimação de coeficientes de validade [acquiescence impact in the estimation of validity coefficients]. Avaliação Psicológica, 19(1), 1–3. http://dx.doi.org/10.15689/ap.2020.1901.ed
Van Sonderen, E., Sanderman, R., & Coyne, J. C. (2013). Ineffectiveness of reverse wording of questionnaire items: Let’s learn from cows in the rain. PloS One, 8(7), e68967. https://doi.org/10.1371/journal.pone.0068967
Vecina, M. L., Chacón, F., & Pérez-Viejo, J. M. (2016). Moral absolutism, self-deception, and moral self-concept in men who commit intimate partner violence: A comparative study with an opposite sample. Violence Against Women, 22(1), 3–16. https://doi.org/10.1177/1077801215597791
Verweij, A. C., Sijtsma, K., & Koops, W. (1996). A mokken scale for transitive reasoning suited for longitudinal research. International Journal of Behavioral Development, 19(1), 219–238. https://doi.org/10.1177/016502549601900115
Vries, R. E. de, Zettler, I., & Hilbig, B. E. (2014). Rethinking trait conceptions of social desirability scales: Impression management as an expression of honesty-humility. Assessment, 21(3), 286–299. https://doi.org/10.1177/1073191113504619
Weijters, B., & Baumgartner, H. (2012). Misresponse to reversed and negated items in surveys: A review. Journal of Marketing Research, 49(5), 737–747. https://doi.org/10.1509/jmr.11.0368
Weijters, B., Cabooter, E., & Schillewaert, N. (2010). The effect of rating scale format on response styles: The number of response categories and response category labels. International Journal of Research in Marketing, 27(3), 236–247. https://doi.org/10.1016/j.ijresmar.2010.02.004
Williams, E. A., Pillai, R., Lowe, K. B., Jung, D., & Herst, D. (2009). Crisis, charisma, values, and voting behavior in the 2004 presidential election. The Leadership Quarterly, 20(2), 70–86. https://doi.org/10.1016/j.leaqua.2009.01.002
Wolf, M. G., Ihm, E., Maul, A., & Taves, A. (2025). The response-process-evaluation method: A new approach to survey-item validation. Advances in Methods and Practices in Psychological Science, 8(3), 1–20. https://doi.org/10.1177/25152459251353135
Woods, C. M. (2006). Careless responding to reverse-worded items: Implications for confirmatory factor analysis. Journal of Psychopathology and Behavioral Assessment, 28(3), 186–191. https://doi.org/10.1007/s10862-005-9004-7
Wu, H., & Estabrook, R. (2016). Identification of confirmatory factor analysis models of different levels of invariance for ordered categorical outcomes. Psychometrika, 81(4), 1014–1045. https://doi.org/10.1007/s11336-016-9506-0
Xia, Y., & Yang, Y. (2019). RMSEA, CFI, and TLI in structural equation modeling with ordered categorical data: The story they tell depends on the estimation methods. Behavior Research Methods, 51, 409–428. https://doi.org/10.3758/s13428-018-1055-2
Zhang, X., & Savalei, V. (2016). Improving the factor structure of psychological scales: The expanded format as an alternative to the likert scale format. Educational and Psychological Measurement, 76(3), 357–386. https://doi.org/10.1177/0013164415596421
Zickar, M. J., & Robie, C. (1999). Modeling faking good on personality items: An item-level analysis. Journal of Applied Psychology, 84(4), 551. https://doi.org/10.1037/0021-9010.84.4.551
Ziegler, M., & Buehner, M. (2009). Modeling socially desirable responding and its effects. Educational and Psychological Measurement, 69(4), 548–565. https://doi.org/10.1177/0013164408324469
Ziegler, M., Maccann, C., & Roberts, R. D. (2011). New perspectives on faking in personality assessment. Oxford University Press.
Zwick, W. R., & Velicer, W. F. (1986). Comparison of five rules for determining the number of components to retain. Psychological Bulletin, 99(3), 432–442. https://doi.org/10.1037/0033-2909.99.3.432