Standardization
- Standardization of assessments refers to standard conditions on the basis that interpreting the scores can be generalized beyond the testing situation.1 2 3
- The process of developing a standardized measure is a rigorous process based on evidence of reliability, validity, and responsiveness to change.4
- Standardized assessments are often perceived to be indicators of professional status in regards to procedures such as scoring and test administration that are fixed and uniform throughout.5
- Limitations of using non-standardized assessments can lead to difficulty defining, documenting, communicating results, changes, and outcomes.6 7 8
- Research published in 2009 found that 35% more of assessments used across geriatrics, pediatrics, physical disability, mental health, and handy therapy were standardized.9
- Research from a survey of 143 Candian pediatric PTs found that less than 50% thought that standardized outcome measures were used in their departments, with the lack of knowledge of available instruments and their measurement properties being the primary barrier to adoption.10
- Lack of time was another frequently mentioned barrier to using standardized measures.11
- Organizational or peers may influence measurement practices among therapists.12
- A study conducted in 2009 of US occupational therapy practitioners found that the majority of assessments addressed body structure and functional impairments (bottom-up) rather than occupational performance – also due to time requirements, as well as ease of use and availability.13
- “Broadening of occupational therapy assessment practice to incorporate more occupation-based measures where appropriate continues to be an important issue for the profession to address…in addition to a receiving string foundation in measurement concepts, occupational therapy students should be taught strategies to incorporate occupation-based assessments in a variety of practice settings. Furthermore, continuing education workshops should emphasize how occupation-based measures may meet the assessment needs of practicing occupational therapists when responding to external requirements such as qualifying a
client for services.”14
Reliability
- The extent to which a measurement produces the same results over a period of time.15
- Internal reliability: the consistency of results within a test.
- Split-half reliability: the extent to which all parts of the test contribute equally to what is being measured.16
- External reliability: the extent to which a measure varies from one to another.
- Threats to reliability: researcher error, environmental changes, participant changes (client factors, e.g., visual acuity), fatigue, and discomfort.18
Validity
- The extent to which the instrument measures what it sets out to measure.15
- Internal validity refers specifically to whether an experimental intervention makes a difference, and whether there is sufficient evidence to support it.
- External validity refers to the generalizibility of the treatment/condition outcomes.
- Threats to validity19 20 21
- History – events which occur between measurements
- Maturation – passage of time of participants
- Testing – effects of taking a test on the outcomes of taking a second test
- Instrumentation – changes in the instrument, observer, or scorers
- Statistical regression – regression to the mean, selection of participants on the basis of extreme characteristics
- Selection of subjects – biases which may result during selection of comparison groups; countered by random assignment
- Experimental mortality – the loss of participants
- Selection-maturation interaction – selection of comparison groups which may lead to confounding outcomes and erroneous interpretation
- Threats to construct validity22 23
- Subject reactivity.- subjects behave in a particular way because they are aware of their role in a study, aka the Hawthrone effect
- Researcher expectancies – influence from the researcher on the participant through the communication of desired outcomes
- Novelty effects – altered behaviors from researchers and participants when treatment is new or novel
- Compensatory effects – compensating for failure to receiving a treatment
- Treatment diffusion (contamination) – alternative treatment is received while in the main study
Fidelity
Fidelity is the faithfulness of an intervention to its underlying therapeutic principles and clinical guidelines.24 High fidelity is necessary to make confident conclusions about the intervention in research.25
Fidelity consists of:26
- Adherence – the extent to which components are delivered as intended
- Quality of delivery – a subjective perspective of how the treatment extends beyond the delivery of prescribed content
- Exposure – the number, length, or frequency of intervention sessions that were implemented
- Participant responsiveness – participant’s judgments about the outcomes and relevance of an intervention27
- Program differentiation – how the intervention is being delivered is different from other interventions28
Research Design
- Nunnally JC (1972). Educational Measurement and Evaluation (2nd edn). New York: McGraw
Hill Book Co[↩] - Angoff WH, Anderson SB (1975). The standardization of educational and psychological tests. In Payne DA, McMorris RF (Eds) Educational and Psychological Measurement: Contributions to Theory and Practice. Morristown, NJ: General Learning Press, p. 397.[↩]
- Wiersma W, Jurs SG (1990). Educational Measurement and Testing (2nd edn). Massachusetts: Allyn and Bacon.[↩]
- Anastasi, A., & Urbina, S. (1997). Psychological testing (7th ed.). Saddle River, NJ: Prentice-Hall,
Inc.[↩] - Glueckauf RL, Sechrest LB, Bond GR, Mcdonel EC (1993). Improving Assessment in Rehabilitation and Health. Newbury Park: Sage[↩]
- Okkema K. Cognition and perception in the stroke patient. A guide to functional outcomes in occupational therapy. Maryland: Aspen Publications; 1993.[↩]
- Unsworth C. Cognitive and perceptual dysfunction. Philadelphia: F.A. Davis; 1999.[↩]
- 41. Matthey S, Donnelly SM, Hextell DL. The clinical usefulness of the Rivermead Perceptual Assessment Battery: statistical considerations. Br J Occup Ther. 1993;56:365/70[↩]
- Mohammed Alotaibi, N., Reed, K., & Shaban Nadar, M. (2009). Assessments used in occupational therapy practice: An exploratory study. Occupational therapy in health care, 23(4), 302-318.[↩]
- Cole, B., Finch, E., Gowland, C., & Mayo, N. (1994). Physical rehabilitation outcome measures. Toronto: Canadian Physiotherapy Association[↩]
- Kay, T., Myers, A., & Huijbregts, M. (2001). How far have we come since 1992? A comparative survey of physiotherapists’ use of outcome measures. Physiotherapy Canada, 53, 268–275.[↩]
- Hanna, S., Russell, D., Bartlett, D., Kertoy, M., Rosenbaum, P., & Wynn, K. (2007). Measurement practices in pediatric rehabilitation: A survey of physical therapists, occupational therapists, and speech-language pathologists in Ontario. Physical & Occupational Therapy in Pediatrics, 27(2), 25–42.[↩]
- Alotaibi, N., Reed, K., & Nadar, M. (2009). Assessments used in occupational therapy practice: An exploratory study. Occupational Therapy in Health Care, 23(4), 302–318.[↩]
- Piernik-Yoder, B., & Beck, A. (2012). The use of standardized assessments in occupational therapy in the United States. Occupational Therapy in Health Care, 26(2-3), 97-108.[↩]
- Michael J. Miller, Reliability and Validity. Western International University, 2000-2009.[↩][↩]
- McLeod, S. A. (2007). What is reliability?. Simply Psychology. https://www.simplypsychology.org/reliability.html[↩][↩]
- Drost E. Validity and reliability in social science research. Educ Res Perspect. 2011;38:105–124.[↩]
- Shaw, R. (1992). Nursing Research: Threats to Reliability and Validity Gerontology.[↩]
- Campbell, D. & Stanley, J. (1963). Experimental and quasi-experimental designs for research. Chicago, IL: Rand-McNally.[↩]
- Cook, T. D., & Campbell, D. T. (1979). Quasi-experimentation: Design and analysis issues for field settings. Boston, MA: Houghton Mifflin Company.[↩]
- Matthay, E. C., & Glymour, M. M. (2020). A graphical catalog of threats to validity: Linking social science with epidemiology. Epidemiology (Cambridge, Mass.), 31(3), 376.[↩]
- Conrad, K. J., & Conrad, K. M. (1994). Reassessing validity threats in experiments: Focus on construct validity. New Directions for Program Evaluation, 1994(63), 5-25.[↩]
- Neo, J. R. J. (2017). Construct validity—Current issues and recommendations for future hand hygiene research. American journal of infection control, 45(5), 521-527.[↩]
- Parham L. D., Cohn E. S., Spitzer S., Koomar J. A., Miller L. J., Burke J. P., . . . Summers C. A. (2007). Fidelity in sensory integration intervention research. American Journal of Occupational Therapy, 61, 216–227. 10.5014/ajot.61.2.216[↩]
- Nelson D. L., & Mathiowetz V. (2004). Randomized controlled trials to investigate occupational therapy research questions. American Journal of Occupational Therapy, 58, 24–34. 10.5014/ajot.58.1.24[↩]
- Dane A. V., & Schneider B. H. (1998). Program integrity in primary and early secondary prevention: Are implementation effects out of control? Clinical Psychology Review, 18, 23–45. 10.1016/S0272-7358(97)00043-3[↩]
- Carroll C., Patterson M., Wood S., Booth A., Rick J., & Balain S. (2007). A conceptual framework for implementation fidelity. Implementation Science, 2, 40 10.1186/1748-5908-2-40[↩]
- Dusenbury L., Brannigan R., Falco M., & Hansen W. B. (2003). A review of research on fidelity of implementation: Implications for drug abuse prevention in school settings. Health Education Research, 18, 237–256. 10.1093/her/18.2.237[↩]