In a previous publication in The Score, Ortiz (2017) highlighted a series of difficulties that characterize assessment of English learners (EL) which are largely rooted in the concept of validity, particularly test score validity. Standardized, norm-referenced tests are problematic for ELs primarily because of the disconnect between age and language development. Whereas age is a reliable indicator of development for individuals who have learned, or been exposed to, only one language, by definition, an EL began learning a second language at some point after having begun development in a different language presumably since birth. Thus, differences in the amount of time two 10-year-old ELs have been learning the same new or second language can vary from as little as one day to more than nine years. For ELs, age no longer controls for differences in language development in EL (or multilingual) populations in the way that it does for monolinguals. Failure to account for such potentially vast differences in development has long been and remains the Achille’s heel in the valid use of tests with ELs.
Differential language development and test performance
Issues that bear upon test score validity in the U.S. have invariably treated ELs as a rather monolithic group, particularly with respect to English and attempts to deal with it have been largely unsuccessful. ELs are often characterized in terms of racial or ethnic categories (e.g., Hispanics, Asian, etc.), by age groupings, or simply as a function of non-native English speaking status in a binary manner regardless of actual development (e.g., limited English proficient vs. fluent English speaker). For these and other reasons, including the idea that native-language testing might prove a viable solution despite the same developmental variability problem in the first language, research on the relationship between language development and test performance is exceedingly scarce. Of note, however, are advancements in the field of speech-language pathology where the topic has received considerably more attention. Publication of the Bilingual English-Spanish Assessment (BESA; Pena et al., 2018) represents the only formal effort to bind these variables and construct a norm sample that accounts for differences among bilinguals. Perhaps the most unique and innovative aspect of the BESA is the use of five categories to accommodate differences in language development for both English and Spanish, which range from functional English monolingual, to bilingual: English-dominant, balanced bilingual, bilingual: Spanish-dominant, functional Spanish monolingual. If one excludes the monolingual categories (which are much like existing test norm samples, albeit none include both English and the heritage language as the BESA does), what is left is a classification of language development that is composed of three general levels of language development and proficiency—from more English than Spanish, to equal English and Spanish, to more Spanish than English. In this way, along with other features of the test, the BESA provides a comprehensive examination of the overall and integrated language ability of pre-school/transition students (i.e., ages 4:0 to 6:11), as that is the age range of the test.
In psychology, these three classifications closely resemble the three classifications that have long been utilized in development and refinement of the Culture-Language Interpretive Matrix (C-LIM; Flanagan, Ortiz, & Alfonso, 2013), which relied not only on empirically-derived differences in subtest performance between English learners (EL) and English speakers (ES) in the available literature, but also on research that sought to examine differences in subtest performance between ELs and other ELs. One example can be seen in the work of Cathers-Schiffman and Thompson (2007) who demonstrated that ELs with average or higher English developmental proficiency (i.e., SS > 90, as measured by a language test), scored equally well across all aspects of the WISC-III and comparable to the mean (i.e., SS = 100) using the test’s norms which are based on monolingual, native-English speaker standards and expectations. However, those ELs with less English language development (SS < 90, as measured by a language test) scored 22 points lower on the Verbal IQ (VIQ) subtests and nine points lower on the Perceptual Organization (POI) subtests, suggesting that language development differentially affects performance relative to the construct being measured. They concluded that rather than treating language as a dichotomous variable, “whenever possible, researchers should conceptualize language fluency as a continuous variable and more closely examine the effects of different levels of English proficiency on intelligence test scores” (p. 51).
Along these same lines, Dynda (2008) and colleagues divided language proficiency into a more continuous variable comprising categories similar to that used in the BESA. In their study, as shown in Figure 1, the Woodcock-Munoz Language Survey-Revised (WMLS-R; Woodcock and Munoz, 2005) was used to assess developmental proficiency in the English language and create categories noted as low, intermediate, and high. Performance across the four subtests from the WMLS-R used to establish language proficiency were evaluated along with the four subtests from the Wechsler Abbreviated Scales of Intelligence (WASI; Wechsler, 1999) by arranging them along the abscissa with respect to the degree of English language development expected or required for responding and the extent to which language ability was the intended construct. The more “language-free” type tasks included Matrix Reasoning (WASI) and Block Design (WASI) on one end and the more “language-based” tasks included Vocabulary (WASI) and Picture Vocabulary (WMLS-R) on the other end. The results closely matched the findings of the Cathers-Schiffman and Thompson (2007) study and demonstrated that as subtests required or utilized more age-based language and acculturative knowledge acquisition, test performance decreased significantly and it decreased further as a function of developmental language proficiency.
In another later study within the same vein, Sotelo-Dynega et al. (2013) replicated the findings very closely by using a test mandated for limited English proficient students called the “New York as a Second Language Achievement Test”(NYSESLAT; NYSTP, 2009) to establish four different levels of developmental language proficiency, proficient, advanced, intermediate, and beginner. All groups were administered the first seven subtests which comprise the General Intellectual Ability (GIA) composite of the Woodcock Johnson III: Tests of Cognitive Ability (Woodcock, 2001). According to the results, not only did the overall ability correlate proportionally to the level of English development (i.e., Proficient SS=101, Advanced SS=89.55, Intermediate SS=82.29, and Beginner SS=71.75), evaluation of performance on the individual subtests demonstrated a similar pattern of decline and with varying levels of declination and attenuation on the basis of the construct being measured as illustrated in Figure 2. Tests that had little to do with language ability or that measured other abilities without relying much on language ability or knowledge (e.g., Spatial Relations, Visual Matching) produced scores that were comparable and well within the average range for all groups of ELs irrespective of developmental proficiency in English. However, as the subtests began to rely more on language ability for measuring the intended construct (e.g., Visual-Auditory Learning, Concept Formation) or when the target of the measurement was language itself (e.g., Verbal Comprehension), performance decreased significantly and still in proportion to English language development. Hence, those with more limited developmental proficiency in English were competitive with their same age EL peers who were found to be more proficient on the NYSESLAT on tests that did not involve much, if any language, but they scored progressively lower on tests that tapped more and more into language ability.
Collectively, these findings indicate that two types of comparisons must be considered in the attempt to measure performance and establish valid normative expectations for ELs. First, it is necessary to establish how ELs, as a general group, perform in comparison to monolingual, native-English speakers because test performance of ELs is moderated by the degree to which a given index or subtest relies on or requires age- or grade-expected English language development and the acquisition of incidental acculturative knowledge. And second, developmental language differences among ELs of the same age must also be accounted for because age-based comparisons no longer accurately reflect equivalent levels of development in either an ELs heritage or new language. A change in focus on norm sample construction that includes attention to variable development in language across all ages should serve as a roadmap that might lead to more effective tools and greater social justice.
Continuous, exposure-based norming
It was in consideration, and on the basis, of validity and fairness issues that the structure of the Ortiz Picture Vocabulary Acquisition Test (Ortiz PVAT; Ortiz, 2018) was predicated. The idea was to create a single, general test of language development and acquisition that could be used with anyone who was learning English, whether from birth or at some point after, no matter the age. Use of English as the target language preserved its relevancy for nearly any purpose with which it might be employed in the U.S. and eliminated the need to consider measurement of the heritage language. Moreover, receptive vocabulary acquisition was chosen as the central construct given its strong relationship to general language and intellectual development, academic outcomes in reading and writing, and the ability to measure language at its earliest stage of acquisition. The hope was that by entrenching fairness into the very conceptualization and foundations of the test, it might prove to be a more equitable measure of English vocabulary acquisition which has historically been subject to the most attenuation in test performance.
To achieve these goals, it was recognized that two distinct and separate norm samples were necessary, one for monolingual, native-English speakers, and one for everyone else. Significant care was taken to ensure that the monolingual, native-English speaking norm sample was composed only of individuals for whom there was no other language present to any significant degree. Similarly, equal care was given to stratifying English-language exposure in the EL sample across the full range of development and across each age level in the test (i.e., 2:6 to 22:11). In addition, in the EL sample, race/ethnicity was dropped in favor of heritage language given that it was believed that race/ethnicity offered little more than an illusory perception of representation and because the focus on language development in the EL sample suggested that variation in language spoken was likely a more relevant factor, albeit it was not clear whether the learning of one first language might interact differentially with learning English than the learning of a different first language. These concerns were laid to rest on both counts in analyses that demonstrated: 1) no significant differences in mean test performance among any of the different racial/ethnic groups in the English speaker sample, see Table 1; and 2) no significant differences in mean test performance among any of the different language groups in the English learner sample, see Table 2. In effect, ensuring true homogeneity of monolingualism in the ES sample eliminated any variance in test performance that might otherwise be ascribed to race/ethnicity and supports the notion that skin color is an irrelevant variable, although it is sometimes indirectly correlated with language differences. Likewise, it does not make any difference what language an individual has learned first when it comes to learning English. While there may be an advantage in the rate of learning or acquisition, depending on the similarity of the heritage language and English, there is no disruption of the developmental path or sequence in learning English. English is acquired (early in life) or learned (later) in essentially the same manner and in the same sequence regardless or age or prior exposure to another language. For example, learners of English initially acquire very common conversational words that are frequently encountered in everyday interactions (e.g., “yes,” “no,” “hello,” “goodbye”). English, like any other language, is not learned in random fashion and although there may be idiosyncratic variation in the specific words one acquires, the process remains quite similar, predictable, and largely invariant regardless of the learners first language or when learning began in that language. And while these analyses certainly attest to the intrinsic validity and fairness of the Ortiz PVAT, a different examination of data may better highlight the importance of language-exposure norms in promoting consequential validity and fairness in outcomes.
Test score interpretation
In 2018, a large suburban district in the southern part of the U.S. provided some preliminary data to the lead author asking for comments which is shown in Table 3. The district was utilizing the Woodcock-Munoz Language Survey-III (WMLS-III; Woodcock & Munoz, 2017) primarily as a way of determining dominance and hence the language for evaluation. The WMLS-III provides a general language score in both English and Spanish that may be used comparatively for evaluating language development in either language. The district had just begun to employ the Ortiz PVAT as an additional measure that might help inform the need and appropriateness of the referral. An examination of the Spanish scores from this small group (N=14) shows that all but two of the students scored quite poorly and well below normal limits with the highest of them being SS=73. Based on comparison to the English language scores, of the 12, only three were considered “dominant” in Spanish. The use of quotes here is to indicate that despite scoring higher than they did in English, these three students had dominance that was being defined as SS=73, SS=61, and SS=59; hardly the performance that one would expect to equate to age-appropriate language development. Assuming the five students who were dominant in Spanish were then evaluated further and comprehensively in Spanish, the outcomes are rather predictable. If all five were to score in the average range on everything else, then at worst three out of the five might be seen as having some type of speech-language impairment. If all five were to score average in some other areas but poorly in others, they all might be viewed as potentially have a learning disability. And if all five scored poorly on everything else that was administered, it would likely raise questions about intellectual disability. Apart from the two students who were average in their Spanish language development, the other three students were very much at risk for identification with some type of disability depending on the nature and pattern of the additional testing.
A similar thing happens when examining the English language scores. Of the 14 students, nine were found to be English “dominant.” Again, the use of quotes is to emphasize that the scores that supported the concept of dominance ranged from a high of SS=69 to a low of SS=45. Even the two Spanish-dominant students scored SS=45 and SS=43 respectively which does not bode well in terms of current academic success. To surmise that such scores might, in any way, render the results from testing in English as being valid is to strain all credulity. It is more likely that once again, if all other testing were within the average range, then each of the nine would be identified as having a speech-language impairment. If the other test data showed variable performances, some average, some not average, then learning disability would be in the discussion. And as before, poor performances across the entirety of the additional test administration would lead examiners to entertain notions of intellectual ability. Taken all together, the use of the WMLS-III which does not have exposure-based norms, would likely have led to 12 of the 14 students as being identified with some type of disability, irrespective of the language in which they were evaluated.
One might think that because this is a referred sample (the students were previously observed to have some type of difficulty with learning in the classroom), an 86 percent identification rate is not surprising or unreasonable. But examination of the concurrent scores from the Ortiz PVAT reveals a vastly different picture. Because the Ortiz PVAT uses exposure-based norms that account for differences in English language development among individuals of the same age, measured performance is startling different despite the fact that both tests measure language. On the Ortiz PVAT, only one student score at a level that would definitively be considered problematic (SS=71) and two others were close to being within normal limits (SS=84) not accounting for measurement error. This means that of the 14 students, only one for certain and possibly two others, are demonstrating any difficulty in their English vocabulary acquisition. This would suggest that if they are having difficulty in the classroom, it is most likely not related to a speech-language impairment and that their acquisition of English is proceeding in accordance with what would be expected of other individuals of the same age and with the same amount of time learning English. This means that at most, three (21 percent) and more likely just one (7 percent) of the individuals who were referred for evaluation might have a disability related to language as opposed to the 86 percent or higher that could have resulted from the use of scores obtained on a test without language-exposure norms.
Norm sample representation and fairness for ELs
The shift in perspective here and the implications for the consequences of using tests on bilingual populations that do not control for language development differences is potentially staggering. The change is not merely a marginally significant increase in fairness but rather a paradigm busting indictment of current practice. That we have been unable to effectively address the kinds of social injustice that we have seen historically and which we continue to see in the present day as related to the use of existing tests, suggests that whatever technical advances that have been promulgated with respect to fairness have been superficial at best and even more discriminatory at worst. For quite some time, many have called for a ban on tests precisely because of their tendencies to support the various forms of social and racial injustice that characterizes many of our systems and organizations in the present day. Yet, it may not be the fault of the tests themselves or any inherent defect which compromises their consequential validity. Rather, it may simply be that we have not endeavored well enough or far enough to recognize and distinguish those factors that have everything to do with how bilinguals perform on tests given to them in any language from those that have little or nothing to do with it. The Ortiz PVAT is not a comprehensive test of all abilities and is thus limited in terms of what it measures, serving primarily as a valid indicator of general English language acquisition. Nevertheless, it also provides a clear demonstration of what is possible if we are not afraid to re-think that which makes a truly representative peer group and distinguish it from that which does not. The future viability and utility of tests and efforts toward achieving some degree of social justice, very likely depends on it.
References
American Educational Research Association, American Psychological Association, & National Council on Measurement in Education (2014). Standards for educational and psychological testing. American Educational Research Association.
Cathers-Schiffman, T. A., & Thompson, M. S. (2007). Assessment of English- and Spanish-speaking Students with the WISC-III and Leiter-R. Journal of Psychoeducational Assessment, 25, 41–52.
Dynda, A. M. (2008). The relation between language proficiency and IQ test performance. (Unpublished manuscript). St. John’s University.
Flanagan, D. P., Ortiz, S. O., & Alfonso, V. C. (2013). Essentials of cross-battery assessment (3rd ed.). John Wiley.
Mercer, J. R. (1979). The system of multicultural pluralistic assessment: Technical manual. The Psychological Corporation.
New York State Testing Program (2009). New York State English as a Second Language Achievement Test (NYSESLAT) Technical Manual. Available at https://www.p12.nysed.gov/assessment/reports/nyseslat/nyseslat-tr-09.pdf
Ortiz, S. O. (2017). Evaluation of English Learners: Issues in measurement, interpretation and reporting. The Score, APA Division 5 (Quantitative and Qualitative Methods) Newsletter, January 2017. Available at https://www.apadivisions.org/division-5/publications/score/2017/01/english-learners.aspx
Ortiz, S. O. (2018). Ortiz Picture Vocabulary Acquisition Test (Ortiz PVAT). Multi-Health Systems.
Ortiz, S. O. (2019). On the Measurement of Cognitive Abilities in English Learners. Contemporary School Psychology, 23(1) 68–86. doi:10.1007/s40688-018-0208-8opens in new window
Peña, E. D., Gutierrez-Clellen, V., Iglesias, A., Goldstein, B. & Bedore, L. M. (2018). Bilingual English-Spanish Assessment (BESA). Brookes Publishing.
Sotelo-Dynega, M., Ortiz, S. O., Flanagan, D. P., & Chaplin, W. (2013). English language proficiency and test performance: Evaluation of bilinguals with the Woodcock-Johnson III Tests of Cognitive Ability. Psychology in the Schools, 50(8), 781–797.
Wechsler, D. (1999). Wechsler Abbreviated Scale of Intelligence Children. Pearson.
Woodcock, R. W., McGrew, K. S., & Mather, N. (2001). Woodcock-Johnson III Tests of Cognitive Abilities. Riverside.
Woodcock, R. W., Munoz-Sandoval, A., Reuf, M. & Alvarado, C. (2005). Woodcock-Muñoz Language Survey-Revised (WMLS-R). Riverside.
Woodcock, R. W., Munoz-Sandoval, A., Reuf, M. & Alvarado, C. (2017). Woodcock-Muñoz Language Survey-Third Edition (WMLS-III). Riverside.