1. Introduction
Test length reduction is not a new concept or rare phenomenon. Since shortened tests require less time and resources (Stanton et al., 2002) and have the propensity to reduce respondent burden (Smits & Vorst, 2007), it’s no surprise that test developers often look for ways to reduce test length. However, these potential benefits also bring potential risks, namely, the possibility of reduced psychometric quality of the assessment. These concerns regarding psychometric quality are not without merit. A literature review of 164 shortened tests indicated that the mean reliability of tests reduced from α=0.84 to α=0.77 when test lengths were reduced and that validity was only assessed in approximately half of the sample (Kruyen et al., 2013). Further, much of the conversation around the shortening of tests or assessments is tied to multiple choice or forced choice formats. There is virtually no literature available on how one may reduce the length of an open-response test.
The goal of this study was to explore optimal test length, which balances practical utility with psychometric integrity (Sitarenios, 2022), for an open-response situational judgement test (SJT).
2. Study Design
To explore this research question, we employed a two-phase sequential design to examine the Casper test; a 14-item open-response SJT that contains both typed-response and video-response scenarios. Each of the 14 items were unique scenarios that each asked two unique questions. Given that this instrument was relatively short already, we identified an opportunity to also explore the impact of changing from scenario-level scoring to question-level scoring. We hypothesized that changing from scenario-level scoring to question-level scoring would have two main advantages: (1) test takers would receive more accurate scores since they were being evaluated for each individual response opposed to holistically for all questions in a scenario and (2) it would allow more flexibility in plausible test length options given that we’d essentially be doubling the number of items in the test.
The study was organized into two main phases. In the first phase, trained raters were asked to provide scenario-level scores and question-level scores for a set of data so that we could assess the extent to which scores differed between the two methods. Using the results of this phase, we conducted a series of simulations across a larger dataset to evaluate the psychometric properties of the test with question-level scoring across various test lengths.
2.1 Phase 1
2.1.1 Procedure
We employed a repeated-measures design in which participants scored the same response twice across the two methods: scenario-level and question-level. The order in which the methods were presented to participants was randomized to control for order effects. Participants were not permitted to score the responses a second time (the alternate approach) until they participated in scoring an active test (i.e., a real test outside the study environment). Scoring an additional set of unique responses between study sessions was intended to reduce the likelihood that participants would recall and reuse scores from the first round when scoring the responses a second time. Response order was also randomized between rounds to further limit recall of prior scores.
2.1.2 Sample
We wanted to ensure that we captured data from a diverse sample of tests since the Casper test is used globally, in different languages, and for admissions to a variety of professional programs. Ultimately, we employed n=23 trained raters to examine data from n=200 test takers across the United States, Australia, and Canada (English and French). Details on the sample data are available in Table 1.
Table 1. Information Regarding Sample Data in Phase 1
Country | Language | Program Type | Number of Raters | Number of Test Takers |
United States | English | Health Sciences | 8 | 50 |
Canada | French | Allied Health | 4 | 50 |
Australia | English | Teacher’s Education | 6 | 50 |
Canada | English | Health Sciences | 5 | 50 |
2.1.3 Results
Descriptive Statistics of Total Scores. Descriptive statistics for both scenario-level and question-level scores are provided in Table 2.
Table 2. Descriptive Statistics of Total Scores Across Scoring Methods
Scoring Method | n | mean | sd | median | min | max | Skew | Kurt | |
US | Scenario-Level | 50 | 5.47 | 0.94 | 5.62 | 2.75 | 7.00 | -0.79 | 0.44 |
US | Question-Level | 50 | 5.31 | 0.90 | 5.41 | 2.69 | 6.88 | -0.49 | -0.06 |
AUS | Scenario-Level | 50 | 4.49 | 0.95 | 4.50 | 2.17 | 7.00 | -0.08 | -0.02 |
AUS | Question-Level | 50 | 4.12 | 0.93 | 4.08 | 1.92 | 6.25 | 0.14 | -0.30 |
CA-E | Scenario-Level | 50 | 5.01 | 1.12 | 5.00 | 2.40 | 7.20 | -0.01 | -0.64 |
CA-E | Question-Level | 50 | 5.02 | 0.96 | 5.00 | 3.00 | 7.20 | 0.26 | -0.39 |
CA-F | Scenario-Level | 49 | 5.96 | 1.03 | 6.00 | 3.50 | 7.75 | -0.34 | -0.45 |
CA-F | Question-Level | 49 | 5.42 | 1.09 | 5.50 | 2.62 | 7.75 | -0.30 | -0.14 |
Note. One CA-F test taker had to be removed as a test taker did not complete the assessment.
Absolute Mean Difference in Total Scores Across Scoring Methods. We examined the difference in total scores between the two methods. To do this, we calculated the mean absolute difference between test takers’ total scores from the scenario-level scoring and their total scores from the question-level scoring. Descriptive statistics are available in Table 3, and a frequency table of absolute differences is available in Table 4.
Table 3. Descriptive Statistics of Mean Absolute Differences in Total Scores
Country | n | mean | sd | median | min | max | Skew | Kurt |
US | 50 | 0.36 | 0.24 | 0.31 | 0.00 | 0.88 | 0.40 | -0.89 |
AUS | 50 | 0.50 | 0.38 | 0.46 | 0.00 | 1.42 | 0.48 | -0.58 |
CA-E | 50 | 0.45 | 0.38 | 0.40 | 0.00 | 1.60 | 0.88 | 0.40 |
CA-F | 49 | 0.64 | 0.46 | 0.50 | 0.00 | 1.75 | 0.49 | -0.86 |
Table 4. Frequency Table of Absolute Differences in Total Scores
US | AUS | CA-E | CA-F | |||||
Absolute Difference | n | % | n | % | n | % | n | % |
0 - 0.10 | 10 | 20 | 11 | 22 | 12 | 24 | 2 | 4.08 |
0.11 - 0.20 | 4 | 8 | 3 | 6 | 7 | 14 | 7 | 14.29 |
0.21 - 0.30 | 7 | 14 | 2 | 4 | 4 | 8 | 7 | 14.29 |
0.31 - 0.40 | 11 | 22 | 4 | 8 | 6 | 12 | 5 | 10.20 |
0.41 - 0.50 | 5 | 10 | 7 | 14 | 2 | 4 | 4 | 8.16 |
0.51 - 0.60 | 3 | 6 | 6 | 12 | 6 | 12 | 0 | 0 |
0.61 - 0.70 | 5 | 10 | 3 | 6 | 2 | 4 | 3 | 6.12 |
0.71 - 0.80 | 2 | 4 | 3 | 6 | 4 | 8 | 3 | 6.12 |
0.81 - 0.90 | 3 | 6 | 2 | 4 | 2 | 4 | 4 | 8.16 |
0.91 - 1.00 | 0 | 0 | 4 | 8 | 2 | 4 | 3 | 6.12 |
1.01 - 1.10 | 0 | 0 | 3 | 6 | 1 | 2 | 0 | 0 |
1.11 - 1.20 | 0 | 0 | 0 | 0 | 0 | 0 | 3 | 6.12 |
1.21 - 1.30 | 0 | 0 | 0 | 0 | 0 | 0 | 5 | 10.20 |
1.31 - 1.40 | 0 | 0 | 0 | 0 | 1 | 2 | 1 | 2.04 |
1.41 - 1.50 | 0 | 0 | 2 | 4 | 0 | 0 | 0 | 0 |
1.51 - 1.60 | 0 | 0 | 0 | 0 | 1 | 2 | 0 | 0 |
1.61 - 1.70 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 2.04 |
1.71 - 7.80 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 2.04 |
Relationship Between Total Scores Across Scoring Methods. The correlation evidenced a strong and statistically significant relationship between the two total scores in the US data (r=0.87, p<.001), Australian data (r=0.86, p<.001), Canadian English data (r=0.85, p<.001), and Canadian French data (r=0.82, p<.001).
2.1.4 Phase 1 Discussion
Across all data, we found that test takers’ total scores differed, on average, by less than 1 point (range: 0.36 - 0.64) between the two methods. More specifically, 60% of test takers’ total scores differed by half a point or less and only 10% of test takers’ total scores differed by more than one point. Correlation analyses evidenced a strong relationship between the two scores (r=0.82 - 0.87, p<.001). Together, these results helped inform the second phase of the study.
2.2 Phase 2
2.2.1 Procedure
In the second phase, we employed a synthetic design approach similar to what is described by Lee and colleagues (2014). Scenario-level scores were collected from the full-length version of the open-response SJT and then used to generate synthetic question-level scores. From here, we were able to evaluate the psychometric properties of various test lengths without the practical and ethical disadvantages of asking participants to take the test repeatedly across different lengths, as would be required with a traditional experimental design (Lee et al., 2014). To generate the synthetic question-level scores, we wrote a script using RStudio Version 2023.3.0.386 (Posit Team, 2023), which generated two item-level scores for each scenario that differed by +/- one point from their original scenario-level score. The choice to allow scores to differ within one point of the original score was determined based on findings in Phase 1.
Prior to conducting the simulations, we identified a number of different test length combinations to explore. Although there were 42 possible combinations in theory, several were not practically feasible. Therefore, we removed several combinations prior to conducting our analyses. Ultimately, we examined 15 possible test length combinations (labeled with letters in Table 5). Data analyses were conducted on each test length combination across each independent test instance.
Table 5. Possible Test Length Combinations
Number of Video Responses | ||||||
1 | 2 | 3 | 4 | 5 | 6 | |
1 | ||||||
2 | ||||||
3 | A | B | C | |||
4 | D | E | F | |||
5 | G | H | I | |||
6 | J | K | L | |||
7 | M | N | O | |||
8 | ||||||
2.2.2 Sample
This phase utilized data from 26,671 test takers from 20 unique test instances across the United States, Australia, and Canada (English and French). See Table 6 for details on the data used for Phase 2.
Table 6. Information Regarding Sample Data in Phase 2
Country | Language | Number of Tests Examined | Number of Test Takers |
United States | English | 7 | 10,545 |
Australia | English | 5 | 5,036 |
Canada | English | 5 | 7,572 |
Canada | French | 3 | 3,518 |
2.2.3 Results
For each combination across each test instance, we examined test-level reliability, the correlation between the total score and the original total score from the 14-item test, and demographic group differences.
Internal Consistency Reliability. Average reliability for the original 14-item test was ɑ=0.82 with an average inter-item correlation (IIC) value of 0.24. Average reliability and IIC values for each combination across all test instances are available in Table 7.
Table 7. Average Alpha and IIC Values Across Test Instances
Test Length Combination | Average Alpha | Average IIC |
Original | 0.82 | 0.24 |
A | 0.77 | 0.22 |
B | 0.77 | 0.20 |
C | 0.78 | 0.18 |
D | 0.80 | 0.23 |
E | 0.80 | 0.20 |
F | 0.81 | 0.19 |
G | 0.83 | 0.23 |
H | 0.83 | 0.21 |
I | 0.83 | 0.20 |
J | 0.85 | 0.24 |
K | 0.85 | 0.22 |
L | 0.85 | 0.21 |
M | 0.87 | 0.25 |
N | 0.87 | 0.23 |
O | 0.87 | 0.22 |
Correlations. Correlation coefficients were used to provide an estimate of the strength of the relationship between the original total score and the total score for each combination. The average correlation coefficient for each test length combination is available in Table 8.
Table 8. Average Correlation Between Original Scores and Various Test Length Combinations Across Test Instances
Test Length Combination | Average Correlation |
A | 0.85 |
B | 0.87 |
C | 0.88 |
D | 0.89 |
E | 0.90 |
F | 0.91 |
G | 0.91 |
H | 0.92 |
I | 0.93 |
J | 0.93 |
K | 0.94 |
L | 0.95 |
M | 0.94 |
N | 0.95 |
O | 0.96 |
Demographic Group Differences. Demographic differences were explored in aggregate for each grouping (US, Australia, Canada English, and Canada French) to ensure a sufficient sample size. We calculated the magnitude of group differences for the total scores of the original 14-item test and compared them to the magnitude observed in the total scores from each test length combination. Generally, we observed very minimal variability in the magnitude of group differences between the original score and scores from the various test lengths. More specifically, the difference in magnitude was no greater than 0.06 across all group comparisons in all geographies.
2.2.4 Phase 2 Discussion
Overall, a majority of the test length combinations performed well. Of the 15 combinations explored, 60% (n=9) produced an average reliability that exceeded that of the original test format. Additionally, 73.3% (n=11) met or exceeded the correlation coefficient threshold of 0.90 that suggests the shortened test adequately represents its full-length counterpart (Lee et al., 2014).
3. Discussion
This study explored whether an open-response situational judgment test (SJT) could be shortened while maintaining psychometric quality. Across both phases, the findings were encouraging. Phase 1 showed that question-level scoring produced scores very similar to scenario-level scoring, with average correlations of r=0.82 to 0.87 (p<.001) between the two methods. At the test taker level, we saw that their total scores differed, on average, by less than 1 point between the two methods. This finding suggests that scoring each question individually does not dramatically change the overall results, but it does provide more flexibility in structuring the test. Phase 2 expanded on the Phase 1 results by simulating various test lengths using synthetic question-level data. The majority of shortened versions maintained reliability greater than what was observed in the original 14-item test. Further, correlations with the original test were high, indicating that the shortened tests adequately represented the longer version. Also, the change in demographic group differences was minimal, suggesting that shortening the test did not introduce bias across groups.
Taken together, these findings have practical implications. Shortening an open-response SJT can reduce test-taking time and respondent burden without sacrificing psychometric quality. Additionally, question-level scoring allows more granular measurement, enabling test developers to explore a wider range of test lengths while maintaining psychometric standards. Of course, there are limitations. Our simulations relied on synthetic data, which, while based on real scoring patterns, may not fully capture all nuances of real-world responses. Further, this study examined only a single instrument. It is recommended that future studies examine different instruments with different formats.
Ultimately, the study demonstrates that open-response tests can be made more efficient without a major trade-off in reliability or validity. This opens the door for more practical and user-friendly assessments in higher-education admissions and other high-stakes testing environments.

