A teacher administers an examination to the same students on two occasions under comparable conditions. The first administration produces substantially different results from the second, even though no meaningful change in students’ learning has occurred. Which conclusion is most justified?
A.
The test demonstrates weak test-retest reliability
B.
The test demonstrates strong content validity
C.
The test has high criterion validity
D.
The test necessarily has strong construct validity