Reliability ER
Validity ER
Name that Evidence
Psychometrics Plot Twist
100

A preschool anxiety questionnaire gives children very different scores when completed again two weeks later, even though their anxiety has not meaningfully changed.

What is poor test-retest reliability? 

100

Children who score highly on a kindergarten school-readiness assessment subsequently tend to perform well when they enter elementary school.

What is strong predictive criterion related validity? 

100

Two independent observers give nearly identical ratings of the same children's behavior.

What is interrater reliability?

100

A new depression measure produces nearly identical scores every time someone completes it. Unfortunately, its items mostly assess physical fitness. Is it reliable, valid, both, or neither?

What is reliable but not valid?

200

Two clinicians independently watch the same video. One records 4 instances of aggression; the other records 17.

What is poor Interrater reliability? 

200
The title of a measure indicates that it's measuring at externalizing behaviors, but all the questions are about parenting and school routines. 

What is poor face validity?

200

Scores on an assessment of anxiety are highly similar when the same people complete it twice under similar circumstances.

What is test-retest reliability?

200

A new measure of empathy correlates strongly with an established empathy measure and has little relationship with a measure of mathematical ability.`

What are convergent validity AND discriminant validity?

300

Form A and Form B were designed as equivalent versions of an assessment. The same people tend to perform quite differently depending on which form they receive.

What is poor alternate forms reliability? 

300

A measure advertised as a comprehensive assessment of preschool social-emotional functioning contains almost exclusively questions about disruptive behavior.

What is poor content validity?

300

Items intended to measure sensory seeking behaviors are strongly related to one another.

What is internal consistency? 

300

Two equivalent versions of an assessment produce nearly identical scores. Clinicians looking at the items say, “I have no idea how these questions are supposed to measure anxiety.” Name the reliability evidence AND validity concern.

What are strong alternate forms reliability and poor face validity?

400

A new preschool anxiety questionnaire has 15 items that are all supposed to measure anxiety. But some items seem to behave very differently from the others—for example, children whose caregivers endorse most of the anxiety items often do not endorse several of the remaining items.


What is poor internal consistency? 

400

Scores on a new depression measure appear completely unreleated to the current school referrals to the counselor. 

What is poor concurrent criterion-related validity? 

400

A new measure of empathy correlates strongly with other measures of empathy.

What is convergent validity?

400

A new behavioral screener correlates highly with the number of referrals received in Kindergarten. Five years later, its original scores also accurately predict which children experience significant behavioral difficulties.

What are concurrent criterion-related validity AND predictive criterion-related validity?

500

A child's behavior rating is extremely consistent across all 30 items on an assessment, but their score changes substantially when they take the same assessment two weeks later. What can you conclude about reliability?

What is strong internal consistency but poor test-retest reliability? 

500

Experts agree that an emotion-regulation measure thoroughly samples all of the important aspects of emotion regulation—but the measure has almost no relationship with other established measures of emotion regulation.

What is strong content validity but poor convergent validity?

500

A measure of empathy has little relationship with a measure of mathematical ability, just as the developers predicted.

What is discriminant validity?

500

An assessment has excellent test-retest reliability, excellent interrater reliability, and excellent internal consistency. Are we good!?

What is reliable but not valid (potentially)?

M
e
n
u