Reliability
Validity A
Validity B
Test Bias
Miscellaneous
100
Reliability tells us whether a test:

a. measures what we think it measures

b. measures something consistently

c. will predict to a future outcome well

d. has test bias

b. measures something consistently

100
The appearance that a measure has validity is:

a. content validity

b. face validity

c. appearance validity

d. criterion validity

b. face validity

100
An experimenter is trying to make sure his test relates to an existing gold standard measure. What type of validity is the experimenter trying to achieve?

a. Construct Validity

b. Criterion Validity

c. Face Validity.

d. Both Criterion and Construct Validity.

b. Criterion Validity

100
If two groups score differently on a test, that means the test is definitely biased - true or false?
FALSE
100
Which of the following is/are a component of item analysis?

A. Item Discriminability

B. Item Distractor

C. Item Difficulty

D. Both a & c

E. Both b & c

D. Both a (discriminability) and c (difficulty).
200
If scores shouldn't change over time, and yet they do, we have a problem with measurement error due to timing. We should address this with what type of reliability calculation?
Test-Retest reliability
200
Which are subtypes of Criterion Validity?

A. Convergent and Discriminant

B. Predictive and Concurrent

C. A and B

D. None of the above

B. Predictive and Concurrent
200
If a test has face validity, we can assume the measure is appropriate to use - true or false?
FALSE!
200
If a test means something different for one group than it does for another, that is called _____________.
Differential validity (can be differential content validity - the content of the test is more relevant/familiar for one group than another - or differential criterion validity - the test predicts differently to the criterion for one from than another).
200
Which does not detract from a test item?

A) Leading questions

B) Double-barreled questions

C) Long items

D) None of the above

D) None of the above

300
All of the following are ways to improve reliability EXCEPT:

a. administer the test to multiple people until you achieve reliability

b. remove inconsistent items

c. determine if subscales exist

d. add consistent items

a. administer the test to multiple people until you achieve reliability

300
Construct underrepresentation (failing to capture important concepts; leaving something out) and construct irrelevant variance (measuring something extra) are subtypes of what type of validity?
Content validity
300
Name one of the three characteristics of an appropriate criterion (i.e., an existing gold standard measure you could compare your new measure with).
It should be: established, relevant, measurable
300
All tests are biased: true or false?
FALSE!
300
_______ indicate(s) the difference in correct responding between high scoring and low scoring participants.

a)Item discriminability

b)Reverse-scored items

c)Lie scales

d)Reliability calculations

a) Item discriminability. Another way of measuring it is the correlation between performance on the test as a whole and performance on that item.
400
If our test items are unclear or inconsistent, it sounds like we have a problem with what specific source of error?
Internal Consistency (of test items)
400
To establish _________ validity, we need as many pieces of evidence as possible.
Construct
400
Which Pearson correlation result best demonstrates divergent evidence for construct validity?

A) -1

B) -.680

C) .480

D) .081

D) .081

400
A graph showing prediction lines for two tests, each with a different slope (one flat, one steep), says what about differential criterion validity?

a) The groups have different results but the test is equally valid for both groups.

b) The criterion is not well predicted for one of the groups and the test should not be used.

c) The test predicts the criterion well for both groups and shows usefulness for both.

d) No evidence of criterion bias - it's a graph common to IQ test results.

b) The criterion is not well predicted for one of the groups and the test should not be used.

400
The set of scores in a particular population against we are comparing an individual's test results is the _____.
Norm. Which should ideally come from a population representative of the population in which we want to use the test/where the individual comes from.
500
Reliability is _____ but not _____ to have validity.
Reliability is NECESSARY but not SUFFICIENT to have validity.
500
To establish construct validity, we need:

a. convergent evidence

b. divergent evidence

c. both

d. neither

c. both

So what does this mean?

500
An aptitude test for a mechanic job that asks questions about preference for cats vs. dogs we can be sure has problems with ______ validity.
Content (Specifically?)
500
In the _______ approach, tests are used to predict but minority group members get a "boost" to adjust for possible bias.
Qualified Individualism. Does this make the test itself more or less valid?
500
The item "Rate the frequency of your use of study guide materials from never (1) to always (5)" is on a ______ scale.
Category
M
e
n
u