Research methods is the third topic on Paper 2, but it accounts for at least 25% of the whole A-level and is assessed on every paper. You need experimental method and design, sampling, ethics, observational and self-report techniques, correlations, reliability and validity, the features of science, levels of measurement, descriptive statistics, probability and significance โ and, above all, how to choose and use the right inferential test.
Research methods is worth more marks than any other topic โ and it is examined across all three papers.
Work through each screen, answer the questions as you go and collect โญ stars. Every claim here is tied to a named study or theory you can quote in an essay. Press Start when you're ready.
Methods ยท experiments
Types of experiment and their variables
Laboratory experiment โ the IV is manipulated in a highly controlled environment. High internal validity and replicability, but risks demand characteristics and low ecological validity.
Field experiment โ the IV is manipulated in a natural setting. Higher mundane realism, but less control over extraneous variables, and there are ethical problems where participants have not consented.
Natural experiment โ the researcher does not manipulate the IV; it varies naturally and would have happened anyway (e.g. Rutter's Romanian orphans). No random allocation, so causal conclusions are risky.
Quasi-experiment โ the IV is a pre-existing characteristic of the participants (age, gender, having a diagnosis). Again, the IV is not manipulated and participants cannot be randomly allocated, so confounding variables are a real threat.
Variables and control. The independent variable is manipulated; the dependent variable is measured; both must be operationalised (defined so that they can be measured). An extraneous variable is any other variable that might affect the DV if not controlled (a nuisance variable); a confounding variable is one that varies systematically with the IV, so we cannot tell which caused the change in the DV.
Threats to validity:demand characteristics (participants guess the aim and change their behaviour โ the 'please-U' or 'screw-U' effect) and investigator effects (the researcher's expectations or behaviour influence the results). Controls: single-blind and double-blind procedures, standardised instructions, and randomisation.
Hypotheses. A directional (one-tailed) hypothesis states the direction of the difference or relationship and is used when previous research points one way. A non-directional (two-tailed) hypothesis states only that there will be a difference or relationship, and is used when previous research is absent or contradictory. The null hypothesis states there is no difference/relationship beyond that expected by chance โ and it is the null hypothesis that a statistical test either rejects or fails to reject.
Methods ยท design & sampling
Experimental design and sampling
Experimental design
Independent groups โ different participants in each condition. Problem: participant variables. Solution:random allocation. Also needs more participants.
Repeated measures โ the same participants in every condition, so participant variables are controlled. Problem:order effects (practice, fatigue, boredom) and greater risk of demand characteristics. Solution:counterbalancing (ABBA) โ half do A then B, half do B then A.
Matched pairs โ different participants, but matched on variables relevant to the study (e.g. IQ, age). Controls participant variables without order effects, but matching is time-consuming and can never be perfect.
Sampling
Random โ every member of the target population has an equal chance of selection (names from a hat, random number generator). Free from researcher bias, but time-consuming, and the sample can still be unrepresentative by chance.
Systematic โ every nth member of the sampling frame. Objective, but again time-consuming.
Stratified โ the composition of the sample reflects the proportions of sub-groups (strata) in the population, with participants randomly selected from each stratum. The most representative method, but the strata cannot capture every difference.
Opportunity โ whoever is available. Quick and cheap, but unrepresentative and open to researcher bias.
Volunteer (self-selected) โ participants respond to an advert. Easy and gives committed participants, but produces a volunteer bias (a particular type of person volunteers).
Ethics (the BPS Code).Informed consent (or presumptive/prior general/retrospective consent where deception is necessary); deception must be justified by an ethics committee weighing the cost-benefit; protection from harm โ participants must leave in the same state they arrived; the right to withdraw; confidentiality and privacy. The key remedies are a full debrief and the right to withdraw data.
Quick check
Choose the design
?A researcher tests reaction times in a quiet room and then in a noisy room, using the same participants in both. She worries that participants will improve simply through practice. What is this problem, and its standard solution?
Methods ยท non-experimental
Observations, self-report, correlations and case studies
Observations โ naturalistic vs controlled; covert vs overt; participant vs non-participant. Behaviour is recorded using behavioural categories, which must be operationalised, observable and mutually exclusive, and sampled by event sampling (count each occurrence) or time sampling (record what is happening at fixed intervals). Time sampling is efficient but may miss behaviours.
Self-report โ questionnaires (open and closed questions; Likert, rating and fixed-choice items) and interviews (structured, unstructured, semi-structured). Both are vulnerable to social desirability bias and acquiescence bias. Design principles: avoid leading questions, double-barrelled questions, jargon and emotive language; pilot the questionnaire first.
Correlations โ measure the strength and direction of an association between two co-variables, expressed as a correlation coefficient from -1 to +1. Correlations cannot demonstrate causation (there may be an untested third variable), and they can be misleading where the relationship is curvilinear.
Case studies โ an in-depth, idiographic investigation of one person, group or event, often using several methods and often longitudinal. Rich, detailed data (HM, Little Hans, Clive Wearing), but cannot be generalised and are prone to researcher subjectivity.
Content analysis and thematic analysis โ a method for turning qualitative material (interviews, media) into data. Coding counts instances of a category, producing quantitative data; thematic analysis identifies recurrent themes and keeps the data qualitative.
Peer review โ independent scrutiny of research by experts before publication, to allocate funding, validate quality and suggest amendments. Criticisms: reviewers may be anonymous and use that to criticise rivals, there is publication bias towards positive and 'headline' findings, and it may suppress opposition to the mainstream view.
Reliability & validity
Reliability and validity โ and how to improve them
Reliability = consistency.
Test-retest โ administer the same test to the same person on two occasions; a correlation coefficient of +0.80 or more indicates good reliability.
Inter-observer reliability โ two or more observers score the same behaviour independently and their totals are correlated (again, aim for +0.80). It is improved by operationalising the behavioural categories more precisely and by training observers.
Improving reliability: standardise procedures in experiments; in questionnaires, replace ambiguous or open questions with fixed-choice items; in interviews, use the same interviewer or train interviewers to avoid leading questions.
Validity = whether we are measuring what we intended to measure, and whether the finding is legitimate.
Internal validity โ did the IV really cause the change in the DV, or was it a confounding variable, demand characteristics or investigator effects?
External validity โ can the finding be generalised? Ecological validity (to other settings and real life), population validity (to other people) and temporal validity (to other eras โ a real problem for Asch's 1950s conformity data).
Face validity โ does the measure look, on the face of it, as though it measures what it claims to? Concurrent validity โ does it correlate closely (again, about +0.80) with an established measure?
Improving validity: use a control group, standardise procedures, use single/double-blind designs, and โ in qualitative research โ use triangulation (comparing evidence from several sources).
Features of science:paradigm and paradigm shift (Kuhn) ยท theory construction and hypothesis testing ยท falsifiability (Popper: a theory that cannot be proved wrong is not scientific) ยท replicability ยท objectivity and the empirical method.
Quick check
Which kind of validity?
?Asch's conformity findings were obtained from American men in the 1950s, and Perrin and Spencer found far lower conformity in 1980s Britain. Which type of validity is threatened?
Data ยท descriptive statistics
Levels of measurement and descriptive statistics
Levels of measurement (you must know these to choose a test):
Nominal โ data in categories (how many people chose A, B or C). Each item appears in only one category. It is a count, not a measurement.
Ordinal โ data that can be ranked, but where the intervals between points are not equal โ for example scores on a rating scale, because one person's '7' is not the same as another's. Ordinal data therefore lack precision.
Interval โ data measured with equal, standardised units: time in seconds, temperature, number of words recalled. This is the most precise level.
Measures of central tendency: the mean (uses every score, so it is most sensitive โ but it is distorted by outliers); the median (unaffected by extreme scores, but ignores much of the data); the mode (the only option for nominal data).
Measures of dispersion: the range (highest minus lowest โ simple, but affected by a single extreme value) and the standard deviation (the average distance of each score from the mean). A large SD means the scores are widely spread and the mean is not very representative; a small SD means the data are tightly clustered and all participants responded similarly.
Distributions: in a normal distribution the mean, median and mode are all at the mid-point and the curve is symmetrical. In a positive skew the long tail is to the right and the mode is to the left of the mean; in a negative skew the long tail is to the left and the mean is pulled below the mode.
Calculate
Calculate the mean
1Six participants recalled the following numbers of words: 7, 9, 6, 12, 8, 6. Calculate the mean number of words recalled.
words
Hint: add the scores (7 + 9 + 6 + 12 + 8 + 6 = 48) and divide by the number of scores (6).
Calculate
Percentages from a table
2In a study, 27 of the 45 participants obeyed the instruction. What percentage of the sample obeyed?
%
Hint: (27 รท 45) ร 100.
Inferential stats ยท significance
Probability, significance and errors
An inferential test tells us how likely it is that our result occurred by chance. Psychology's usual level of significance is p โค 0.05 โ a 5% (or less) probability that the result is due to chance. If the test is significant, we reject the null hypothesis.
A more stringent level (p โค 0.01) is used where a Type I error would be especially serious โ for example in drug trials, or where the study is a replication of a controversial finding.
Type I error โ a false positive: we reject the null hypothesis when it was in fact true (we claim a difference that isn't there). More likely with a lenient significance level such as 0.10.
Type II error โ a false negative: we accept the null hypothesis when it was actually false. More likely with a stringent level such as 0.01.
Using critical value tables. To find the critical value you need three things: whether the hypothesis is one-tailed or two-tailed; the N (or degrees of freedom); and the level of significance. Then compare:
R-cubed: Rank tests โ Result must be ReducedSign test, Wilcoxon, Mann-Whitney: calculated value must be EQUAL TO OR LESS THAN the critical value. Chi-square, Spearman, Pearson, both t-tests: calculated value must be EQUAL TO OR GREATER THAN the critical value.
Degrees of freedom for chi-square: df = (number of rows โ 1) ร (number of columns โ 1).
Calculate
Degrees of freedom
3A chi-square test is carried out on a contingency table with 3 rows and 3 columns. Calculate the degrees of freedom.
The sign test is used when you are looking for a difference, the design is repeated measures (related), and the data are nominal (or can be reduced to a simple direction of change).
Step 1. For each participant, record the sign of the change: a + if the score went up, a โ if it went down, and 0 if there was no change.
Step 2.Discard all the zeros (participants who showed no change). N is the number of participants left.
Step 3. Count the pluses and the minuses. The calculated value of S is the number of the less frequent sign.
Step 4. Look up the critical value of S for that N, the chosen significance level, and whether the test is one- or two-tailed.
Step 5. The result is significant if the calculated value of S is equal to or less than the critical value.
Worked example. 15 participants are tested before and after therapy. Nine improve (+), four get worse (โ), and two show no change (0). We discard the two zeros, so N = 13. The less frequent sign is the minus, so S = 4. The critical value of S for N = 13 at p โค 0.05, two-tailed, is 2. Because 4 is greater than 2, the result is not significant and we fail to reject the null hypothesis.
Calculate
Calculate S
4Twenty participants are tested before and after an intervention. Fourteen improved (+), three got worse (โ) and three showed no change (0). Calculate the calculated value of S for a sign test.
= S
Hint: discard the zeros. S is the number of the LESS frequent sign โ here, the minuses.
Calculate
What is N?
5Using the same data (14 pluses, 3 minuses, 3 no-change), what is the value of N that you would use to look up the critical value?
= N
Hint: N is the number of participants left after the zeros have been discarded: 20 โ 3.
Inferential stats ยท choosing a test
Choosing the right statistical test
Three questions decide the test. (1) Am I testing a difference or a correlation? (2) If a difference, is the design related (repeated measures or matched pairs) or unrelated (independent groups)? (3) What is the level of measurement?
Learn this grid cold โ 'Carrots Should Come Mashed With Swede Under Roast Potatoes' gives the nine cells reading across each row.
The parametric tests (the two t-tests and Pearson's r) are more powerful, but require interval data, a normal distribution and homogeneity of variance.
Quick check
Choose the test (1)
?A researcher compares the number of words (counted in whole words) recalled by an independent groups design: one group revised with music, the other in silence. Which test?
Quick check
Choose the test (2)
?A researcher tests whether there is a relationship between the number of hours slept and scores on a 10-point rating scale of mood. Which test?
Quick check
Significant or not?
?A Mann-Whitney test gives a calculated value of U = 21. The critical value of U for this N, at p โค 0.05 two-tailed, is 23. Is the result significant?
Calculate
Type I or Type II?
6A researcher uses a significance level of p โค 0.05. Out of every 100 significant results obtained at this level, how many, on average, would be expected to be a Type I error (a false positive) purely by chance?
in 100
Hint: p โค 0.05 means a 5% probability that the result is due to chance.
Scientific processes ยท reporting & impact
Pilot studies, data types and the scientific report
Pilot studies โ a small-scale trial run of an investigation, carried out before the real thing. The aim is to find out whether aspects of the design do not work โ ambiguous questionnaire items, instructions that confuse participants, a task that is too easy โ so that they can be modified and time and money saved. In observational research, a pilot allows observers to be trained and behavioural categories to be checked.
Primary data โ information collected first-hand by the researcher for the purpose of the investigation. It is authentic and fits the aim exactly, but it takes time and money to gather.
Secondary data โ information already collected by someone else (government statistics, journal articles, existing datasets). It is inexpensive and immediately available, but it may be outdated, incomplete or not quite fit the researcher's aim.
Meta-analysis โ a form of research using secondary data, in which the findings of many studies on the same question are combined and analysed together, often producing an effect size. It creates a much larger, more representative sample and greater confidence in conclusions โ but it is vulnerable to publication bias, because researchers may omit studies with negative or non-significant findings.
Presenting data: a bar chart shows the frequency of discrete categories, with gaps between the bars; a histogram shows continuous data, with no gaps and the area of the bar representing frequency; a scattergram shows the relationship between two co-variables in a correlation.
Sections of a scientific report:abstract (a short summary of the whole study) โ introduction (a literature review ending in the aims and hypotheses) โ method (design, sample, apparatus, procedure, ethics โ detailed enough for replication) โ results (descriptive and inferential statistics) โ discussion (what the findings mean, limitations, implications) โ referencing.
Implications for the economy โ psychological research has real economic value. Attachment research supports flexible working and shared parental leave; research into eyewitness testimony saves police time and prevents costly miscarriages of justice; effective treatments for depression and OCD reduce absence from work, which the economy pays for.
Sort it
Which test?
Tap a scenario, then the level of data it uses. Getting the level of measurement right is the first step to choosing the correct test.
๐ข Nominal
๐ Ordinal
๐ Interval
Match it
Match the design to the test
Tap an item on the left, then its partner on the right.
The study
The correct test
Recap
The big ideas to know
Experiments: lab ยท field ยท natural ยท quasi; IV, DV, extraneous and confounding variables