โ† Back to subjects
โญ 0
AQA A-level Psychology (7182) ยท Research methods
Mini-Lesson

Research methods

Research methods is the third topic on Paper 2, but it accounts for at least 25% of the whole A-level and is assessed on every paper. You need experimental method and design, sampling, ethics, observational and self-report techniques, correlations, reliability and validity, the features of science, levels of measurement, descriptive statistics, probability and significance โ€” and, above all, how to choose and use the right inferential test.

methods & design data & descriptive stats inferential tests Paper 2 ยท at least 25% of the whole A-level, and it is assessed on every paper
Research methods is worth more marks than any other topic โ€” and it is examined across all three papers.

Work through each screen, answer the questions as you go and collect โญ stars. Every claim here is tied to a named study or theory you can quote in an essay. Press Start when you're ready.

Methods ยท experiments

Types of experiment and their variables

  • Laboratory experiment โ€” the IV is manipulated in a highly controlled environment. High internal validity and replicability, but risks demand characteristics and low ecological validity.
  • Field experiment โ€” the IV is manipulated in a natural setting. Higher mundane realism, but less control over extraneous variables, and there are ethical problems where participants have not consented.
  • Natural experiment โ€” the researcher does not manipulate the IV; it varies naturally and would have happened anyway (e.g. Rutter's Romanian orphans). No random allocation, so causal conclusions are risky.
  • Quasi-experiment โ€” the IV is a pre-existing characteristic of the participants (age, gender, having a diagnosis). Again, the IV is not manipulated and participants cannot be randomly allocated, so confounding variables are a real threat.

Variables and control. The independent variable is manipulated; the dependent variable is measured; both must be operationalised (defined so that they can be measured). An extraneous variable is any other variable that might affect the DV if not controlled (a nuisance variable); a confounding variable is one that varies systematically with the IV, so we cannot tell which caused the change in the DV.

Threats to validity: demand characteristics (participants guess the aim and change their behaviour โ€” the 'please-U' or 'screw-U' effect) and investigator effects (the researcher's expectations or behaviour influence the results). Controls: single-blind and double-blind procedures, standardised instructions, and randomisation.

Hypotheses. A directional (one-tailed) hypothesis states the direction of the difference or relationship and is used when previous research points one way. A non-directional (two-tailed) hypothesis states only that there will be a difference or relationship, and is used when previous research is absent or contradictory. The null hypothesis states there is no difference/relationship beyond that expected by chance โ€” and it is the null hypothesis that a statistical test either rejects or fails to reject.

Methods ยท design & sampling

Experimental design and sampling

Experimental design

  • Independent groups โ€” different participants in each condition. Problem: participant variables. Solution: random allocation. Also needs more participants.
  • Repeated measures โ€” the same participants in every condition, so participant variables are controlled. Problem: order effects (practice, fatigue, boredom) and greater risk of demand characteristics. Solution: counterbalancing (ABBA) โ€” half do A then B, half do B then A.
  • Matched pairs โ€” different participants, but matched on variables relevant to the study (e.g. IQ, age). Controls participant variables without order effects, but matching is time-consuming and can never be perfect.

Sampling

  • Random โ€” every member of the target population has an equal chance of selection (names from a hat, random number generator). Free from researcher bias, but time-consuming, and the sample can still be unrepresentative by chance.
  • Systematic โ€” every nth member of the sampling frame. Objective, but again time-consuming.
  • Stratified โ€” the composition of the sample reflects the proportions of sub-groups (strata) in the population, with participants randomly selected from each stratum. The most representative method, but the strata cannot capture every difference.
  • Opportunity โ€” whoever is available. Quick and cheap, but unrepresentative and open to researcher bias.
  • Volunteer (self-selected) โ€” participants respond to an advert. Easy and gives committed participants, but produces a volunteer bias (a particular type of person volunteers).

Ethics (the BPS Code). Informed consent (or presumptive/prior general/retrospective consent where deception is necessary); deception must be justified by an ethics committee weighing the cost-benefit; protection from harm โ€” participants must leave in the same state they arrived; the right to withdraw; confidentiality and privacy. The key remedies are a full debrief and the right to withdraw data.

Quick check

Choose the design

?A researcher tests reaction times in a quiet room and then in a noisy room, using the same participants in both. She worries that participants will improve simply through practice. What is this problem, and its standard solution?
Methods ยท non-experimental

Observations, self-report, correlations and case studies

  • Observations โ€” naturalistic vs controlled; covert vs overt; participant vs non-participant. Behaviour is recorded using behavioural categories, which must be operationalised, observable and mutually exclusive, and sampled by event sampling (count each occurrence) or time sampling (record what is happening at fixed intervals). Time sampling is efficient but may miss behaviours.
  • Self-report โ€” questionnaires (open and closed questions; Likert, rating and fixed-choice items) and interviews (structured, unstructured, semi-structured). Both are vulnerable to social desirability bias and acquiescence bias. Design principles: avoid leading questions, double-barrelled questions, jargon and emotive language; pilot the questionnaire first.
  • Correlations โ€” measure the strength and direction of an association between two co-variables, expressed as a correlation coefficient from -1 to +1. Correlations cannot demonstrate causation (there may be an untested third variable), and they can be misleading where the relationship is curvilinear.
  • Case studies โ€” an in-depth, idiographic investigation of one person, group or event, often using several methods and often longitudinal. Rich, detailed data (HM, Little Hans, Clive Wearing), but cannot be generalised and are prone to researcher subjectivity.
  • Content analysis and thematic analysis โ€” a method for turning qualitative material (interviews, media) into data. Coding counts instances of a category, producing quantitative data; thematic analysis identifies recurrent themes and keeps the data qualitative.
  • Peer review โ€” independent scrutiny of research by experts before publication, to allocate funding, validate quality and suggest amendments. Criticisms: reviewers may be anonymous and use that to criticise rivals, there is publication bias towards positive and 'headline' findings, and it may suppress opposition to the mainstream view.
Reliability & validity

Reliability and validity โ€” and how to improve them

Reliability = consistency.

  • Test-retest โ€” administer the same test to the same person on two occasions; a correlation coefficient of +0.80 or more indicates good reliability.
  • Inter-observer reliability โ€” two or more observers score the same behaviour independently and their totals are correlated (again, aim for +0.80). It is improved by operationalising the behavioural categories more precisely and by training observers.
  • Improving reliability: standardise procedures in experiments; in questionnaires, replace ambiguous or open questions with fixed-choice items; in interviews, use the same interviewer or train interviewers to avoid leading questions.

Validity = whether we are measuring what we intended to measure, and whether the finding is legitimate.

  • Internal validity โ€” did the IV really cause the change in the DV, or was it a confounding variable, demand characteristics or investigator effects?
  • External validity โ€” can the finding be generalised? Ecological validity (to other settings and real life), population validity (to other people) and temporal validity (to other eras โ€” a real problem for Asch's 1950s conformity data).
  • Face validity โ€” does the measure look, on the face of it, as though it measures what it claims to? Concurrent validity โ€” does it correlate closely (again, about +0.80) with an established measure?
  • Improving validity: use a control group, standardise procedures, use single/double-blind designs, and โ€” in qualitative research โ€” use triangulation (comparing evidence from several sources).

Features of science: paradigm and paradigm shift (Kuhn) ยท theory construction and hypothesis testing ยท falsifiability (Popper: a theory that cannot be proved wrong is not scientific) ยท replicability ยท objectivity and the empirical method.

Quick check

Which kind of validity?

?Asch's conformity findings were obtained from American men in the 1950s, and Perrin and Spencer found far lower conformity in 1980s Britain. Which type of validity is threatened?
Data ยท descriptive statistics

Levels of measurement and descriptive statistics

Levels of measurement (you must know these to choose a test):

  • Nominal โ€” data in categories (how many people chose A, B or C). Each item appears in only one category. It is a count, not a measurement.
  • Ordinal โ€” data that can be ranked, but where the intervals between points are not equal โ€” for example scores on a rating scale, because one person's '7' is not the same as another's. Ordinal data therefore lack precision.
  • Interval โ€” data measured with equal, standardised units: time in seconds, temperature, number of words recalled. This is the most precise level.

Measures of central tendency: the mean (uses every score, so it is most sensitive โ€” but it is distorted by outliers); the median (unaffected by extreme scores, but ignores much of the data); the mode (the only option for nominal data).

Measures of dispersion: the range (highest minus lowest โ€” simple, but affected by a single extreme value) and the standard deviation (the average distance of each score from the mean). A large SD means the scores are widely spread and the mean is not very representative; a small SD means the data are tightly clustered and all participants responded similarly.

Distributions: in a normal distribution the mean, median and mode are all at the mid-point and the curve is symmetrical. In a positive skew the long tail is to the right and the mode is to the left of the mean; in a negative skew the long tail is to the left and the mean is pulled below the mode.

Calculate

Calculate the mean

1Six participants recalled the following numbers of words: 7, 9, 6, 12, 8, 6. Calculate the mean number of words recalled.
words
Hint: add the scores (7 + 9 + 6 + 12 + 8 + 6 = 48) and divide by the number of scores (6).
Calculate

Percentages from a table

2In a study, 27 of the 45 participants obeyed the instruction. What percentage of the sample obeyed?
%
Hint: (27 รท 45) ร— 100.
Inferential stats ยท significance

Probability, significance and errors

An inferential test tells us how likely it is that our result occurred by chance. Psychology's usual level of significance is p โ‰ค 0.05 โ€” a 5% (or less) probability that the result is due to chance. If the test is significant, we reject the null hypothesis.

  • A more stringent level (p โ‰ค 0.01) is used where a Type I error would be especially serious โ€” for example in drug trials, or where the study is a replication of a controversial finding.
  • Type I error โ€” a false positive: we reject the null hypothesis when it was in fact true (we claim a difference that isn't there). More likely with a lenient significance level such as 0.10.
  • Type II error โ€” a false negative: we accept the null hypothesis when it was actually false. More likely with a stringent level such as 0.01.

Using critical value tables. To find the critical value you need three things: whether the hypothesis is one-tailed or two-tailed; the N (or degrees of freedom); and the level of significance. Then compare:

R-cubed: Rank tests โ†’ Result must be ReducedSign test, Wilcoxon, Mann-Whitney: calculated value must be EQUAL TO OR LESS THAN the critical value.
Chi-square, Spearman, Pearson, both t-tests: calculated value must be EQUAL TO OR GREATER THAN the critical value.

Degrees of freedom for chi-square: df = (number of rows โˆ’ 1) ร— (number of columns โˆ’ 1).

Calculate

Degrees of freedom

3A chi-square test is carried out on a contingency table with 3 rows and 3 columns. Calculate the degrees of freedom.
df
Hint: df = (rows โˆ’ 1) ร— (columns โˆ’ 1) = (3 โˆ’ 1) ร— (3 โˆ’ 1).
Inferential stats ยท the sign test

The sign test, step by step

The sign test is used when you are looking for a difference, the design is repeated measures (related), and the data are nominal (or can be reduced to a simple direction of change).

  • Step 1. For each participant, record the sign of the change: a + if the score went up, a โˆ’ if it went down, and 0 if there was no change.
  • Step 2. Discard all the zeros (participants who showed no change). N is the number of participants left.
  • Step 3. Count the pluses and the minuses. The calculated value of S is the number of the less frequent sign.
  • Step 4. Look up the critical value of S for that N, the chosen significance level, and whether the test is one- or two-tailed.
  • Step 5. The result is significant if the calculated value of S is equal to or less than the critical value.

Worked example. 15 participants are tested before and after therapy. Nine improve (+), four get worse (โˆ’), and two show no change (0). We discard the two zeros, so N = 13. The less frequent sign is the minus, so S = 4. The critical value of S for N = 13 at p โ‰ค 0.05, two-tailed, is 2. Because 4 is greater than 2, the result is not significant and we fail to reject the null hypothesis.

Calculate

Calculate S

4Twenty participants are tested before and after an intervention. Fourteen improved (+), three got worse (โˆ’) and three showed no change (0). Calculate the calculated value of S for a sign test.
= S
Hint: discard the zeros. S is the number of the LESS frequent sign โ€” here, the minuses.
Calculate

What is N?

5Using the same data (14 pluses, 3 minuses, 3 no-change), what is the value of N that you would use to look up the critical value?
= N
Hint: N is the number of participants left after the zeros have been discarded: 20 โˆ’ 3.
Inferential stats ยท choosing a test

Choosing the right statistical test

Three questions decide the test. (1) Am I testing a difference or a correlation? (2) If a difference, is the design related (repeated measures or matched pairs) or unrelated (independent groups)? (3) What is the level of measurement?

Level of data Difference UNRELATED design Difference RELATED design Correlation NOMINAL Chi-square Sign test Chi-square ORDINAL Mann-Whitney Wilcoxon Spearman's rho INTERVAL Unrelated t-test Related t-test Pearson's r
Learn this grid cold โ€” 'Carrots Should Come Mashed With Swede Under Roast Potatoes' gives the nine cells reading across each row.

The parametric tests (the two t-tests and Pearson's r) are more powerful, but require interval data, a normal distribution and homogeneity of variance.

Quick check

Choose the test (1)

?A researcher compares the number of words (counted in whole words) recalled by an independent groups design: one group revised with music, the other in silence. Which test?
Quick check

Choose the test (2)

?A researcher tests whether there is a relationship between the number of hours slept and scores on a 10-point rating scale of mood. Which test?
Quick check

Significant or not?

?A Mann-Whitney test gives a calculated value of U = 21. The critical value of U for this N, at p โ‰ค 0.05 two-tailed, is 23. Is the result significant?
Calculate

Type I or Type II?

6A researcher uses a significance level of p โ‰ค 0.05. Out of every 100 significant results obtained at this level, how many, on average, would be expected to be a Type I error (a false positive) purely by chance?
in 100
Hint: p โ‰ค 0.05 means a 5% probability that the result is due to chance.
Scientific processes ยท reporting & impact

Pilot studies, data types and the scientific report

  • Pilot studies โ€” a small-scale trial run of an investigation, carried out before the real thing. The aim is to find out whether aspects of the design do not work โ€” ambiguous questionnaire items, instructions that confuse participants, a task that is too easy โ€” so that they can be modified and time and money saved. In observational research, a pilot allows observers to be trained and behavioural categories to be checked.
  • Primary data โ€” information collected first-hand by the researcher for the purpose of the investigation. It is authentic and fits the aim exactly, but it takes time and money to gather.
  • Secondary data โ€” information already collected by someone else (government statistics, journal articles, existing datasets). It is inexpensive and immediately available, but it may be outdated, incomplete or not quite fit the researcher's aim.
  • Meta-analysis โ€” a form of research using secondary data, in which the findings of many studies on the same question are combined and analysed together, often producing an effect size. It creates a much larger, more representative sample and greater confidence in conclusions โ€” but it is vulnerable to publication bias, because researchers may omit studies with negative or non-significant findings.
  • Presenting data: a bar chart shows the frequency of discrete categories, with gaps between the bars; a histogram shows continuous data, with no gaps and the area of the bar representing frequency; a scattergram shows the relationship between two co-variables in a correlation.
  • Sections of a scientific report: abstract (a short summary of the whole study) โ†’ introduction (a literature review ending in the aims and hypotheses) โ†’ method (design, sample, apparatus, procedure, ethics โ€” detailed enough for replication) โ†’ results (descriptive and inferential statistics) โ†’ discussion (what the findings mean, limitations, implications) โ†’ referencing.
  • Implications for the economy โ€” psychological research has real economic value. Attachment research supports flexible working and shared parental leave; research into eyewitness testimony saves police time and prevents costly miscarriages of justice; effective treatments for depression and OCD reduce absence from work, which the economy pays for.
Sort it

Which test?

Tap a scenario, then the level of data it uses. Getting the level of measurement right is the first step to choosing the correct test.

๐Ÿ”ข Nominal

๐Ÿ“ˆ Ordinal

๐Ÿ“ Interval

Match it

Match the design to the test

Tap an item on the left, then its partner on the right.

The study
The correct test
Recap

The big ideas to know

Experiments: lab ยท field ยท natural ยท quasi; IV, DV, extraneous and confounding variables

Design: independent groups (random allocation) ยท repeated measures (counterbalancing) ยท matched pairs

Sampling: random ยท systematic ยท stratified ยท opportunity ยท volunteer

Reliability: test-retest and inter-observer; correlate at +0.80 or above

Validity: internal ยท external (ecological, population, temporal) ยท face ยท concurrent

Levels of data: nominal (categories) ยท ordinal (ranks) ยท interval (equal units)

Significance: p โ‰ค 0.05; Type I = false positive, Type II = false negative

The nine tests: Chi-square/Sign/Chi-square ยท Mann-Whitney/Wilcoxon/Spearman ยท Unrelated t/Related t/Pearson

Critical values: Sign, Wilcoxon and Mann-Whitney โ€” calculated must be โ‰ค critical. All the others โ€” calculated must be โ‰ฅ critical.

Also on spec: pilot studies ยท primary/secondary data and meta-analysis ยท graphs ยท report sections ยท implications for the economy

You have covered the whole of AQA 4.2.3 Research methods โ€” the highest-value topic on the course. Press Finish to see your score.

๐Ÿ†

Mini-lesson complete!

โญโญโญ

You've worked through Research methods for AQA A-level Psychology (7182). ๐ŸŽ‰

Your stars: 0 / 0

Next: test yourself in the Evaluate stage Confidence Quiz, then lock it in with Verify.

๐Ÿ“ฃ Smashed it? Share your score

Challenge a mate to beat your stars, or show a parent how you got on.

โ†’ Back to all subjects