1  Foundations

Intensive longitudinal methods involve repeatedly measuring the same people multiple times, usually dozens to hundreds of times, over a relatively short period. The data let you ask questions that you can’t answer with a single survey, e.g., How does someone’s mood change from hour to hour? When someone is more stressed than usual, do they also feel worse? When someone uses social media more than they usually do, do they feel more connected to others?

1.1 Terminology

Several labels overlap, and different fields prefer different ones. The distinctions are mostly about emphasis and origin, not really about fundamentally different data.

Term Origin and emphasis Typical form
Experience sampling method (ESM) Larson & Csikszentmihalyi (1983) paged adolescents at random moments to capture subjective experience in context. Several random prompts per day, self-report of thoughts, feelings, activity, company.
Ecological momentary assessment (EMA) Stone & Shiffman (1994) coined the term. It generally emphasises real-world settings, momentary reports, and repeated sampling. Includes symptoms, behaviours, physiology. Time- or event-based self-report; often paired with devices.
Ambulatory assessment (AA) Umbrella term (see the Society for Ambulatory Assessment). Covers self-report plus physiological, behavioural and environmental monitoring. Smartphone surveys, accelerometers, ECG, GPS.
Daily diary Also dates to Csikszentmihalyi, but Bolger et al. (2003) provide an early overview. One report per day, usually in the evening, about the day as a whole.
Intensive longitudinal data (ILD) Methodological umbrella term for any data with many closely spaced repeated measures. All of the above, plus passive sensing.

In this primer, ESM and EMA are used interchangeably for momentary self-report protocols, and intensive longitudinal refers to the broader family. A single questionnaire occasion is called a prompt, beep, signal or assessment.

1.2 Why sample experience?

A single questionnaire asks people to summarise themselves at a single point in time. That is, it asks them to compress weeks of shifting thoughts, feelings and behaviour into a single response that stands for how things generally are for them. Such cross-sectional values can be useful, especially when we need to measure general traits or tendencies, but there is a lot that it obscures. Two people who report the same average stress may live very different lives — one steady, the other fluctuating between calm and crisis — and with a single aggregate summary number we would not be able to assess and study these differences.

The case for sampling experience rests on (at least) five overlapping ideas.

  • Less recall bias. Retrospective reports (“How anxious were you last month?”; “How much time do you typically spend on social media?”) ask people to reconstruct the past, and this reconstruction is can innacurate and biased. People do not mentally average their experience; they over-weight its peaks and its most recent moments, and they colour the whole by how they feel at the moment of asking and by their beliefs about the kind of person they are. Momentary reports shorten the recall window to “right now” or “since the last prompt”, so the answer generally suffers less memory degradation. This is especially important for aspects like mood, craving, pain, effort, which are momentary in nature.
  • Ecological validity. Measurement happens where life happens–at work, at home, while commuting, with other people, rather than in a lab. The behaviour you record is the behaviour of interest in its natural setting, not a proxy produced under observation. This also widens the range of situations you capture, because you sample the ordinary moments a participant would never think to mention in an interview.
  • Within-person processes. This is often the primary reason to sample experience. Many theories are really claims about what happens inside a person over time — when this person uses social media more than usual, their mood worsens — but they are often tested with between-person data, which can only show that people who use social media more tend to have worse mood than other people. These are different questions, and their answers can differ in magnitude and even in direction (the within-person and between-person associations do not necessarily have to align; see Section 6.1 and Section 6.9). Repeated measurement of the same person allows you to separate the two, letting you study the process the theory is actually about while still describing how people differ.
  • Dynamics. With closely spaced measures you can study change itself. That is, you can study how a state carries over from one moment to the next (inertia), how much it fluctuates around a person’s own average (instability), how quickly someone returns to baseline after a disturbance (recovery), and whether one variable predicts later change in another (lagged effects).
  • Context. Because each report is time-stamped and made in situ, depending on the study design, it can be linked to where the person was, what they were doing, and who they were with, and to sensor data from the same moment — location, movement, phone use, physiology.
NoteExample: the same average, two different weeks

Two participants both report an average stress of 4 on a 7-point scale. Sampled six times a day for two weeks, the first turns out to respond 4 at almost every prompt, while the second alternates between 1 on quiet days and 7 when they are close to a deadline. A single end-of-study survey would score them identically. The intensive longitudinal data separate a person’s typical level (their between-person value) from their moment-to-moment variability and reactivity (their within-person dynamics).

1.3 What ESM cannot do

  • It does not give causal answers by itself. Temporal order helps, but unmeasured time-varying confounders, measurement error, and the choice of time interval can all create or hide lagged effects (see Section 6.12).
  • It is not free of bias. Participants who are sick or busy respond less often, so missing prompts are rarely random. Repeatedly answering the same questions can change how people answer and possibly also their behaviour too (Section 3.8). Not everyone is willing to participate in this type of research.
  • It cannot capture rare or fast events well unless the design targets them (event-contingent sampling, sensors).
  • It is burdensome and can be expensive to run. Every design choice trades information against participant burden, and burden can erode data quality.
  • Group findings do not automatically describe individuals. A within-person effect averaged across people may not hold for any given person (Section 6.9).