2 Research design
Design starts from the research question and the timescale of the process you care about. Sampling frequency, study length, and questionnaire content should all follow from those two things, not from what the software makes easy or what other studies have done.
2.1 Research questions that fit ESM
When thinking about research questions that fit ESM it can help to name which level and which kind of effect your question is about.
| Question type | Example |
|---|---|
| Descriptive, momentary | How much time do students spend alone, and how do they feel then? |
| Within-person association | When people are more stressed than usual, is their negative affect higher? |
| Between-person association | Do people who are stressed on average have higher average negative affect? |
| Cross-level moderation | Is the stress–affect link stronger in people high in neuroticism? |
| Temporal (lagged) | Does stress at one prompt predict affect at the next, controlling for prior affect? |
| Dynamics as individual differences | Is emotional inertia higher in people with depression? |
| Person-specific structure | For this client, which symptoms precede which? |
| Intervention effect in daily life | Does a prompt to take a walk reduce craving in the next hour? |
A useful exercise is to write the question as a sentence that says who changes and relative to what. For example, “when a person is more stressed than their own average…” is a within-person question. “People who are more stressed than other people…” is a between-person question. Interestingly, Hamaker (2012) argues that most psychological theories are implicitly about within-person questions, but most data collected tend to be at the between-person level.
2.2 Match the design to the timescale
Every process has a natural tempo. Affect can change within minutes; sleep quality is a nightly event; self-esteem may drift over weeks. Your sampling interval sets the lag at which you can see effects. With prompts every two hours you cannot detect a process that plays out in 20 minutes. As a general rule of thumb:
- Fast-changing states (mood, craving, pain, social interaction): several prompts per day.
- Daily events and summaries (sleep, conflict, alcohol use, daily goals): one or two diaries per day.
- Slow processes (symptom change over treatment, adjustment to a life event): measurement-burst designs, with a week of ESM repeated every few months.
2.3 Sampling schemes
Wheeler & Reis (1991) distinguished three basic protocols by what triggers a report: interval-contingent, in which participants report at fixed, regular intervals; signal-contingent, in which they report when a device prompts them (at random or semi-random times); and event-contingent, in which they report whenever a designated event occurs.
Time-based schemes (interval- and signal-contingent) aim to sample representative moments of everyday life, so they suit questions about typical states and their moment-to-moment dynamics. Event-based schemes deliberately oversample particular occasions, so they suit questions about specific, often rare, episodes.
Time-based (signal-contingent and interval-contingent)
Here the schedule, not the participant, decides when a report happens. There are two common variants to this approach: interval-contingent approaches and signal-contingent approaches. In a meta-analysis of 477 EMA studies, signal-contingent schedules were used in 58% of studies and interval-contingent schedules in 18% (Wrzus & Neubauer, 2023).
- Interval-contingent (fixed-interval) designs place prompts at set clock times: 10h00, 12h00, 14h00, and so on. Equal spacing is convenient because it keeps the lag between observations constant, which simplifies lagged and time-series models. The cost of this convenience, however, is predictability. Participants can anticipate a fixed schedule and, consciously or not, change their behaviour or bring a report forward, which undermines the “momentary” logic of the method.
- Signal-contingent designs prompt at random or semi-random times, so a report cannot be anticipated. Pure random scheduling, however, can bunch several prompts into a short window and leave long gaps elsewhere. Most studies therefore use stratified (semi-)random sampling in which the waking day is divided into equal blocks (e.g., six two-hour blocks) and one prompt is placed at a random time within each block, usually with a minimum gap between consecutive prompts (e.g., 30–45 minutes).
Event-based (event-contingent)
Here the participant, not the schedule, initiates the report whenever a defined event occurs (e.g., after each social interaction lasting more than ten minutes, each cigarette, each argument, each binge episode). This approach captures rare or brief events that time-based prompts would almost always miss, and it records them close to when they happen rather than at the next scheduled beep. But there are still tradeoffs to this approach. You cannot distinguish a moment when no event occurred from one where an event occurred but went unreported, so compliance is harder to audit than with time-based prompts. The event definition must be very clear and easy to recognise in the moment. Self-initiated reporting can also make people more attentive to the target behaviour, a form of reactivity (Section 3.8). Two fixes to these issues are typically used. The first involves pairing event reports with time-based prompts so you also have non-event moments for comparison. The second involves letting the software detect the event and trigger the prompt automatically (for example, a logged app-open, a completed call, or a sensor threshold), which removes the burden of remembering to self-report.
Continuous and passive
Sensors record continuously in the background, without direct participant effort (e.g., accelerometry for movement and sleep, GPS for mobility, phone logs for screen time and communication, wearables for heart rate or skin conductance, etc.). Passive sensing lowers the participant burden and self-report bias, but it does raise issues of privacy, battery and storage. It also remains a challenge to link the passive event stream to what a moment actually meant to the person.
Hybrid and adaptive schemes
Combining schemes lets you measure processes that run at different tempos, or target specific events, within one design:
- Morning and evening anchors (sleep report on waking, day summary at night) combined with random daytime prompts.
- Sensor-triggered prompts: ask about mood when the phone detects a period of unusual inactivity or a spike in heart rate not explained by movement.
- Planned missingness: each prompt shows a random subset of items, lowering burden while all items are still measured often enough across the study.
2.4 Design parameters and the burden trade-off
Once the scheme is chosen, you need to decide how long to run it for, how many prompts, how many items per prompt, how far apart they should be and so on. There is no universally correct combination. The table below provides some common ranges and the main considerations.
| Parameter | Common range | Considerations |
|---|---|---|
| Study length | 7–14 days (up to 30+) | Include weekdays and weekends. Longer studies give more reliable person-specific estimates but more fatigue and drop-out. |
| Prompts per day | 3–10 | Set by the timescale of the process. Typical: about 6 per day for 7 days for ESM (Wrzus & Neubauer, 2023) and 1 per day for a diary study. |
| Items per prompt | 5–30 | Aim for under 2–3 minutes to complete. It is typical to have 5 - 15 items in an ESM survey. |
| Response window | 10–60 min | How long a prompt stays open. Short windows keep reports momentary; long windows raise compliance but blur timing. |
| Reminders | 0–2 | Re-notify after 5–10 min if unanswered. Helpful, but annoying if overused. |
| Sampling window | e.g. 09:00–22:00 | Ideally tailored to the target population’s waking hours. |
| Minimum gap | 15–60 min | Prevents consecutive prompts so close that they measure the same moment twice. |
Keep the per-prompt questionnaire short and focused on your primary question. It is usually better to add a few more prompts than to add more items per prompt. If you need many constructs, consider planned missingness or a separate morning/evening questionnaire for slower-changing variables.
2.5 How big should the study be?
Sample size in ESM is a function of two numbers: how many people you recruit, and how many prompts each person answers.
- More prompts per person increases precision within a person. Within-person associations (does this person’s affect rise when they are more stressed than usual?) and person-specific dynamics (one participant’s inertia or variability) are powered mainly by the number of observations per person. Within-person effects are often well powered with only a moderate number of people, because each person contributes many observations.
- More people increases precision between persons. Between-person associations, and cross-level moderation (is the stress–affect link stronger in people high in neuroticism?), depend mainly on the number of participants; once each person has a reasonable number of prompts, adding more prompts helps these effects very little.
- Plan for the prompts you will actually get. Power depends on the answered prompts, not the scheduled ones, so multiply your design by an expected compliance rate (compliance averages around 80%; Wrzus & Neubauer (2023)) before judging whether it is enough.
Because closed-form power formulas rarely fit the models ESM studies use, the standard approach is simulation in which you generate data under plausible assumptions, fit the model you plan to run, and repeat many times to see how often it recovers the effect. Later in this primer after introducing multilevel models in Chapter 6, a worked example of this simulation-based approach is presented in Section 6.11.
2.6 Specialised designs
Measurement-burst designs
Intensive “bursts” (e.g., 7 days of ESM) are repeated at wider intervals (e.g., every three months). They connect short-term dynamics to long-term change. For example, does emotional reactivity to stress decrease over a course of therapy?
Ecological momentary interventions and micro-randomised trials
Ecological momentary interventions deliver support in daily life. A just-in-time adaptive intervention (JITAI) decides when to intervene based on momentary data (Nahum-Shani et al., 2018). The micro-randomised trial (Klasnja et al., 2015) randomises, at each decision point, whether a person receives an intervention prompt. Because treatment is randomised within person many times, it gives unconfounded estimates of proximal (short-term) effects and how they change over time.
N = 1 and small-N clinical designs
If the aim is to describe one person, you need many more observations from that person (often 50–100+ per variable) and roughly equal spacing. Group size matters less.
2.7 Preregistration
ESM studies carry an unusually large number of “researcher degrees of freedom”– which prompts count as valid, which compliance cut-off to apply, whether to lag across nights, how to centre predictors, which random effects to include, what to do when a model will not converge. Making these decisions after seeing the data, even in good faith, inflates false positives and makes confirmatory claims hard to trust. Preregistration — writing the decisions down and time-stamping them before you look — is the main safeguard against p-hacking and HARKing.
General-purpose preregistration templates were not built with intensive longitudinal data in mind, and leave out much of what makes an ESM study reproducible. Kirtley et al. (2021) fill that gap with a template and tutorial designed specifically for ESM; it adapts the general OSF preregistration template and is freely available at osf.io/2chmu. The paper walks through each field with worked examples. A good ESM registration covers the sampling scheme, the items (with exact wording and response scales), planned exclusions and compliance criteria, the preprocessing pipeline, the models (random-effects structure, centering, and a convergence fallback), and the power rationale.