3  Measurement and questionnaire design

ESM questionnaires are answered dozens of times, often in a few seconds, on a small screen, in the middle of something else. Because of the limited space and different temporal scope, scales and items that work well in a one-off survey tend not to be compatible with ESM designs.

3.1 Writing momentary items

  • Fix the time frame and state it. “Right now I feel…” measures a state at the prompt. “Since the last prompt, I…” measures the interval in between and supports questions about events and behaviour. Do not mix frames within a block without clear visual cues.
  • One idea per item. “I feel tense and irritable” cannot be answered by someone who is tense but calm.
  • Make every item applicable at every prompt. “I enjoy what I am doing” is odd when someone is doing nothing. Use branching: ask about the activity first, then rate it.
  • Keep wording short and concrete. Items are read on a phone, often dozens of times. Avoid negations and double negatives.
  • Expect within-person variation. An item that barely changes within people (e.g., a trait-like statement) adds burden and little information.
  • Order deliberately. Put the most time-sensitive items (e.g., current affect) first, before items about context or events that may change how people rate their mood. Randomising item order within a block can reduce patterned responding but makes the questionnaire less predictable; many studies keep a fixed order.
  • Reuse where possible. The ESM Item Repository is an open collection of items used in published ESM studies, with wording, response scales and translations. ESM-Q, a consensus-based quality assessment tool for ESM items (Eisele et al., 2025) can help to evaluate new or existing items.

3.2 Response formats

Format Strengths Watch out for
Likert (e.g., 1–7) Fast on touch screens, familiar, ordinal categories are clear. Coarse; few categories limit detection of small within-person changes; ceiling and floor effects.
Visual analogue slider (0–100) Fine-grained; good for detecting small fluctuations; treated as continuous. Default handle position anchors answers. Require the participant to touch the slider, and do not use a midpoint default that looks like a real answer.
Binary/categorical Very fast; ideal for context (alone yes/no, location). Needs logistic or multinomial models; rare categories give sparse data.
Open text Rich descriptions of events. Slow to type; low completion; needs coding.

3.3 Single- or multi-item measures?

ESM rewards brevity, so many constructs are measured with a single item (“Right now, how stressed do you feel?”) rather than a multi-item scale. The cost is that you cannot estimate its reliability in the usual way — with one item, measurement error is confounded with genuine within-person change, and there is no internal-consistency check (Section 3.5).

Prefer a multi-item measure when the construct is abstract or multi-faceted, when clinical content coverage matters, or when you need to estimate within-person reliability or fit a measurement model. Prefer (or accept) a single item when the construct is simple and unambiguous, when there is strong precedent for the exact wording, and when burden is the binding constraint.

3.4 Baseline and trait measures

Not everything is measured repeatedly. Alongside the momentary questionnaire, most studies collect a baseline (intake) battery once. This can include things like demographics, and trait or dispositional questionnaires (e.g., neuroticism, trait rumination) that change slowly or not at all over the study. These become person-level (level-2) variables in analysis — they describe how people differ, and they are the natural moderators of within-person effects (the cross-level moderation questions of Section 2.1). In the running example, trait neuroticism is measured once and used to test whether stress reactivity is stronger in more neurotic people.

3.5 Reliability and validity

Because the data are nested, reliability has two different meanings:

  • Between-person reliability: how consistently a scale ranks people by their average level. It increases with the number of prompts, because person means average over many occasions.
  • Within-person reliability: how well the scale detects a person’s change from one moment to the next. This is usually lower.

Two common approaches estimate both: generalizability theory, which partitions variance into persons, occasions, items and their interactions (Cranford et al., 2006), and multilevel confirmatory factor analysis with separate within and between omegas (Geldhof et al., 2014). Both need at least two or three items per construct; single items cannot be evaluated this way. A worked example is in Section 5.8.

Validity also has two levels. A scale’s factor structure within persons (how items covary across moments) need not match its structure between persons (how people’s averages covary).

3.6 Context and passive data

Self-reported context (location, activity, company) is cheap to collect but has its limits. Passive sensing goes further, using the sensors and logs already on a smartphone or wearable to record behaviour continuously, with greater accuracy, and with almost no burden on the participant. It fills the long gaps between prompts and captures things people cannot report accurately from memory: how much they moved, how they slept, how far they travelled, how much they used their phone.

Source Example features Considerations
GPS/location Time at home, number of places visited, location entropy, distance travelled Highly identifying; battery drain; indoor gaps; strict privacy protections needed
Accelerometer Steps, activity intensity, sedentary time, sleep estimates Phone may not be carried; wrist devices better for sleep and activity
Phone use logs Screen-on time, unlocks, app categories, night-time use Operating system restrictions (especially iOS) limit what can be logged
Communication Call and message counts (not content) Often restricted by app stores; consent from third parties is an ethical question
Audio Ambient sound snippets, time spent talking (e.g., the Electronically Activated Recorder) Records non-consenting bystanders; transcription burden
Wearable physiology Heart rate, heart rate variability, skin conductance, skin temperature Movement artefacts; must separate physical from psychological arousal

Passive streams are powerful, but they shift the work from data collection to data processing, and they carry risks that self-report does not.

From raw signal to feature. A sensor does not provide you with a measure for “time at home”, “total hours of smartphone use”, or “sleep duration”; it provides a stream of coordinates, smartphone usage events, or accelerations that you must turn into a feature. Every step — the window length, the aggregation, the threshold, the algorithm — is a researcher decision. This is the garden of forking paths again (Section 2.7). I recommend preregistering the features that you will use in your analyses before you look, and build on validated pipelines rather than bespoke ones where you can. A dedicated preregistration template for passive smartphone measures is available (Langener et al., 2024).

Aligning sensors to prompts. To link a passive feature to a momentary report you must choose a time window — steps in the 60 minutes before each prompt, mobility since the previous prompt, sleep the night before. Fix these windows in advance, for the same reason you fix the sampling scheme.

Missingness is the norm, and it is not random. Phones are switched off, left on a desk, or run out of battery; operating systems (especially iOS) restrict background logging; wearables are taken off to charge. The result is gaps that differ by platform and by person, often correlated with the very states you study — a flat, withdrawn day can also be a low-data day. Treat passive missingness as seriously as prompt non-compliance, and report it per stream.

3.7 Piloting the questionnaire

Pilot before you launch. A useful pilot does three things:

  • Cognitive pretesting. Have a handful of people from the target population complete a few prompts and think aloud, or interview them afterwards: does each item mean what you intended, at a glance, on a phone, mid-activity? This catches ambiguous wording, items that do not apply, and confusing branching before they cost you thousands of responses.
  • Check the numbers each item produces. Run a short pilot burst and inspect, per item, completion time, missingness, and — importantly — within-person variance. An item that barely moves within people (Section 3.1), or that everyone answers at the ceiling, is not useful. This is also where you confirm that a single item behaves well enough to stand alone (Section 3.3).
  • Test the whole procedure end to end. Time the full questionnaire on the actual device and app, and check that prompts fire on schedule, that notifications and reminders work, that data are recorded in the right format, and that response windows and branching behave as intended.

Treat the pilot as a rehearsal of the entire protocol, not just a proofread of the items, and budget time to act on what it tells you. General guidance on building and testing daily-life questionnaires is available in Myin-Germeys & Kuppens (2026).

3.8 Reactivity and response fatigue

Repeated measurement can change the very thing it measures, and can decline in quality as a study proceeds.

Measurement reactivity means that being asked changes what is measured. Self-monitoring is itself an intervention for some behaviours and repeatedly rating a symptom may heighten attention to it. Reactivity is a threat to validity when your estimate of daily-life experience is shifted by the act of measuring.

Response fatigue (and careless responding) means the answers themselves deteriorate with repetition. This appear as less time per item, less variability, more straightlined or identical responses, or more skipped prompts.

The evidence suggests reactivity is modest for many momentary states (Shiffman et al., 2008; Trull & Ebner-Priemer, 2020), but do not assume it — check it in your own data:

  • Model the effect of study day and prompt number on item means, within-person variances, and completion times; a systematic trend is a warning sign.
  • Compare early and late days, and inspect careless-responding indicators (long strings of identical answers, implausibly fast completions).
  • Ask participants at debriefing whether taking part changed how they felt or behaved (this doubles as the exit check of Section 3.4).

Design choices are the main defence against reactivity and response fatigue. Keep each prompt short and focused (Section 2.4), avoid an over-long study, and invest in onboarding and engagement so that compliance and effort hold up.