An analysis is never better than the data it rests on. If the data were collected badly — with biased questions, from a sample that represents nobody, or in violation of customers' privacy — no later technique can fix it. In this lesson you will learn where data come from (primary and secondary sources), how to design surveys and questionnaires that measure what you actually want to measure, when a census is worthwhile and when a sample is, the main sampling methods (probability and non-probability), the biases lurking at every step, and the ethical and legal obligations of working with personal data. The connecting thread is a real assignment from Marta: designing NovaMarket's annual satisfaction survey.
Contents
- The assignment: the annual satisfaction survey
- Primary and secondary sources
- Designing surveys and questionnaires
- Census or sample
- Probability sampling
- Non-probability sampling
- Biases: the silent enemies
- Representativeness and sample size: an intuitive view
- Ethics and data protection
The assignment: the annual satisfaction survey
Monday morning. Marta calls the team together:
"Management wants to know how satisfied our customers are this year: with the stores, with the e-commerce and with Club Nova. I need you to design the complete study: which data we use, whom we ask, what we ask, and how we guarantee that the results are reliable and legal. Limited budget, as always."
This assignment contains, in compressed form, every decision in this lesson. Let's take it apart piece by piece.
Primary and secondary sources
The first decision: do we collect new data, or make use of what already exists?
- Primary sources: data you collect yourself, expressly for your study. The satisfaction survey you are about to design is a primary source: it did not exist, and you create it tailored to your question.
- Secondary sources: data that already exist, collected by others or for another purpose. For NovaMarket: the purchase receipts (generated for billing, not for analysis), the delivery records, the Club Nova sign-ups, and also external sources such as INE (Spain's national statistics institute) figures on household consumption or retail-sector industry studies.
| Aspect | Primary | Secondary |
|---|---|---|
| Cost and time | High | Low (they already exist) |
| Fit to your question | Total (you design them) | Partial (collected for something else) |
| Quality control | In your hands | You inherit the flaws of the source |
| NovaMarket example | Satisfaction survey | Transaction history; INE data |
In professional practice you combine both: for Marta's assignment, the receipts (secondary) will tell you what customers do (how much they spend, whether they come back, whether they leave), but only the survey (primary) will tell you what they think and why. The golden rule: exhaust the secondary sources first — they are cheap and often answer half the question — and collect primary data only for what is left.
Designing surveys and questionnaires
Designing a questionnaire looks trivial ("they're just questions") and is among the hardest things to do well: every word can alter the answers.
Question types
- Closed-ended: the respondent chooses among predefined options. Easy to analyze (they produce clean nominal or ordinal variables). Example: "How often do you shop at NovaMarket? ☐ Several times a week ☐ Weekly ☐ Every two weeks ☐ Less often".
- Open-ended: free-text answers ("What would you improve about your usual store?"). Rich in nuance, but costly to analyze. Use them sparingly: one or two at the end.
- Semi-closed: predefined options plus an "Other: ______".
Likert scales
The go-to tool for measuring attitudes. A statement is presented and the respondent indicates their level of agreement on an ordered scale, typically with 5 points:
"NovaMarket's fresh produce is of good quality." 1 = Strongly disagree · 2 = Disagree · 3 = Neither agree nor disagree · 4 = Agree · 5 = Strongly agree
Good practices with Likert scales:
- Label every point (not just the endpoints), so that all respondents interpret the scale the same way.
- Decide deliberately whether to include a neutral midpoint (odd-numbered scale) or force respondents to take a side (even-numbered). Including it is the usual choice.
- Keep the same direction throughout the block (5 should always be "best"), or, if you alternate positive and negative statements to catch automatic responding, remember to reverse the codes before analyzing.
- Remember from the previous lesson: a Likert item produces an ordinal variable; treating its values as numbers is a convention to be handled with care.
Wording: the rules that keep you from measuring garbage
| Rule | Bad example | Corrected version |
|---|---|---|
| One idea per question (avoid "double-barreled" questions) | "Are you satisfied with the price and quality of our products?" | Two questions: one on price, one on quality |
| Do not lead the respondent | "Do you agree that our excellent delivery service is fast?" | "How would you rate the speed of the delivery service?" |
| Plain language, no jargon | "Do you view our omnichannel proposition favorably?" | "Do you find it easy to combine shopping in-store and online?" |
| Avoid confusing negatives | "Don't you think we shouldn't change the opening hours?" | "Which opening hours would you prefer?" |
| Exhaustive, mutually exclusive options | Age: "18-30, 30-45, 45-60" (where does a 30-year-old go? and a 65-year-old?) | "18-29, 30-44, 45-59, 60 or over" |
| Sensitive questions at the end, and only if necessary | Opening by asking for household income | Income in brackets, at the end, with a "Prefer not to say" option |
Professional tip: always run a pilot test (10-20 people) before launch. You will catch ambiguous questions, missing options and how long it takes to respond (more than 8-10 minutes sends abandonment soaring).
Census or sample
You have the questionnaire. Whom do you send it to? Recall lesson 01-01: the target population for the assignment is NovaMarket's customers (realistically, we will start with the 380,000 Club Nova members, for whom we have contact details). Two options:
- Census: ask everyone. It eliminates sampling error, but it is expensive, slow and often unnecessary. It makes sense for small, accessible populations: for example, surveying the 42 store managers.
- Sample: ask a well-chosen part. With a few thousand well-selected responses you can estimate the club's satisfaction precisely enough to make decisions.
A nuance that surprises people: a badly executed census can be worse than a good sample. If you send the survey to all 380,000 members and 9,000 reply (whoever feels like it — usually the very happy and the very angry), you do not have a census: you have an enormous, biased, self-selected sample. Better a sample of 2,000 randomly chosen members with good follow-up (reminders, an incentive) that achieves a high response rate.
Decision for the assignment: sample. Now, how do we choose it?
Probability sampling
Sampling is probabilistic when every unit in the population has a known, non-zero probability of being selected, and the selection is made by chance, not by a person. It is the only family of methods that later allows you to compute margins of error on solid ground (Module 5). It requires a sampling frame: the complete list of the population (here, the Club Nova database — an enormous advantage).
Simple random sampling
Every member has the same probability of being drawn, as in a raffle. Operationally: the 380,000 members are numbered and a random number generator picks, say, 2,000.
- Advantage: it is the conceptual gold standard, simple and unbiased.
- Drawback: by pure chance, it can underrepresent small groups (you might end up with very few online-only customers).
Systematic sampling
A random starting point is chosen and then 1 in every \( k \) units is taken, where \( k \) is the sampling interval:
\[ k = \frac{N}{n} = \frac{380{,}000}{2{,}000} = 190 \]
A number between 1 and 190 is drawn (say it comes out 37) and members 37, 227, 417, 607... of the list are selected. Very convenient with long lists. Caution: if the list has a periodic pattern that coincides with \( k \), the sample becomes biased (rare, but check).
Stratified sampling
The population is divided into strata (homogeneous groups, relevant to the study) and sampling is done at random within each one. For NovaMarket, a natural stratum is the purchase channel. With proportional allocation (each stratum weighs in the sample what it weighs in the population):
| Stratum | Members (N) | Proportion | Sample (n = 2,000) |
|---|---|---|---|
| In-store only | 266,000 | 70% | \( 2{,}000 \times 0.70 = 1{,}400 \) |
| Mixed (in-store + online) | 76,000 | 20% | \( 2{,}000 \times 0.20 = 400 \) |
| Online only | 38,000 | 10% | \( 2{,}000 \times 0.10 = 200 \) |
| Total | 380,000 | 100% | 2,000 |
- Advantage: it guarantees that every important group is represented and improves precision; it also lets you report results by stratum.
- When to use it: when the population contains clearly different groups with respect to what you are measuring (online customers' satisfaction most likely differs from in-store customers') and you know which stratum each unit belongs to. (If a small stratum is of particular interest, it can be oversampled and reweighted afterwards; just keep the idea in mind.)
Cluster sampling
The population is divided into clusters (groups that are internally varied "mini-populations", typically defined by physical proximity), a few whole clusters are drawn at random, and the study is carried out within them. Example: to survey customers in person at the store door, instead of sending interviewers to all 42 stores, 8 stores are drawn and the surveying happens there.
- Advantage: major logistical savings.
- Drawback: less precision for the same \( n \) (customers of the same store resemble one another), which calls for somewhat larger samples.
Strata and clusters are often confused. The key:
| Stratified | Cluster | |
|---|---|---|
| Internally, the groups are... | Homogeneous (similar to each other) | Heterogeneous (varied, like the population) |
| How many groups enter the sample? | All of them (you sample within each) | Only some (drawn whole) |
| Motivation | Precision and group representation | Saving logistical cost |
| NovaMarket example | Purchase channel (in-store/mixed/online) | Stores for the in-person survey |
Non-probability sampling
When there is no sampling frame, no time and no budget — or the goal is exploratory — you resort to methods where chance does not govern the selection. They are legitimate as long as you use them knowing what they do not allow: generalizing with margins of error.
- Convenience sampling: you survey whoever is easy to reach. Example: asking the customers who happen to walk through the Madrid-Centro store on Tuesday morning. Fast and cheap; the sample is probably biased (who shops on a Tuesday morning? Not many people who work office hours).
- Quota sampling: the "disciplined" version of convenience: quotas are set to mimic the population (for example, 60% women / 40% men, so many per age bracket) and the interviewer fills each quota with whomever they find. It improves the appearance of representativeness, but within each quota it is still the interviewer choosing, not chance.
- Snowball sampling: each participant recruits others. Useful for populations that are hard to locate without a list — for example, if NovaMarket wanted to interview celiac customers to design its gluten-free range, it would start with a few identified ones and ask them for contacts of others.
When to use each family?
| Situation | Recommended method |
|---|---|
| A population list exists and you want estimates with a margin of error | Probability sampling (whichever fits logistically) |
| Relevant groups that are very different and identifiable | Stratified |
| Strong logistical constraints, units physically grouped | Cluster |
| A long list ordered in a "neutral" way | Systematic |
| Quick exploratory study, questionnaire pilot | Convenience |
| No sampling frame, but you want to mimic the population structure | Quota |
| Hidden population, no list possible | Snowball |
For Marta's assignment, the reasoned decision is: stratified sampling by channel over the Club Nova database, delivered by email, with reminders and a small incentive (a coupon) to lift the response rate.
Biases: the silent enemies
A bias is a systematic error: it pushes the results in the same direction every time, and — unlike random sampling error — it cannot be fixed by enlarging the sample. The three most dangerous:
Selection bias
The way the sample is chosen excludes or underrepresents part of the population. Examples in our assignment:
- Surveying only Club Nova members leaves out customers without a loyalty card (perhaps the least loyal and most critical). The honest conclusion: the results speak "about the members", not "about the customers".
- Surveying only by email excludes members with no email address on file (often older ones).
Non-response bias
Even if the selection is perfect, those who respond differ from those who do not. If 25% respond, and dissatisfied customers respond more (they want to complain) or less (they have already gone to a competitor), the measured mean satisfaction drifts away from the real one. Defenses: a short questionnaire, reminders, incentives, and comparing the respondents' profile with the population's (do they match on age, channel, spend? — if not, bad sign, and it can partly be corrected by reweighting).
Social desirability bias
People tend to answer what sounds good. Asked face to face by an interviewer in a NovaMarket vest, a customer will soften their criticism; asked "do you buy healthy products?", they will exaggerate their virtue. Defenses: genuine anonymity (and communicating it), self-completion (web beats face-to-face for sensitive topics), neutral wording that legitimizes every answer.
Other biases worth knowing by name: survivorship bias (analyzing only current customers while forgetting those who left — the receipts only contain the people who are still shopping) and coverage bias (the sampling frame does not match the population, as in the email case).
Representativeness and sample size: an intuitive view
A sample is representative when it reproduces, on a small scale, the composition and diversity of the population in everything relevant to the study. Two key ideas, both counterintuitive:
- Representativeness depends on HOW the sample is chosen, not on how many are in it. A random sample of 1,500 members tells you more about the club than 40,000 self-selected voluntary responses. Size does not cure bias: it just makes you measure the biased value more precisely.
- What matters is the absolute size, not the percentage of the population. For a given precision, you need roughly the same number of respondents to estimate the satisfaction of 380,000 members as of 4 million. That is why national polls work with 1,000-2,500 interviews.
On size, three intuitions we will develop formally in the lesson Errors, Power and Sample Size:
- More sample ⇒ more precision (a smaller margin of error), but with diminishing returns: to double the precision you must quadruple the sample, not double it.
- If you want reliable results by subgroup (by channel, by region), each subgroup needs a sufficient size of its own — this is what drives study costs up the most.
- Rough orders of magnitude: with about 400 random responses, the margin of error of a percentage is around ±5%; with about 1,100, around ±3%. (Where these numbers come from you will see in Module 5; for now, use them as a pocket reference.)
With all this, the team proposes to Marta: 2,000 invitations stratified by channel, targeting at least 800-1,000 responses, with results reported overall and by channel.
Ethics and data protection
Collecting data about people carries ethical and legal obligations. In Spain and the EU, the framework is the GDPR (General Data Protection Regulation) and Spain's LOPDGDD. Principles every analyst must internalize:
- Data minimization: collect only what is necessary for the declared purpose. If the survey will not analyze by income, do not ask about income.
- Purpose and transparency: respondents must know who is collecting the data and what for, and give their consent where required. Club Nova data may only be used for purposes compatible with those accepted at sign-up.
- Anonymisation and pseudonymisation: whenever the analysis allows it, separate the responses from the identity. Anonymising means irreversibly breaking the link to the person; pseudonymising means replacing the identity with a code (reversible via the lookup table, which must be kept separately and securely). Beware: removing the name is not enough — the combination of postal code + age + usual store can re-identify someone in a small town. With very small groups, aggregate (for example, do not publish results for strata below a minimum number of responses).
- Security and access: microdata (individual responses) are stored encrypted with restricted access; what circulates around the company are aggregates.
- Storage limitation: define how long responses are kept, and delete them afterwards.
Important caveat: this section offers general principles of good statistical practice, not legal advice. Before launching any collection of real personal data, the design (consent wording, legal basis, retention periods) must be reviewed by your organization's data protection officer or a compliance/GDPR professional. At NovaMarket, Marta launches no survey without the DPO's sign-off — do the same.
Common Mistakes and Tips
- Believing that a huge sample makes up for bad sampling. This is conceptual error number one. Large bias + large sample = a wrong conclusion held with great (false) confidence.
- Double-barreled or leading questions. Review every question looking for the conjunction "and" and for evaluative adjectives ("our excellent service...").
- Ignoring non-response. Always report the response rate and whether the respondents resemble the population; a 15% response rate with no bias analysis should set off every alarm.
- Confusing strata with clusters. Remember: strata = homogeneous groups, all of them are in, you draw within each; clusters = heterogeneous groups, some are drawn whole.
- Generalizing from convenience samples. "We asked people at the Madrid-Centro store" may do for exploring ideas, never for headlining "80% of our customers think that...".
- Neglecting privacy. Publishing tables with cells of 2-3 identifiable people, or keeping named responses in a shared spreadsheet, are serious incidents. When in doubt, ask the DPO.
- Professional tip: document the design (population, frame, sampling method, response rate, questionnaire text) in a technical datasheet, the way polling institutes do. It is your study's quality guarantee, and it protects you when someone questions the results months later.
Exercises
Exercise 1
Classify the sampling method in each NovaMarket scenario (simple random, systematic, stratified, cluster, convenience, quota or snowball) and justify it in one sentence:
- To audit the quality of the receipt data, 1 receipt in every 500 in the month's database is taken, starting from one drawn at random among the first 500.
- For a quick test of the new questionnaire, 15 head-office employees who happened to be in the cafeteria are asked for their opinion.
- To estimate satisfaction with deliveries, online customers are split into "urban area" and "rural area" and 300 are drawn from each group.
- To study the in-store shopping experience, 6 of the 42 stores are drawn at random and every customer who agrees is surveyed at the exit of those 6 stores for a week.
- To interview customers who left Club Nova years ago and no longer have valid contact details on file, the first few who are located are asked to provide contacts of other former members they know.
Exercise 2
This draft question has slipped into the satisfaction questionnaire:
"Don't you agree that NovaMarket's prices and product variety have improved thanks to our hard work? ☐ Yes ☐ No"
Identify at least three wording problems and rewrite the block properly (you may split it into more than one question and change the response format).
Exercise 3
The survey was emailed to 2,000 members (stratified by channel) and 640 responded. Comparing profiles reveals: the respondents' mean age is 51 versus 44 for the club population, and 78% of those who gave a score are members with more than 5 years' tenure (who make up 45% of the club). The mean satisfaction came out at 8.1.
- What response rate was achieved?
- Which specific biases threaten the 8.1 figure? Name at least two using the terminology of this lesson.
- Would expanding the mailing to 20,000 members solve the problem? Why?
Solutions
Solution 1
- Systematic: a fixed interval (1 in 500) with a random start.
- Convenience: you pick whoever is at hand; valid only as an exploratory pilot.
- Stratified: homogeneous groups defined in advance (urban/rural) with a random draw within each; moreover, 300 per stratum is non-proportional allocation, designed to allow comparing the groups.
- Cluster: a few whole stores are drawn (heterogeneous groups defined by logistics) and the study happens inside them. Common mistake: answering "stratified"; the giveaway is that only some stores enter the sample. (Note also that "customers who agree" adds a self-selection component within each store.)
- Snowball: each person located recruits the next ones; appropriate because no sampling frame exists for that population.
Solution 2
Problems (three are enough):
- Double-barreled question: it mixes prices and variety in a single question.
- Leading wording: "thanks to our hard work" and the framing "Don't you agree that...?" push towards a positive answer (and trigger social desirability).
- Confusing negative: "Don't you agree...?" makes it ambiguous what answering "No" means.
- Poor response format: a Yes/No captures no degrees; a scale is called for.
A possible rewrite:
"Please indicate how much you agree with the following statements (1 = Strongly disagree, 2 = Disagree, 3 = Neither agree nor disagree, 4 = Agree, 5 = Strongly agree): a) NovaMarket's prices are reasonable. 1 · 2 · 3 · 4 · 5 b) NovaMarket's product variety is adequate. 1 · 2 · 3 · 4 · 5"
Solution 3
- Response rate: \( 640 / 2{,}000 = 0.32 \), that is, 32%.
- Mainly non-response bias: the respondents do not resemble the population (older, and far more long-tenured/loyal: 78% versus 45%); if veteran members are more satisfied, the 8.1 will be inflated. You can also point to the underlying coverage/selection bias: only members with a valid email address — on top of the fact that the whole study speaks about club members, not customers in general.
- No. Expanding the mailing multiplies the responses but does not change who tends to respond: the imbalance (older, veteran members) would reproduce itself at a larger scale. The bias is systematic and size does not correct it; the right levers are raising the response rate (reminders, incentives, a short questionnaire), reweighting the results according to the known population profile, and complementing the email channel. Common mistake: believing that "more data" is always the answer — this is exactly the confusion between random error (which does shrink with \( n \)) and bias (which does not).
Conclusion
You have closed the introductory module by learning the step that precedes every analysis: getting good data. You now know how to combine secondary sources (cheap, already in existence) with purpose-built primary sources; how to write questionnaires with clean questions and well-constructed Likert scales; how to decide between a census and a sample; how to choose sensibly among the probability sampling methods (simple random, systematic, stratified, cluster) and the non-probability ones (convenience, quota, snowball); how to recognize selection, non-response and social desirability bias, with the lesson etched in: size does not cure bias; and how to work with personal data under GDPR principles, always with review by a data protection professional. NovaMarket's satisfaction survey is now designed and in the field; when the responses come in — along with the receipts, the deliveries and the rest of the company's data — they will need to be summarized and made sense of. That is what Module 2 is about, and it opens with Measures of Central Tendency: the mean, the median and the mode, and the art of knowing which one to give Marta on each occasion.
Statistics Course
Module 1: Introduction to Statistics
Module 2: Describing Data
- Measures of Central Tendency
- Measures of Dispersion
- Measures of Position and Outliers
- Graphical Representation of Data
Module 3: Probability
Module 4: Probability Distributions
- The Binomial Distribution
- The Normal Distribution
- Other Important Distributions
- The Central Limit Theorem
Module 5: Statistical Inference
Module 6: Data Analysis
- Correlation Analysis
- Regression Analysis
- Analysis of Variance (ANOVA)
- Categorical Data Analysis: Chi-Square
