English:Experimental Design

Experimental Design
Introduction
Experimental design is the careful planning of a study in which you deliberately change one or more conditions and measure what happens. A strong design helps you separate a real treatment effect from coincidence, bias, measurement error, and other sources of variation. Experimental design is therefore central to Statistics, the scientific method, Research methods, Biology, Chemistry, Physics, Psychology, engineering, medicine, agriculture, and many other fields.
For learners in Grades 11–13, the key question is not simply "Did two groups have different averages?" but "Does the design justify a causal conclusion?" Good experiments answer that question by defining variables precisely, using fair comparisons, controlling unwanted influences, assigning treatments appropriately, collecting enough independent evidence, and analyzing uncertainty.

A design of experiments can be thought of as a plan that connects a research question to data. You decide what will be changed, what will be measured, which experimental units will receive which conditions, what will be held constant, and how random variation will be handled. In statistical design, planning happens before the data are collected because weaknesses in the design often cannot be repaired later.
Learning Goals
By the end of this aiMOOC, you should be able to distinguish an experiment from an observational study, write a testable hypothesis, identify explanatory and response variables, choose appropriate controls, explain randomization and replication, use blocking when it improves precision, recognize confounding and bias, interpret factorial designs and interactions, plan ethical experiments, and justify what conclusions a design can support.
You should also be able to evaluate an experiment critically. That means asking whether the comparison is fair, whether the experimental units are truly independent, whether the measurement is valid and reliable, whether the analysis matches the design, and whether the claimed conclusion goes beyond the evidence.
From Questions to Testable Experiments
Experiments and Observational Studies
In an observational study, researchers measure variables without deliberately assigning the explanatory condition. Observational studies can reveal patterns and associations, but an observed association can be explained by confounding variables. In a true experiment, the researcher deliberately imposes at least one treatment or condition on experimental units.
This distinction matters because controlled random assignment can balance both known and unknown background characteristics across treatment groups. When random assignment is carried out properly and the rest of the design is sound, differences in outcomes can be attributed more confidently to the assigned treatment.
Random assignment is not the same as random sampling. Random sampling concerns how units are selected from a population and mainly affects how well results can be generalized. Random assignment concerns how selected units are allocated to treatments and is the main design tool for supporting causal inference.
Research Questions and Hypotheses
A useful experimental question identifies a possible cause and a measurable outcome. For example: "Does studying with retrieval practice rather than rereading improve quiz performance after one week?" This question can be turned into a hypothesis because the study method can be assigned and delayed quiz performance can be measured.
A hypothesis is a testable prediction. It should specify the expected relationship between variables without claiming certainty. A statistical hypothesis test may later compare a null model with an alternative, but the scientific hypothesis should be understandable before any calculations are performed.
Avoid vague questions such as "Is music good for learning?" A better question defines the population, treatment, comparison, outcome, and time frame, such as "For students in this class, does 20 minutes of instrumental background music during a vocabulary task change the number of correctly recalled words five minutes later compared with silence?"
Variables, Treatments, and Experimental Units
Independent, Dependent, and Controlled Variables
The explanatory variable or independent variable is the condition that is deliberately changed. Its possible settings are called levels. A particular combination of factor levels is a treatment. The response variable or dependent variable is the outcome you measure.
A controlled variable is a condition intentionally kept the same across groups. In a plant-growth experiment, light intensity might be the factor, plant height after fourteen days the response, and soil volume, pot size, watering schedule, and temperature controlled variables.

Variables need operational definitions. Instead of saying "stress," define how stress will be measured, such as a validated questionnaire score. Instead of "plant growth," specify height change in millimetres, dry biomass in grams, or leaf count. An operational definition makes a concept observable and helps another researcher reproduce the measurement.
Experimental Units and Subjects
The experimental unit is the smallest unit independently assigned to a treatment. If 40 students are individually assigned to two study methods, each student is an experimental unit. If four whole classes are assigned as intact groups, the class rather than the individual student may be the experimental unit for treatment assignment.
This distinction is essential because repeated measurements from the same experimental unit are usually not independent replications. Measuring one plant ten times does not create ten independent plants. Treating non-independent observations as independent is a form of pseudoreplication and can make uncertainty look smaller than it really is.
Core Principles of Strong Experimental Design
Control and Fair Comparison
A control condition provides a meaningful baseline. In a simple laboratory experiment, the control may receive no treatment. In other contexts, the control may receive the standard treatment, a placebo, or a comparison treatment. The best control depends on the scientific question.
A fair comparison means that the groups differ systematically in the factor being tested while other relevant conditions are held constant, randomized, balanced, or otherwise handled by the design.

A placebo is an inactive or comparison treatment designed to resemble the active treatment. Placebos are especially useful when participants' expectations could influence outcomes. A placebo is not appropriate in every field and must be ethically justified.
Randomization
Random assignment uses a chance mechanism to allocate experimental units to treatments. Examples include a random-number generator, shuffled labels, or another transparent chance procedure. Randomization reduces systematic allocation bias and tends to distribute background differences across groups.
Randomization should be planned, not improvised. The allocation rule should be specified before outcomes are known. In many studies, researchers also randomize the order of runs so that gradual changes such as temperature drift, fatigue, or equipment warming do not consistently favour one condition.
Datei:Randomization in clinical trials.webm
Replication
Replication means applying treatments to multiple independent experimental units. Replication allows you to estimate natural variation and determine whether an observed difference is large compared with the variation expected among units.
More replication usually increases precision, but sample size should be planned rather than chosen by a rule such as "more is always better." The needed number of units depends on expected variability, the size of effect worth detecting, the design, the planned analysis, available resources, and ethical constraints.
Replication should not be confused with repeated measurement. Repeated measurements can be useful because they show change over time or reduce measurement noise, but repeated observations on the same unit do not automatically increase the number of independent units.
Blocking
Blocking groups experimental units that are similar with respect to an important source of variation and then randomizes treatments within each block. Suppose you compare two teaching methods and prior achievement strongly predicts the response. You could form blocks of students with similar prior scores and randomly assign methods within each block.
Blocking can improve precision because comparisons are made among more similar units. A matched-pairs design is a special form of blocking for two treatments. In one version, similar units are paired and the two treatments are randomly assigned within each pair. In another version, each unit receives both treatments in a randomized order when carryover effects can be controlled.
Blinding
Blinding keeps participants, assessors, or both from knowing which treatment was assigned when that knowledge could influence behaviour or measurement. Blinding can reduce expectancy effects and observer bias.
In a single-blind study, one relevant group is unaware of treatment identity. In a double-blind study, both participants and the people assessing or administering outcomes may be unaware, depending on the study structure. Not every experiment can be blinded. For example, a student usually knows whether they studied with flashcards or a textbook. When blinding is impossible, objective measurement rules and standardized procedures become even more important.
Confounding, Bias, and Validity
Confounding
A confounding variable changes together with the explanatory variable and provides an alternative explanation for the outcome. Imagine comparing plant growth under two fertilizer types while all plants receiving fertilizer A are placed by a sunny window and all plants receiving fertilizer B are placed farther away. Light and fertilizer are confounded, so the experiment cannot tell which factor caused a difference.
Random assignment is powerful because it prevents the experimenter from deliberately or unconsciously assigning particular kinds of units to one treatment. However, randomization does not guarantee perfectly identical groups in every realized experiment. It provides a defensible chance mechanism and a basis for statistical inference.
Sources of Bias
Selection bias occurs when the units studied differ systematically from the intended population. Measurement bias occurs when the measurement process systematically favours particular results. Observer bias can occur when expectations influence judgments. Attrition bias can arise when dropout differs across treatment groups.
A well-designed experiment anticipates these problems before data collection. Standardized instructions, calibrated instruments, pre-specified inclusion criteria, blinded outcome assessment, and careful handling of missing data can all reduce bias.
Internal and External Validity
Internal validity asks whether the observed treatment difference can reasonably be attributed to the treatment rather than to confounding, bias, or flawed measurement. External validity asks how far the findings can be generalized to other people, settings, materials, or times.
A tightly controlled laboratory experiment may have strong internal validity but limited generalizability. A field experiment may better reflect real conditions but can be harder to control. Good research often uses multiple complementary studies rather than expecting one experiment to answer every question.
Common Experimental Designs
Completely Randomized Design
In a completely randomized design, all eligible experimental units are assigned to treatments using one randomization process. This is often appropriate when units are reasonably similar and there is no major known variable that should be blocked.
A simple example is randomly assigning identical seedlings to three light levels and then measuring biomass after a fixed period. If the environment is homogeneous and there are enough independent seedlings, complete randomization is efficient and easy to analyze.
Randomized Block and Matched-Pairs Designs
In a randomized block design, units are first grouped using a characteristic expected to affect the response, and treatments are randomized within each block. This design deliberately separates variation associated with the blocking factor from the treatment comparison.
A matched-pairs design is useful when two treatments are compared and strong matching is possible. For example, two similar samples from the same material can receive different treatments, or one participant can experience both conditions in random order if practice, fatigue, and carryover are addressed.
Factorial Designs
A factorial design studies two or more factors at the same time. A full factorial design includes every combination of the selected factor levels. For two factors with two levels each, there are four treatment combinations.
Factorial designs can estimate main effects and interactions. A main effect is the average effect of one factor across the levels of another. An interaction occurs when the effect of one factor depends on the level of another factor. Interactions are scientifically important because systems often behave differently when factors act together.

For example, suppose you study the effects of fertilizer level and watering frequency on plant growth. Fertilizer may have little effect under low watering but a large effect under high watering. That pattern is an interaction and would be missed by studying each factor separately in isolated one-factor-at-a-time experiments.
Repeated-Measures and Crossover Designs
In a repeated-measures design, the same experimental unit is measured under multiple conditions or at multiple times. This can control for stable differences among units because each unit acts partly as its own comparison.
The design must account for order, learning, fatigue, and carryover. Randomizing or counterbalancing the treatment order can reduce these problems. A washout period may be needed when one condition can continue to affect later measurements.

Planning an Experiment Step by Step
Step One: Define the Objective
State the scientific objective before choosing a design. Decide whether your purpose is to compare treatments, estimate an effect, screen many factors, investigate interactions, or optimize a process. A clear objective prevents the experiment from becoming a collection of measurements without a coherent question.
Step Two: Specify Units, Factors, Levels, and Responses
Define the population of interest and identify the actual experimental units. List the controllable factors, choose scientifically meaningful levels, and define each response variable with units and measurement procedures.
Also identify possible nuisance variables: factors that are not the main focus but could affect the response. Decide whether to hold them constant, block on them, randomize over them, record them for analysis, or redesign the experiment to reduce their influence.
Step Three: Choose Controls and an Assignment Method
Choose a comparison condition that answers the research question. Then select a randomization method that is practical and auditable. If an important source of variation is known in advance, consider blocking before randomization.
A simple plan should be written before data collection. For example: "Within each prior-achievement block, assign students to study method A or B using a random-number generator, keep study time equal, use the same test, and score responses using a fixed answer key."
Step Four: Plan Replication and Measurement Quality
Decide how many independent units are needed and justify the choice. In advanced work, sample-size planning can use a power analysis based on a meaningful effect size, expected variability, significance level, and desired power.
Check the measurement system before the experiment begins. A highly randomized experiment cannot rescue a response variable that is unreliable, badly calibrated, or unrelated to the intended construct.
Step Five: Pre-Specify the Analysis
The analysis should match the design. A two-group randomized experiment may compare group means, medians, proportions, or another suitable summary. A randomized block design should preserve the block structure. A factorial design should consider main effects and interactions.
Pre-specifying the main outcome and analysis reduces the temptation to search through many alternatives until one looks impressive. Exploratory analyses are valuable, but they should be identified as exploratory rather than presented as if they were the original plan.
Step Six: Run, Document, and Monitor
Follow the plan consistently. Record treatment assignments, measurement times, deviations, missing observations, equipment problems, and any changes made during the experiment. Do not quietly remove inconvenient data points. Any exclusions should follow justified rules and be reported.
A run sheet or laboratory notebook makes the experiment reproducible. Documentation is part of experimental design because another researcher should be able to understand what actually happened rather than only what was intended.
Step Seven: Analyze, Interpret, and Report
Start with graphs and descriptive summaries. Look for unusual values, spread, skewness, missingness, and patterns connected with time or blocks. Then use an inferential method suited to the design.

A p-value, confidence interval, or other inferential summary should not replace scientific reasoning. Report the estimated effect, its uncertainty, and its practical importance. A statistically detectable effect can be too small to matter, while a scientifically important effect may remain uncertain if the experiment is small or noisy.
Statistical Reasoning After the Experiment
Effect Size and Uncertainty
An effect size describes how large a treatment difference is. Examples include a difference in means, a risk difference, a ratio, or a standardized difference. A confidence interval describes a range of values compatible with the data and statistical model under the method used.
Do not reduce a result to "significant" or "not significant." Ask how large the estimated effect is, how precise the estimate is, whether the assumptions are reasonable, and whether the result is meaningful in context.
Random Variation and Reproducibility
Even a perfectly executed randomized experiment will not produce identical group results every time because random assignment and natural variability create chance differences. Statistical inference quantifies this uncertainty under stated assumptions.
Reproducibility also depends on transparent methods. A result is easier to evaluate and repeat when the research question, design, data-processing rules, analysis, and deviations are documented clearly. Repeating a study in new settings can show whether an effect is robust beyond one sample or one laboratory.
Correlation Is Not Automatically Causation
An association between variables does not by itself establish causation. Experiments strengthen causal reasoning because the explanatory condition is assigned rather than merely observed. However, even experiments can fail if the assignment is compromised, if groups receive different additional treatments, if outcomes are biased, or if the analysis ignores the actual experimental unit.
The phrase "correlation does not imply causation" is therefore a warning to examine design, not a claim that causal knowledge is impossible.
Ethics and Responsible Experimentation
Experiments involving people, animals, hazardous materials, or significant environmental impact require appropriate ethical and safety oversight. For classroom work, choose low-risk questions, follow school rules, obtain required consent, protect personal data, and avoid coercion or unnecessary deception.
Participants should understand what they are agreeing to when informed consent is required. Sensitive information should be minimized and handled securely. In medical and psychological research, formal ethics review may be required before recruitment begins.
Responsible research also includes honest reporting. Do not fabricate data, alter observations to fit a hypothesis, conceal important exclusions, or present exploratory findings as pre-planned confirmations. Scientific credibility depends on transparent methods and accurate records.
A Worked Example
Suppose you want to test whether short retrieval-practice sessions improve delayed vocabulary recall compared with rereading. Your experimental units are students who volunteer under the school's approved procedures. The factor is study method with two levels: retrieval practice and rereading. The response is the number of correctly recalled target words on the same delayed test.
A stronger design might block students by a short pretest, then randomly assign the two study methods within each block. All students receive the same word list, total study time, instructions, testing delay, and scoring rules. The person scoring written responses can be given anonymous codes so the scorer does not know which study method each student used.
The main comparison is the difference in delayed recall between the two assigned methods. The design can support a causal conclusion for the participating students if treatment assignment and procedures were followed well. Generalizing to all students would require additional reasoning about how the participants relate to the wider population.
A flow diagram can make recruitment, assignment, follow-up, and analysis transparent:

Design Checklist
Before collecting data, ask yourself whether the research question is specific and testable, the experimental unit is correctly identified, treatments are clearly defined, the response is measured consistently, the control condition is meaningful, randomization is appropriate, replication is truly independent, blocking is useful, blinding is possible, ethical requirements are met, and the analysis matches the assignment structure.
During data collection, check whether the planned conditions are being followed, instruments are functioning, assignment is preserved, missing data are documented, and unplanned changes are recorded.
After data collection, examine the data graphically, estimate treatment effects, quantify uncertainty, check assumptions, distinguish confirmatory from exploratory analysis, discuss limitations, and state conclusions only as strongly as the design allows.
Reliable References and Further Study
For a deeper introduction to statistical design, you can use the NIST/SEMATECH e-Handbook section on Process Improvement and Design of Experiments, which covers objectives, randomization, blocking, factorial designs, assumptions, and analysis.
The Penn State STAT 503 Design of Experiments course notes provide open course material on basic principles, randomized blocks, factorial designs, fractional factorial designs, and advanced topics.
The Khan Academy study design materials offer accessible practice on experiments, observational studies, random assignment, matched pairs, and causal conclusions.
Interactive Tasks
Quiz: Test Your Knowledge
Which feature most directly supports a causal conclusion in a controlled experiment? (Random assignment) (!Random sampling) (!Large population size) (!Colorful graphs)
What is the response variable in an experiment? (The measured outcome) (!The assigned treatment) (!The randomization method) (!The blocking rule)
Why is replication important? (It estimates natural variation) (!It guarantees a large effect) (!It removes all bias) (!It replaces randomization)
What is the main purpose of blocking? (To compare treatments within similar groups) (!To eliminate the response variable) (!To avoid measuring outcomes) (!To replace every control group)
When does confounding occur? (When treatment differences mix with another factor) (!When all units use the same instrument) (!When treatments are assigned by chance) (!When outcomes are measured twice)
What distinguishes random sampling from random assignment? (Sampling concerns selection and assignment concerns treatment allocation) (!Sampling proves causation and assignment proves generalization) (!Sampling is only for laboratories) (!Assignment is only for surveys)
What can a factorial design reveal that separate one-factor studies may miss? (Interactions between factors) (!The identity of every population member) (!A guarantee of normal data) (!Perfect measurement reliability)
What is pseudoreplication? (Treating nonindependent observations as independent replicates) (!Repeating an experiment in a new laboratory) (!Using a placebo in a control group) (!Randomizing the order of treatments)
What is an interaction effect? (The effect of one factor depends on another factor) (!Every factor has exactly the same effect) (!The response is measured without units) (!The experiment has no control condition)
What should a scientific conclusion report in addition to statistical uncertainty? (The size and practical meaning of the effect) (!Only whether a threshold was crossed) (!Only the largest observation) (!Only the number of variables)
Memory Game
| Randomization | Allocation of experimental units to treatments by chance |
| Replication | Use of multiple independent experimental units |
| Blocking | Grouping similar units before treatment assignment |
| Placebo | Comparison condition designed to resemble an active treatment |
| Interaction | Situation in which one factor changes the effect of another |
| Confounding | Mixing of a treatment effect with another explanatory influence |
Drag and Drop
| Match the correct terms. | Topic |
|---|---|
| Experimental unit | Smallest unit independently assigned to a treatment |
| Response variable | Outcome measured after treatment |
| Control condition | Baseline used for comparison |
| Randomized block design | Treatment assignment by chance within similar groups |
| Factorial design | Experiment that studies combinations of two or more factors |
...
Crossword Puzzle
| Randomization | What process assigns experimental units to treatments by chance? |
| Replication | What principle uses multiple independent experimental units? |
| Blocking | What design strategy groups similar units before random assignment? |
| Placebo | What inactive comparison can resemble an active treatment? |
| Response | What type of variable records the measured outcome? |
| Confounding | What problem mixes a treatment effect with another influence? |
LearningApps
Cloze Text
Open-Ended Tasks
Easy
- Variable Map: Choose a safe classroom experiment and create a labeled diagram showing its factor, levels, response variable, experimental units, and at least three controlled variables.
- Fair Comparison Audit: Find a simple experiment described in a textbook, news article, or classroom handout and explain whether its control condition creates a fair comparison.
- Random Assignment Simulation: Use shuffled cards or a random-number generator to assign twenty fictional subjects to two treatments, document the procedure, and explain why the method reduces allocation bias.
- Measurement Protocol: Write a one-page protocol for measuring a response such as paper-airplane flight distance, cooling time, or seedling height so that different researchers would measure it consistently.
Standard
- Paper Airplane Experiment: Design and carry out a safe experiment comparing two paper-airplane designs while controlling launch position, measuring flight distance repeatedly on independent planes, and graphing the results.
- Blocking Investigation: Design a study in which an important background variable is used to create blocks, then explain how treatment comparisons within blocks could be more precise than complete randomization.
- Factorial Mini Project: Plan a two-factor experiment with two levels per factor, create the four treatment combinations, predict possible main effects and an interaction, and present the design in a table.
- Research Interview: Interview a teacher, laboratory technician, engineer, or researcher about how they control variability in real experiments, then summarize at least three design practices and connect them to course concepts.
Advanced
- Power and Sample Size Exploration: Use a statistical simulator or approved software to explore how sample size, effect size, and variability influence power, then explain why design choices should be made before seeing the final outcomes.
- Replication Study Proposal: Choose a safe published classroom-scale experiment and write a proposal to replicate it, including the hypothesis, assignment method, measurement protocol, planned exclusions, and analysis.
- Interaction Video Explanation: Produce a short educational video or animation explaining a factorial interaction with your own example, graph, and explanation of why separate one-factor experiments could miss the pattern.
- Experimental Design Critique: Locate a published experiment in a scientific or reputable educational source, reconstruct its design, identify strengths and limitations, and judge which causal and generalization claims are justified.
Learning Assessment
- Causal Reasoning Assessment: Given a study description, decide whether a causal conclusion is justified and defend your answer by referring to assignment, controls, confounding, and the actual experimental unit.
- Design Repair Assessment: Analyze a flawed experiment with confounded treatment groups and redesign it so that randomization, replication, and measurement are improved.
- Blocking Decision Assessment: Compare a completely randomized design with a randomized block design for the same research question and justify which design is more appropriate.
- Factorial Interpretation Assessment: Interpret a two-factor results table or interaction graph and explain the main effects, possible interaction, and a scientifically meaningful follow-up question.
- Inference and Scope Assessment: Explain separately what random assignment and random sampling contribute to the scope of a study's conclusions.
- Ethics and Transparency Assessment: Evaluate an experimental plan for consent, risk, privacy, pre-specified outcomes, data exclusions, and honest reporting, then propose concrete improvements.
Evidence of Learning
Evidence of learning includes accurate use of terms such as experimental unit, factor, treatment, response, control, randomization, replication, blocking, blinding, confounding, and interaction. You should be able to identify these elements in unfamiliar examples rather than only repeat definitions.
Strong evidence also includes a complete experimental plan with a focused question, operational definitions, a defensible assignment procedure, independent replication, appropriate controls, attention to nuisance variables, ethical safeguards, and an analysis that matches the design.
Your products may include a design diagram, laboratory protocol, data table, graph, short research report, critique of a published study, interview summary, simulation, poster, or explanatory video. High-quality work makes assumptions and limitations visible instead of hiding them.
Transfer is demonstrated when you can apply the same principles across subjects. For example, randomization and blocking can be useful in biology, psychology, agriculture, manufacturing, and educational research even though the measurements and treatments differ.
OERs on the Topic
Linked Learning Areas
aiMOOC Projects
MOOCwiki · Deutsch
Nach dem Lernen ist vor dem Lernen
Entdecke direkt den nächsten Lernkurs. Weitere Inhalte erscheinen, wenn Du weiter nach unten scrollst.
Zur MOOCwiki-HauptseiteMediathek
Mediathek
Mediathek wird aus dem Wiki geladen ...
Keine passenden Inhalte gefunden. Bitte ändere Suche oder Filter.
NEWSLernweltNOAH fragen