Zum Inhalt springen

English:Assessment and Evaluation

Aus MOOCsWiki Staging
Die Druckversion wird nicht mehr unterstützt und kann Darstellungsfehler aufweisen. Bitte aktualisiere deine Browser-Lesezeichen und verwende stattdessen die Standard-Druckfunktion des Browsers.
aiMOOC-Siegel

Assessment and Evaluation



Introduction

Assessment and Evaluation are central to learning, teaching, curriculum design, quality assurance, and educational decision-making in higher education. They influence what students study, how they study, what teachers notice, which achievements are recognized, and what institutions decide to improve.

In this aiMOOC, assessment means the systematic collection and interpretation of evidence about learning, performance, or progress. Evaluation means making a reasoned judgment about quality, value, effectiveness, or achievement by comparing evidence with explicit purposes or criteria. In practice, universities and academic disciplines sometimes use these terms differently or even interchangeably. You should therefore check the definitions and regulations used in your own institution.

Assessment is broader than testing. Evidence can come from an examination, a laboratory performance, a design, a portfolio, a case analysis, an oral defense, a simulation, peer review, fieldwork, or a professional product. Evaluation is also broader than assigning a grade: you can evaluate an assignment design, a course, a teaching innovation, a curriculum, or a whole academic program.

The following video introduces questions about what assessment measures and why authentic evidence matters in higher education.


Learning Goals

By the end of this aiMOOC, you should be able to explain major purposes and forms of assessment, distinguish assessment from evaluation in context, design assessment evidence that aligns with learning outcomes, apply validity and reliability principles, use rubrics and feedback productively, interpret quantitative and qualitative evidence, evaluate a course or learning intervention, recognize fairness and accessibility issues, and propose improvements based on evidence.


Core Concepts


Assessment, Measurement, Grading, and Evaluation

These four ideas are related but should not be treated as synonyms.

Assessment asks: What evidence do we have about learning or performance? It includes gathering evidence, interpreting it, and using it for decisions.

Measurement assigns numbers according to defined rules. A score of 18 out of 25, a response time, or a rating on a scale is a measurement. Numbers can be useful, but they do not interpret themselves.

Grading converts judgments about achievement into a reporting category such as a mark, percentage, pass/fail decision, or letter grade. A grade is a compact summary and may hide important information about strengths, weaknesses, uncertainty, and improvement.

Evaluation asks: How good, effective, valuable, suitable, or successful is something for a stated purpose? Evaluation requires criteria, evidence, interpretation, and judgment. For example, a department might evaluate whether a redesigned course improved student learning, accessibility, and progression.

A useful academic habit is to ask three questions whenever you see a score: What construct was measured? What evidence produced the score? What decision will be made from it?


Purposes of Assessment

Assessment can serve several purposes, and one task can sometimes serve more than one.

Diagnostic assessment identifies prior knowledge, misconceptions, readiness, or learning needs before or early in instruction.

Formative assessment produces evidence that can still influence learning or teaching. Its defining feature is not that it is ungraded, but that the evidence is used to decide what to do next.

Summative assessment supports a judgment at the end of a learning period, such as certification, progression, or a final grade.

Assessment for learning emphasizes using evidence and feedback to improve learning. Assessment of learning emphasizes judging achievement. Assessment as learning emphasizes the learner's active role in monitoring and regulating learning through self-assessment, reflection, and metacognition.

Ipsative assessment compares current performance with the same learner's earlier performance. Criterion-referenced assessment compares performance with stated criteria or standards. Norm-referenced assessment compares performance with the performance of a reference group. These approaches answer different questions and should not be confused.


Learning Outcomes and Cognitive Demand

A learning outcome should state what you are expected to know, understand, produce, perform, analyze, or create. Assessment should provide evidence at the same level of cognitive demand.

Bloom's revised taxonomy is one useful framework for thinking about increasing cognitive complexity. It is not a complete theory of learning and should not be used mechanically, but it can help you notice a mismatch between an outcome and an assessment task.

For example, if the outcome says that you should evaluate competing policy options, a recall-only multiple-choice test cannot provide complete evidence of that outcome. You need an opportunity to compare evidence, apply criteria, justify a judgment, and respond to counterarguments.


Constructive Alignment

Constructive alignment links intended learning outcomes, learning activities, and assessment. The central design question is whether students practice the kind of performance that the assessment later asks them to demonstrate.

A constructively aligned course begins with clear outcomes, selects assessment evidence that can demonstrate those outcomes, and then provides learning activities that prepare students for the required performance. Alignment does not mean making assessment easy. It means making expectations coherent.

Consider this example. If an engineering outcome requires you to diagnose faults in a complex system, then lectures about fault diagnosis are not enough. You need guided practice with cases or simulations, and the assessment should require an actual diagnosis with reasoning. If a literature outcome requires interpretation, then reproducing definitions alone is insufficient evidence.

Constructive alignment also helps you detect under-assessment and over-assessment. Under-assessment occurs when important outcomes are barely tested. Over-assessment occurs when a course demands many tasks without a clear connection to distinct learning outcomes.


Quality of Assessment Evidence


Validity

Validity concerns the strength and appropriateness of the interpretation and use of assessment evidence. A useful practical question is: Does this assessment give us defensible evidence for the conclusion we want to make?

If you want to assess clinical communication, a written recall test alone has weak alignment with the performance claim. If you want to assess statistical reasoning, a task dominated by complex language may unintentionally measure reading proficiency as well as statistics.

Validity is not a permanent property of a test. The same assessment may support one interpretation but not another. You should therefore consider the intended construct, task design, scoring method, consequences, and alternative explanations for performance.


Reliability and Consistency

Reliability concerns consistency. If performance is judged repeatedly under comparable conditions, how stable and reproducible is the result?

Sources of inconsistency include ambiguous questions, poorly specified criteria, inconsistent administration, rater differences, random guessing, fatigue, and too few observations. Reliability can be strengthened through clearer tasks, sufficient sampling of the domain, explicit scoring criteria, marker training, moderation, and multiple sources of evidence.

Reliability is necessary for many high-stakes uses, but consistency alone does not guarantee validity. A perfectly consistent assessment can still measure the wrong thing.


Fairness, Accessibility, and Transparency

A fair assessment gives students a meaningful opportunity to demonstrate the intended learning without irrelevant barriers. Fairness does not always mean identical conditions. Reasonable accommodations, accessible formats, flexible modes of demonstration, and clear instructions can reduce construct-irrelevant barriers while preserving academic standards.

Transparency requires students to understand the task, criteria, weighting, permitted resources, academic integrity expectations, and consequences. Exemplars and annotated samples can make quality visible, but they should be used to illuminate standards rather than encourage imitation.

Bias can enter through topic selection, language, examples, scoring expectations, technology access, time limits, cultural assumptions, and rater judgments. Good assessment design therefore includes an explicit review for potential disadvantage and unintended measurement.


Authenticity and Transfer

Authentic assessment asks you to apply knowledge and skills in a meaningful context that resembles, models, or connects with practices beyond a conventional test setting. Examples include a policy brief, clinical simulation, engineering design, legal argument, data dashboard, lesson plan, public communication product, or research proposal.

Authenticity is not automatically superior to every traditional form. The best method depends on the learning outcome. A short quiz may be efficient for checking foundational knowledge, while a complex performance task may be needed to assess integration, judgment, or professional practice.


Assessment Methods


Selected-Response and Short-Answer Tasks

Multiple-choice, true/false, matching, and short-answer tasks can efficiently sample a broad content domain. High-quality items can test application and reasoning, not only recall. Weak items often contain clues, unnecessary difficulty, implausible distractors, or language unrelated to the intended construct.

For multiple-choice questions, the stem should present a clear problem, the correct answer should be defensible, and distractors should represent plausible misconceptions rather than tricks. Item review should consider both content quality and statistical evidence after use.


Essays, Reports, and Extended Responses

Extended writing can assess argumentation, synthesis, evidence use, disciplinary communication, and judgment. The task prompt should make the purpose, audience, evidence expectations, scope, and criteria clear.

Because scoring complex work involves judgment, rubrics, marker calibration, exemplars, and moderation can improve transparency and consistency. However, overly mechanical rubrics can fragment a complex performance into disconnected parts. Criteria should reflect what genuinely matters in the discipline.


Performance, Projects, Portfolios, and Oral Assessment

Performance assessment can show what you can actually do. Projects can integrate knowledge over time. Portfolios can demonstrate development, selection, reflection, and a body of work. Oral assessment can test explanation, reasoning, spontaneous response, and authorship.

These methods can generate rich evidence, but they require careful design. You need clear criteria, manageable workload, appropriate support, secure procedures, and attention to accessibility. Group work also requires decisions about how individual and collective contributions will be recognized.


Rubrics and Criteria

A rubric describes important criteria and levels of performance. An analytic rubric separates dimensions such as evidence, reasoning, organization, and communication. A holistic rubric gives an overall judgment based on an integrated description.

A useful rubric does more than make marking faster. It communicates what quality looks like, supports self-assessment, guides feedback, and can improve consistency across markers. Criteria should be observable, relevant to the learning outcomes, and understandable to students.

Before using a rubric, test it on sample work. Ask whether two informed readers interpret the descriptors similarly, whether the levels distinguish meaningful differences in quality, and whether the rubric rewards the intended learning rather than surface compliance.


Feedback, Feedforward, and Learning

Feedback becomes educationally useful when it helps you understand the current quality of your work, the desired quality, and a feasible next step. Feedforward emphasizes using feedback to improve future work.

A general feedback loop illustrates the idea that information from an outcome can become an input for the next action.

Effective feedback is usually specific, selective, timely enough to be used, connected to criteria, and oriented toward action. More comments are not necessarily better. A learner needs enough information to make a decision and another opportunity to apply that information.

Feedback is also a learner capability. Feedback literacy includes understanding standards, making judgments about quality, managing emotional responses, seeking information, and taking action. Peer review and self-assessment can strengthen these capabilities when students are trained to use criteria and justify judgments.

The following higher-education resource explores assessment for learning and the relationship among teaching, grading, and feedback.


Measurement and Data Interpretation


Scores and Scales

Scores are representations of performance, not the performance itself. Before interpreting any number, identify the scale, its possible range, the scoring rule, and the intended comparison.

A percentage can mean very different things depending on the assessment. Sixty percent on an easy recall quiz and sixty percent on a difficult professional simulation are not directly comparable. A cut score should therefore be justified in relation to standards, not assumed to have universal meaning.

The image above visualizes learning evaluation through repeated measurements before and after instruction and self-study. Such designs can be useful, but a difference between two scores should not automatically be attributed to one cause. Practice effects, task differences, motivation, timing, and measurement error can also influence change.


Descriptive Statistics

For numerical assessment data, the mean summarizes the arithmetic average, the median identifies the middle observation, and the mode identifies the most frequent value. Measures of spread such as the range, interquartile range, and standard deviation show how dispersed results are.

A single average can hide important patterns. Always inspect the distribution, missing data, outliers, subgroup patterns, and sample size before making conclusions. In small classes, a few unusual values can strongly influence the mean.


Rating Scales and Surveys

Course evaluations and questionnaires often use Likert-type items in which respondents indicate degrees of agreement or another ordered response.

A single Likert item is ordered categorical data. A multi-item scale may be summarized in more complex ways when its construction and assumptions support that use. Regardless of analysis, a numerical rating should be interpreted together with the wording of the item, response rate, sampling process, and qualitative evidence.

Student evaluation of teaching can provide useful information about student experiences, but it is not a complete measure of teaching quality. Strong program evaluation combines multiple sources such as student work, peer observation, curriculum evidence, learning analytics, interviews, focus groups, and reflective teaching documentation where appropriate.


Item Analysis

After a test, item analysis can help identify questions that may need review. Common indicators include the proportion of students answering an item correctly and the relationship between item performance and overall test performance.

Statistical flags do not prove that an item is good or bad. A difficult item may be entirely appropriate, and an unusual discrimination pattern may reflect ambiguous wording, a miskeyed answer, content that was not taught, or a genuinely complex concept. Expert review should accompany statistics.


Educational Evaluation


From Evidence to Judgment

Evaluation extends beyond collecting data. It connects evidence to a purpose, criteria, and a decision. A strong evaluation makes the reasoning visible: what was being evaluated, for whom, according to which criteria, with what evidence, and with what limitations.

For a university course, possible evaluation questions include: Did students achieve the intended outcomes? Which students benefited and which faced barriers? Did the assessment design support the intended learning? Was workload proportionate to credit? How did students use feedback? What should be retained, changed, or investigated further?


An Evaluation Cycle

  1. Clarify the purpose: Define what is being evaluated and why the evaluation is needed.
  2. Identify stakeholders: Determine who is affected by the program and who will use the findings.
  3. Form evaluation questions: Write answerable questions about quality, outcomes, implementation, equity, or value.
  4. Specify criteria and indicators: Decide what evidence would count as success, concern, or improvement.
  5. Collect evidence: Use suitable quantitative and qualitative methods while protecting privacy and consent.
  6. Analyze and triangulate: Look for patterns, contradictions, plausible explanations, and limitations across sources.
  7. Make a judgment: Compare the evidence with criteria and state the reasoning and uncertainty.
  8. Act and follow up: Decide what will change, who is responsible, and how the effect of the change will be reviewed.

Evaluation should be iterative. Findings influence redesign, and redesign creates new questions for later evaluation.


Quantitative, Qualitative, and Mixed Evidence

Quantitative evidence is useful for questions about frequency, distribution, change, relationships, and patterns across larger groups. Qualitative evidence is useful for questions about experience, reasoning, meaning, process, and context.

A mixed-methods evaluation combines both when the evaluation question requires breadth and depth. For example, course completion rates can reveal where a problem occurs, while interviews may help explain how students experienced the barriers behind that pattern.

Triangulation does not mean that all sources must agree. Disagreement can be informative. If survey ratings are positive but assignment performance declines, the contradiction becomes a question to investigate rather than a reason to discard one source automatically.


Ethical and Responsible Practice

Assessment and evaluation can affect progression, confidence, professional opportunities, workload, and institutional decisions. Ethical practice therefore requires proportionality, confidentiality, data minimization, secure handling of records, transparent decision rules, and procedures for review or appeal.

High-stakes decisions should not depend on weak evidence. Where possible, use multiple observations or methods, especially when performance is complex. Document uncertainty and avoid claiming more than the evidence supports.

When evaluating teaching, programs, or people, distinguish developmental uses from accountability uses. Participants should understand how data will be used. Anonymous and confidential data are not the same, and institutional policies may impose specific legal and ethical requirements.


Assessment in the Age of Generative AI

Generative AI changes what some assignments can demonstrate. If an assignment can be completed by a tool with little student reasoning, the task may no longer provide the evidence that the instructor intended.

A responsible response is not simply to make every assessment invigilated. Instead, assessment can be redesigned to gather richer evidence of process and judgment. Examples include staged drafts, annotated source choices, oral follow-up questions, local or novel data, reflective decision logs, authentic professional tasks, and supervised performances where appropriate.

Rules for AI use should be explicit. Students need to know which tools are permitted, what uses are allowed, what must be disclosed or cited, which data may not be uploaded, and which parts of the work must be demonstrably their own. Requirements should align with institutional policy and disciplinary norms.

AI-generated output should itself be evaluated critically. You should verify claims, check sources, inspect reasoning, identify missing perspectives, and remain accountable for submitted work. Automated AI-detection scores should not be treated as self-sufficient proof of misconduct because detection systems can produce uncertain or erroneous results.


Designing a Strong Assessment Plan

A coherent assessment plan balances learning, evidence quality, student workload, staff workload, feedback opportunities, academic integrity, accessibility, and decision consequences.

A practical design sequence is to start with the learning outcome, identify the performance that would demonstrate it, choose an assessment method, define criteria, plan learning activities and practice, decide where feedback will occur, review accessibility and integrity risks, pilot or moderate the task, and then use evidence from implementation to improve the next version.

The key question is not Which assessment method is best? but Which combination of evidence is most defensible for this learning outcome, this context, and this decision?


Research and Teaching Resources

  1. Constructive Alignment at the University of Tasmania: A higher-education guide to aligning outcomes, assessment, and learning activities.
  2. Creating and Using Rubrics at the University of Minnesota Duluth: Guidance on criteria and performance-level descriptions.
  3. Authentic Assessments at the University of Illinois Chicago: Guidance on real-world application, validity, and design.
  4. Constructive Alignment at Queen Mary University of London: An explanation of curriculum alignment and assessment for learning.
  5. Formative Assessment resources from Edutopia: Examples of using evidence during learning to adjust instruction and support progress.


Interactive Tasks


Quiz: Test Your Knowledge

Which statement best describes formative assessment? (It produces evidence that can still be used to improve learning or teaching) (!It is any task that does not receive a grade) (!It is always completed before teaching begins) (!It is used only to calculate a final course grade)




What is the central question of assessment validity? (Whether the evidence supports the intended interpretation and use) (!Whether every student receives the same score) (!Whether the assessment contains many questions) (!Whether the final grade is reported as a percentage)




What does reliability primarily concern? (Consistency of assessment results or judgments) (!Popularity of an assessment method) (!Difficulty of the course content) (!Length of the feedback comments)




Which example is criterion-referenced assessment? (Comparing a performance with explicit competency standards) (!Ranking students from highest to lowest) (!Comparing one cohort with a national sample) (!Awarding grades according to a fixed curve)




What is the main idea of constructive alignment? (Align learning outcomes activities and assessment evidence) (!Use the same assessment method in every course) (!Assess only knowledge that can be scored automatically) (!Replace all examinations with group projects)




What is a major purpose of a rubric? (To communicate criteria and levels of performance) (!To eliminate all professional judgment from marking) (!To guarantee that every marker gives identical scores) (!To convert every complex task into a multiple choice test)




Which practice best supports useful feedback? (Give actionable information that students can apply to later work) (!Provide as many comments as possible after the course ends) (!Focus only on the numerical grade) (!Avoid showing students the assessment criteria)




What does authentic assessment usually emphasize? (Application of knowledge and skills in meaningful contexts) (!Memorization without a context) (!Ranking students against one another) (!Using only standardized test questions)




Why should student survey ratings be combined with other evidence when evaluating teaching? (Because one data source cannot capture every dimension of teaching quality) (!Because survey responses can never provide useful information) (!Because qualitative evidence is always more accurate than numbers) (!Because course evaluation should ignore student experience)




What is triangulation in evaluation? (Using multiple sources or methods to examine an evaluation question) (!Converting every result into a single percentage) (!Removing data that disagree with the expected result) (!Using three graders for every assessment task)





Memory Game

Validity Strength of the interpretation and use supported by assessment evidence
Reliability Consistency of scores or judgments under comparable conditions
Rubric Criteria with descriptions of different levels of performance
Formative Evidence used while there is still an opportunity to improve learning
Summative Evidence used to make an end-point judgment about achievement
Authenticity Connection between an assessment task and meaningful disciplinary or real-world practice





Drag and Drop

Match the correct terms. Topic
Evidence used to guide next learning steps Formative assessment
Judgment made at the end of a learning period Summative assessment
Comparison with explicit performance standards Criterion referencing
Consistency across comparable judgments Reliability
Fit between evidence and intended interpretation Validity




...


Crossword Puzzle

Validity What concept asks whether assessment evidence supports the intended interpretation?
Reliability What concept describes consistency across comparable measurements or judgments?
Rubric What scoring tool describes criteria and levels of performance?
Feedback What information can help a learner improve a current or future performance?
Authenticity What quality connects an assessment with meaningful disciplinary or real-world practice?
Criterion What word completes the phrase referenced assessment when performance is compared with a stated standard?





LearningApps


Cloze Text

Complete the text.
Assessment gathers and interprets

about learning or performance. A

use of assessment allows evidence to influence what happens next. A

use supports an end-point judgment about achievement. Assessment

concerns whether evidence supports the intended interpretation and decision. Assessment

concerns consistency under comparable conditions. A

communicates criteria and performance levels for complex work. Constructive

connects learning outcomes, learning activities, and assessment. Evaluation should combine evidence with explicit

before making a judgment.




Open-Ended Tasks


Easy

  1. Assessment Glossary: Create a one-page glossary that distinguishes assessment, measurement, grading, and evaluation, then add one university example for each term.
  2. Feedback Audit: Select feedback from a past learning task, anonymize it, and identify which comments are actionable, which are unclear, and what next step each useful comment suggests.
  3. Outcome Match: Choose three learning outcomes from a university module and identify an assessment method that could provide suitable evidence for each outcome.
  4. Rubric Critique: Find a rubric used in your discipline and write a short critique of its clarity, relevance, and usefulness for self-assessment.


Standard

  1. Assessment Redesign: Redesign a conventional recall-heavy assessment so that it also captures application or judgment, and explain how your redesign changes the evidence collected.
  2. Mini Evaluation Interview: Interview a student or instructor about one assessment experience, summarize the main themes, and distinguish reported experience from your own interpretation.
  3. Survey Design: Create a short course-evaluation questionnaire with clearly worded rating items and two open questions, then explain what each item can and cannot tell you.
  4. Authentic Assessment Prototype: Produce a prototype professional or disciplinary task such as a policy brief, case response, design, lesson, simulation, or public communication product with assessment criteria.


Advanced

  1. Validity Argument: Build a written validity argument for one high-stakes university assessment by stating the intended claim, evidence, alternative explanations, risks, and safeguards.
  2. Mixed Methods Evaluation: Design a small mixed-methods evaluation of a course innovation using at least one quantitative and one qualitative source, and explain how you would integrate the findings.
  3. Marker Calibration Study: Create two sample responses and a rubric, ask several peers to score them independently, compare the judgments, and propose changes that could improve consistency.
  4. Assessment Portfolio Video: Produce a short academic video presenting an assessment system for a university module, including alignment, feedback points, accessibility, academic integrity, and a plan for evaluating the design.



Learning Assessment

  1. Alignment Case Analysis: Given a set of learning outcomes, teaching activities, and assessments from a hypothetical course, identify mismatches and justify a redesigned assessment plan.
  2. Validity and Reliability Diagnosis: Analyze a scenario in which students receive inconsistent marks on a task that does not fully match the learning outcome, then separate the validity problem from the reliability problem and recommend remedies.
  3. Evidence Interpretation: Examine a small dataset of grades, survey ratings, and student comments, identify patterns and limitations, and produce a cautious evidence-based conclusion.
  4. Rubric Construction: Design an analytic rubric for a complex disciplinary performance, justify each criterion, and explain how you would test the rubric before high-stakes use.
  5. Program Evaluation Proposal: Write an evaluation plan for a new university teaching intervention with stakeholders, evaluation questions, criteria, data sources, ethics, analysis, and follow-up.
  6. AI Responsive Assessment: Redesign an assignment for a context in which generative AI is available, specify permitted use, identify evidence of student learning, and explain how the design protects validity and accessibility.




Evidence of Learning

Important evidence of learning includes knowledge, skills, products, and transfer achievements.

  1. Knowledge: You can accurately explain assessment purposes, criterion and norm referencing, validity, reliability, authenticity, constructive alignment, rubrics, feedback, and evaluation.
  2. Analytical skills: You can identify mismatches between outcomes and assessment evidence, distinguish measurement problems from interpretation problems, and recognize limitations in data.
  3. Design skills: You can create assessment tasks, criteria, rubrics, feedback processes, and evaluation plans that are coherent, transparent, and suitable for university learning.
  4. Evidence products: You can produce a validity argument, an assessment blueprint, a rubric, an evaluation proposal, a data interpretation, or an authentic assessment prototype.
  5. Ethical judgment: You can identify fairness, accessibility, privacy, bias, academic integrity, and high-stakes decision risks and propose proportionate safeguards.
  6. Transfer achievement: You can apply assessment and evaluation principles to a new discipline, workplace, research, or professional learning context and justify your decisions with evidence.




OERs on the Topic

The English Wikipedia article on educational assessment provides an open starting point for further study.



Linked Learning Areas


aiMOOC Projects