Showing posts with label classroom assessment. Show all posts
Showing posts with label classroom assessment. Show all posts

Wednesday, February 8, 2023

AI generated unit plans and assessment

 Charles Ungerleider, Professor Emeritus, The University of British Columbia

[permission to reproduce granted if authorship is acknowledged]  

Last week I reported the results of using AI as a foundation for planning lessons. In general, I thought the bot I used produced a workable framework and, if thoroughly reviewed, some useful content and examples for a first draft plan. That experience prompted me to want to explore the bot’s contribution to unit planning and classroom assessments of various types.  

I asked the bot to prepare an outline for a teaching unit about the French Revolution. Here’s the output:  

I. Introduction

A. Definition and overview of the French Revolution

B. Historical context of the Revolution

C. Significance of the Revolution for France and the world

 II. Causes of the Revolution

A. Economic crisis and poverty

B. Political corruption and abuse of power

C. Enlightenment ideas and the influence of philosophers

D. International events and wars

 III. Major events of the Revolution

A. The fall of the Bastille

B. The Reign of Terror

C. The rise of Napoleon Bonaparte

D. The abolition of the monarchy and the establishment of the First Republic

 IV. Impact of the Revolution

A. Political and social changes in France

B. The spread of revolutionary ideals to other countries

C. The influence on later political and social movements

D. Legacy of the Revolution in France and the world

 V. Conclusion

A. Summary of key points

B. Reflection on the significance of the French Revolution

C. Final thoughts and open discussion.

 VI. Assessment

A. Written tests or essays

B. Group projects or presentations

C. Class discussions or debates

D. Quizzes or exams.

Like the lesson plans, the unit outline is a useful starting point upon which a teacher might build a detailed unit plan. In developing my plan, I would want to place greater emphasis on the Declaration of the Rights of Man and of the Citizen and the contradiction between advocacy of those rights and slavery and religious intolerance. Nonetheless, the outline is a starting point. It, as is the case with everything produced by the bot or even one’s colleagues, needs to be carefully reviewed.  

My next experiment was to ask the bot to produce a ten-item multiple choice test on the French Revolution. I stipulated that each successive item should be more difficult than the preceding item. I asked the bot to mark the correct answer with an asterisk. Here’s the output:  

What was the main cause of the French Revolution?

a. The Enlightenment

b. The American Revolution

c. The Industrial Revolution

d. Economic crisis and poverty *


Who was the first ruler of France during the Revolution?

a. King Louis XVI *

b. Napoleon Bonaparte

c. Maximilien Robespierre

d. Charles De Gaulle


What was the fall of the Bastille symbolize for the French people?

a. The end of the Revolution

b. The end of the monarchy

c. The start of the Reign of Terror

d. The start of their freedom *

 

What was the Reign of Terror?

a. A period of intense violence during the Revolution *

b. A period of peace and prosperity

c. The name of Napoleon's army

d. A period of intellectual enlightenment

 

Who was Maximilien Robespierre?

a. A philosopher of the Enlightenment

b. A military leader during the Revolution

c. The leader of the Reign of Terror *

d. A supporter of the monarchy

 

What did the Declaration of the Rights of Man and Citizen do?

a. Established the First Republic

b. Granted citizenship to all men *

c. Abolished the monarchy

d. Established a dictatorship

 

Who was Napoleon Bonaparte?

a. A philosopher of the Enlightenment

b. A military leader during the Revolution

c. The leader of the Reign of Terror

d. The first ruler of the First Republic *

 

What was the impact of the French Revolution on other countries?

a. It led to the spread of revolutionary ideals *

b. It had no impact on other countries

c. It led to the decline of revolutionary ideals

d. It led to increased political corruption

 

What was the long-term legacy of the French Revolution?

a. The decline of revolutionary ideals

b. The end of the monarchy

c. The establishment of the First Republic *

d. The rise of dictatorship

 

What was the significance of the French Revolution?

a. It changed the course of European history *

b. It had no significant impact

c. It led to the decline of revolutionary ideals

d. It led to increased political corruption.

Over my career I’ve read many multiple-choice items. I’d put the items produced by the bot somewhere near the 80th percentile among all the items I’ve read. Without a prompt, the bot avoided items that asked “NOT-type” items, ones in which the respondent is to pick out the one item that is NOT an instance of the phenomenon addressed by the question. The bot also avoided tipping off the answer by keeping all plausible answers the same length.  

I am not happy with all items. The item about Louis XVI is confusing at best (Louis XVI was the last king before the revolution). I wouldn’t use it as written by the bot. The item that asks what was the fall of the Bastille symbolize for the French people?  is worded awkwardly. It should read what did the fall of the Bastille symbolize for the French people?  I wouldn’t use the one asking, “who was Maximilien Robespierre?” because it places greater emphasis on his part in the reign of terror than it does on his advocacy of universal male suffrage. Nor is it evident to me that the questions become progressively difficult. Notwithstanding those shortcomings,, the output is a useful, time-saving starting point for the construction of a quiz.  

To my previous caution about scrutinizing the bot’s output for accuracy I would add nuance is not a feature of the bot’s output. Yet, when judged against the output produced by novice item-writers – and I would put most teachers in that category – the bot scored about 80% in my mental calculus.  

In a future blog, I will report on the bot’s production of marking rubrics, group project ideas, and discussion questions. In the meantime, you might want to evaluate the bot’s production of multiple-choice items on other topics and in other fields.

Wednesday, March 10, 2021

All teachers teach to the test (or at least they should)

 Charles Ungerleider, Professor Emeritus, The University of British Columbia

[permission to reproduce granted if authorship is acknowledged]

Opponents of large-scale assessments say that they prompt teachers to teach to the test. As I have said in an earlier blog, if the assessments measure performances on things that are important to be able to do or know, teaching to the test is not a bad thing.

But all teachers “teach to the test.” What I mean by that is all instruction begins with establishing a clear objective or destination. Although they may not (should not) wait for external assessments to determine whether the students have achieved the objective or reached the destination, teachers use classroom assessments that they have devised to monitor student progress along the way and, often, at the end of a unit.

Sometimes called backward design or backward planning, the process starts with the end point of the unit or lesson and works backward toward the beginning. Once the goals (and big ideas) are established, teachers develop the instructional sequence that they think will help the students reach the goals, including the ‘way points’ or indicators of progress that teachers assess along the path.

The assessments may entail simple observation on the part of the teacher, the completion of a task or problem the teacher has set, a quiz, a demonstration, a dramatization, a graphic representation, essays, presentations, project work, portfolios, etc. All teacher assessments are high stakes in the sense that cumulatively they will figure in a teacher’s overall appraisal of a student’s performance.

Teacher-made assessments are very “costly.” They are more costly than large scale assessments when you add up the amount of time teachers spend creating and marking the various assessments that they use.

With both teacher assessments and large-scale assessments, validity is a big deal. Are the judgments or decisions made based on the assessment(s) justified by the data generated by them? For instance, I do not eat at [restaurant name] because the first time I ate there I had food poisoning. The tenuous connection between my decision and the data is insufficient to justify my conclusion. Can a recommendation for a gifted program be made because the student produced an imaginative science-fair project? Does a two-minute audition justify a director’s decision to offer or deny a leading role in a school play?

Reliability is another important consideration in assessment. Does the assessment accurately measure whatever it is trying to measure? As I have mentioned in an earlier blog, when I was a child our furnace was fueled with oil from a large tank in our basement. Each fall my father would climb a step ladder and put a stick into the tank to determine how much oil we needed. This was important because the truck that delivered the oil served many customers. The driver, to ensure that there was sufficient oil in the truck, would ask in advance, “how much do you need?” My father would tell him. But my father’s measurement of the oil in the tank was very inconsistent. Sometimes he would insert the stick in the tank on an angle. In the low light in the basement, he would sometimes misread the level on the stick. His appraisal of the volume of oil we needed was often mistaken because of the measurement errors he made. The driver would be annoyed when my father’s assessment was too low, and the driver would have to return to finish filling the tank.

Individual teacher assessments are often unreliable. Fatigue, the pressure of time, ambiguous instructions, and many other factors detract from the reliability of teacher assessments. It is fortunate, however, that teachers make (or should make) many individual assessments before arriving at a judgement about performance (assigning a grade for the year, for example). Although notoriously imprecise, the use of multiple assessments throughout the year is assumed to average out the measurement error.

Interpretation (by parents and other teachers, for example) of the judgements that teachers make about student performance in an area of study is very challenging because teachers vary in what they consider in making an assessment. For some, the assessment reflects work habits, punctuality, task persistence, engagement, etc. There is little consistency across teachers and, often, little consistency from one assessment to another for the same teacher.

Lack of consistency about the features that a teacher is taking into account and inconsistency across teachers compromise the validity of the judgements and decisions made on the basis of assessments. These undisclosed aspects of a teacher’s assessment might be one reason why some teachers complain about “teaching to the test.” The “test” is only one of several things that a teacher uses to gauge student achievement.

That all teachers “teach to the test” is definitely not an indictment of what they do. Teaching to the test is a crucial element of all instruction and should be celebrated.  

Enjoy Your Spring Break, Charles

I will post again on March 31st

Wednesday, March 3, 2021

The Challenge of Classroom Assessment

 

Charles Ungerleider, Professor Emeritus, The University of British Columbia

[permission to reproduce granted if authorship is acknowledged]

Teachers’ classroom assessments are high stakes. They are the basis for decisions about promotion of students from one grade to the next, the awarding of graduation diplomas and scholarships, and they are often consequential for admission to post-secondary studies.

High stake decisions are not the only purposes to which classroom assessments are put. Teachers use classroom assessments to inform students about whether they have mastered the knowledge they are supposed to have acquired, to plan and modify the teachers’ instructional plans, to provide opportunities for students to practice and apply what they have learned, to determine the knowledge and understanding students need to progress to the next level, to communicate to the student and the student’s parents how the student is doing, to motivate students and more.

Teacher classroom assessment occurs daily and is labour intensive. Observation of student performance happens frequently albeit unsystematically throughout the day. Teachers create numerous kinds of assessments. Assessments include quizzes, tests, and opportunities for demonstration and oral presentations. Appraising performance on each of the teacher-created assessments is time-consuming.

The challenge with teacher assessment is that there isn’t agreement among teachers about what should be assessed. Teacher assessments are driven by a teacher’s professional judgment and, thus, by the teacher’s values. Some teachers believe that assessment should be confined to performance in the subject being assessed. Others believe that assessment should consider student attitude, motivation, or work ethic. Most teachers are aware that a student’s background, gender orientation, and/or personality are not relevant considerations for assessment.

Some teachers argue that assessment should be confined to the student’s performance in the classroom. Others say that homework is an extension of the classroom. Those opposed to assessing homework point out that homework is a form of practice – the most important facet of which is the teacher’s feedback. Others take a developmental perspective, arguing that a formal assessment of homework (beyond just the teacher’s feedback) is acceptable, but that teachers should take a developmental perspective by weighting the assessments more heavily the closer they are to the end of the unit. Some teachers are opposed to assessing homework because it is difficult to determine whether the student worked independently or with the assistance of others.

Teachers sometimes say that colleagues place too much value on print-based assessments (quizzes, tests, essays, reports, etc.) penalizing students who understand the material but have difficulty demonstrating what they know in writing. These teachers supplement written assessment with other demonstrations of understanding (oral presentations, dramatizations, diagrams, use of manipulatives, etc.).

Educational researcher Susan Brookhart and her colleagues reviewed the literature devoted to grading over the past 100 years with a particular emphasis on the meaning and value associated with grades. Among the many important findings of this review were:

·       Teachers primarily assess achievement using tests.

·       Teachers assess many non-academic factors in assigning grades, including effort, improvement, perceptions of student ability, completion of work, and other student behaviour.

·       The relative emphasis teachers assign to academic and non-academic factors differs markedly across teachers.

·       Assessment varies considerably by grade level.

The conclusion that Brookhart and her colleague draw is worth quoting at length:

This review suggests that most teachers’ grades do not yield a pure achievement measure, but rather a multidimensional measure dependent on both what the students learn and how they behave in the classroom. This conclusion, however, does not excuse low quality grading practices or suggest there is no room for improvement. One hundred years of grading research have generally confirmed large variation among teachers in the validity and reliability of grades, both in the meaning of grades and the accuracy of reporting.

Their conclusion is troubling. Inconsistency in the assessment process means that teachers other than the one who made the assessment find it difficult to know how to interpret their colleague’s assessment. It takes a leap of faith to believe that, in the aggregate, teacher assessments can support the high stakes decisions on which they are based.

Some ministries and school boards try to reduce the variability of teacher assessments by encouraging the use of performance standards and separating the appraisal of academic and non-academic factors. While those efforts are to be commended, those practices are also inconsistent across jurisdictions.

Notwithstanding these efforts, much more work is needed to improve classroom assessment, including clarifying the purpose of assessment, distinguishing between academic and non-academic factors, improving teacher preparation in assessment and in communicating results of assessment to students, parents, and others who make use of assessment results.

_______________

Brookhart, S. M., Guskey, T. R., Bowers, A. J., McMillan, J. H., Smith, J. K., Smith, L. F., Stevens, M.T., Welsh, M.E. (2016). A Century of Grading Research: Meaning and Value in the Most Common Educational Measure. Review of Educational Research, 86(4), 803-848.