Showing posts with label assessment. Show all posts
Showing posts with label assessment. Show all posts

Wednesday, November 1, 2023

Some mistakes are too important to ignore

 Charles Ungerleider, Professor Emeritus, The University of British Columbia

[permission to reproduce granted if authorship is acknowledged]  

There are mistakes in Geoff Johnson’s article “FSA not a good way of assessing Indigenous students” in the Times Colonist posted on October 29th that deserve correction for several reasons. One is that the Times Colonist is widely read. The other is that the designation of “former superintendent” adds authority to the errors he has made.  

As noted in his article, “the Foundations Skills Assessment is administered to all Grade 4 and Grade 7 students, both Indigenous and non-Indigenous.” The mistakes arise from his claim that the Foundation Skills Assessment is not a good way of assessing Indigenous students. One of them is the logical fallacy in his statement. He writes, “turns out that the four ‘lowest performing’ elementary schools on the FSA, according to the Fraser Institute’s one-shot ranking system, have significant populations of First Nations students.”  

In making that statement Johnson is committing the logical fallacy “after this, therefore because of this.” The logical fallacy occurs when someone assumes that because one event or set of conditions preceded another, the first event must have caused the second. In saying that the lowest performing elementary schools have significant populations of First Nations students, he is implying that the poor performance is attributable to the student composition of those schools.  

Another mistake deserving correction is drawing the connection between the performance of those schools and the assertion that “assessing a child in a way that does not seem meaningful or relevant to their life and culture is inauthentic and therefore meaningless, because it does not respect the learning of the whole child.” Indigenous children live in a society in which and literacy in the dominant language and numeracy figure prominently and where both have been used to deny Indigenous peoples their rights. Knowing how well Indigenous students perform on such assessments is essential to ensuring that they are being educated to the same standard as their non-Indigenous peers and from equipping them with the knowledge they need to defend their rights.  

I agree with Johnson that ensuring that the measures used to assess Indigenous youngsters are fair is essential. Johnson strongly asserts that “a central problem with the lack of validity of the FSA, as far as Indigenous students are concerned, is that the tests often contain items expressed in a way not obvious to an Indigenous student who might have a worldview and experiences that differ from the dominant Western culture.”  

This is something about which evidence can be brought to bear. Differential Item Functioning (DIF) is a technique used in assessment to determine if a question is fair to different groups of people. Imagine, for example, a math assessment, and you find out that for some reason, boys are more likely to get Question 5 right than girls, even when both groups are equally good at math. If that's the case, then Question 5 has "differential item functioning." Question 5 is not measuring math skills equally for boys and girls. DIF helps us figure out if a particular question on a test is easier or harder for one group of people compared to another, helping to ensure that tests are fair and unbiased.  

Before strongly asserting that the FSA is inappropriate for Indigenous students, one should consider the evidence. Johnson’s reference to the 2013 article by Jane P. Preston and Tim R. Claypool and the invocation of the imprimatur of The Canadian Council on Learning does not substitute for evidence about the Foundation Skills Assessment.  

Johnson quotes a B.C. government media release saying, “the redesign of curriculum maintains a focus on sound foundations of literacy and numeracy while supporting the development of citizens who are capable thinkers and communicators, and who are personally and socially competent in all areas of their lives.” He follows the quotation with the claim that the government statement ignores research about Indigenous ways of learning and is “dangerously close to being as colonial as you can get.”  

Could it not be argued that Johnson’s assertion is colonial? Johnson seems to imply that Indigenous students are by reason of ancestry or circumstance unable to demonstrate that they are capable thinkers and communicators. I doubt that was his intention. However, it seems very similar to the incorrect inferences drawn about women; namely that they are by constitution incapable of being pilots, surgeons, entrepreneurs, etc.  

There are three crucial omissions from Johnson’s article. One is the fact that the First Nations Leadership Council supports the use of the Foundation Skills Assessment as “. . . one among many tools necessary to address the ‘racism of low expectations’ experienced by First Nations learners as identified by the BC Auditor General in their 2015 report.” The second is that Indigenous educators are involved in item development for the assessment. The third is that the assessments contain First Peoples content written by and about First Peoples, and their development is guided by the First Peoples Principles of Learning.  

It is for these reasons that I hope Johnson will revise his article in the Times Colonist.

Wednesday, September 6, 2023

Making sense of British Columbia’s proficiency scales

Charles Ungerleider, Professor Emeritus, The University of British Columbia

[permission to reproduce granted if authorship is acknowledged]

Traditional letter grading systems have been ingrained in our thinking about education. But when subjected to closer examination, their limitations become glaringly apparent. For example, an A in Social Studies might give an impression of overall mastery, but it may mask specific areas of struggle, like the interpretation of primary sources. The focus on letter grades has led us to assign tremendous value to a system that falls short in providing comprehensive insights about students’ academic performance or about the curriculum they are trying to master. 

Letter grading systems are appealing. Their simplicity, familiarity, and directness have made the “A” a universally accepted symbol of achievement. There's a certain comfort in being able to categorize performance into neat compartments from A to F. However, the core of performance assessment lies in the interpretation of the letter grades.  

A significant portion of our society tends to view these grades as an indicator of a student's standing relative to their peers. We've invested so much meaning into the symbol of an A, for instance, that it's taken as a benchmark for excellence, even outside of academia; corporations strive for AAA bond ratings and consumers purchase Grade A beef. Yet, what these grades fail to communicate is the extent of a student's learning and areas that require further improvement.  

The transition to proficiency scales has been met with skepticism, often born from the fear that the scales might dilute academic competitiveness and the standards we are accustomed to. Yet, these scales present a nuanced picture of a student's abilities, growth, and areas for improvement. By breaking down performance into categories of 'emerging,' 'developing,' 'proficient,' and 'extending,' teachers can communicate more information about student progress.  

In this system, 'emerging' isn't synonymous with failure but signifies the early stages of understanding. 'Developing' means the student is applying their learning more consistently. The aim is for every student is to be 'proficient,' where they can reliably demonstrate their learned skills. 'Extending' students go beyond, showing a deeper understanding than is typically expected.  

This framework differentiates between a student who is "just starting to demonstrate learning" and one who is "showing an initial understanding". These phrases might seem similar, but they capture crucial differences in a learner's progress. Proficiency scales that are properly aligned with well-defined competencies enable parents, educators, and students to identify specific areas requiring further focus, a feature traditional letter grades do not offer.  

Proficiency scales are a departure from the punitive notions of 'passing' or 'failing,' signaling that learning is a continuum. The purpose of the proficiency scale is to facilitate progression, to help students navigate from 'emerging' to 'developing,' and eventually to 'proficient' or 'extending.'  

Skepticism and resistance are inevitable results of systemic changes. But I would counsel those who are skeptical to approach proficiency scales with an open mind. It's crucial to look beyond the symbol of a grade and focus more on the depth of learning and areas for growth. Proficiency scales promise a broader, more comprehensive understanding of a student's educational progress that letter grades cannot convey.

 

Wednesday, October 27, 2021

What is measured matters

 Charles Ungerleider, Professor Emeritus, The University of British Columbia

[permission to reproduce granted if authorship is acknowledged]

 There are variations on a common theme in the discussion of large-scale student assessment. One version is “what is measured matters.” Another is “what matters is measured.” Those who argue that confining large-scale student assessment to literacy and numeracy gives prominence to those capacities and diminishes the other contributions that schools make. 

It is important to give prominence to literacy and numeracy because they are so fundamental to learning in school and out. There are, however, many important contributions of schooling that are not systematically measured across the education system.

Consider the school system where I live. British Columbia’s school system is designed to “enable learners to develop their individual potential and to acquire the knowledge, skills, and attitudes needed to contribute to a healthy society and a prosperous and sustainable economy.” To that end, it strives to develop educated citizens who are:

·         thoughtful, able to learn and to think critically, and who can communicate information from a broad knowledge base;

·         creative, flexible, self-motivated and who have a positive self image;

·         capable of making independent decisions;

·         skilled and who can contribute to society generally, including the world of work;

·         productive, who gain satisfaction through achievement and who strive for physical wellbeing;

·         cooperative, principled, and respectful of others regardless of differences; and

·         aware of the rights and prepared to exercise the responsibilities of an individual within the family, the community, Canada, and the world.[1]

In recent years, the British Columbia Ministry of Education has revised the provincial curriculum to better reflect these goals. Now it is time for the Ministry to revise its assessments to align with its vision of the educated citizen and the curricula designed to help students realize that vision. To that end, the Ministry should develop and implement a new suite of provincial assessments:

Print and media literacy: Literacy is the foundation for school success and success later in life. Literacy is essential for developing numeracy, critical thinking, problem-solving, and almost every other human capacity. When students do not acquire a strong literacy foundation early in their school careers, they are more likely to experience failure in school and lack the foundation for productive, adult citizenship.

Using communication technologies is ubiquitous. Misinformation and dis-information are major societal problems. Being media literate is as important as being print literate and is as crucial to critical thinking.

Numeracy: Understanding and working with numbers is fundamental to everyone’s life. Thought and action depend on understanding and using numbers. Deciphering a recipe, reading a climate graph, computing interest, sequencing an argument, dancing, playing an instrument, and constructing an historical timeline are illustrative activities that require an understanding of numbers and the ability to apply them.

Critical thinking: The ability to formulate a question, analyze an argument, ask and answer challenging questions, judge the credibility of sources, make inferences, and identify unstated assumptions are among the abilities that critical thinkers possess and use in every aspect of life.

Communicating: Representing and presenting ideas, arguments, and emotions in ways that are coherent and understandable to others are essential to effective communication.

Social and personal competence: We use our abilities to self-regulate, empathize, motivate, read social situations, and develop relationships to work with others productively, settle disputes, and cooperate with others every day.

 These are examples of assessments that the revised BC curriculum requires to realize its promise to society. These assessments should have NO consequences for individual students or teachers; in other words, they will be no stakes assessments. They should be designed to provide information:

                    about how well students have mastered the curriculum.

                    about equity among sub-populations of students.

                    to parents about the progress their children are making.

                    for developing policy, allocating resources, and providing opportunities for professional learning.

                    about how well the education system is fulfilling its mandate.

                    to improve public confidence in the education system.

Provincial assessments that give teachers information about their students provide an opportunity for a rich discussion among educators about their own expectations and those of others.  In these discussions teachers can learn from one another about their instructional practices, what Andy Hargreaves calls the “derivatization” of the classroom. If each teacher operates within her/his own bundle of expectations for students, with no reference to others inside and outside the school, there is no reason for the teacher to challenge her or his assumptions and expectations. Such discussions are essential to the collaboration among professionals that can lead to greater equity of performance and, ultimately, outcomes. 

There is much more to schooling than what we currently measure on a system-wide basis. Schools develop capacities for thinking critically, for communicating ideas and emotions in a variety of media, and for developing us as human beings and teaching us to relate to others respectfully . . . and much more. Those capacities matter and they should be measured systematically.



[1] Statement of Education Policy Order, OIC 1280/89. British Columbia Ministry of Education. https://www.llbc.leg.bc.ca/public/PubDocs/bcdocs/365524/oic_1280-89.pdf

Wednesday, March 3, 2021

The Challenge of Classroom Assessment

 

Charles Ungerleider, Professor Emeritus, The University of British Columbia

[permission to reproduce granted if authorship is acknowledged]

Teachers’ classroom assessments are high stakes. They are the basis for decisions about promotion of students from one grade to the next, the awarding of graduation diplomas and scholarships, and they are often consequential for admission to post-secondary studies.

High stake decisions are not the only purposes to which classroom assessments are put. Teachers use classroom assessments to inform students about whether they have mastered the knowledge they are supposed to have acquired, to plan and modify the teachers’ instructional plans, to provide opportunities for students to practice and apply what they have learned, to determine the knowledge and understanding students need to progress to the next level, to communicate to the student and the student’s parents how the student is doing, to motivate students and more.

Teacher classroom assessment occurs daily and is labour intensive. Observation of student performance happens frequently albeit unsystematically throughout the day. Teachers create numerous kinds of assessments. Assessments include quizzes, tests, and opportunities for demonstration and oral presentations. Appraising performance on each of the teacher-created assessments is time-consuming.

The challenge with teacher assessment is that there isn’t agreement among teachers about what should be assessed. Teacher assessments are driven by a teacher’s professional judgment and, thus, by the teacher’s values. Some teachers believe that assessment should be confined to performance in the subject being assessed. Others believe that assessment should consider student attitude, motivation, or work ethic. Most teachers are aware that a student’s background, gender orientation, and/or personality are not relevant considerations for assessment.

Some teachers argue that assessment should be confined to the student’s performance in the classroom. Others say that homework is an extension of the classroom. Those opposed to assessing homework point out that homework is a form of practice – the most important facet of which is the teacher’s feedback. Others take a developmental perspective, arguing that a formal assessment of homework (beyond just the teacher’s feedback) is acceptable, but that teachers should take a developmental perspective by weighting the assessments more heavily the closer they are to the end of the unit. Some teachers are opposed to assessing homework because it is difficult to determine whether the student worked independently or with the assistance of others.

Teachers sometimes say that colleagues place too much value on print-based assessments (quizzes, tests, essays, reports, etc.) penalizing students who understand the material but have difficulty demonstrating what they know in writing. These teachers supplement written assessment with other demonstrations of understanding (oral presentations, dramatizations, diagrams, use of manipulatives, etc.).

Educational researcher Susan Brookhart and her colleagues reviewed the literature devoted to grading over the past 100 years with a particular emphasis on the meaning and value associated with grades. Among the many important findings of this review were:

·       Teachers primarily assess achievement using tests.

·       Teachers assess many non-academic factors in assigning grades, including effort, improvement, perceptions of student ability, completion of work, and other student behaviour.

·       The relative emphasis teachers assign to academic and non-academic factors differs markedly across teachers.

·       Assessment varies considerably by grade level.

The conclusion that Brookhart and her colleague draw is worth quoting at length:

This review suggests that most teachers’ grades do not yield a pure achievement measure, but rather a multidimensional measure dependent on both what the students learn and how they behave in the classroom. This conclusion, however, does not excuse low quality grading practices or suggest there is no room for improvement. One hundred years of grading research have generally confirmed large variation among teachers in the validity and reliability of grades, both in the meaning of grades and the accuracy of reporting.

Their conclusion is troubling. Inconsistency in the assessment process means that teachers other than the one who made the assessment find it difficult to know how to interpret their colleague’s assessment. It takes a leap of faith to believe that, in the aggregate, teacher assessments can support the high stakes decisions on which they are based.

Some ministries and school boards try to reduce the variability of teacher assessments by encouraging the use of performance standards and separating the appraisal of academic and non-academic factors. While those efforts are to be commended, those practices are also inconsistent across jurisdictions.

Notwithstanding these efforts, much more work is needed to improve classroom assessment, including clarifying the purpose of assessment, distinguishing between academic and non-academic factors, improving teacher preparation in assessment and in communicating results of assessment to students, parents, and others who make use of assessment results.

_______________

Brookhart, S. M., Guskey, T. R., Bowers, A. J., McMillan, J. H., Smith, J. K., Smith, L. F., Stevens, M.T., Welsh, M.E. (2016). A Century of Grading Research: Meaning and Value in the Most Common Educational Measure. Review of Educational Research, 86(4), 803-848.

Tuesday, February 23, 2021

The fallacies of a sampling approach to student assessment

 

Charles Ungerleider, Professor Emeritus, The University of British Columbia

[permission to reproduce granted if authorship is acknowledged]

My previous blog examined some of the claims made about large scale student assessment: that they prompt teachers to teach to the test; waste valuable time and resources; do not assess everything that is important; are stressful for students and teachers; do not take into account differences among students; and allow some individuals and organizations to make invidious comparisons among schools.

Another argument that opponents make about large scale student assessment is that they should be administered to a percentage of students (a sample) rather than all the students at a particular level (census approach). That, it is argued, would save resources, and prevent those who misuse the results from doing so. It is doubtful that a sample approach to large scale student assessment would achieve those desired outcomes. But there is a more serious problem: sampling will not work if one is really concerned about equity of outcomes for all students because it reduces the ability to identify factors that impede success for each and every student.

 A sampling approach (as opposed to a census approach) has other deficiencies. Sampling prevents examining the trajectories of students. In other words, with a sample, it is almost impossible to tell whether the students who performed poorly at grade 3 had improved by grade 6. Sampling doesn’t allow the system to see what has happened to transient students, a particularly vulnerable population that is easy to overlook. It also increases imprecision by increasing measurement error. If you want to know how students’ prior performance relates to their subsequent performance, you need to survey all students.

I can illustrate the benefits of a census approach from a study in British Columbia. In 2006, the Minister of Education described high school completion rates of students for whom English was a second language (ESL) to be better than any other group the Ministry assessed (Victoria, Parliamentary Debates, p. 3530). After a closer examination of the data, Bruce Garnett and I (Garnett and Ungerleider, 2008) confirmed the accuracy of the Minister’s statement, but found that the high achievement of the numerous Chinese speakers masked the fact that smaller subgroups of the ESL population were faring poorly in the school system. We found that the strong performance of Chinese speakers (the largest ESL group in the data set) pulled the aggregate ESL graduation rates upwards. The graduation rates of all groups except Chinese speakers were very low, generally below 60%. The worst outcomes were among ESL speakers of Spanish, Vietnamese and Filipino languages. Identifying the differential success of various ESL groups would not have been possible if the data set had been generated on a sample basis. You cannot systematically address problems that you cannot identify.

Another advantage of a census approach is that it permits analyses that can help to identify factors over which the system exerts influence that facilitate or impede educational progress of groups of learners (for example, First Nations, Metis, and Inuit students, second language learners, students with special needs, etc.).

Most important, if equity among students is a priority, samples simply do not work. Even with carefully drawn samples it is difficult to detect small sub-populations of students to support meaningful analyses. I can illustrate this with reference to British Columbia’s student population which, at the time of the calculations below, was about 40,848 students at grade 4. With a student population of that size, we could draw a sample of 1,481 student, a number sufficient to meet the requirements for sound statistical analyses of the grade level population.

However, if you wanted to break down results by school board or wanted to study the performance of sub-groups of students, the sample would not work. Here is an illustration of why it doesn’t. The illustration assumes that the assessment is intended to produce results that fall within a confidence interval of +/- 2.5% 95 times out of 100 at the school board level.

School Boards

Number of Students available for assessment in 2019/20

Number of students required for a sample with a confidence interval of 2.5 at a 95% level of confidence

Abbotsford

1524

765

Alberni

294

247

Arrow Lakes

34

33

Boundary

99

93

Bulkley Valley

144

132

Burnaby

1740

816

Campbell River

397

316

Cariboo-Chilcotin

305

255

Central Coast

26

26

Central Okanagan

1690

805

Chilliwack

1022

614

Coast Mountains

271

230

Comox Valley

631

448

Conseil Scolaire francophone

601

432

~~~~~~~~~~~~~~~~~~~~~~~~

Table truncated*

~~~~~~~~~~~~~

Richmond

1377

726

Rocky Mountain

288

243

Saanich

484

368

Sea to Sky

409

323

Sooke

868

555

Southeast Kootenay

466

358

Stikine

15

15

Sunshine Coast

235

204

Surrey

5459

1199

Vancouver

3493

1067

Vancouver Island North

94

89

Vancouver Island West

35

34

Vernon

610

437

West Vancouver

501

378

Grand Total

40,848

22,200

*FULL TABLE AVAILABLE UPON REQUEST

         

In very large boards such as the Surrey School Board, the number of students sampled (1199) would be a relatively small proportion of the grade 4 student population (5,459). In a smaller board such as Vancouver Island North, the sample required (89) would encompass almost all the 94 grade 4 students in the Board. In the Conseil Scolaire Francophone Board the sample required would be more than two-thirds (432) of the 601 grade 4 students.

Overall, if you wanted to break down results by school board or wanted to study the performance of sub-groups of students, you would need to sample more than half of the total number of students at grade 4.

It is inconsistent for those concerned about the role education plays in helping to achieve social justice to want to restrict large scale student assessments by arguing in favour of sampling students. Those determined to achieve social justice ought to want to shine a light on discrepancies, not obscure them. My hunch is that, when considering the evidence of the limitation posed by sampling, advocates for social justice will see the benefits of a census approach to large scale student assessment.