Popular Post
Showing posts with label Testing. Show all posts
Showing posts with label Testing. Show all posts

Sunday, February 6, 2011

Testing, Scoring and Trusting the Data

A fascinating case study is the NY Regents, scoring, and the unintended (or purposeful) consequences of a line in the instructions to the scorers.

First, here's the graph of the number of kids getting each score (from WSJ).  The issue can be seen quite clearly. The graph overall is a typical left-skewed distribution, as you'd expect from this type of test. The trouble comes when you explore that jump in the middle at the passing mark.

Of course, the nattering class is all up in arms over this, claiming fraud and misconduct.  You can almost hear jeers of "Union Bastards trying to save their jobs by lying on the tests."  Unfortunately for those people, the reason comes down to one sentence:
"[State officials] note that the state actually requires teachers to regrade certain Regents tests where the student barely fails in order to check for grading errors."
When you are singling out tests for special consideration, and the stated focus is to look for "scoring errors" on "barely failing tests", the only way for the scores to change is up.  Since it's easiest to give one or two points to a 64 or 63, it's logical that those would be the most effected.


Looking at the graphs, you can see that some teachers (probably a school at a time) set their cut off at 50 while the majority set the cutoff at 55.

Why the negative slope in that region? When looking for ambiguous answers which could be scored higher, you have to "rescore" each problem until you get an appropriate number of points. It's easier to find one such than 10 such.

Similarly, since there was no reason to check more answers than just enough to get the kid to pass, the uptick was only to 65, though it seems that many teachers weren't keeping close track of the extra points and brought the kid up to a 66 or 67.

Conversely, if the teachers had been instructed to rescore all students within ten points of the cutoff instead of just those who were below it, then you would have seen some students adjusted downward, balancing out much of the upward movement.  If there were a similar uptick at the 65 mark in this hypothetical, then and only then can you claim that teachers are deliberately mis-scoring to jigger their VA measures. (and even then, I'd put the reason as teachers wanting to help students rather than being so coldly self-interested.)

The place where I found the link to this article had this comment: "Teachers don't want to flunk kids that just barely miss the passing score. Until all responsibility for creating and scoring state exams is given to an independent body with no interest in the results of the tests, the results reported should be viewed skeptically."

Ummm, no. Scoring tests is really complicated.  Pearson, the biggest company, uses part-time, barely out of college, minimum wage people to do the scoring. Getting the "right" score is more a matter of whether or not the scorer speaks English and actually knows the material.

Frankly, given the mess that the testing industry is in when it comes to scoring, I have a feeling that the teachers are doing a more conscientious job. If you want a nasty introduction to the follies of testing company scoring sessions, check out Todd Farley's "Making the Grades." It's well-written but damn depressing if counting on accurate test scores because you're stuck in the hell of value-added and merit pay.

Thursday, January 20, 2011

PISA scores - another look.

What do you see? When you look solely at schools with fewer than 10% of students on FRL (i.e., poor), US schools would be better than those of the top of the chart, Korea. When you include the whole spectrum of US schools and their students, the US is much lower.

The US has a much bigger spread than any other country (the largest standard deviation of wealth in the developed world).
Our overall scores are unspectacular because we have a high percentage of children living in poverty, over 20%. This is the highest among all industrialized countries. In contrast, child poverty in high-scoring Finland is less than 4%. For Cleveland and for the US as a whole, the major problem is poverty. Before we worry about teacher quality, institute longer school days, and increase testing, we need to make sure that all children are protected from the effects of poverty: This means adequate health care and nutrition, and access to books. When we do this, American test scores will be at the top of the world.-- Stephen Krashen

Saturday, December 25, 2010

That'll improve the school

Think they'll pass muster?
Forget about replacing the teachers ... replace the students. Bring in the superstars, the superheroes and the supernerds.Then our scores are sure to go up.  New Trier High School, Fairfax High School -- we're looking at you. Let's trade yours for ours!

Never mind.

Funny idea, though.

Monday, December 13, 2010

What's it worth if it isn't scored accurately?

SCores are being held consant.
The next time your principal complains about how the scores aren't going up, think about this.
From the Loneliness of the Long Distance Test Scorer: For some mysterious reason, unbeknownst to test scorers, the scores we are giving are supposed to closely match those given in previous years. So if 40 percent of papers received 3s the previous year (on a scale of 1 to 6), then a similar percentage should receive 3s this year. Lest you think this is an isolated experience, Farley cites similar stories from his fourteen-year test-scoring career in his book, reporting instances where project managers announced that scoring would have to be changed because “our numbers don’t match up with what the psychometricians [the stats people] predicted.”
Let's hear it for testing!

The whole article is interesting, but that line caught my eye.

Saturday, December 11, 2010

Looking for the Education

We're below average on test scores.
U.S. students are below average in math skills, according to PISA, while Asian countries excel.
So somebody decided to look at why. Family attitude seems to be the key: working your butt off and getting extra help seems to be the key to doing better on the test.
source: Just as the latest international testing data once again highlight the relatively poor performance of U.S. students in math, a new report has come out to further explore why the United States may be struggling, with a focus on the math attitudes, beliefs, and behaviors of parents, and their children's out-of-school activities. Among the key findings: Parents in Singapore are far more likely than those in the United States and England to engage a math tutor to help their child, they're more likely to get assistance from teachers and others in how to help their child, and their children more often take part in math competitions and math/science camps.
I'm okay with that.  If you want to do better on something, you practice it. If you want to test better, you need to practice on the test.

Just don't call it an education.

Call it "Education Hero."

Just as Guitar Hero ruins the ability to play an actual guitar, Education Hero shouldn't be used on actual students.

Sunday, December 5, 2010

What happened to "Education"?

Joanne Jacobs has a piece on evaluation programs that are looking to videotape teachers as they give lessons and then "Go Look at Tape." The issue seems to be that
More than 99 percent of teachers are rated satisfactory by their principals, reports a study on “the widget effect” by the New Teacher Project.
and this is somehow "bad," hence the need for overhaul. While I would argue that you need to improve the system so that middle-managers (Admin) can appropriately rate, measure and evaluate their assembly-line workers (teachers), I'm not sure that going to videotape is much of an improvement. If the Principal can't see the "errors" while sitting in the class, how is he supposed to make anything out of the 17th such review, even if he did have the time. Well, they have an answer for that ... Gates figures we'll pay outside evaluators who will somehow be able to see what the highly paid admin can't or won't.
The Gates Foundation is developing a new model, with the help of social scientists and teachers, reports the New York Times. Outside evaluators analyze videotapes to determine whether teachers are teaching well.
We've got money for outside evaluators? Pretty cool. I'll sign up. What gives you the idea that I'll be any more or biased/petty/superficial than the current system? The fact that I'll be able to rewind the tape? More likely I'll fast forward through it. The fact that I don't know the teacher so I'll be more honest? More likely, I'll make a snap judgment and move on.

You know we'll pay for training out of our RTTT money or something, hire bunches of consultants to teach teachers to evaluate other teachers who are trying to teach students. Pretty amazing, ain't it?
Hundreds of teachers will be trained to review 64,000 hours of classroom video. They will look “for possible correlations between certain teaching practices and high student achievement, measured by value-added scores.”
Let's say it's 800 teachers - that's 80 hours of tape to watch, per person. Asking a lot? I can't even spend that much time watching TV or movies in a month while sitting mindless and half-comatose. I'd need to be unemployed to look at 80 hours of tape multiple times and evaluate someone.

But here's the crux of my complaint ... those correlations. I've had students who completely and utterly bombed in class but did well on testing. I've also had plenty who did well in my class, well in the next couple of classes, had a great high school career, and then a great college career, and are now cranking out 6 figure salaries ... but didn't do well on testing. How do I know they did well? They told me.

Since when did value-added scores on NEAP mean an education? Since when could anyone figure out what I do that "works" by watching 10, 20 hours of videotape? Really? Do they know me? Since when did someone watching a few hours of video actually have a clue as to what went on and who and how the students were effected? Ten hours of video? That's two days. Does anyone here think that I couldn't fake it for two days and mess around the rest of the year if I was of a mind to?

I didn't think so.

Friday, December 3, 2010

Awards for all?

Joanne Jacobs has a piece about a parent complaining that only some kids were recognized for their performance on state testing. I actually agree with the parent on this one.
This, he said, was unfair to students who traditionally score lower on standardized tests and might not reach proficiency no matter how hard they try — mainstreamed special education students, for example.
Some kids will never reach proficiency. It's just a fact of life. If you make your standards low enough for all, then the achievement is meaningless and all of the kids know it and blow you off.The better response, for me, is to simply thank everyone publicly, en masse, and announce the barbeque for all. Then, reward or thank the good students separately.

Really, the "Always Praise in Public" rule isn't always your best course of action. I am against, for example, the tactic of an academic assembly during the first two periods of the day for which every class trudges down to the gym and sits by class.

"Everyone is sure to be a winner
with these fun 4" trophies."
What happens next is the whole point and is also the most excruciating part: the Guidance counselor, feeling all very important because she gets to "honor" the "good" students and bask in the reflected glory, reads the names of the high honor roll (20 kids), honor roll (140 kids), and merit (15 kids). She asks them to stand while this endless list is being read. Do you know how long it takes to read 175 names?

A little subtraction shows that there's maybe 30 kids in that grade who couldn't manage to get anything ... administration is proud that they didn't publicly shame any of them by saying their names. Except that they are still sitting down while everyone around them is standing up. Is it any wonder that they feel like shit? Most memorable student quote about the assembly (in informal geometry afterward): "Here are all the smart people in the school and none of them are YOU."

Back to the fun. Guidance has them sit ... and does the 11th grade. And repeats for the 10th. And the 9th. Can't have anyone left out, can we? An hour and something later, you've managed to humiliate as many people as possible, so you send the school back to class. "Don't make any comments to the Dweeb in the hallway, now."


Maybe the proper response is to not require everyone to recognize them. Give them their own awards night and invite them and their parents to come or not, as they choose.  You know, like the sports awards night, where the non-athletic can avoid having to sit through endless coaches' attempts at public speaking.

God knows there is nothing worse than a coach with limited vocabulary and no experience speaking to a crowd who's attempting to appear smart, clever, witty and interesting ... for each of his 54 football players, naming and praising the "spectacular work ethic" of every member of the 1-6 team, including the kids who lost eligibility for drinking and fighting.


It was also interesting that the rest of the fall teams had much better seasons, one winning a state championship, but the football team spent the most time congratulating itself.  But I digress ...

Maybe the takeaway from all this is simple:  The people who attend an awards night should be the ones who were there to watch the achievement itself.  Anyone else is excused. Those who wish to attend can do so.  I'd much rather have an awards night with the rest of my team and the spectators who were at the games. Everyone else feels like "Johnny come lately" hangers-on.  The same is true for academic awards:
If you weren't there when we did it, why would you want to be there to celebrate it?
If you can answer that question, then you can come to the ceremony and we'll all have a blast. If you can't, then you shouldn't be required to be there and we probably would feel uncomfortable if you did show up.

Tuesday, November 23, 2010

Looking for the Value in Value-Added.

In all of the argument and debate in the merit-pay/ teachers suck arena, the topic of value-added assessment comes up often. I wonder if anyone can actually delineate in which fashion or way that a multiple-choice, standardized test that is taken by students who have no pressure on them at all and who have no immediate personal need to pass the thing can possibly measure the educational state of said students, let alone any "added value" imparted by me.

Standard-based, constructivist, traditional ... none of that seems to matter.

Friday, October 1, 2010

Highly Qualified or Effective?

Mike Petrilli on Education Gadfly writes about HQT and how terrible it is:
Everyone knows it’s a meaningless designation. Nobody will defend its focus on paper credentials. The conversation has moved on to teacher “effectiveness” as measured by student learning and other meaningful indicators. Yet in the real world of real schools, HQT is still the law of the land, wreaking havoc every day. It continues to make teachers jump through unnecessary hoops. It continues to tie the hands of charter schools that have to demonstrate that their teachers have requisite “subject matter knowledge”—never mind the autonomy charters are supposed to receive. And now it’s causing material harm to Teach For America, one of the best things our education system has going.
Awesome. and Stupid. and Inconsistent.

HQT makes teachers hop or skip through low-lying hoops, like certification and demonstrating that you actually know what you are trying to teach. The praxis test is required. Not much more. The checkboxes for HQT are so incredibly easy to tick off that 93.8% of Vermont teachers are HQT. The obvious response is "You want to be a teacher. Teachers have paperwork. This is easy. Get over yourself."

Mike wants "teacher 'effectiveness' measured by student learning and other meaningful indicators." So, if I understand this, he wants the teachers' effectiveness measured by students taking a meaningless high-stakes test instead of the teachers taking a meaningful one. That's silly. The teachers should take the tests. After all, if they fail, it'll be their college professors' fault.

Saturday, June 19, 2010

Merit Pay and Testing

Campbell's Law: “the more any quantitative social indicator is used for social decision making, the more subject it will be to corruption pressures and the more apt it will be to distort and corrupt the social processes it is intended to monitor.”
h/t: Jim Horn

Saturday, April 24, 2010

This is how testing should be used.

From Joanne Jacobs' Community College blog is this note about California's Early Assessment Program, meant to ascertain whether students are ready for college. Instead of waiting until after admission, they've decided to do it early, in the 11th grade. (DUH!) This allows students to pick up the pace in high school while there's still time to get these skills.
Most juniors who take the optional exam are told they’re not ready for college math and English, giving them an incentive to use senior year to boost their skills.

Forget NECAP, indirectly measuring teachers and punishing schools. This is testing that is useful and appropriate.
All agree on the value of eap (at Educated Guess Blog)
Here are some numbers:
"In 2009, 79 percent of juniors took the English exam, but only 16 percent of them were deemed ready. Last year, 36 percent of juniors took the math EAP. Of those who took the Algebra II test, only a quarter were deemed college-ready (20 percent conditionally ready, 5 percent fully ready; of those who took Summative Math, with of bits of Algebra 1, Geometry and Algebra II, the results were better: 88 percent were ready (67 conditionally and 21 percent fully). The combined result for students who took either test: 57 percent (34 percent conditionally ready, 13 percent fully ready)."

Are you on track for college? (at CCBlog)
California’s community colleges want 11th graders to take the Early Assessment Program – EAP–  exam developed by the California State University system. Most juniors who take the optional exam are told they’re not ready for college math and English, giving them an incentive to use senior year to boost their skills. Until now, the state’s 112 community colleges have given their own exams to decide who needs remedial coursework, writes John Fensterwald on Educated Guess. Now 10 community colleges have agreed to use the EAP and 15 more are considering it. It’s a lot cheaper to catch up on academic skills in high school than to wait till college. In fact, higher education leaders want to get students on the college track in eighth and ninth grade.

Teacher Quality Study

Via Joanne Jacobs, via Reuters, from Science comes a twin study on teacher effectiveness.

I’d be interested in the criteria for “strong” teacher and “weak” teacher. If the “strong” one is defined (as the linked article seems to imply) as someone whose 20 students did well on a single reading test, then perhaps it isn’t surprising that one of those students did better on a test.

At least this study seems to be comparing teachers within the same building, using twins. That’ll help control for income and such. I’m not as convinced that a single measure of just one skill is a particularly good way to differentiate, however, though my expertise is certainly not at the k-3 level. Assuming that you can even separate the discipline at this level I have to wonder: if one teacher is better at science or math while the first is only good at teaching the reading skills needed for that one test, the differences could be enough to label one or the other as “deficient” and that’s not a good way to do this.

I’d also like to know how the kids were placed in the various classes. If it’s random, then this study has more credibility, but if the placements are done for reasons of ability or behavior, then that test is measuring too many variables to definitively say one teacher is better than the other.

I'm sure the researchers tried to control for this, but schools do tend to separate kids on purpose and that will skew the study. If a weaker student is intentionally placed with teacher #1, then it shouldn't be surprising if the teachers show a difference.

All in all, a good study. Hopefully it's not taken directly to the Congress as the article suggests - it's way too early for that and the Congress would probably try to apply its lessons to high or middle schools and it's way to specific for that.

BTW, the photo that accompanied the article showed a child using an abacus during a national math test in India. If one of our teachers used an abacus during teaching and the other used a calculator, that would definitely change the results.

Article below the jump

Poor teachers may hamper good students: U.S. study
Thu Apr 22, 2010 4:45pm EDT
WASHINGTON (Reuters) - An unusual genetic study supports the argument that good teachers make a difference and shows that poor teachers may do damage, even to gifted students, researchers said on Thursday.
The study, published in the journal Science, showed that effective teachers help kids with the best genes read better, while poor teachers brought down all the children in a classroom to the same mediocre level.
The findings by behavioral geneticist Jeanette Taylor at Florida State University and colleagues could influence the debate in Congress, the White House and school districts across the United States about measuring the quality in schools.
"In circumstances where the teachers are all excellent, the variability in student reading achievement may appear to be largely due to genetics. However, poor teaching impedes the ability of children to reach their potential," Taylor and colleagues concluded.
To tease out the effects of genes and environment, the researchers turned to the time-tested model of twins. Identical twins share all their DNA, while fraternal twins share about half, or as much as any brother or sister.
Their theory: if one identical twin does better than his or her sibling in a different classroom, much of the difference must be due to the teacher.
They studied 280 identical twin pairs and 526 fraternal twin pairs in the first and second grades from a diverse selection of Florida schools.
To determine teacher quality, they used Oral Reading Fluency test scored for the entire classroom of each twin.
"(It's) a timed measure of how many words children can read in a passage," Taylor said in a Science podcast.
They checked to see how many more words children could read at the end of the year compared to the beginning.
"We felt that was a reasonable estimate of teacher effectiveness," Taylor said.
When teachers were good, the genes really mattered. If one identical twin excelled with a good teacher, the other did too. But if one twin had a strong teacher and the other twin had a weak teacher, twins with strong genetic potential did just so-so in the poor teacher's class.
"Better teachers provide an environment that allows children to reach their potential," Taylor said.
"As state and national policy increasingly focuses on teacher quality, the effect that teachers have on the strong documented genetic foundation of reading is an important question," the researchers wrote.
While great teachers do not guarantee success, they said, policymakers need to realize that good teaching is important, even for gifted children.
One weakness of the study -- the researchers threw out data from identical twins who happened to be in the same classroom who scored differently from one another on the reading test. Taylor said she did not know how many such cases there were.
(Reporting by Maggie Fox, editing by Stacey Joyce)

Sunday, April 18, 2010

Incentives and Achievement

"A multicity experiment to test the effect of paying students for performance succeeded in increasing achievement when the payments were tied to specific behaviors related to learning, such as reading books, but not when the awards depended directly on test scores, new findings show."
“Providing incentives for achievement-test scores has no effect on any form of achievement we can measure,” wrote Harvard University economist Roland G. Fryer in a working paper published last week by the National Bureau of Economic Research."
I can't say that I am surprised by this. Tests seem to be out of the student's control and out of mind quickly. The feedback on tests is too late to be tied to the achievement, while the monetary payoff for reading and such is much more direct and seemingly more in the control of the student.

It may also be an additional explanation why high school students don't score as well on these tests as might be expected from all the other measures of their ability - the reward payoff is delayed so far that the effort simply isn't put into the test. NECAP tests were in November, scores are returned in March/April and school AYP determinations are delayed until May.

One more nail in the coffin of "Merit Pay Should be Based on Test Scores."

Monday, April 12, 2010

Testing, Testing ... 1, 2, 3?

Up here in the great white north, the math & English tests are given in Nov of junior year and the scores are returned to us by march - maybe. We (and many other schools in this state) are on a block schedule so any student who has English or math during first semester is done by the time scores are returned. For the 40% of the students who had English and/or math in the fall semester of sophomore year then not again until spring of junior year, there is a gap of nearly a year between their most recent relevant class and the tests that measure "the students" - go figure the intelligence of that idea.

The science tests are given in May of the junior year and the scores returned to us in the fall of the senior year. Not very helpful either.

The upshot is that the scores cannot be used to measure the students, only the school and the teachers. Likewise, we cannot use the scores to help the students who earned them by giving them remedial classes or somehow making guidance decisions based on the results. Only the curriculum and the teachers can be affected.

Of course, the idea of basing anything on a test that so few take seriously is pretty silly, too. Why should they care, after all? No one gets to see these scores and nothing goes on a transcript.

It's obvious that neither the students' benefit nor their measurement is what is at stake here. This is, purely and simply, an anti-public school, pro-voucher, let's-find-something-to-hang-the-teachers-with idea.

An interesting, though only partially related, opinion in the mid-Vermont newspaper the other day. Full text below the jump of an English teacher's take on some of these issues. One of the hot-button topics up here is the consolidation of districts. We have one or two towns per district - making them all pretty small. The Commish wants to consolidate to save money but there isn't any real chance of saving money this way.


Imperial overreach

By CONRAD TUERK - Published: April 8, 2010

Though I agree with your strong denunciation of Vermont Commissioner Armando Vilaseca's proposal for school consolidation ("Time to get real" April 1), you fail to apply the same skepticism toward his department's pernicious meddling in local schools. In questioning Vilaseca's lust for power at the expense of local control, you ask, "Who is the commissioner of education to supersede this tradition of community involvement and local democracy?"

A good question, yet your enthusiasm for Obama's Race to the Top ignores the same potential dilemma. You urge Vilaseca toward "actual reforms that focus on teachers in classrooms." But why would Vilaseca's imperial overreach be any different there than in his consolidation scheme? Armed with a bushel of federal money, won't Vilaseca and his minions wield a heavier club? More importantly, why should we, at the local level, trust in the Vermont Department of Education's leadership?

If local trust is the heart of the issue, as I think it is, why has the Vermont Department of Education been so tone deaf and condescending in dealing with its constituents? The department's opaque and moronic ranking system, disgraced by bungled computations, is only one public example of the department's ineptitude and arrogance.

In a stunning display of hubris, Rae Ann Knopf, the department's deputy commissioner, distributed a memo on March 30 to state superintendents, principals and teachers chastising them for questioning her department's omnipotence. Primarily, she defends the department's reliance on one single test, the NECAP, to determine a school's effectiveness. She writes: "Our NECAP testing is based on the countless hours of work educators in Vermont and other northeast states put into creating high-quality standards for teaching and learning for our children."

As an experienced educator myself, who works with actual students, I take exception to her claim. I too value standardized test scores, but recognize them for what they are: a snapshot of student learning. If I based my classroom grading on one exam, as the DOE does, I would be compromising my professional responsibilities. Any serious educator understands the importance of using a variety of assessments to determine student learning.

The DOE should know better, especially since they themselves have trained teachers to differentiate their instruction to accommodate various learning styles. The practice arises from Gardner's theory of multiple intelligences, which identifies seven types of intelligences. Since only two of them, linguistic and logical-mathematical, have been traditionally served by schools, teachers have been urged to alter their methods and move away from instruction that focuses entirely on "academic" intelligences. But then they are being judged by a standardized test that measures only those two traits. In essence, teachers are being trained for a practice that runs contrary to how they are judged.

What makes it more maddening is that the NECAP test has no relevance to students. Their results do not affect them in any way. Any reasonable person with the slightest understanding of teenagers will recognize the hazards in such an enterprise. As a monitor for past NECAP exams, I can assure you that students have viewed the test, at best, as a needless distraction or, in most cases, as a complete waste of time. To judge the entirety of student learning, and school effectiveness, based on a single test that most students do not take seriously is beyond ludicrous.

Knopf also takes teachers and schools to task for not caring about all Vermont children: "There is no question that Vermont has a tremendous reputation for its strong educational system, where our young people learn and excel and most go on to lead meaningful, productive adult lives. The key here is most, but not all." She then backs up her claim by relying, you guessed it, on NECAP scores. "If only 18.5 percent of our students eligible for free and reduced lunch reach proficiency in math by eleventh grade, this means 29,000 kids in our state have little hope of gaining the skills necessary to be meaningful participants in today's global economy."

So the only way to be a "meaningful participant in today's global economy" is by reaching proficiency on the NECAP test? What about those young people who possess intelligences that cannot be measured by a standardized test? Do we not tap their interests and aptitudes, for fear of them flourishing, or do we insist they squeeze into our narrow definitions of success? Does a skilled laborer have a less secure future than a college-degreed cubicle jockey?

Not only do I, as a foot soldier, resent having my work in the trenches impugned by distant bureaucrats, but by setting up a strict dichotomy of winners and losers, based on faulty logic and evidence, the DOE may harm the very students it aspires to help. My years of teaching have taught me that every young person is a genius, but genius need not be academic. Rather than stigmatize students who struggle with the NECAP, and relegate them to less-meaningful-citizen status, we should create viable alternatives for students whose interests and aptitudes do not conform to traditional "schooling." Our goal should be to help all young people capitalize on their strengths, not just those who reach proficiency on standardized tests. Utopian thinking may play well in theory, but it fails to address real issues.

What the results of the first round of Race To The Top reveal is that consensus matters. Both winners, Tennessee and Delaware, showed strong stakeholder support: from school districts, business leaders, superintendents, and union reps. Unfortunately, from my vantage point, the Vermont DOE has a long way to go in garnering this level of support. To start, they could spend less time with data and more time with students and teachers. The human world might humble them.

Conrad Tuerk is an English teacher at Rutland High School.

Wednesday, February 10, 2010

I'm for testing - sort of.

Schools matter quotes a letter to the editor of the NYTimes:
"All educators understand the necessity of assessment, but it is our obligation to do the minimum amount of testing necessary, and no more. Every minute spent testing that is not necessary bleeds time from learning, and every dollar spent on testing that is not necessary is stolen from investments that really need to be made in schools. Any new education law should result in less testing, not more. - Stephen Krashen"

Apologies to Stephen but state-wide testing, done right, doesn't bleed anything or "steal money" from anything, really. Using loaded words does get you published but it doesn't make your argument any better.

If I had my druthers, I'd have testing that would happen mid-course and end of course. These would be called Midterm Exam and Final Exam. The exams would not be written by the teacher but by a group made up from the district. Exams should at least be department-wide. Scoring would be be done by the teacher, but with other teachers being able to review the materials. Multiple choice is half. Clearly defined scoring for the student-constructed response section. All teachers aware of the curriculum and of the topics on the exam.

That's it. Two tests. Every grade at the same time.
No assemblies. No illnesses. No field trips. No sports dismissals. No bullying seminars. No peer mediations. No suspensions. No doctor's appointments. No guidance appointments. No excuses.

If you want a nationwide test, that would take place at the same time for every student tested and would count somehow. Like the SAT.

Now that I think about it, giving everyone an SAT in October of their senior year would be cheaper and more accurate than the silly state-written things I've seen so far.

Tuesday, February 9, 2010

What's a failing school?

These truths I hold to be self-evident:
1. You can't fix a failing school by converting it into a charter school, replacing its administration with a new administration that had been fired from another failing school, or by testing until the kids go bat-crack crazy. (h/t to Ritchie for that expression!)

2. You can't measure a school based on tests that aren't taken seriously.

3. You can't reliably identify a failing school by testing because the criteria were so poorly defined in the first place and because the Law is looking from the wrong perspective.

4. The educational experience of a small fraction of one ethnic group doesn't represent the experience of all of the students in that school any more than my abilities as a teacher can be "averaged" with those of the loser next door and the PhD on the other side.

More:Why should the failure or marginal failure of one subgroup (a significant portion of whom passed) lower the boom on a school?

Looking from the top down, you can measure how the "school" performed, but each student has a different experience. This group may have done poorly because their particular teachers didn't "get" them or were stupid (someone has to fit the stereotype), but those other groups had a great education.

Same building. Same "faculty." Different teachers. Different families. Different education.

As any parent knows, there are good teachers and bad ones but far more good than bad. The bad ones just don't stay - if they do, there's a reason. What not everyone understands is that often it's not a matter of good and bad teaching but of good or weak connection between teacher and student.

If your kid is not capable of "meshing with" or learning from a particular teacher, you ask for a change of section. Every teacher will have kids who won't or can't learn well from them but who magically blossom with someone else -- but there's always a vice-versa. I know, for instance, that some students just enjoy being in my class -- for whatever reason -- and try to set up their schedules accordingly. Often, the older brother counsels the younger to take my class. Others prefer other teachers, whether for their teaching style or gender or height or discipline policy. Many either don't care or know enough to have an opinion. I don't take it personally.

You all have had the experience of sitting in an IEP, SAT, EST, 504, MVP, IPod, or faculty meeting and listening to the others complain about KidX. When it comes to you, you shrug and say "I've never had a problem with him. I called his mother and everything just worked out." Judging by the looks on everyone's faces, you've dropped a bombshell.

The real problem with NCLB is that an entire school feels the punitive effects of the law when as little as 2% - 5% actually experience the school environment that produced those low scores.

All schools have some students who receive a great education. Those schools have a majority of students who get a decent education. The school certainly hasn't failed them. All schools have the mouth-breathing, drug-using lowlifes who really need a life-altering experience before they have a life-ending one - why should the results of the latter reflect on the school of the former?

Thursday, November 26, 2009

Portfolios Inflate Scores - Who Knew?

Once more into the Breach, Dear Friends ...

According to the Washington Post, Portfolios inflate scores. Color me surprised.

When one type of assessment fails too many students, the response is "Let's change the teaching." When too many still fail, it's "Let's blame the teachers for not changing." When the scene doesn't improve, we then try to game the system and teach to the test "Let's teach test-taking skills using the released questions." If we are STILL not making the grade, we change the test and measure the students differently: "Testing without testing."

Portfolios.

We claim it's "more authentic" and a "21st century skill" and all that, but it's just misdirection. They're "fairer" and "more meaningful" only because they artificially raise the scores.

Portfolios are not necessarily the students' work (parents and teachers help, books are consulted), aren't as structured or as difficult, are usually the four or fifth rewrite (with so many specific corrections that the student's voice is lost), and most importantly aren't definite -- anything that looks good goes in, anything that doesn't, won't. How can one fail under those conditions?

"Teachers document learning throughout the year in a binder of class work, including worksheets, quizzes and writing samples." When you have such an obvious selection bias, don't be surprised if the final numbers are off. In this case, higher than they probably should be if you are actually expecting that the students know the same things as students in other jurisdictions.

Let's focus on this paragraph:
Last year, students tested with portfolios outperformed classmates who took multiple-choice tests in Fairfax. Students with disabilities surpassed schoolwide pass rates in reading or math tests in more than a dozen schools. Students learning English were far more likely to score in the highest performance tier on the reading test, which measures knowledge of language arts concepts such as metaphor and plot, than their native-speaking peers. Overall, English-learners and students with disabilities charted 20- and 18-point gains respectively in reading pass rates, compared to a six-point gain for the division.

Hummmmm.



Reposted here because it'll disappear from there.

Alternative test may inflate score gains
'Portfolio' exams spread in Va.
'How do you know we are closing the . . . gap?'

By Michael Alison Chandler
Washington Post Staff Writer
Thursday, November 19, 2009

Lynbrook Elementary School, which serves one of the poorest communities in Fairfax County, seems to be a model for reform. Three years ago, the Springfield school failed to meet state testing goals in English. Since then, it has charted double-digit gains in passing rates for every one of its closely monitored racial and ethnic groups of students.

But the success at Lynbrook and other schools throughout the state is not only due to better teaching. More and more, students who have struggled to pass Virginia's Standards of Learning exams are taking different tests.

The trend dates to 2007, when federal officials approved an alternative assessment after the Fairfax School Board threatened to defy a mandate to give multiple-choice reading tests to students who were destined to fail -- students who, like many at Lynbrook, were just beginning to learn English.

The Virginia Grade Level Alternative, like the multiple-choice test, assesses students' understanding of the state academic standards. Teachers document learning throughout the year in a binder of class work, including worksheets, quizzes and writing samples. Some special education students and non-native speakers in early stages of learning English are eligible for the portfolio, but final decisions are made by committees of educators and often parents.

Educators say the "portfolio" tests are valuable teaching tools and fairer and more meaningful than multiple-choice tests. With more time and flexibility, students have seen their passing rates soar.

Since 2007, Lynbrook's reading passing rate for students learning English shot from 52 to 94 percent. Among special education students, the rate went from 34 to 100 percent. At the same time, the number of portfolios increased from a handful to more than 100, including nearly half of the English learners and 78 percent of students with disabilities. All passed. The school had more than 460 students last year.

With more students taking the new test, many schools are showing sudden surges in performance. And some parents are concerned the portfolios are muddling scores the public relies on to see how racial and ethnic groups of students are performing and how they compare.

"How do you know we are closing the achievement gap, because thousands of our kids are not being tested the same way?" said Maria Allen, a Fairfax parent and longtime advocate for minority students.
Success at a cost

The remarkable gains at Lynbrook fit into a picture of ever-greater success in the region's largest school system. Fairfax Superintendent Jack D. Dale announced record highs in test scores and impressive progress in narrowing achievement gaps this fall. He attributed the progress to "a powerful shift" toward more personalized instruction systemwide.

Dale, who helped lead the fight to provide an alternative test for those beginning to learn English, said portfolios produce more accurate results that are consistent with how non-native speakers perform on multiple-choice tests once they master English. "We are seeing the same great improvement in our kids and our teachers no matter what instrument you look at," Dale said.

In an era of high-stakes testing, school leaders walk a tightrope. They must balance a lofty mandate to measure all students according to the same high expectations with a reality of classrooms filled with children who have trouble processing basic information or who recently arrived from another country. Every state makes some allowances for students who cannot meet testing requirements.

Maryland officials permit students who fail an exit exam required for graduation to do a project instead. District schools offer a "read aloud" accommodation for students with disabilities during reading tests, but began to dial back the program this spring after education officials found it was being overused. Most states offer alternative tests for students with serious cognitive disabilities.

Virginia's move to expand its use of portfolios to include students who are learning grade-level skills is unusual. It's costly. Fairfax spent more than $500,000 to train teachers and score portfolios last year, not to mention thousands of hours of teacher time compiling them. It's also risky. Experts say blending the results of different tests is very difficult. Closely watched trend lines and the accountability system's credibility are at stake.

"Schools or districts that are administering more of these alternative assessments may look better than those who are using fewer, and it may not have anything to do with the quality of the program," said Joan Herman, director of the National Center for Research on Evaluation, Standards and Student Testing at UCLA.

Virginia education officials say they have worked hard to make the tests comparable in rigor and scoring. A Virginia Commonwealth University study found that both tests are "well aligned" to the same academic standards, and the federal government has scrutinized and approved the alternative test.

But rollout has been uneven as the number of portfolios in Virginia has more than doubled to 47,000 in the past three years. Richmond, a district with about 23,000 students, administered nearly 3,800 portfolios last year; Loudoun, a district of 57,000, collected fewer than 1,000.

Fairfax, with 169,000 students last year, compiled 9,440 portfolios, up from 700 three years ago. The number represents about 2 percent of the total assessments given in Fairfax last year and about 6 percent of reading and math tests given in elementary and middle school. High school students are not eligible for the portfolio.
Students excel

Last year, students tested with portfolios outperformed classmates who took multiple-choice tests in Fairfax. Students with disabilities surpassed schoolwide pass rates in reading or math tests in more than a dozen schools. Students learning English were far more likely to score in the highest performance tier on the reading test, which measures knowledge of language arts concepts such as metaphor and plot, than their native-speaking peers. Overall, English-learners and students with disabilities charted 20- and 18-point gains respectively in reading pass rates, compared to a six-point gain for the division.

At Weyanoke Elementary School near Annandale, a third of students were tested with reading portfolios last year, up from none three years ago. Passing rates jumped from 41 to 100 percent for students with disabilities, from 69 to 97 percent for English learners, and from 66 to 91 percent for black students (more than a quarter of whom were tested with portfolios).

Principals at Weyanoke and Lynbrook say that the boost in scores has gone hand in hand with improvements in instruction and that portfolios help teachers focus on students' unique learning styles.

Weyanoke teacher Candy Kwiecinski is assembling about 10 portfolios for students in her fourth-grade class this year. One October afternoon, she taught a lesson on dictionary skills and how to use guide words at the top of the page. Some students might see a question on guide words next spring on a multiple-choice test. Others were tested that day.

A work sheet asking for examples of guide words could go in the portfolio. Or if it that proves too challenging, Kwiecinski can ask a student to explain what they are or whether they can select examples of guide words from an assortment of flashcards. Her job is to find the right way to teach and to test each student.

Last year, 100 percent of the portfolios at Weyanoke received passing scores. That does not mean the students who took them are the school's top performers, Kwiecinski said; it means they all learned the curriculum.

The portfolios show that her students "are learning the exact same things in different ways," she said.

Tuesday, November 17, 2009

Facts and International Tests

21st Century Skills, as promoted, are not effective without a command of some basic factual information. If you don't know any factual material, a few minutes on Google may or may not get you relevant facts. Then, of course, the student who has no knowledge cannot evaluate what facts are found, cannot compare the new knowledge to the old, cannot "stand on the shoulders of giants". An hour's worth of Powerpoint and some multimedia cribbed from Google images won't hide that lack, which is why I refer to it as Glitzenbullshit.

One commenter replied to an earlier post that
"There is a reason why the U.S. scores 35th in the world in math. It is because the rest of the world is embracing inquiry and prioritizing 'flexible' thinking...while we are still preparing to compete in the industrial age."

I'd really like to know how he arrived at this cause-effect relationship.

The US averages must naturally include many schools that have made these changes to a 21st Century Focus -- I would maintain that it is precisely because so many of our schools have changed to 21stCS that our averages have declined on tests like PISA and TIMSS. These tests don't measure 21stCS sets; they measure factual information and basic skills.

It is not, perhaps, indicative of a 'good education', but until you define THAT and determine what such should entail, you cannot use a fact-based international test to determine superiority or otherwise.

I would also point out that the whole concept of comparison conveniently ignores the fact (there's that dirty word again) that many nations to which we compare ourselves have much more homogenous culture and demographics, and most have a national curriculum and province-wide, if not national, control over schools. Almost all of the "great" successes have nationwide testing programs to finish off school. The final product is much more amenable to testing like PISA.

When I compare students and programs, talk with ex-students, visit colleges and speak with professors and admissions, converse with tradesmen and businessmen, I find one thing across the board. Success in college, life and careers is correlated more closely with schools that stressed facts and knowledge first and then blended in communication, information, literacy and critical thinking skills. Any school that tried to do it the other way invariably and ultimately did a disservice to its students.

Saturday, October 17, 2009

A Message for Educational Critics.

I have heard plenty about how bad our students are compared to the rest of the world, state, nation. Our test scores are stagnant. Kids don't know anything. Teachers are bums.

If anyone is interested, the released questions from the 11th grade test are here:
http://education.vermont.gov/new/pdfdoc/pgm_assessment/necap/released_items/07/grade_11_math.pdf

Take an hour and see how you would do. On certain questions, a calculator is not allowed. Look for the icon. Getting 50% would mean you are proficient.

Think about how difficult this set of questions is for you and consider whether you need to know this information or be able to use these skills to do your job.

Then you can rant all you like about how bad our students are and what lazy overpaid teachers we have here in this state.

Damned lazy teachers. Kids aren't passing tests.

As a response to some of the folks who feel that kids should be able to get every question correct ...

First, remember that the questions are vetted. If any question is answered correctly by more than 90%, it is removed from the test.

Not every kid will have taken the math courses necessary to understand every question. Some kids are not "math types". They might be artists or writers or skateboarders.

Most importantly, none of the schools is allowed to make the test count towards graduation or even report all scores on transcripts, so the kids have little incentive to succeed.

I'm not sure that the readers of this paper realize how difficult some of these questions are. How many adults can answer the following released question? What does that say about you? Why are we expecting the students to understand things that most of the adults in this population cannot?

Here's one of the "easy" questions (easy if you know it):

There is a line drawing of a pyramid with a dihedral angle of 52°. The length of each side of the square base is 230 meters. Which equation represents the height, h, of the pyramid?
A. h =115 sin52°
B. h =115 tan52°
C. h =115/sin52°
D. h =115/tan52°

Anyone? Anyone? Beuller?

This one is easy, too ...
A square with a side length of 8.0 cm is rolled up, without overlap, to form the lateral surface of a cylinder. What is the radius of the cylinder to the nearest tenth of a centimeter?

Anyone?