Ignore Effect Size at Your Peril

This article in The Economist entitled “Removing Phones from Classrooms Improves Academic Performance” (https://papers.ssrn.com/sol3/papers.cfm?abstract_id=5370727) has been cited by Chris Ferguson as an example of how ignoring effect sizes can lead to false conclusions. Here is the abstract of the paper, followed by Ferguson’s critique:

    Widespread smartphone bans are being implemented in classrooms worldwide, yet their causal effects on student outcomes remain unclear. In a randomized controlled trial involving nearly 17,000 students, we find that mandatory in-class phone collection led to higher grades — particularly among lower-performing, first-year, and non-STEM students — with an average increase of 0.086 standard deviations. Importantly, students exposed to the ban were substantially more supportive of phone-use restrictions, perceiving greater benefits from these policies and displaying reduced preferences for unrestricted access. This enhanced student receptivity to restrictive digital policies may create a self-reinforcing cycle, where positive firsthand experiences strengthen support for continued implementation. Despite a mild rise in reported fear of missing out, there were no significant changes in overall student well-being, academic motivation, digital usage, or experiences of online harassment. Random classroom spot checks revealed fewer instances of student chatter and disruptive behaviors, along with reduced phone usage and increased engagement among teachers in phone-ban classrooms, suggesting a classroom environment more conducive to learning. Spot checks also revealed that students appear more distracted, possibly due to withdrawal from habitual phone checking, yet, students did not report being more distracted. These results suggest that in-class phone bans represent a low-cost, effective policy to modestly improve academic outcomes, especially for vulnerable student groups, while enhancing student receptivity to digital policy interventions.

    Ferguson says: This is a great example of why study authors relying on p-values and ignoring the likelihood of weak effect sizes being statistical noise can do great harm. The study looked at 17000 students in a study of cellphone bans in schools…still found null effects for most outcomes, but focus on one positive outcome with a d = .086 (about r = .043). Aside from the likelihood the one finding (against the other null findings) is due to chance, this effect size is well within the range of those indistinguishable from statistical noise (Ferguson & Heene, 2021). It should never have been interpreted as hypothesis supportive. Yet the Economist (falsely) portrays it as a “ringing endorsement.” It is, in fact, far better evidence against the use of cellphone bans than for it.

    A Creative Research Design for Studying the Origins of Hunger

    The following example can be used simply to illustrate creative research design but can also be useful in a section or course on motivation to illustrate the role of memory in influencing people’s decisions about when to start and stop eating. Suppose you have just finished a big lunch at your favorite restaurant when a server gets mixed up and brings you another plate of that same lunch that was meant for someone else. You would almost certainly send it away, but why? Decisions to start eating or stop eating are affected by many biological factors, including signals from your blood that tell your brain how much “fuel” you have available, but Paul Rozin and his colleagues were interested in how these decisions are affected by psychological factors, such as being aware that you have already eaten. What if you didn’t remember that you had just had lunch? Would you have started eating that second plate of food?

      To explore this question, Rozin conducted a series of tests with R.H. and B.R., two men who had anterograde amnesia, a kind of brain damage that left them unable to remember anything for more than a few minutes. The men were tested individually, on three different days, in a private room where they sat with a researcher at lunchtime and were served a tray of their favorite food. Before and after eating, they were asked to rate their hunger on a scale from 1 (extremely full) to 9 (extremely hungry). Once lunch was over, the tray was removed, and the researcher continued chatting, making sure that each man drank enough water to clear his mouth of food residue. After ten to thirty minutes, a hospital attendant reentered with an identical meal tray and announced “Here’s lunch.” These men had no memory of having eaten lunch already, but would signals from their stomachs or their blood be enough to keep them from eating another one?

      Apparently not. In every test session, R.H. and B.R. ate all or part of the second meal and, in all but one session, ate at least part of a third lunch that was offered to them ten to thirty minutes after the second one. Rozin conducted similar tests with J.C. and T.A., a woman and a man who had also suffered brain damage but who still had normal memory for recent events. In each of two test sessions, these people finished their lunch but refused the opportunity to eat a second one. These results suggest that the memory of when we last ate can indeed be a factor in guiding decisions about when to eat again. They also support the conclusion that eating is controlled by a complex combination of biological, social, cultural, and psychological factors. As a result, we may eat when we think it is time to eat, regardless of what our bodies tell us about our physical need to eat.

      [Reference: Rozin, P., Dow, S., Moscovitch, M., & Rajaram, S. (1998). The role of memory for recent eating experiences in onset and cessation of meals. Evidence from the amnesic syndrome. Psychological Science, 9, 392–396.]

      An Example of Limits on Test Validity

      An article found here provides a good example of how a test can be valid for one purpose but invalid for another. In this case, the test is the Wonderlic Contemporary Cognitive Ability Test and it has been used for years by the NFL to help teams decide which players to draft from the college ranks. The problem is that cognitive ability is not as predictive of athletic performance as it is of academic performance. As a case in point, quarterback Johnny Manziel achieved a high score on the test and had an undistinguished two-year NFL career whereas repeat SuperBowl champion quarterback Tom Brady scored quite low but played for more than 20 years with two different teams and will be inducted into the NFL Hall of Fame.

      Hearing Loss and Dementia

      A recent JAMA study, found here describes a strong correlation between hearing loss and dementia in older adults. A commentator interpreted this correlation to mean that people should do all they can to mitigate hearing loss as soon as it appears so as to reduce the risk of developing dementia. You may want to use this study, and particularly the commentator’s assertion as an example of how easy it is to over-interpret the meaning of a correlation.

      Nightmares and Dementia

      An article found here describes a study of middle-aged and older adults showing that men who have nightmares at least once a week are significantly more likely than other men to develop dementia. This correlation did not reach statistical significance for women. Speculating on the basis for this association, the researchers suggest that when dementia is in its undetectable early stages, it creates negative emotions that may be expressed as depression while awake and as nightmares while asleep. They hypothesize that medication that can reduce nightmares might also help to slow cognitive decline. Is the association between nightmares and development of dementia causal or correlational? If the latter, what third factor(s) might be responsible for it and why did it only appear in men? The study on which this article is based provides a good example of research methods whose interpretation depends on critical thinking about correlation and causation. It should provide a useful target for classroom discussion as well as a topic for student writing assignments aimed at promoting critical thinking.

      Spotting Examples of Pseudoscience

      This chapter emphasizes the importance of helping students to think critically, including by alerting them to these ten warning signs of pseudoscientific assertions:

      Lack of falsifiability and overuse of ad hoc hypotheses

      Lack of self-correction

      Emphasis on confirmation

      Evasion of peer review

      Overreliance on testimonial and anecdotal evidence

      Absence of connectivity

      Extraordinary claims

      Ad antequitem fallacy (an appeal to tradition as an argument for validity)

      Use of hypertechnical language

      Absence of boundary conditions

      Source: Lilienfeld, S. O., Ammirati, R., & David, M. (2012). Distinguishing science from pseudoscience in school psychology: Science and scientific thinking as safeguards against human error. Journal of School Psychology, 50(1), 7-36.