In a previous post, I showed that Reddit entries were slightly, but significantly, more likely to contain spelling errors and other unusual word choices.
I had not discounted the most common slang and swear words, which one could say contain proper meaning and are not misspelled even while not in the dictionary. In this post, I will account for them.
Dropping common slang and swear words ("ok", "lol", "bs", "gg", "awol", "4Chan", "nsfl", "nsfw", "ama", "s%#&t*", "f&%k*", etc. and ":)", ":(", ";)", ";(", ":/", ":\", etc., and their main capitalization variants) as well as links, the difference between Mercury Retrograde (MR) and Non Mercury Retrograde (NMR) in incidence of word spelling errors divided by entry word length increased to 2.14%.
Moreover, this difference of 2.14% is quite statistically significant with an ANOVA p-value for equality at 6.4x10^-14. The 95% mean confidence interval for NMR is ~{0.0528, 0.0531} and for MR it is ~{0.540 , 0.542}.
Here is the histogram showing the separation of MR (yellow) from NMR (blue) as well as the quartile-quartile plot comparing MR to NMR upon removing slang, emojis, links, and non-English entries.
Again, the difference could still be accounted for by other strange words, like unusual last names, that suddenly became fashionable in the MR period, but keep in mind that these results are from scanning evenly across 53 million entries. That is something like two thousand front page posts with full comments per day in both NMR and MR. This huge number is likely to dilute away any such fashionable blip with a similarly fashionable blip in NMR, although more data is always better. (Looking at all or at least more MRs and NMRs would be nice, but this is all the data I have right now.) [Edited: If you would like to see such a study across many Mercury retrograde seasons, see here.]
To give you a sense of perspective, the average number of words in an entry only increased by 0.4% during MR despite wide variation, and that increase is not statistically significant.
So, again I say, a >2% increase is huge, a statistically and culturally significant result.
Consider this: if you are just 2 percent more likely to make a mistake in some small thing during Mercury Retrograde, but you build up on hundreds (if not thousands) of those small things per day, that effectively implies that your rate of making some bigger mistake skyrockets.
For example, if you do 10 related small things in a MR day and each step is two percent more likely to cause an error than in NMR, then you are at least 21 percent more likely to commit a compounded error that day, and within the three weeks of MR, you are almost sure to make at least four such serious errors.
The Mathematica notebook is available for download.
Postscript: Re-running the file but also dropping "OMG" and "(͡ ͜ºʖ ͡º)" increased the difference between NMR and MR misspelling rates further to 2.83 percent. Mean 95% confidence intervals are {0.061656, 0.0619098} and {0.0632795, 0.0635948}.
I wrote the piece above in October 2015, a week after the one it builds on. I have left every word of it as it was. What follows is what I think of it now.
I have appended a long note to the earlier article setting out what is wrong with the design both pieces share: one boundary in one month, 1.6 million comments counted as though they were 1.6 million independent facts, and an effect too small to mean anything. All of it applies here and I will not repeat it. This piece has four faults of its own.
The compounding argument was wrong, and nearer to right than I deserved
This is the part of the article people remember, and it is the part that took me longest to think through properly. I wrote that if you do ten small things in a retrograde day and each is two percent more likely to go wrong, you are at least 21 percent more likely to commit a compounded error that day.
The 21 percent is 1.0210 = 1.219.
As a derivation that is wrong. The factor is only correct if a compounded error
means all ten steps failing at once, and at a five percent per-step error rate that event has a probability of about one in ten trillion, retrograde or not. Read the way any reader would read it — the chance of making at least one mistake across ten steps — the arithmetic goes the other way: 40.13 percent ordinarily against 40.75 percent in retrograde. An increase of 1.6 percent, not 21. With only ten steps there is no honest route from a two percent nudge to a twenty-one percent anything.
But the instinct underneath it was better than the arithmetic on top of it, and since the better argument is the one I failed to make in 2015, it should be here.
Do not take ten steps. Take a thousand, which is nearer to a real day in any case. Say fifty of them ordinarily go wrong, and ask not will I make a mistake today
— of course I will — but will today be a bad day?
Call a bad day one carrying fifty-seven errors or more, which happens about one day in six. Now nudge the per-step rate by my two percent. That bad day becomes twenty-three percent more likely. Draw the line at sixty errors instead and it is thirty percent more likely; at seventy, fifty-six percent. That is real. A small shift in a rate moves the tail of a distribution a great deal more than it moves the middle, and the tail is what people actually notice, because nobody remembers the average Tuesday. If Mercury retrograde did nudge the rate, then bad days genuinely would bunch up, and by something not far off the factor I claimed.
Three qualifications, and I need all three. First, that amplification arrives by a road I never travelled: it is the behavior of a binomial tail over a thousand trials and has nothing whatever to do with raising 1.02 to the tenth power. My figure and the defensible figure share their digits by coincidence, which is the least respectable way for a number to be nearly right. Second, the size of it depends entirely on where I draw the line at bad
— 1.19 at fifty-five errors, 1.23 at fifty-seven, 3.5 at a hundred and ten. Twenty-one percent is therefore not a fact about the world but a fact about a threshold, and one I would have to choose in advance and defend rather than discover afterwards at the value I liked. Third, what never amplifies, for any number of steps whatsoever, is the average: expected errors per day rise by exactly two percent and not a hair more. So the ordinary day gets two percent worse while the rare bad day gets disproportionately worse — and my sentence quietly needed both readings at once. Within the three weeks of MR, you are almost sure to make at least four such serious errors
requires the serious error to be common. The twenty-one percent requires it to be a conjunction that essentially never occurs. No single event can oblige both.
And every word of this is downstream of the two percent. When I finally put that number to fourteen years instead of one month, it came back p = 0.45. An amplified tail of a figure that cannot survive its own null is still a figure that cannot survive its own null. The compounding was the most interesting thing in this article, and I built it on the one quantity I had never checked.
My effect grew every time I deleted more words
Look at the sequence honestly. In the first article, leaving the slang in: 1.3 percent. Drop the slang and swear words and links: 2.14 percent. Then, in the postscript, drop OMG
and the lenny face too: 2.83 percent. Each deletion was chosen after I had seen what the previous one did, and each one pushed the number the way I wanted it to go. I presented that as the signal getting cleaner. It is at least as consistent with my having found a knob and turned it until the reading pleased me. There is no principled place to stop deleting words, which is exactly the problem: I could have kept going. What I ought to have done was fix the exclusion list before looking at a single result, and publish it whatever came out. To my credit I did publish the trail, so you can see the knob turning. That is how I know to distrust it now.
The figures are not pictures of 53 million entries
I leaned hard on 53 million entries
and on those two cleanly separated histograms. But count the bars: there are about thirty values in each group, not millions. The notebook confirms it — the intervals come from MeanCI applied to fractionsNMR and fractionsMR, which are lists of roughly thirty aggregates, not the per-post fractions. So the separation a reader sees is the separation of about thirty summary numbers, and its cleanness is a statement about how precisely I pinned down an average, not about how differently people spell. My own earlier article has the honest picture: plot the actual entries and the two windows lie almost perfectly on top of one another. Same data, same month, opposite impression. I showed the flattering one.
And a decimal point
I wrote that the 95% mean interval for NMR is ~{0.0528, 0.0531} and for MR ~{0.540, 0.542}. The MR figure is missing a zero; it should read ~{0.0540, 0.0542}. Taken at face value my own sentence claims retrograde raises misspellings by 922 percent, which would have been quite a finding. I have left the text as published rather than quietly repair it, since the error is part of the record.
What the answer turned out to be
The bracket I added at the time — if you would like to see such a study across many Mercury retrograde seasons, see here
— was pointing in the right direction, and I have now actually run the test on that longer series: 5,295 days, 2000 to 2014, 49 retrograde periods, with the retrograde flag slid to every possible position so I could see whether the true alignment stands out from arbitrary ones. It does not. p = 0.45. The working is at the foot of the earlier article. In that fourteen-year series the gap between Monday and Thursday is three times the gap between retrograde and direct.
Where that leaves me. I called two percent huge, a statistically and culturally significant result.
It is not huge; it is a third of the effect that the day of the week has, and it does not survive being asked whether it prefers Mercury to any other rhythm on the calendar. The care in this piece went into the wrong place. I was scrupulous about which words to throw out and careless about whether the comparison could mean anything at all, and the more scrupulous I was about the first, the more the second slipped away from me. I would rather leave this here, with its faults marked, than take it down. It is a decent record of how a person talks themselves into a result — and I was the person.