• Авторизация


Наука и результаты выборов 06-03-2012 17:34 к комментариям - к полной версии - понравилось!


Закон Бенфорда, теорема Вейля и фальсификация на выборах.
Скопировала без графиков и таблиц, позднее вставлю.

http://www.guardian.co.uk/news/datablog/2012/mar/0...a-putin-voter-fraud-statistics

http://kvant.mirror1.mccme.ru/pdf/1998/01/kv0198arnold.pdf


Russian election: does the data suggest Putin won through fraud?

Vladimir Putin's landslide vistory in the 2012 presidential elections has been marred by allegations of fraud. Does the data support his detractors?
• Get the data

*
o Tweet this
o Share
o reddit this
* Comments (…)

Vladimir Putin casts his vote at a polling station in Moscow on 4 March 2012.
Vladimir Putin casts his vote at a polling station in Moscow on 4 March 2012. Photograph: KeystoneUSA-ZUMA / Rex Features

To the surprise of virtually no-one, Vladimir Putin has won a landslide vistory in Russia's presidential elections. As also seemed inevitable, electoral observers and Putin's opponents alike have reported allegations of widespread electoral fraud, as the Guardian russian correspondent Miriam Elder reports:

Two women hover over a ballot box in the industrial Russian city of Cherepovets, stuffing in ballot after ballot. On the streets of Moscow, an independent election monitor armed with an iPhone trails a van full of "carousel" voters – people bussed from polling site to polling site in order to cast multiple votes for Vladimir Putin.

Three months after Moscow exploded in a storm of fury over allegedly widespread electoral fraud during the country's parliamentary vote, Russians went to the polls to vote against or, mostly, for Vladimir Putin in his quest to return to the presidency.

Putin quickly claimed victory, waiting until just over 20% of votes were counted, but his opponents just as quickly cried foul, armed with reels of evidence of alleged fraud. They uploaded them by the thousands to their Twitter accounts and LiveJournal blogs, helping the indignation go viral.

Measuring the scale of electoral fraud is always a challenge. Confirming some irregularities have taken place is (relatively) straightforward, but proving whether fraud is systemic or isolate is vastly more difficult.

In many aspects of life, however, there is a statistical trick that can find large-scale fraud, particularly in finance.

In short, it's that far more numbers start with a "1" than you'd think: in a normal ledger book without fraud, around a third of the figures (whether £17.20 or £1.16bn) would be expected to begin with a 1.

This pattern continues throughout the digits, and is known as Benford's law (a fuller explanation can be found here)

We've taken results from just over 2,150 polling stations submitted by Russia's election observers to a Russian-language site here. These aren't the verified official results, but give us a much bigger dataset to check than those would.

We then grouped together the results for all candidates. Did the results comply to Benford's law? The short answer is no, as shown in the graph below:

The results for Putin's vote share alone were even more striking, and it's worth noting the difference was statistically significant in both cases:

Does this mean the statistics show Russia's results were subject to fraud? Unfortunately, as my colleague Ben Goldacre – who helpfully did some of the number-crunching for this piece (though all errors are mine) – is fond of saying, it's a bit more complicated than that.

Each polling station within this set of data has a relatively tight range of votes: between three and around 2,600. This makes it unlike most data in, say, financial accounts, due to lack of variance, and can mean data doesn't comply to the Benford pattern even when totally legitimate.

Whether Benford's Law has potential for finding political fraud is a matter of contention. A study by Joseph Deckert, Mikhail Myagkov and Peter C. Ordeshook published in August 2011 concluded not:


It is not simply that the Law occasionally judges a fraudulent election fair or a fair election fraudulent. Its "success rate" either way is essentially equivalent to a toss of a coin, thereby rendering it problematical at best as a forensic tool and wholly misleading at worst.

However, this piece quickly received a response from a professor at the University of Michigan, Walter Mebane:

The paper mistakenly associates such a test with Benford's Law, considers a simulation exercise that has no apparent relevance for any actual election, applies the test to inappropriate levels of aggregation, and ignores existing analysis of recent elections in Russia ...

Whether the tests are useful for detecting fraud remains an open question, but approaching this question requires an approach more nuanced and tied to careful analysis of real election data than one sees in the discussed paper.

Data-driven investigations clearly offer intriguing potential for catching widespread fraud, whether financial or political. But it also seems the scope and potential of this method and others are far from a subject of consensus.

We've published the full data from the electoral observers below. If you've picked up on anything interesting within, or run more sophisticated tests,
let us know in the comments below, by email to james.ball@guardian.co.uk, or through twitter @jamesrbuk.

http://www.badscience.net/2012/03/is-there-statist...-in-the-russian-election-data/


Is there statistical evidence of fraud in the Russian election data?

March 5th, 2012 by Ben Goldacre in bad science, data, structured data | 6 Comments »

James Ball sent me the data for the Russian election vote counts this morning and asked me to test whether it deviates from Benford’s law, a test that can give a hint at whether numbers are the product of fraud. Posted below is my analysis, and also a check for last digit preference, which is another method for spotting sneakiness.

You might remember Benford’s Law from this post, it’s a good way to check if data has been faked, and was used on shady Greek economic data submitted to the EU:

www.badscience.net/2011/09/benfords-law-using-stat...entire-nation-for-naughtiness/

Imagine you have the data on, say, the population of every country in the world. Now, take only the “leading digit” from each number: the first number in the number, if you like. For the UK population, which was 61,838,154 in 2009, that leading digit would be “six”. Andorra’s was 85,168, so that’s “eight”. And so on.

If you take all those leading digits, from all the countries, then overall, you might naively expect to see the same number of ones, fours, nines, and so on. But in fact, for naturally occurring data, you get more ones than twos, more twos than threes, and so on, all the way down to nine. This is Benford’s law: the distribution of leading digits follows a logarithmic distribution, so you get a “one” most commonly, appearing as first digit around 30% of the time, and a nine as first digit only 5% of the time.

The next time you’re waiting for a bus, you can think about why this happens (bear in mind what leading digits do when quantities repeatedly double, perhaps) but reality agrees with this theory pretty neatly, and if you go to the website testingbenfordslaw.com you’ll see the proportions of each leading digit from lots of real-world datasets, graphed alongside what Benford’s law predicts they should be, with data from Twitter users’ follower counts to the number of books in different libraries across the US.

Benford’s law doesn’t work perfectly: it only works when you’re examining groups of numbers that span several orders of magnitude, for example, so for age, in years, of the graduate working population, which goes from around 20 to 70, it wouldn’t be much good, but for personal savings, from nothing to millions, it should work fine. And of course, Benford’s law works in other counting systems, so if three-fingered sloths ever develop numeracy, and count in base-6, or maybe base-12, the law would still hold.

I analysed for election fraud using the stats package Stata, beloved of economists and epidemiologists. James has posted graphs from the analysis I did on his blog, where you can also (I assume!) download the data he sent me:

www.guardian.co.uk/news/datablog/2012/mar/05/russia-putin-voter-fraud-statistics

The bottom line is this: the data massively do not conform to Benford’s Law, which would make you think that fraud is at work; but I think we shouldn’t expect the data to conform to Benford’s law, as the numbers don’t span several orders of magnitude. Essentially, this tool won’t work on these numbers. Anyway, here are the graphs and tables. Firstly, a graph showing the leading digit for all vote counts, in blue, against Benford’s distribution in red.

But here’s the distribution of the vote counts, not spanning multiple orders of magnitude:

Here are the tables with the Benford figures, in case you want them, and the second table reports a chi-squared test for whether the real figures deviate from Benford’s distribution (they do, but we could have guessed that from eyeballing the data!):

Lastly, I looked at last digit preference, which is a crude tool for looking at whether someone has made up figures. If they’re all made up by the same human, you might expect them to have some preference for specific digits, as humans are quite bad at generating random numbers. There’s no evidence of digit preference in the data. You might say that’s not surprising, as if there was fraud it would not have been centralised. In any case, I’m only really posting this as an illustration of how you can take data and play around with it.

Lastly, bear in mind that all these tests are only to check if people have made up numbers after the votes have been counted, and won’t detect people stuffing dodgy voting slips into ballot boxes.

For completeness, here’s the Stata code I used. Opensource dorks: in the future, I’m planning to do more public data analyses and detailed walk-throughs, for which I’ll use R, the open source statistics package, but because I did this in a hurry between work this morning I had to use Stata (I know it well so can code without thinking too hard).

Here’s the Stata code:


* housekeeping
clear all
cd /Users/bens/Documents/academia/projectscompleted/benford/
insheet using CandCount.csv, names

* eyeball
summ votecount, detail

* benford action
benford votecount
firstdigit votecount
firstdigit votecount, by(candidate)
digituse votecount

* digit preference
gen lastdigit = mod(votecount,10)
tab lastdigit

* nice pics
hist votecount, bin(100) freq
eqprhistogram votecount, bin(100)

* install packages if needed
ssc install benford
ssc install firstdigit
ssc install digits

вверх^ к полной версии понравилось! в evernote


Вы сейчас не можете прокомментировать это сообщение.

Дневник Наука и результаты выборов | Ивалон - Possunt qui posse videntur (Может тот, кто думает, что может) | Лента друзей Ивалон / Полная версия Добавить в друзья Страницы: раньше»