• Security incident: ISF was recently accessed by intruders. Please change your password, and change it anywhere else you used it. Read more

Just how much does google know about you?

The scope of Google's apparent ambition is breathtaking. I like it. They're not there yet, but I swear they've got the best chance of being the first real competitor Microsoft has ever had.
 
If this was addressed to me I think I did not make myself clear. The examples that you give about what you can say about an individual seem to make my point: of course we are predictable in our behaviour. But that says nothing at all about what we might do in odd circumstances.

For example it is absolutely true that if my shopping pattern changes something has changed. So what? There are loads of reason that will happen and there are loads of people at any given time it will happen to.

And yes, if you want to you can tell my politics from my posting. So what? They are not a secret. If they were I woud not be posting about them, would I?

You don't need to analyse my shopping records to find out who my friends are: we have CCTV for that. I object to CCTV but the vast majority of people in this country seem to actively welcome it. But again, I have school records and employment records and all sorts of things: It would not be hard to find out who I hang out with if you wanted to: and that is no different from any period in living memory.

What I do not see is any evidence that all this increase in information makes the use of that informaion practical: if anything it has the opposite effect. I am fairly sure that much more was known about people in rural and feudal societies than at any time since.
Of course it makes uses of the information practical. This is AUTOMATED. I do not need to analyze the data personally.

Now lets make it practical and real. I have the data for, say, a small city. 50,000 people. This is about the size of Harrisburg, PA, a medium sized capital city.

I know from exit polls the demographics of people likely to vote. I know from information gathering, the political opinion of most of the inhabitants of the city, 95% certainty. I have a political candidate I wish to win, who is behind by 7%, average voter turnout is about 30% in municipal elections.

I can absolutely locate the EXACT PEOPLE in the 15,000 voters necessary to swing the election, and carefully demographically target the opposition voters with various information-related measures to disrupt their lives so they can't vote. For instance, if people have been working a certain amount of overtime and are X% less likely to vote for each hour of overtime they've worked in the preceding two weeks leading up to the election (total information, remember? We do exit polls) and target the specific 5-7 employers necessary to create the necessary disruption to engineer a surprise victory. Given a million dollars, 2025 information levels, and a bit of unethical behavior, this is easy.

For a smaller town, say, 3,000 people, this is TRIVIAL. Given that small towns can still award multi-million dollar contracts, this is both trivial AND profitable.

I assume you can see why this is something to consider. Information, used properly, is power.
 
For instance, if people have been working a certain amount of overtime and are X% less likely to vote for each hour of overtime they've worked in the preceding two weeks leading up to the election (total information, remember?

But that value of X is just an average (even in 2025). One cannot calculate X for each voter. Furthermore the standard deviation on that distribution will be huge. Lastly given that the data is based on only a few data points for each person, the amount of noise in that statistical function will be so large as to make the calculation meaningless when trying to apply the average value of X to a single individual or even a small group of individuals. Lastly, trying to take action so that a specific person is required to work overtime is absurdly complicated. You have not convinced me.
 
GreyICE,



There need to be laws and treaties to regulate this.



Agreed, but either nobody seems to care, the law hasn't caught up with the technology, and ethical standards as a result are not being kept in line with scientific advancement.

Science without morality, as anybody with a brain knows, is extremely dangerous.
Hi INRM:

I suggest that you go to Wikipedia and the look at things like Internet Protocol Suite WP IP Address WP, DHCP WP, now the point of this is not to overwhelm you with information but to help you understand the amazing thing that is called the Internet.

First point, we need computers to make it work, it is such a goofy system it would not work otherwise at all. In a sort of analogy that is not related to the reality but just trying to give you an idea of how it works.

You log onto your computer, it sends out a message (unless you have a static address) that screams out to everyone on the local network Give Me a Name!, something in the network, a server or your Internet Provider say Here Is a Name, and seriously on a local network it does go slamming around until the server responds.

Now you have Your Name (DHCP gives you an IP address) then you type something into your web browser like www.randi.org, you know what is goofy? There is no www.randi. org there is an IP address that is called www.randi.org but really today it is 67.228.115.46. So your machine (actually a bunch of machines) uses a DNS (Domain Name Server) to find out What Does It Mean, so it goes off to the DNS and asks What Does It Mean [www.randi.org] and the DNS says 67.228.115.46.

Now you machine sends out a request to find 67.228.115.46 and guess what it is not like a letter getting delivered, it is not like a car driving. It is more like the system says Are You 67.228.115.46 and if the answer is no then it just throws your request to somebody else, and so it goes bouncing around with no particular notion of where it is heading but getting closer until it does.

Now today my request went through 13 steps to get to the forum, I went through Indianapolis, Chicago, Denver and Seattle, some days as the trunks get busier it will take an even stranger route.


And what is even crazier is that the message often get broken down, take different routes and arrive in no particular order, so when your machine says 'waiting for http://www.randi.org' sometimes it is waiting for packages to arrive, it has some already but not the one that says what to do with them, then it assembles the packages and you see the screen of material on your computer.

But it is so crazy, it all happens in no particular order, it just sort of whizzes around until it gets there and things just bounce around in cyber space. And each stop on the trip your package has this label, where it is coming from and where it wants to go, to get to www.randi today my message to 13 steps and eleven of those are intermediate servers. So those are places that my message went, sat on a server and was then passed on.

That does not even look at how the actual hosting of web pages works and how something like a search engine works.

But here is the thing, when you use the Internet, a lot of data is labeled and passed around, and it is all data, Google is a service, they track what people do, first off that is what makes Google so strong, their searches are sorted by the data they gather and the order on the page reflects the aggregate usage. Second they get your IP address and will track what that specific IP address it doing. Now if you use Google Desktop they will get a whole lot more information about you.

But here is the thing, because of DHCP you basically gets a new IP address each time you log on to your machine, so it is hard for them to know exactly who you are. Which is why tracking cookies and spy ware and important to keep off your machine, because they do a much better job.

So user beware!

Every message you send in cyber space is stored somewhere and many places at the same time. Over and over and over. Even on your home machine files that you delete can be recovered.

So yes the laws should try to keep up with the technology but to some extent that is never going to be possible.


BTW if you want to see the steps your message takes and you have Windows
Go to Start, select RUN
type cmd, this will open the DOS box for cmd
type tracert www.randi.org and press enter.
 
Just remember as far as Google's CEO is concerned if you have concerns about your on-line privacy it's because you are up to no good!

"If you have something that you don't want anyone to know, maybe you shouldn't be doing it in the first place, if you really need that kind of privacy, the reality is search engines, including Google, do retain that information." - Schimdt
 
Just remember as far as Google's CEO is concerned if you have concerns about your on-line privacy it's because you are up to no good!

"If you have something that you don't want anyone to know, maybe you shouldn't be doing it in the first place, if you really need that kind of privacy, the reality is search engines, including Google, do retain that information." - Schimdt
He's basically right, except society hasn't caught up to the point we need to be where google search info won't matter if it's private.

the thing about embarrassing google search information is that it only holds power over you if society maintains a puritanical double standard. For as long as we pretend that porn is for the other guy or that my interest in categorizing poop smells is shameful, we will have a need for private info.
 
I suggest that you go to Wikipedia and the look at things like Internet Protocol Suite WP IP Address WP, DHCP WP, now the point of this is not to overwhelm you with information but to help you understand the amazing thing that is called the Internet.
Hi Dancing David,

That's a good overview of IP. I want to comment on a couple of points.

(much stuff snipped)

But here is the thing, when you use the Internet, a lot of data is labeled and passed around, and it is all data, Google is a service, they track what people do, first off that is what makes Google so strong, their searches are sorted by the data they gather and the order on the page reflects the aggregate usage. Second they get your IP address and will track what that specific IP address it doing. Now if you use Google Desktop they will get a whole lot more information about you.
One word: cookies. Google tracks you using a cookie they sent to your browser. Every site that subscribes to Google Analytics (and there are a LOT of them) causes your browser to make a connection to Google so that Analytics can update their information. And the unique cookie that Google assigned to your browser, which is permanent unless you clear it, gets sent with that request. The cookie is independent of your current IP address, and it remains across internet connects, browser restarts, and system reboots. (But if you use two different browsers, they will have two different cookies.)

But here is the thing, because of DHCP you basically gets a new IP address each time you log on to your machine, so it is hard for them to know exactly who you are. Which is why tracking cookies and spy ware and important to keep off your machine, because they do a much better job.
That's true of ADSL. I'm using a cable provider, and I have had the same IP address for over a year.

(stuff snipped)

type tracert www.randi.org and press enter.
I'm on Linux, and prefer using mtr :p
 
Last edited:
I have cookies denied for Google. The only downside is that I have to re-establish my Search Settings every time I use it and that's no big deal since I keep it running all the time.

The other thing I've thought of is just flooding Google with irrelevant searches. I partially do this by looking up spellings for words that Mozilla has problems with. ;)
 
He's basically right, except society hasn't caught up to the point we need to be where google search info won't matter if it's private.

the thing about embarrassing google search information is that it only holds power over you if society maintains a puritanical double standard. For as long as we pretend that porn is for the other guy or that my interest in categorizing poop smells is shameful, we will have a need for private info.

Or you are trying to plan a surprise get together for someone one, or looking for advice on how to save your marriage or...

There are many, many things that many people would wish to be private that do not fall into the either the social taboo nor illegal camp.
 
Last edited:
Hi Dancing David,

That's a good overview of IP. I want to comment on a couple of points.


One word: cookies. Google tracks you using a cookie they sent to your browser. Every site that subscribes to Google Analytics (and there are a LOT of them) causes your browser to make a connection to Google so that Analytics can update their information. And the unique cookie that Google assigned to your browser, which is permanent unless you clear it, gets sent with that request. The cookie is independent of your current IP address, and it remains across internet connects, browser restarts, and system reboots. (But if you use two different browsers, they will have two different cookies.)


That's true of ADSL. I'm using a cable provider, and I have had the same IP address for over a year.


I'm on Linux, and prefer using mtr :p

Cool, this is something that has been dsicussed here such as if you go to Stormfront they can read out your vBulletin cookie for the JREF.

Maybe I do have the same IP each time, I have digital cable, but I can repair the connection, I will have to check.
 
Last edited:
But that value of X is just an average (even in 2025). One cannot calculate X for each voter. Furthermore the standard deviation on that distribution will be huge. Lastly given that the data is based on only a few data points for each person, the amount of noise in that statistical function will be so large as to make the calculation meaningless when trying to apply the average value of X to a single individual or even a small group of individuals. Lastly, trying to take action so that a specific person is required to work overtime is absurdly complicated. You have not convinced me.
Really? Why can one not calculate X for each voter? One merely needs sufficient amounts of information, and sufficient processor time. Neither is a scarce resource. The standard deviation is hardly as huge as you think - people are reasonably predictable, and we have a good thousand of them or so to play with. That's more than enough for standard dev to be minimal.

Finally, your entire assumption is based around limitations - you assume that the action is limited to one single thing. You assume that the number of data points is limited. You assume the data gathered is too poor to have a good reliability. You assume that no one has the processor time to compute this.

These are limitations. I see no reason why anyone would operate under them. This becomes relevant if and when those limitations are exceeded - and for all, it is simply a matter of time (with time, data improves, processors grow faster, and reliability increases).
 
Or you are trying to plan a surprise get together for someone one, or looking for advice on how to save your marriage or...

There are many, many things that many people would wish to be private that do not fall into the either the social taboo nor illegal camp.

And there are many things that society would prefer to remain private, because of negative consequences.

See health insurance discrimination on the basis of genetic codes, for instance.
 
Really? Why can one not calculate X for each voter? One merely needs sufficient amounts of information, and sufficient processor time. Neither is a scarce resource. The standard deviation is hardly as huge as you think - people are reasonably predictable, and we have a good thousand of them or so to play with. That's more than enough for standard dev to be minimal.

Yes, all one needs is sufficient amounts of information. The catch is that people vote once every four years. Assume your target is 33 years old. He has had the chance to vote in in four elections. Say he voted in three of them and missed one of them. That's already public record and easily obtainable. The reason he didn't vote in that one election is not deducible no matter how much electronic information you have. Maybe he had a sore throat that day, maybe he simply forgot, maybe he drove there and the line was too long, maybe he got a flat tire, maybe inclement weather kept him away, maybe he saw the difference in polls going into election day and decided his vote wasn't needed, maybe it really was the amount of overtime he worked in the past seven days (although, I am unsure how unlimited internet tracking will allow a third party to determine how much overtime was worked in the past week). I'll agree that if you tracked his every movement on the internet, you could infer with a certain reliability who he would vote for, but determining if he is going to vote in an upcoming election requires more than four data points spread out over 16 years.



Finally, your entire assumption is based around limitations - you assume that the action is limited to one single thing.

No, I don't. You were the one who asserted that the amount of overtime worked was the deciding factor.

You assume that the number of data points is limited.

The fact that the number of data points related to election behavior is limited is beyond question.

You assume the data gathered is too poor to have a good reliability.

Yes, I do.

You assume that no one has the processor time to compute this.

No, I do not. I will agree that processor time is not the limiting factor.
 
Yes, all one needs is sufficient amounts of information. The catch is that people vote once every four years. Assume your target is 33 years old. He has had the chance to vote in in four elections. Say he voted in three of them and missed one of them. That's already public record and easily obtainable. The reason he didn't vote in that one election is not deducible no matter how much electronic information you have. Maybe he had a sore throat that day, maybe he simply forgot, maybe he drove there and the line was too long, maybe he got a flat tire, maybe inclement weather kept him away, maybe he saw the difference in polls going into election day and decided his vote wasn't needed, maybe it really was the amount of overtime he worked in the past seven days (although, I am unsure how unlimited internet tracking will allow a third party to determine how much overtime was worked in the past week). I'll agree that if you tracked his every movement on the internet, you could infer with a certain reliability who he would vote for, but determining if he is going to vote in an upcoming election requires more than four data points spread out over 16 years.
You are going about this the wrong way. You are insisting that we must act like it's the 20th century, as if computers had not been invented yet.

This is nonsense. Last presidential election, we got exit polls. Assuming a nice spread of data points, we most likely got anywhere from 50,000 to 200,000 data points. Improved surveying methods might drive this up.

We in fact have very good data about how a 26 year old male who voted democratic, worked 55 hours a week on average, was college educated with a liberal arts major, and regularly followed the news voted. We also have demographics of the average 26 year old male, the number of hours a week they worked, their education level, etc.

In fact, we can isolate many factors. The number of recent positive versus negative articles they remember about their candidate. The effect increased work hours has on that demographic. The effect that different issues had.

You act like this is impossible. It is already occurring. Our methods simply, in terms of sophistication, are primitive. You are staring at a sharp rock, thinking about metalworked blades, and determine that no one would dig up a bunch of rock, melt it, purify out different metals, and forge it just to get something sharper than that sharp rock (it would take years, if not decades for someone to make themselves this new sharper rock (call it a knife).

Of course you are correct. No one would go through all that trouble for one knife. But the fact that you are correct does not mean that you have entirely thought the matter through.

No, I don't. You were the one who asserted that the amount of overtime worked was the deciding factor.

The fact that the number of data points related to election behavior is limited is beyond question.

No, I do not. I will agree that processor time is not the limiting factor.
Do you see what is happening here? Overtime is not the ONE DECIDING factor. On this we both agree. But it is a factor. On this we most likely probably agree (if your argument is that people decide to vote by rolling dice, I suppose we'll have to agree to disagree - but I would be agreeing to be right, and you would not be).

You further agree that overtime's effect on voting can probably be analyzed. Finally you note that processor time is not a limiting factor.

But we are at a quandary here. If election behaviors are based on a number of factors, then the factors can be analyzed. All we need is good data. You have agreed to this. You have limited the number of data points to one per year (rather than one per person per election year, which seems a tad more likely). Then you have concluded the data is poor because we only get one point per year rather, than, say, 250 million or so.

It's somewhere in-between the two. Each year it gets closer to the latter than the former. So the quality and granularity of the data is improving.

Do we agree or disagree on this basic premise?
 
You are going about this the wrong way. You are insisting that we must act like it's the 20th century, as if computers had not been invented yet.

This is nonsense. Last presidential election, we got exit polls. Assuming a nice spread of data points, we most likely got anywhere from 50,000 to 200,000 data points. Improved surveying methods might drive this up.

We in fact have very good data about how a 26 year old male who voted democratic, worked 55 hours a week on average, was college educated with a liberal arts major, and regularly followed the news voted. We also have demographics of the average 26 year old male, the number of hours a week they worked, their education level, etc.

In fact, we can isolate many factors. The number of recent positive versus negative articles they remember about their candidate. The effect increased work hours has on that demographic. The effect that different issues had.

You act like this is impossible. It is already occurring. Our methods simply, in terms of sophistication, are primitive. You are staring at a sharp rock, thinking about metalworked blades, and determine that no one would dig up a bunch of rock, melt it, purify out different metals, and forge it just to get something sharper than that sharp rock (it would take years, if not decades for someone to make themselves this new sharper rock (call it a knife).

Of course you are correct. No one would go through all that trouble for one knife. But the fact that you are correct does not mean that you have entirely thought the matter through.

Do you see what is happening here? Overtime is not the ONE DECIDING factor. On this we both agree. But it is a factor. On this we most likely probably agree (if your argument is that people decide to vote by rolling dice, I suppose we'll have to agree to disagree - but I would be agreeing to be right, and you would not be).

You further agree that overtime's effect on voting can probably be analyzed. Finally you note that processor time is not a limiting factor.

But we are at a quandary here. If election behaviors are based on a number of factors, then the factors can be analyzed. All we need is good data. You have agreed to this. You have limited the number of data points to one per year (rather than one per person per election year, which seems a tad more likely).

I am not sure how you arrived at that statistic. I said a 33-year-old who voted in three of the last four elections produces four data points spread out over 16 years.

Then you have concluded the data is poor because we only get one point per year rather, than, say, 250 million or so.

It's somewhere in-between the two. Each year it gets closer to the latter than the former. So the quality and granularity of the data is improving.

Do we agree or disagree on this basic premise?

I disagree. You may get to the point of saying exactly 351 of the 500 33-year-old voters in this city will vote in the upcoming election. You could even say that number drops to 298 if it rains that day.1 But that is a very far cry from saying I know that this 33-year-old will vote and that 33-year-old will not vote; that distinction is necessary for your scenario. The candidate must know how to prevent a specific voter from turning up.2
I gave four virtually untrackable reasons why a person might not have voted in the last election. I can give a dozen more. That's where the variability comes in. I don't care if you can track a thousand variables related to why a person votes or not. There are still a hundred untrackable reasons why a person does or does not show up on election day. In large numbers these factors tend to cancel each other out, but we are not talking about large numbers, we are talking about one specific voter.



.. . . . . . . . . .

1) I'll agree that defining this number will become more and more accurate in the future.
2) Isn't there an issue here related to what happens if it comes out that a candidate spent money trying to prevent a specific voter from turning out? More accurately, a candidate who did not use phone calls, mailings, or visits from volunteers to change this person's mind, but rather made an attempt to make that person work extra hours on election day in order to prevent that voter from going to the polls.
 
Last edited:
But why? It's rather useless to make sure that a specific voter does not turn up. It's rather a tad more useful to make sure that five hundred, or one thousand voters do not turn up. If an industry employs, say, one thousand workers, and you can calculate they are 70% likely to vote and 80% likely to vote for your opponent, you can use that for informed decision making.

If one can identify 3,000 voters who are likely to turn up and make them 3,000 voters who are unlikely to turn up, that's a very big margin of change. And there is beginning to be enough information out there that a computer can automatedly generate the highest probability methods of making sure those people do not vote - tailored for each individual.

Your argument that one must treat each individual differently is based on a concept that somehow each person is unique and unpredictable. They are not. You are not a unique and beautiful snowflake.
 
But why? It's rather useless to make sure that a specific voter does not turn up. It's rather a tad more useful to make sure that five hundred, or one thousand voters do not turn up. If an industry employs, say, one thousand workers, and you can calculate they are 70% likely to vote and 80% likely to vote for your opponent, you can use that for informed decision making.

If one can identify 3,000 voters who are likely to turn up and make them 3,000 voters who are unlikely to turn up, that's a very big margin of change. And there is beginning to be enough information out there that a computer can automatedly generate the highest probability methods of making sure those people do not vote - tailored for each individual.

But if you are tailoring it to the individual, then you must understand what motivates and demotivates each one of those 3000 voters.

Give me an example of what actionable information this computer can generate on these 3000 individuals. What exactly is the candidate going to do to change these people from the likely-to-vote category into the unlikely-to-vote category?
 
Last edited:
I wonder if we are talking at cross purposes here.

At first I thought the concern was information about the individual and I do not think that kind of thing is possible in principle for reasons like Ladewig's, and the ones I mentioned earlier

If the concern is gross manipulation on the basis of statistics, then I think that is possible: but it has always been possible without the statistics and, as an example, we have seen voter manipulation before. I see no reason to suppose some people will not wish to do it again: and I see no reason why most will not oppose it. Nor do I see any reason that that opposition will not be effective: starting a rumour that overtime is offered to prevent people voting will render it useless all by itself
 

ISF - Join now!

Every member here is approved by hand. No bots, no spam, just people who care about evidence and honest debate.

Membership is free!

Create your free account

Back
Top Bottom