Sunday, July 6, 2008

Facebook Lexicon


I discovered the Facebook Lexicon very recently. It basically counts up words and
phrases over the walls (like orkut's scrap book) of all the facebook users. It then plots the frequency of these words over time on a graph.
I found the following hilarious, "party tonight, hangover" which compares the
frequencies of both phrases.


Notice the near perfect phase shift ;)

Sunday, April 27, 2008

Social Networks, the death knell for Stand-alone IM?



I always suspected that with the advent of Social Networks like Facebook, orkut, Myspace etc. users were using Instant Messaging (IM) less and less. I recall in my high school years instant messaging was the in-thing. Especially for the 13-25 demographic instant messaging was extremely popular. Now it seems that Social
Networks have taken over which seems like an interesting throw-back to the asynchronous nature of communication like email.


Now, with Facebook's Chat Application and Google integrating Gtalk gently into Orkut, I suspect companies whose core business is IM, like Meebo will feel the pinch. Instant Messaging is still useful but might now be just a feature on a Social Network. Maybe not all demographics but the 13-30 year old web user (which is the demographic on a typical social network) will definitely not be too keen on maintaining a different friends list for Meebo (or other Web based IM clients) and one for Facebook. Sure, Meebo might develop a chat client geared for facebook like social.im but I would think that it will be hard to compete with Facebook on Facebook's turf. And a tug of war with Facebook for user visits seems like a scary prospect.


I am sure Meebo, Social.im and others have thought long and hard about this. I am curious to see how these companies evolve and adapt to the changing landscape.I had an opportunity to talk with Seth Sternberg (CEO, Meebo) whom I had invited to talk at the Stanford I don't know to CEO Business Conference. I learnt that Meebo does other things as well, like Meebo rooms which plug in nicely to Myspace, so they definitely have other tricks up their sleeve and I am sure we might see more in the months to come.

As an aside, on the topic of time spent by users online on Social Networks, I heard Max Levchin (Founder and CEO, Slide & PayPal) remark that most of this time comes from the bucket that users spend on email and other entertainment on TV. I would add stand-alone IM clients to that list.

Sunday, April 6, 2008

Beauty & Intelligence


Here is a fun probability result which has to do with conditional probability. Its not very hard to see what's going on but its a fun result nonetheless. I just made up a dummy example to make it more fun :)

Lets take 2 fairly independent attributes like beauty and intelligence (It can be argued that they are not independent by appealing to genetics and preferential/unequal selection rights for species perpetuation but lets ignore that for now and think simple). Lets assume beauty and intelligence are independent.
i.e if
B=beautiful
I=intelligence

P(I)=P(I|B) ->(1)

Now, think of all the people who stick in your memory (as opposed to you forgetting them after a few days of meeting them). Lets assume (simplistically)
that the people who stick in your memory are ones who are either intelligent or beautiful.

Lets throw in some numbers.
Lets assume the prior probability of someone being beautiful is
0.4
and for someone being intelligent is 0.1
Lets assume 10% of the beautiful people are intelligent.

Say, you've met 200 individuals in your lifetime.

If we go with the assumption of you being able to recollect only people who
are intelligent or beautiful, you will remember
80+12=92 individuals

Now when you look at this sample set, lo and behold, it looks like intelligence
and beauty are negatively correlated!

This looks counterintuitive since it looks like,
P(I)=0.21
but P(I|B)=0.1
(thus violating the independence assumption in (1))

whereas in reality, the conditional probability eqns are,
P(I|B,P) < P(I|P)
where P=I U B
which is another case of selection bias at work..

Hmmm..I wonder how many people feel this way about beauty and intelligence ;)

Monday, February 18, 2008

Infoaxe - Stealth Search Engine out of Stanford


I have some BIG news in this post. I have just started my company, Infoaxe along with my long time friend and classmate from Stanford, Vijay Krishnan.

Infoaxe is the next generation search engine searching a very different kind of Web.
We are in stealth mode currently and hence my lips are sealed. But stay tuned for updates. Infoaxe is short for 'Information Access' with some liberties taken :)
It could also be a reference to the Stanford Axe ;).

We developed the Infoaxe Search Engine while we were graduate students at Stanford.
We are very excited to have Prof. Hector Garcia Molina on our technical advisory board.

We are really excited about Infoaxe which has at its core many innovations in applications of data mining and machine learning.

Thursday, December 27, 2007

Roll with the punches..tomorrow is another day..


The title of today's post has only a vague connection with the substance of today's post. Today's post is about a betting strategy for roulette. This strategy has an intrinsic flaw and I am going to leave it to my readers to ponder as to what the flaw is. I'll explain the flaw in a later post.
In roulette, the casino houses typically offer odds such that the expected return is negative (which is how they stay in business). This can be shown as a direct consequence of the law of large numbers.
The betting strategy is remarkably simple. If you lose a round, you double your bet. This is also called 'doubling up' or the Martingale strategy. The intuition is that it is very unlikely that you will have a string of losses and you hope to cover your losses when you eventually win. You have to win sometime right? And at that time you would have recovered all your losses.
Do you see the problem with this strategy?

Wednesday, December 19, 2007

The Monty Hall Puzzle

Here's a fun probability puzzle. This is called the Monty Hall Puzzle since its based off of an American Game Show, Let's make a deal.

Say you're in a game show where the host invites you to open one of 3 doors. One of the doors has a prize(a car) and the other 2 doors have goats. The host knows which door has the prize. After you pick a door, the host then opens one of the remaining 2doors which do not have the prize and reveals the goat. Now, he offers you the option of switching your choice from your current pick to the other remaining door.

The question is, what should you do?
A. Doesn't matter if you switch, the probability of winning is still the same
B. Always switch

Tuesday, December 18, 2007

The Challenge of Duplicates on the Web


A few weeks back, Chad Walters, the Search Architect at Powerset told me an interesting anecdote from the time he headed the Runtime efforts at Yahoo!.
There is a fundamental problem in estimating the number of hits for search queries because of duplicates in Web Search results. Near Duplicates & Duplicates occur because of many reasons like mirroring, RSS feeds, track backs etc. These add a fair amount of noise to the estimates of Web Search Engines about the size of their indexes which is always a matter of some pride to the big guys (Google, Yahoo! & Microsoft) and rightly so.
Anyway, a french researcher, Jean Véroni did some analysis on the number of hits reported by search engines and writes an interesting article here.
Here are some charts from his blog (http://aixtal.blogspot.com/2006/07/search-crazy-duplicates-1.html),


Google with similar pages





Google without similar pages






Yahoo! with similar pages






Yahoo! without similar pages






The fluctuation in the estimates is apparent when similar pages are included.

To quote Jean Véronis from his blog,

It’s interesting to note that:
* once duplicates are removed, Google and Yahoo’s figures are about the same;
* Yahoo’s curves are much more stable than Google’s.


I found this interesting in the context of the battle for supremacy in Web Search.
Maybe Yahoo! does a better job than its given credit (or market share) for :)

I did some work on near duplicate detection on the Web as part of my graduate research while at Stanford with Andreas Paepcke of the Stanford InfoLab and that's one reason I found this interesting. My work can be found here.
(Thanks to Chad for this interesting tidbit)