الاثنين، 27 أبريل 2009

The Continuing Metamorphosis of the Web



I just returned from giving a talk at the 18th World Wide Web Conference in Madrid and was pleased to see a healthy and dynamic conference despite difficult economic conditions. Madrid had beautiful spring weather, and a magnificent modern architecture abounds throughout the city. I will say, though, that the Madrid subway does not vibrate (shake, rattle, and roll) one’s soul quite as much as does our local NYC subway.

My talk was entitled The Continuing Metamorphosis of the Web. In it, I noted that the initial web standards were so simple and sensible that they engendered a path of stepwise innovations, which taken together have aggregated into amazing accomplishments. Metaphorically, I feel our community has been on a kind of pseudo-random walk that has taken us to remarkable places. The truly great results have included the creation of a virtual Library of Alexandria, the creation of the search engine (to be that library’s super-card catalog), the empowerment of the long tail (in diverse communities), and great innovations to doing business. I argued that the bottom up evolution is continuing (perhaps even accelerating) today, and that the current stepwise improvements are still leading to broad innovations, which we will come to view as extraordinary as any that have occurred to-date.

Here are three great achievements currently a-brewing:
  1. “Totally Transparent Processing.” By this, I argued that our use of the web (whether for search, communication, or information access) can increasingly occur in a fluid manner that is independent of the device we are using, independent of the human language we prefer, independent of the modality of the data, and independent of the corpus of information on which our interaction is based. In effect, processing can be transparent ∀d∈D, ∀l∈L, ∀m∈M, ∀c∈C. Our barriers to using information technology are fading away and becoming transparent.
  2. “Ideal Distributed Computing.” While we have known the fundamentals of distributed computing for many decades, only today are we reaching a state where we can achieve a powerful and efficient balance of computation between all end-user devices and a vast collection of shared storage and computational resources. Cloud computing is today’s term d’arte, but I talked more generally about systems with the flexibility that computation and data can move across computers within a cluster, across clusters of computers and—of course—between clusters and all other (say, end user) devices. The result is the efficient, even awesome, capability to provide communication, computation and data to a vast collection of people and applications.
  3. “Hybrid, Not Artificial, Intelligence.” Systems are regularly augmenting the capability of all of us in day-to-day life, and our collective use of those systems is, in turn, augmenting the capabilities of those systems in a beneficial virtuous circle. The virtuous circle is operating already in the search engine, voice recognition systems, recommendation systems, and more. There is every reason to think the effect will become ever more potent as computers are applied to more domains and and used by larger populations. The result may not be artificially intelligent machines that pass the Turing Test, but instead systems that will be ever more capable of helping us achieve our goals in life -- in a kind of partnership. For a related take on this, you might look at a Google Official Blog post, “The Intelligent Cloud,” which Franz Och and I posted last Fall.
More explanation and many examples, based on Google research and services, are available in the slides I used with my talk. A PDF file of those slides is available on the WWW2009 website under the papers and presentations link.

الخميس، 23 أبريل 2009

Congratulations to NSF CLuE Grant awardees



The first goal of the Academic Cluster Computing Initiative was to familiarize the academic community with the methods necessary to run very large datasets on massive distributed computer networks. By expanding that program to include research grants through the National Science Foundation's Cluster Exploratory (CLuE) program, we're also hoping to enable new and better approaches to data-intensive research across a range of disciplines.

Now that the NSF has announced the 2009 CLuE grants in addition to some previous Small Grant for Exploratory Research (SGER) grants, we're excited to congratulate the recipient researchers and wish them the best as they bring new projects online and continue to run existing SGER projects on the Google/IBM cluster.

The NSF selected projects based on their potential to advance computer science as well as to benefit society as a whole, and researchers at 14 institutions are tackling ambitious problems in everything from computer science to bioinformatics. The institutions receiving CLuE grants are Purdue, UC Santa Barbara, University of Washington, University of Massachussetts-Amherst, UC San Diego, University of Virginia, Yale, MIT, University of Wisconsin-Madison, Carnegie Mellon, University of Maryland- College Park, University of Utah and UC Irvine. Florida International University, Carnegie Mellon and University of Maryland will continue other projects with exiting SGER grants. These grantees will run their projects on a Google/IBM-provided cluster running an open source implementation of Google's MapReduce and File System.

We're excited to help foster new approaches to difficult, data-intensive problems across a range of fields, and we can't wait to see more students and researchers come up with creative applications for massive, highly distributed computing.

الخميس، 16 أبريل 2009

Socially Adjusted CAPTCHAs



Unfortunately, there is a war going on between humans and 'bots. Software
'bots are attempting to generate massive numbers of computer accounts
which are then sold in bulk to spammers. Spammers use these accounts to
inundate emails and discussion boards. Meanwhile humans are trying to
simply create an account and don't want to spend a lot of time proving
that they are not a program.

Typically we use CAPTCHAs -- we present an image of some distorted text
and then ask the applicant to type in the letters. As image processing gets
more sophisticated, these letter sequences tend to get longer and more
distorted, sometimes to the point where humans fail too.

So we switched the game. We show an image, say an airplane, but it
is randomly rotated and we ask the applicant to rotate it to "up." This
is generally hard for computers but easy for people. Well, for the most
part.

Since computers are good at faces, skies, text, etc. we sift
through our database of images running state-of-the-art up detectors to
remove those images. But of the images that remain, some are too hard
for people to figure out. What is up for a plate or a piece of
abstract art?

So here is where it gets interesting. We show people several images, one
of which is a "candidate" and we see how people do. If everyone rotates
it the same way, it is a keeper. If there is a lot of variation, we
discard it. As extra credit it turns out that even if the original image were
taken at an angle, it does not matter, since people, in large numbers,
socially adjust the CAPTCHA.

Read the full paper here (posted with the permission of WWW'09).

الأربعاء، 15 أبريل 2009

The Grill: Google's Alfred Spector on the hot seat



Alfred Spector, Google's VP of Research, tells COMPUTERWORLD the ins and outs of Research at Google and where it's headed for the future. Read the complete interview here.

الخميس، 2 أبريل 2009

Predicting the Present with Google Trends



Can Google queries help predict economic activity?

The answer depends on what you mean by "predict." Google Trends and Google Insights for Search provide a real time report on query volume, while economic data is typically released several days after the close of the month. Given this time lag, it is not implausible that Google queries in a category like "Automotive/Vehicle Shopping" during the first few weeks of March may help predict what actual March automotive sales will be like when the official data is released halfway through April.

That famous economist Yogi Berra once said "It's tough to make predictions, especially about the future." This inspired our approach: let us lower the bar and just try to predict the present.

Our work to date is summarized in a paper called Predicting the Present with Google Trends. We find that Google Trends data can help improve forecasts of the current level of activity for a number of different economic time series, including automobile sales, home sales, retail sales, and travel behavior.

Even predicting the present is useful, since it may help identify "turning points" in economic time series. If people start doing significantly more searches for "Real Estate Agents" in a certain location, it is tempting to think that house sales might increase in that area in the near future.

Our paper outlines one approach to short-term economic prediction, but we expect that there are several other interesting ideas out there. So we suggest that forecasting wannabes download some Google Trends data and try to relate it to other economic time series. If you find an interesting pattern, post your findings on a website and send a link to econ-forecast@google.com. We'll report on the most interesting results in a later blog post.

It has been said that if you put a million monkeys in front of a million computers, you would eventually produce an accurate economic forecast. Let's see how well that theory works ...

الأربعاء، 25 مارس 2009

The Unreasonable Effectiveness of Data



Alon Halevy, Peter Norvig, and I argue that we should stop acting as if our goal is to author extremely elegant theories, and instead embrace complexity and make use of the best ally we have: the unreasonable effectiveness of data. See the full article here (IEEE Intelligent Systems, March/April 2009).

الخميس، 19 مارس 2009

Google and WPP Marketing Research Awards: Improving industry understanding and practices in online marketing



Google and the WPP Group have teamed up to create a new research program with the goal of improving industry understanding of digital marketing. Eleven research awards have been given to universities through the Google and WPP Marketing Research Awards Program, announced by both companies in the fall of 2008. The academic studies will harness WPP client data to explore how online media influences consumer behavior, attitudes, and decision making. The research provides an opportunity for very innovative thinking in an area that is at the crossroads of marketing, computer science, economics, and various mathematical disciplines.

More than 120 entries were received by the deadline for proposals. The awards represent the first round of grants in the three-year program towards which WPP and Google will commit up to $4.6 million in an effort to support research around digital marketing. Hal Varian, Google's Chief Economist, participated on the decision committee.

The winning projects offer convincing designs for exploring how online and offline marketing influence consumer attitudes, decisions, and purchase behavior. As marketing continues to become more digital and more measurable, the results of these studies will also advance our understanding of how advertising investment should be allocated among media channels.

The researchers and affiliated academic institutions participating in this first round of awarded projects are:

• “Effect of Online Exposure on Offline Buying: How Online Exposure
Aids or Hurts Offline Buying by Increasing the Impact of Offline
Attributes”; Amitav Chakravarti, New York University, Stern School of
Business, Department of Marketing

• “The Interaction Between Digital Marketing Tactics and Sales
Performance Online and Offline”; Elie Ofek, Associate Professor
Marketing, Harvard Business School and Zsolt Katona, Associate
Professor of Marketing, UC Berkeley, Haas School of Business

• ”Are Brand Attitudes Contagious? Consumer Response to Organic
Search Trends”; Donna L. Hoffman, Professor, A. Gary Anderson
Graduate School of Management, University of California Riverside and
Thomas P. Novak, A. Gary Anderson Graduate School of Management,
University of California Riverside

• “Does internet advertising help established brands or niche ("long
tail") brands more? Catherine Tucker, Assistant Professor of
Marketing, MIT Sloan School of Marketing and Avi Goldfarb, Associate
Professor of Marketing, Joseph L. Rotman School of Management
University of Toronto

• “Marketing on the Map: Visual Search and Consumer Decision Making”;
Nicolas Lurie, Assistant Professor of Marketing, College of
Management, Georgia Institute of Technology, College of Management and
Sam Ransbotham, Assistant Professor of Information Systems, Carroll
School of Management, Boston College

• “Methods for multivariate metric analysis; identifying change
drivers”; Trevor J. Hastie, Professor, Department of Statistics,
Stanford University

• “Unpuzzling the Synergy of Display and Search Advertising: Insights
from Data Mining of Chinese Internet Users”; Hairong Li, Department of
Advertising, Public Relations, and Retailing, Michigan State
University and Shuguang Zhao, Media Survey Lab, Tsinghua University

• “Optimal Allocation of Offline and Online Media Budget”; Sunil
Gupta, Professor of Business Administration, Harvard Business School;
Anita Elberse, Associate Professor, Harvard Business School; and
Kenneth C. Wilbur, Assistant Professor of Marketing, Marshall School
of Business, University of Southern California

• “Targeting Ads to Match Individual Cognitive Styles: A Market
Test”; Glen Urban, Professor, MIT Sloan School of Management

• “How do consumers determine what is relevant? A psychometric and
neuroscientific study of online search and advertising effectiveness”;
Antoine Bechara, Professor of Psychology and Neuroscience, Department
of Psychology/Brain & Creativity Institute, University of Southern
California and Martin Reimann, Fellow, Department of Psychology/Brain
& Creativity, University of Southern California

• “A Comprehensive Model of the Effects of Brand-Generated and
Consumer-Generated Communications on Brand Perceptions, Sales and
Share”; Douglas Bowman and Manish Tripathi, Professors of Marketing,
Goizueta Business School, Emory University.

You can find more information about the Google and WPP Marketing Research Awards Program on the website.