الجمعة، 28 أبريل 2006

Statistical machine translation live



Because we want to provide everyone with access to all the world's information, including information written in every language, one of the exciting projects at Google Research is machine translation. Most state-of-the-art commercial machine translation systems in use today have been developed using a rules-based approach and require a lot of work by linguists to define vocabularies and grammars.

Several research systems, including ours, take a different approach: we feed the computer with billions of words of text, both monolingual text in the target language, and aligned text consisting of examples of human translations between the languages. We then apply statistical learning techniques to build a translation model. We have achieved very good results in research evaluations.

Now you can see the results for yourself. We recently launched an online version of our system for Arabic-English and English-Arabic. Try it out! Arabic is a very challenging language to translate to and from: it requires long-distance reordering of words and has a very rich morphology. Our system works better for some types of text (e.g. news) than for others (e.g. novels) -- and you probably should not try to translate poetry ... but do stay tuned for more exciting developments.

Update: We've just opened a discussion forum for all topics related to machine translation.

Update: Fixed broken link to NIST results.

الخميس، 27 أبريل 2006

Our conference on automated testing



Automated testing is one of my passions: it has hard problems to be solved, and they get harder every day. Over the past few years, I've had the opportunity to work on several automation projects, and now I'm getting a chance to combine my passion for automation with my love for the city of London.

I'm happy to announce that Google will be hosting a Conference on Test Automation in our London office on September 7 and 8, 2006. Our goal is to create a collegial atmosphere where participants can discuss challenges facing people on the cutting edge of test automation, and evaluate solutions for meeting those challenges.

Call for Presentations
We're looking for speakers with exciting ideas and new approaches to test automation. If you have a subject you'd like to talk about, please send me email at londontestconf@google.com that includes a description of your 60- or 90-minute talk in 500 words or less. Deadline for submissions is June 1.

We're planning to have 10 people give presentations at the conference followed by adequate time for discussion. If you'd like to attend as a non-speaker, watch this space. Once we've got a slate of speakers, we'll post it along with details on attending.

الأحد، 23 أبريل 2006

See you at CHI



The raison d’etre for our user experience research team is driven by Google's keen interest in focusing on the user. So we help many product teams provide the best possible experience to everyone around the world, primarily by inviting thousands of people to take part in usability tests in our labs, and by analyzing our logs to identify problems which need fixing. From this we get the data we help our engineers make Google products as easy as possible to use for the millions of people out there who think computers are far too complicated. People like my Mum, Dad, girlfriend, Gran — and pretty much everyone I know!

We’re one of several Google teams that publish research at academic and industry conferences, and this week a number of us will be attending the CHI (Computer-Human Interaction) conference in Montreal, the world's premiere gathering for CHI researchers and practitioners. Googlers from several teams will take part in eight sessions, each focusing on different aspects of human-computer interaction. (The full program is here – it’s a PDF file.)

A Large Scale Study of Wireless Search Behavior: Google Mobile Search – In a session on Search and Navigation: Mobiles and Audio, we'll present the first large-scale study of search behavior for mobile users, highlighting some shortcomings of wireless search interfaces.

Scaling the card sort method to over 500 items: Restructuring the Google AdWords Help Center – Here we adapt the popular card-sorting research methodology to large information sets where the traditional approach is impractical and discuss how we've applied this technique.

No IM Please, We’re Testing – During the Usability Evaluations: Challenges and Solutions session we’ll discuss the use of instant messaging tools like Google Talk in usability tests, and the benefits of this technique for enabling live collaboration between test moderators and observers.

Add a Dash of Interface: Taking Mash-Ups to the Next Level – Here we contribute to the discussion of how extendable interfaces like Google Maps are enabling exciting new online innovation through the combining of data sources.

Why Do Tagging Systems Work? – This panel will address the design challenges of scaling tagging systems to meet their recent surge in popularity. Gmail is an example of email tagging that offers more flexibility than traditional hierarchical systems.

Design Communication: How Do You Get Your Point Across? – A key challenge for UI designers is communicating solutions and challenges within product teams. This panel focuses on effective ways to do that.

“It’s About the Information, Stupid!” Why We Need a Separate Field of Human Information Interaction – This interdisciplinary panel will discuss arguments for and against a distinct field focusing on information rather than computing technology. One for the theoreticians? (-;

Incorporating Eyetracking into User Studies at Google – In this Eyetracking in Practice workshop, we’ll talk about some of the challenges we’ve encountered in studies of eyetracking in our labs.

If you work in, or study, the area of human-computer interaction, the user experience team is hiring. Right now we’re looking for user experience researchers (including those with specialized quantitative skills), UI designers, and more.

الأربعاء، 22 مارس 2006

First Robots



With 4 seconds left to go, the Team Cheesy Poofs robot shouldered its way onto the 3 foot platform, pivoted 90 degrees into scoring position, and rapid-fired 10 balls directly into the 3-point goal. They won the match, and the Google Silicon Valley Regional Championship for US FIRST, a non-profit "For the Inspiration and Recognition of Science and Technology" (FIRST).

Google jumped at the opportunity to sponsor this organization after Dean Kamen (inventor of the Segway and the first implantable dialysis pump) spoke to a packed Google audience about his lifelong crusade to improve education in the United States. Dean founded US FIRST over 15 years ago, and from humble beginnings in the Northeast, FIRST has now grown to involve over 60,000 high school students all over the United States and the world.

FIRST was a natural partner for Google, given their focus on science and technology, their passion for changing the world for the better, and their single-minded focus on making education fun for students. When the final buzzer rang at the recent championship match the students jumped and hugged like they'd won the Superbowl. And in a way, they had. This event has all the excitement, tension, and drama of a major sporting event and then some.

Beyond sponsoring the FIRST tournament, Google also funded half a dozen teams in the Bay Area, ranging from East Palo Alto High School to Notre Dame High School. Several dozen employees also served as team mentors, meeting the students once a week to help construct the competition robots over the frantic six-week design/build cycle. Others volunteered at the Regional event as judges, coordinators, and referees, and plenty of Googlers were on hand to spectate the exciting matches.




We congratulate all the teams at the regional tournament for their hard work and innovation. We wish the six bay-area teams who qualified for the finals in Atlanta the best of luck . Bring home the gold!

السبت، 11 مارس 2006

Hiring: The Lake Wobegon Strategy



You know the Google story: small start-up of highly-skilled programmers in a garage grows into a large international company. But how do you maintain the skill level while roughly doubling in size each year? We rely on the Lake Wobegon Strategy, which says only hire candidates who are above the mean of your current employees. An alternative strategy (popular in the dot-com boom period) is to justify a hire by saying "this candidate is clearly better than at least one of our current employees." The following graph compares the mean employee skill level of two strategies: hire-above-the-mean (or Lake Wobegon) in blue and hire-above-the-min in red. I ran a simulation of 1000 candidates with skill level sampled uniformly from the 0 to 100th percentile (but evaluated by the interview process with noise of ±15%) starting from a core team of 10 employees with mean 75 and min 65. You can see how hire-above-the-min leads to a precipitous drop in skill level; one we've been able to avoid.



Another hiring strategy we use is no hiring manager. Whenever you give project managers responsibility for hiring for their own projects they'll take the best candidate in the pool, even if that candidate is sub-standard for the company, because every manager wants some help for their project rather than no help. That's why we do all hiring at the company level, not the project level. First we decide which candidates are above the hiring threshold, and then we decide what projects they can best contribute to. The orange line in the graph above is a simulation of the hiring-manager strategy, with the same candidates and the same number of hires as the no-hiring-manager strategy in blue. Employees are grouped into pools of random size from 2 to 14 and the hiring manager chooses the best one. We're pleased that these little simulations show our hiring strategy is on top. You can learn more about our hiring and working philosophy.

الثلاثاء، 7 مارس 2006

An experimental study of P2P VoIP



VoIP (Voice-over-IP) systems are one of the fastest growing means of communication on the Internet, enabling free or low-cost phone calls. But to date, researchers have had little data to work with to learn how to build VoIP systems better. Some of these systems are proprietary, and obtaining data about their operational characteristics has been particularly challenging. For instance, even though the Skype network has tens of millions of users, it has been hard for researchers to benefit from its commercial success.

Data was collected from a Skype 'supernode' running at Cornell. Skype is a Peer-to-Peer (P2P) system in which clients (for example, a home user's PC) communicate directly to exchange voice packets with other clients (also called peers). However, their communication is facilitated by special peers called supernodes that can allow the peers to connect even if they are behind firewalls or other network elements such as NATs (Network Address Translators). P2P in Skype already connects millions of users behind NATs today. Prior to our research, not much has been known about how Skype users and clients behave, and how supernodes are selected or what kinds of demands they place on the network they reside in.

We learned a couple things from the data. For example, we found that Skype users typically keep their client software open during the workday, as opposed to users of file-sharing P2P systems (such as KaZaa) where users typically join and leave the network with much greater frequency. In further contrast to P2P file-sharing applications, which typically tend to be bandwidth hogs, Skype clients and supernodes use relatively little bandwidth and CPU even when they relay VoIP calls. So this means you can run Skype without having it slow down your Internet connection.

You'll find even more results discussed in the paper. In addition to better P2P systems, researchers can use the data to design a better Internet. Based on what we've learned, perhaps researchers can design a next-generation P2P-friendly Internet that is commercially viable.

السبت، 4 مارس 2006

Teamwork for problem-solving



Google Research is about teamwork with outstanding engineers to solve novel and challenging problems that have an impact. But it's also about being at the forefront of scientific innovations. We're an active part of the research community, and we like to interact with researchers and scientists in academia. We're happy to serve as a hub for researchers to come and discuss their latest findings and get exposed to the large-scale problems and challenges that we face. Robert Tarjan, John Lafferty, and Brian Kernighan are among the professors that have spent time here.

We host world-renowned scientists spanning diverse areas including neuroscience, climatology, internet security and e-commerce -- and of course, computer science. In the fall, our Research Seminars attracted such prominent figures as John Hopcroft and Michael Rabin. This spring we're welcoming Christos Papadimitriou and Vladimir Vapnik, to name just a few.

So if you're curious about the latest meteor findings in Antarctica or interested in high-end computing and scientific visualization at NASA, do check out our "tech talks" on Google Video. You don't actually need to work at Google to "attend" the talks -- but if you're interested, we're always looking.