Thursday, July 15, 2010

Google PhD Fellowships go international

We introduced the Google Fellowship program last year in the United States to broaden our support of university research. The students who were awarded the 2009 fellowships were a truly impressive group, many having high profile internships this past summer and even a few with faculty appointments in the upcoming year.

Universities continue to be the source of some of the most innovative research in computer science, and in particular it’s the students that they foster who are the future of our field. This year, we’re going global and extending the fellowship program to Europe, Israel, China and Canada. We’re very happy to be continuing our support of excellence in graduate studies and offer our sincere congratulations to the following PhD students for receiving Google Fellowships in 2010:

Google European Doctoral Fellowships
  • Roland Angst, Google Europe Fellowship in Computer Vision (Swiss Federal Institute of Technology Zurich, Switzerland)
  • Arnar Birgisson, Google Europe Fellowship in Computer Security (Chalmers University of Technology, Sweden)
  • Omar Choudary, Google Europe Fellowship in Mobile Security (University of Cambridge, U.K.)
  • Michele Coscia, Google Europe Fellowship in Social Computing (University of Pisa, Italy)
  • Moran Feldman, Google Europe Fellowship in Market Algorithms (Technion - Israel Institute of Technology, Israel)
  • Neil Houlsby, Google Europe Fellowship in Statistical Machine Learning (University of Cambridge, U.K.)
  • Kasper Dalgaard Larsen, Google Europe Fellowship in Search and Information Retrieval (Aarhus University, Denmark)
  • Florian Laws, Google Europe Fellowship in Natural Language Processing (University of Stuttgart, Germany)
  • Cynthia Liem, Google Europe Fellowship in Multimedia (Delft University of Technology, Netherlands)
  • Ofer Meshi, Google Europe Fellowship in Machine Learning (The Hebrew University of Jerusalem, Israel)
  • Dora Spenza, Google Europe Fellowship in Wireless Networking (Sapienza University of Rome, Italy)
  • Carola Winzen, Google Europe Fellowship in Randomized Algorithms (Saarland University / Max Planck Institute for Computer Science, Germany)
  • Marek Zawirski, Google Europe Fellowship in Distributed Computing (University Pierre and Marie Curie / INRIA, France)
  • Lukas Zich, Google Europe Fellowship in Video Analysis (Czech Technical University, Czech Republic)
Google China PhD Fellowships
  • Fangtao Li, Google China Fellowship in Natural Language Processing (Tsinghua University)
  • Ming-Ming Cheng, Google China Fellowship in Computer Vision (Tsinghua University)
Google United States/Canada PhD Fellowships
  • Chong Wang, Google U.S./Canada Fellowship in Machine Learning (Princeton University)
  • Tyler McCormick, Google U.S./Canada Fellowship in Statistics (Columbia University)
  • Ashok Anand, Google U.S./Canada Fellowship in Computer Networking (University of Wisconsin)
  • Ramesh Chandra, Google U.S./Canada Fellowship in Web Application Security (Massachusetts Institute of Technology)
  • Adam Pauls, Google U.S./Canada Fellowship in Machine Translation (University of California, Berkeley)
  • Nguyen Dinh Tran, Google U.S./Canada Fellowship in Distributed Systems (New York University)
  • Moira Burke, Google U.S./Canada Fellowship in Human Computer Interaction (Carnegie Mellon University)
  • Ankur Taly, Google U.S./Canada Fellowship in Language Security (Stanford University)
  • Ilya Sutskever, Google U.S./Canada Fellowship in Neural Networks (University of Toronto)
  • Keenan Crane, Google U.S./Canada Fellowship in Computer Graphics (California Institute of Technology)
  • Boris Babenko, Google U.S./Canada Fellowship in Computer Vision (University of California, San Diego)
  • Jason Mars, Google U.S./Canada Fellowship in Compiler Technology (University of Virginia)
  • Joseph Reisinger, Google U.S./Canada Fellowship in Natural Language Processing (University of Texas, Austin)
  • Maryam Karimzadehgan, Google U.S./Canada Fellowship in Search and Information Retrieval (University of Illinois, Urbana-Champaign)
  • Carolina Parada, Google U.S./Canada Fellowship in Speech (Johns Hopkins University)
The students will receive fellowships consisting of full coverage of tuition, fees and stipend for up to three years. These students have been exemplary thus far in their careers, and we’re looking forward to seeing them build upon their already impressive accomplishments. Congratulations to all of you!

Wednesday, July 14, 2010

Translating Wikipedia

(Cross-posted from the Google Translate Blog)

We believe that translation is key to our mission of making information useful to everyone. For example, Wikipedia is a phenomenal source of knowledge, especially for speakers of common languages such as English, German and French where there are hundreds of thousands—or millions—of articles available. For many smaller languages, however, Wikipedia doesn’t yet have anywhere near the same amount of content available.

To help Wikipedia become more helpful to speakers of smaller languages, we’re working with volunteers, translators and Wikipedians across India, the Middle East and Africa to translate more than 16 million words for Wikipedia into Arabic, Gujarati, Hindi, Kannada, Swahili, Tamil and Telugu. We began these efforts in 2008, starting with translating Wikipedia articles into Hindi, a language spoken by tens of millions of Internet users. At that time the Hindi Wikipedia had only 3.4 million words across 21,000 articles—while in contrast, the English Wikipedia had 1.3 billion words across 2.5 million articles.

We selected the Wikipedia articles using a couple of different sets of criteria. First, we used Google search data to determine the most popular English Wikipedia articles read in India. Using Google Trends, we found the articles that were consistently read over time—and not just temporarily popular. Finally we used Translator Toolkit to translate articles that either did not exist or were placeholder articles or “stubs” in Hindi Wikipedia. In three months, we used a combination of human and machine translation tools to translate 600,000 words from more than 100 articles in English Wikipedia, growing Hindi Wikipedia by almost 20 percent. We’ve since repeated this process for other languages, to bring our total number of words translated to 16 million.

We’re off to a good start but, as you can see in the graph below, we have a lot more work to do to bring the information in Wikipedia to people worldwide:

Number of non-stub Wikipedia articles by Internet users, normalized (English = 1)

We’ve also found that there are many Internet users who have used our tools to translate more than 100 million words of Wikipedia content into various languages worldwide. If you do speak another language we hope you’ll join us in bringing Wikipedia content to other languages and cultures with Translator Toolkit.

We presented these results last Saturday, July 10, at Wikimania 2010 in Gdańsk, Poland. We look forward to continuing to support the creation of the world’s largest encyclopedia and we can’t wait to work with Wikipedians and volunteers to create more content worldwide.

Google Books goes Dutch

(Cross-posted from the European Public Policy blog)

In recent months, I’ve got to know a group of people in the Hague who are working on an ambitious project to make the rich fabric of Dutch cultural and political history as widely accessible as possible - via the Internet.

That team is from the National Library of the Netherlands, the Koninklijke Bibliotheek (KB), and as of today, we’ll be working in partnership to add to the library’s own extensive digitisation efforts. We’ll be scanning more than 160,000 of its public domain books, and making this collection available globally via Google Books. The library will receive copies of the scans so that they can also be viewed via the library’s website. And significantly for Europe, the library also plans to make the digitised works available via Europeana, Europe’s cultural portal.

The books we’ll be scanning constitute nearly the library’s entire collection of out-of-copyright books, written during the 18th and 19th centuries. The collection covers a tumultuous period of Dutch history, which saw the establishment of the country’s constitution and its parliamentary democracy. Anyone interested in Dutch history will be able to access and view a fascinating range of works by prominent Dutch thinkers, statesmen, poets and academics and gain new insights into the development of the Netherlands as a nation state.

This is the third agreement we've announced in Europe this year, following our projects with the Italian Ministry of Cultural Heritage and the Austrian National Library. The Dutch national library is already well underway with its own ambitious scanning programme, which will eventually see all of its Dutch books, newspapers and periodicals from 1470 onwards being made available online. By any measure, this is a huge task, requiring significant resources, and we’re pleased to be able to help the library accelerate towards its goal of making all Dutch books accessible anywhere in the world, at the click of a mouse.

It's exciting to note just how many libraries and cultural ministries are now looking to preserve and improve access to their collections by bringing them online. Much of humanity's cultural, historical, scientific and religious knowledge, collected and curated over centuries, sits in Europe's libraries, and its great to see that we are all striving towards the same goal of improving access to knowledge for all.

Google and other technology companies have an important role to play in achieving this goal, and we hope that by partnering with major European cultural institutions such as the Dutch national library, we will be able to accelerate the rapid growth of Europe's digital library.

Our commitment to the digital humanities

(Cross-posted on the Google Research Blog)

It can’t have been very long after people started writing that they started to organize and comment on what was written. Look at the 10th century Venetus A manuscript, which contains scholia written fifteen centuries earlier about texts written five centuries before that. Almost since computers were invented, people have envisioned using them to expose the interconnections of the world’s knowledge. That vision is finally becoming real with the flowering of the web, but in a notably limited way: very little of the world’s culture predating the web is accessible online. Much of that information is available only in printed books.

A wide range of digitization efforts have been pursued with increasing success over the past decade. We’re proud of our own Google Books digitization effort, having scanned over 12 million books in more than 400 languages, comprising over five billion pages and two trillion words. But digitization is just the starting point: it will take a vast amount of work by scholars and computer scientists to analyze these digitized texts. In particular, humanities scholars are starting to apply quantitative research techniques for answering questions that require examining thousands or millions of books. This style of research complements the methods of many contemporary humanities scholars, who have individually achieved great insights through in-depth reading and painstaking analysis of dozens or hundreds of texts. We believe both approaches have merit, and that each is good for answering different types of questions.

Here are a few examples of inquiries that benefit from a computational approach. Shouldn’t we be able to characterize Victorian society by quantifying shifts in vocabulary—not just of a few leading writers, but of every book written during the era? Shouldn’t it be easy to locate electronic copies of the English and Latin editions of Hobbes’ Leviathan, compare them and annotate the differences? Shouldn’t a Spanish reader be able to locate every Spanish translation of “The Iliad”? Shouldn’t there be an electronic dictionary and grammar for the Yao language?

We think so. Funding agencies have been supporting this field of research, known as the digital humanities, for years. In particular, the National Endowment for the Humanities has taken a leadership role, having established an Office of Digital Humanities in 2007. NEH chairman Jim Leach says: "In the modern world, access to knowledge is becoming as central to advancing equal opportunity as access to the ballot box has proven to be the key to advancing political rights. Few revolutions in human history can match the democratizing consequences of the development of the web and the accompanying advancement of digital technologies to tap this accumulation of human knowledge."

Likewise, we’d like to see the field blossom and take advantage of resources such as Google Books that are becoming increasingly available. We’re pleased to announce that Google has committed nearly a million dollars to support digital humanities research over the next two years.

Google’s Digital Humanities Research Awards will support 12 university research groups with unrestricted grants for one year, with the possibility of renewal for an additional year. The recipients will receive some access to Google tools, technologies and expertise. Over the next year, we’ll provide selected subsets of the Google Books corpus—scans, text and derived data such as word histograms—to both the researchers and the rest of the world as laws permit. (Our collection of ancient Greek and Latin books is a taste of corpora to come.)

We've given awards to 12 projects led by 23 researchers at 15 universities:
  • Steven Abney and Terry Szymanski, University of Michigan. Automatic Identification and Extraction of Structured Linguistic Passages in Texts.
  • Elton Barker, The Open University, Eric C. Kansa, University of California-Berkeley, Leif Isaksen, University of Southampton, United Kingdom. Google Ancient Places (GAP): Discovering historic geographical entities in the Google Books corpus.
  • Dan Cohen and Fred Gibbs, George Mason University. Reframing the Victorians.
  • Gregory R. Crane, Tufts University. Classics in Google Books.
  • Miles Efron, Graduate School of Library and Information Science, University of Illinois. Meeting the Challenge of Language Change in Text Retrieval with Machine Translation Techniques.
  • Brian Geiger, University of California-Riverside, Benjamin Pauley, Eastern Connecticut State University. Early Modern Books Metadata in Google Books.
  • David Mimno and David Blei, Princeton University. The Open Encyclopedia of Classical Sites.
  • Alfonso Moreno, Magdalen College, University of Oxford. Bibliotheca Academica Translationum: link to Google Books.
  • Todd Presner, David Shepard, Chris Johanson, James Lee, University of California-Los Angeles. Hypercities Geo-Scribe.
  • Amelia del Rosario Sanz-Cabrerizo and José Luis Sierra-Rodríguez, Universidad Complutense de Madrid. Collaborative Annotation of Digitalized Literary Texts.
  • Andrew Stauffer, University of Virginia. JUXTA Collation Tool for the Web.
  • Timothy R. Tangherlini, University of California-Los Angeles, Peter Leonard, University of Washington. Northern Insights: Tools & Techniques for Automated Literary Analysis, Based on the Scandinavian Corpus in Google Books.
We have selected these proposals in part because the resulting techniques, tools and data will be broadly useful: they’ll help entire communities of scholars, not just the applicants. We look forward to working with them, and hope that over time the field of digital humanities will fulfill its promise of transforming the ways in which we understand human culture.

Tuesday, July 13, 2010

Introducing our Google Fiber for Communities website

In February we announced our plans to build experimental, ultra-high speed broadband networks. Over the past several months, our team’s been hard at work reviewing the nearly 1,100 community responses to our request for information—not to mention the nearly 200,000 responses from individuals across the U.S.

Throughout this process, one message has come through loud and clear: people are hungry for better and faster Internet access. With that in mind, today we’re launching a new site called Google Fiber for Communities, where you can learn more about fiber networks and keep up-to-date on our project. You’ll also be able to advocate for common-sense federal and local policies that would help fiber deployments nationwide.

We also wanted to thank every community and individual that submitted a response, posted a YouTube video, started a website, joined a rally or otherwise let their voice be heard. We were so honored by the grassroots enthusiasm across the country for this project that we put together a short video to say thank you:



As we explained back in March, we plan to name our target community or communities by the end of the year. We still have some work ahead of us before we’re ready to make that announcement, but in the meantime, we hope this site helps to keep the conversation going.

BLOG STAR © 2008. Design by :Yanku Templates Sponsored by: Tutorial87 Commentcute