Showing posts with label Google Translator Toolkit. Show all posts
Showing posts with label Google Translator Toolkit. Show all posts

Wednesday, June 1, 2011

Google Chokes on McLocalization’s Perpetual Garbage Creation Machine


(Of course, there would be a cost to Google for the Biblical amounts of resources I have wasted on my little antisocial effort, but that is irrelevant for this thought experiment. Let us further assume that Google can’t track me down even with those Google Maps cars that go around the world with little cameras on them.)

(I go AWOL with a new identity and end up eating Fritos while watching Sopranos re-runs in some fleabag motel.)

It turns out, apparently, that the McLocalization Brigade was busy doing this exact same thing on an even more massive scale. To the point that they were putting a considerable strain on Google’s infrastructure.

Thursday, February 17, 2011

Why the Machine Translation Crowd Hates Google

Jack: You have more sexual hang ups than an adult chat line run by Gilbert Gottfried.
Liz: What?
Jack: That was written by a computer program we're working on to replace you.
(30 Rock)

To get a measure of how insignificant the localization industry is, witness the chasm that separates the amoeba in the machine translation (MT) sector and the Behemoth from Mountain View. The Google MT people are regularly quoted in major news outlets and then you witness the carping that ensues in the blogosphere and twittersphere. You would think that Google people would be lauded as pioneers in the field. By rolling out their engine for free, they have at least placed computerized translation at the forefront of the public’s imagination. You would be wrong.

It is telling to see that not a single Google guy or gal showed up at last November’s Denver Conference of the Association for Machine Translation in the Americas (AMTA). The usual suspects were there: Asia Online, SDL-Language Weaver and a veritable Star Trek convention of academic geeks. Google was not even a corporate sponsor. None of the high-level Google MT execs was a guest speaker. The keynote address was delivered by Paul Bremer, for Chrissake. It is a measure of how far out in the wilderness these people are that they have to pay a former Bush Administration official to visit.

Even more telling is how the slightest critical piece on Google’s translation engine circulates at the speed of light through the blogosphere, especially by agencies and MT specialists. This, of course, is aided and abetted by slightly misinformed freelance translators who have nightmares at least once a week in which Google forecloses on their mortgage.

Google’s people, of course, are blissfully unaware of this. As in so many other fields, they are Moses on the mountain while the Israelites are in the valley building shrines to stones that look a little bit like sheep. I guess Google researchers going to visit the AMTA would be a little like a twenty-first century Homo sapiens going to a Cro-Magnon convention to find viable alternatives to fossil fuel. “So, Mr. Ughhho, what do you have along the lines of fuel cells?” “Ughhho have fire! Fire good! Ughhho powerful!”

To get a feel of the little love that the Big G gets in Machine Translation Island, see the reaction to this article in The Guardian. One member of the Google MT team is quoted as saying the following:

Andreas Zollmann, who has been researching in the field for many years and working at Google Translate for the last year, suggests, along with Blunsom, that the idea that more and more data can be introduced to make the system better and better is probably a false premise. "Each doubling of the amount of translated data input led to about a 0.5% improvement in the quality of the output," he suggests, but the doublings are not infinite. "We are now at this limit where there isn't that much more data in the world that we can use," he admits. "So now it is much more important again to add on different approaches and rules-based models."

What this means is that once Google Translate achieves a certain degree of quality (and pay no heed to any bulls**t to the contrary: Google Translate is the state of the art in SMT), the rate of progress reaches a plateau. Improvements are still achieved, but they are painfully gradual compared to the pioneering years of the technology. Whereas five years ago doubling a one-billion-word corpus brought a 100% improvement in quality, nowadays a doubling of a ten-billion-word MT corpus only brings a marginal rate of improvement.

To analyze another example, take the reaction when another Google executive remarked to an Australian newspaper that "I'd be really careful about having any kind of a sensitive debate with someone either spoken or written using these translations." Whew, that prompted a firestorm!

Despite the myriad things that are questionable about the company, one thing you have to love about The Google is its intellectual honesty. Of course, that honesty is enabled by the fact that its interest in setting up the MT engine in the first place isn’t commercial, but rather strategic. Instead of making pennies selling their MT application to corporates and Joe Schmoe, they decided to make the application available for free in order to expand their corpus. This is typical of the company’s long-term vision. Instead of licensing their invention, they decided to put it out there because any narrowing of the language barrier will broaden the reach of the Internet, which is their real core business. So they simply devote a tiny sliver of their R&D budget to initiatives such as GT.

The Machine Translation Sector Has Been Googled

The thing is that a tiny squirt of Google’s R&D is like Gargantua flooding Paris. And the Parisians can get hopping mad and resentful! Zollmann’s unexceptional pronouncement can trigger a lot of sniping. From the purely clueless to the slyly disingenuous. Or take the whispering campaign about Google Translate and confidentiality issues (you know who you are).

Frankly, I find it all slightly smug and more than a little infuriating. Because it is typical of the intellectual dishonesty that pervades the vulgar push to drive down in standards in the translation industry through Web 2.0 idiocy. You see, folks, the thing is that when a Google dude says “our MT engine has reached its limit” what he really means is “machine translation has reached its limit.”

And why is that? Because none of the pygmies sniping at Google can match its content aggregation capabilities. Or its raw (human) brain power.

The other strategy is to insinuate snidely the following: “Well, Google might have reached its limits, but we know a better shortcut down the Yellow Brick Road.” In response, one should entertain the following thought experiment. Let us imagine that there is an Albert Einstein of machine translation. Let’s imagine that he is 23 years old and just graduated from MIT with a double Ph.D. in physics and linguistics. What is more likely: That he works for Google or for a mega-agency that haggles for nickels and dimes with freelancers? I have my own (incredibly biased) answer to that question. I leave you to draw your own.

Another, slightly bizarre, instance of this tendency is when Google began to warn that it would exclude from its search results any pages that had been machine translated, even those that had a smattering of post-editing. This, of course, is a huge bummer for the MT crowd, because its main objective isn’t to provide better machine translation technology. Its real objective is to convince the industry to crowdsource the proofreading of MT drivel by non-professionals.

Google’s epistemological modesty prompts this sort of conspiratorial reaction: “Google just wants us to use only their MT tool.”

Which, in my view, is slightly bizarre. What does Google care about whether its free tool is used or not to machine translate a website? “Aha,” our conspiracy theorist will riposte, “if Google translate is not used, it cannot enrich its corpus. Right?” Wrong! While the translated website doesn’t enrich the Google corpus immediately, eventually it will make it out into the Internet and the company’s crawlers will eventually catch it in its nets. After identifying it as multilingual content, the site will ultimately make its way to further contaminate the SMT corpus.

Google Translate: Massive Party Pooper

Why all this animus? After all, Google is not in direct competition with the MT Lilliputians. “Free” is not in competition with companies that want to license their own SMT applications.

No, it’s not the competition. In a sense, it is something much, much worse. You see, Google squelched in one fell swoop the opportunity for another one of those rounds of incredibly wasteful capital misallocation to which Silicon Valley has treated us over the years. By creating the best MT application and putting it out there for free, many years of free-spending (and competing) venture-capital-financed startups failed to get off their ground. In a world without Page and Bryn, all of these MT dotcoms would have gotten exactly where we are now in double the time and at several times the cost to society.

Wasteful as they are, these manias are yummy because they create VC-funded millionaires that take seed money and sprinkle it on sportscars and trophy wives. And, lo, how many advisory fees were foregone by investment banks!

No, Virginia, there will never be a machine translation IPO. Or, to put it another way, it already happened. In 2003, when the U.S.S. Google floated on the wide open market seas and sank a lot of paper boat dreams.

The reverberations of the Google bomb are still felt to this day. Any presentation to a venture capitalist of the new, new thing in computerized translation is doomed to fail. The prospect goes home, takes off his tie, opens his laptop and applies the Beta version to his daughter’s French homework. His conclusion: The Next Big MT Thing actually performs pretty much along the lines of something that is free (yuck!).

“What are these clowns peddling? Next!”

Miguel Llorens is a freelance financial translator based in Madrid who works from Spanish into English. He is specialized in equity research, economics, accounting, and investment strategy. He has worked as a translator for Goldman Sachs, the US Government's Open Source Center and H.B.O. International, as well as many small-and-medium-sized brokerages and asset management companies operating in Spain. To contact him, visit his website and write to the address listed there. Feel free to join his LinkedIn network or to follow him on Twitter.

Thursday, August 19, 2010

A Financial Translator Dabbles in Machine Translation: Can Google Translator Toolkit Replace a Licensed CAT tool?

After a spell in which I reinstalled Windows several times in the space of a few months because of an array of hardware/software problems and the purchase of a new PC, I basically decided not to trust computer hard drives with long-term storage. I now consider my main PC’s hard drive as an empty shell for temporary storage of non-essential items and some of the most recent files I use. The heavy lifting of my real storage is done by my Carbonite account (which backs up my data in the cloud, for a fee after 2GB) and external hard drives. Redundant? Yes, but I now no longer worry about my backup failing.
Now, every time I encounter intractable PC problems (at least for a user with middling computer literacy), I simply reformat my HD and reinstall my OS, which, of course, erases any data on it. This, however, creates the inconvenience of also erasing all my programs and settings, forcing me to spend an entire day re-teaching my PC to think like me. This is particularly uncomfortable for my CAT tools, which feature relatively roundabout ways of reinstalling licenses in order to avoid piracy. Wordfast forces you to reinstall, obtain an install number, visit their website, punch in an account name and password, and “re-license” your product by obtaining a new license number for your install number. Trados’s softkey procedure similarly implies logging in to their support site using an account password, downloading a new .txt document, saving it on your hard drive and then navigating to it from Workbench. The program then identifies the file as a new license and unlocks the program for use with large TMs.
Needless to say, these procedures, while simple, are never hassle-free. There are always hiccups along the way and several restarts are necessary before the software starts to work.
This led me to wonder about CAT tools in the cloud-computing sky. After all, cloud-based software is the wave fo the future. Eventually, our hard drives will not contain hardly any software. Our PCs will simply be a hub to connect with all the programs we need, lodged in the servers of software providers. Many of us use e-mail that way, never downloading messages to our PCs but rather reading and storing messages via our browsers.
To avoid the mess of reinstalling software on the latest reinstallation of my Windows environment, I decided to try out Google Translator Toolkit (GTT). I heard of Lingotek a few years back, which featured sharing of large TMs among thousands of translators, but they have since adopted a B2B business model, apparently, and only provide maintenance of the tool to people who originally signed up for it.
I was familiar with Google’s machine translation through the plug-in in Wordfast. And I knew the expanded Toolkit featured some limited CAT capabilities. I decided to give it a spin as a CAT tool to see whether it could supplant paid programs that require fussy re-licensing procedures. As many know, GTT combines certain aspects of machine translation with CAT capabilities. You can upload a small translation memory (TM) of less than 50 GB. The downside is that this material automatically becomes the property of Google Inc., with the attendant confidentiality problems. The work of the users uploading TMs and processing sentences in the Google environment is then used by the company to enrich the quality of its own machine translations (MT).
The conclusion, sadly, is that GTT is not ready for prime time as a replacement of CAT tools. The main drawback: the way it scans tables and visual elements of Word docs. Instead of leaving visual elements in the same state as in the original Word document, it scans it partially or whollly and then translates elements that do not need translation (such as letterheads), forcing the user to backtrack over the document at the end of the project and manually insert translations, reinsert originals, make little tweaks here and there, copy-paste, paste-copy, etc.
The key word here is “manual”. Everyone knows that anything that has to be done manually on a computer augments exponentially the amount of mistakes.  So that is a strike against the use of GTT as a replacement of heftier CATs (or TEnTs, as they are also known).
Another major drawback: the tool translates everything in the text, as opposed to one sentence at a time the way CAT tools do, which usually work segment by segment. This can create a lot of headaches. GTT doesn’t contain an “Insert original” (like say CTRL+C or CTRL+O) option. Which forces the professional translator to cut and paste (a lot in some cases).
Another (rather bizarre) aspect of Google Translator Toolkit: the Help info in other languages is… well… (how shall I put this?)… er… translated by Google Translate. Which means two out of three sentences are complete gibberish. Of course, this makes sense in some sort of twisted Silicon-Valley way only a Sheldonian computer engineer would understand. “If we’re designing a machine translation tool, wouldn’t it be hypocritical to use a human editor to translate the output in our Help files spit out by that very MT tool?” Well, not if you want your users to actually embrace the tool. This kind of laudable intellectual honesty mixed with utter idiocy is probably a major reason why more translators don’t embrace projects such as Google’s GTT.
Finally, I hesitate to mention this… After all, it is a free tool. You get what you pay for, right? I (rather naively) sent a message making suggestions to improve the tool to the contact address of the team devoted to maintaining GTT and promptly got a “Delivery Status Notification (Failure)” reply from the server. The developer’s contact addresses no longer exist on the Google servers (!), which may mean that any further improvement or expansion of GTT has been either called off or postponed indefinitely. Hardly encouraging.
In any case, TEnT developers do not need to lose any sleep… at least for now. I guess the option for me would be to try out open source CAT tools, since by definition they do not require messy re-licensing. The problem is I am not familiarized with any, so self-training from ground zero would be a requisite. Furthermore, I do not even know if they require the use of Linux, with which I am totally unfamiliarized. We shall see…