Big Leap in Machine Language Translation

 Language Translation
(1) Statistical machine translation – in which computers essentially learn new languages on their own instead of being “taught” the languages by bilingual human programmers – has taken off. The new technology allows scientists to develop machine translation (MT) systems for a wide number of obscure languages at a pace that experts once thought impossible.

(2) Experts in the field say the progress and accuracy of statistical machine translation have recently surpassed that of the traditional machine translation programs used by Web sites like Yahoo and BabelFish. In the past, such programs were able to compile extensive databanks of foreign languages that allowed them to outperform statistics-based systems.

(3) Traditional machine translation relies on painstaking efforts by bilingual programmers to enter the vast wealth of information on vocabulary and syntax that the computer needs to translate one language into another. But in the early 1990s, a team of researchers at International Business Machines (IBM) devised another way to do things: feeding a computer an English text and its translation in a different language. The computer then uses statistical analysis to “learn” the second language.

(4) Compare two simple phrases in Arabic: rajl kabir and rajl tawil . If a computer knows that the first phrase means “big man” and the second means “tall man,” the machine can compare the two and deduce that rajl means “man,” while kabir and tawil mean “big” and “tall,” respectively. Phrases like these, called N-grams (with “N” representing the number of terms in a given phrase), are the basic building blocks of statistical machine translation, or MT. Although in one sense it was more economical, this kind of machine translation was also much more complex, requiring powerful computers and software that did not exist for most of the 1990s. But a workshop at John Hopkins University in Baltimore changed all that in the summer of 1999. A team led by Kevin Knight, the head of machine translation research at the Information Sciences Institute at the University of Southern California, came up with a software application package, Egypt/Giza, that made statistical translation accessible to researchers across the United States.

(5) Today, researchers are racing to improve the quality and accuracy of the translations. The final translations generally give an average reader a solid understanding of the original meaning but are far from grammatically correct. While not perfect, statistics-based technology is also allowing scientists to crack scores of languages in a fraction of the cost, that traditional methods involved. A team of computer scientists at John Hopkins led by David Yarowsky is developing machine translations of such languages as Uzbek, Bengali, Nepali—and one from “Star Trek.”

(6) “If we can learn how to translate even Klingon into English, then most human languages are easy by comparison,” Yarowsky said. “All our techniques require is having texts in two languages. For example, the Klingon Language Institute translated ‘Hamlet’ and the Bible into Klingon, and our programs can automatically learn a basic Klingon – English MT system from that.” Yarowsky hopes to have working translation systems for as many as 100 languages within five years. Although the grammatical structures of languages like Chinese and Arabic make them hard to analyze statistically, he said, it will only be a matter of time before such hurdles are overcome.
(7) In addition to the release of Egypt/Giza in 1999, the spread of the Internet has led to an explosion of translated text in far-flung languages, greatly aiding the team’s research. Researchers have also benefited from a much faster means of evaluating the outcome of translation experiments: a computerized technique developed by IBM enables researchers to test 10 to 100 new approaches for cracking languages each day. The technique, known as the Bleu Metric, compares machine translations with a “gold standard” based on human translations. Instead of waiting for human beings to assign a score to the quality of a machine translation, the Bleu Metric does so almost instantly through a statistical comparison. This provides scientists with a fast, objective measurement that they can use to note improvement and saves them from having to review every unsuccessful experiment.

(8) Despite the progress being made in statistical machine translation, some researchers remain skeptical, preferring to focus their efforts on language-specific translation techniques. Knight acknowledges that statistical machine translation is far from perfect. His team has sought to combine the statistical and traditional approaches to achieve maximum accuracy and to produce translations the average computes user can understand.

(9) The best machine translation system today, while capable of yielding a passage’s general meaning, are better known for their muddled syntax than their accuracy. By applying the principles of statistical translation to varying grammatical structures, Knight hopes to solve some of these problems. “N-grams are one of those things where you don’t know how much you need it until you take it away,” he said. “The way our imagination work, we need help.” Google gives 714 results for the Klingon Language Institute.

Part A – Multiple Choice Questions