Showing posts with label machine translation. Show all posts
Showing posts with label machine translation. Show all posts

Saturday, September 12, 2026

There and Back Again: Bidirectional Diffusion Bridges for Multimodality Translation

This could be an interesting new paper by Stefano Ermon and his team!

From the abstract:
"Multimodality translation (e.g., text-to-image) is a core generative AI task. However, existing approaches 
(1) follow generative paths that do not directly represent the source modality, limiting the flexibility of some sampling algorithms; and
(2) are unidirectional, preventing inversion (e.g., image-to-text).
We propose BIT: Bidirectional Image-Text Diffusion Bridges.
In contrast to previous approaches, BIT starts directly from text and interpolates into images, providing
(1) a source-aware generative path that enables diverse and flexible sampling algorithms; and
(2) an endpoint-conditioned process that can be traversed from image to text, providing a unified, bidirectional generative framework.
BIT is derived through stochastic calculus, yielding SDE forms amenable to simulation and tractable loss functions that scale to high dimensions.
Our experiments show that BIT is competitive with denoising-diffusion and deterministic-flow baselines, and outperforms them on several vision--language and natural-science evaluations."

[2608.27885] There and Back Again: Bidirectional Diffusion Bridges for Multimodality Translation






Saturday, July 25, 2026

Language diversity: A historical context for language loss for the past 6,000 years or more

Over human history, many languages have disappeared and continue to disappear!

Why should the evolution of human languages be so different than the biological evolution? Do we need about 7,000 existing and actively spoken languages on our planet?

Hopefully, the Tower of Babel will go away in the 21st century. The world would be a much better place if we all speak the same language (with or without machine translation)!

Is this article another way of trying to indoctrinate with the DEI (Diversity, Equity, and Inclusion) ideology?

The approach taken in this study, i.e. a kind of extrapolation from present-day hunter-gatherers, is certainly questionable if not junk! This assumption is possibly a very strong inductive bias! What if e.g. the first humans spoke only one or a few languages and that for a long time before language diversity emerged maybe with the geographic spread of humans?

Notice that this study does not blame colonialism, but the preceding multinational empire building and expansion for the loss of languages.

"Humanity’s linguistic diversity is in trouble [???]. Roughly half of the 7500 languages signed or spoken today are endangered [???], and at least four go extinct every year. But while preserving the languages that do persist is a top priority, a new study in Science suggests that we are already more than one millennium beyond humanity’s linguistic “golden age.”

Linguists have long debated when the most languages were in use: Some have argued that the total number of languages remained roughly steady from the last ice age until modern-day European colonialism, others suggested that diversity has steadily declined since the dawn of agriculture about 12,000 years ago, and a third camp theorized a high point sometime before large states became established 4000 years ago.
Since writing was not developed until around 6000 years ago, however, understanding dynamics of early language remained challenging.
So, a team of anthropologists applied data on the languages, distributions, and behaviors of present-day hunter-gatherers to estimates of the total human population size and the size and number of individual cultures over the last 12,000 years. This allowed them to create a model of the total number of languages in use at any point since the last ice age ended.

The researchers showed that linguistic diversity rose steadily for around 10,000 years, reaching a “golden age” of tens of thousands of languages around 1000 to 3000 years ago before entering a language loss free-fall that continues to this day. The results reveal that the language crisis is not the result of modern colonialism but rather the culmination of millennia of cultural exchange and homogenization."

"... The golden age was followed by a period of rapid decline in linguistic diversity that coincided with the rise of large states and multinational empires, such as the Roman Empire, the researchers found. The finding challenges a commonly held view that widespread language extinction began later, about 500 years ago, with the onset of European colonial expansion. ..."

"To the Point
  • Fewer languages before agriculture: The researchers estimate that between approximately 4,500 and 6,200 languages were spoken at the beginning of the Holocene, around 12,000 years ago - probably fewer than the roughly 7,500 languages spoken and signed today.
  • A linguistic “golden age”: As the global human population grew, language diversity also increased. The models suggest that tens of thousands of languages may have existed between 1,000 and 3,000 years ago.
  • A deep history of language loss: The decline in linguistic diversity appears to have begun with the expansion of large states and empires, long before modern European colonialism.
  • Today’s languages are survivors [???]: The languages spoken today represent a small and historically biased sample of past linguistic diversity [???], with important implications for how researchers explain global patterns in language and culture.
...
They began with ethnographic data from 171 hunter-gatherer and fisher societies whose traditional subsistence and mobility had not been profoundly transformed by contact with food-producing populations. These data allowed the researchers to estimate the likely distribution of ethnolinguistic group sizes near the beginning of the Holocene.

The team then combined these estimates with reconstructions suggesting that the global human population 12,000 years ago was between approximately 4.4 and seven million. Assuming that ethnolinguistic groups at that time generally corresponded to distinct languages, the models produced an early-Holocene estimate centred on roughly 4,500 to 6,200 languages, although broader plausible estimates ranged from around 3,300 to 7,800. ..."


From the editor's summary and abstract:
"Editor’s summary
More than four languages are lost every year, and this rate is predicted to increase. Blasi et al. investigated how language diversity has changed over the Holocene using a Bayesian modeling approach, ethnographic data, and estimates of the total human population size over time.
They assumed a maximum size of ethnolinguistic groups and that the number of different ethnolinguistic groups is related to the number of languages. Their findings suggested that language loss is not a new phenomenon, and that language diversity has been decreasing rapidly over the past one to three millennia following a peak of an order of magnitude higher diversity than we see today. ...

Abstract
Characterizing the factors that have shaped linguistic diversity is fundamental for understanding human history, culture, and cognition.
In this study, we combined statistical and social computational modeling, ethnographic data, and paleodemographic inference to model trajectories of global linguistic diversity.
Before the onset of plant and animal domestication, the number of languages was smaller than it is today (4500 to 6000 compared with 7500).
Subsequent increases in global population precipitated increased linguistic diversity.
We uncovered a linguistic “golden age” with tens of thousands of languages 3000 to 1000 years ago. Great loss of linguistic diversity did not begin with recent colonial expansion but as multinational empires first spread along with their languages, pathogens, and cultures. Thus, extinction has likely played a much greater role in shaping linguistic and cultural diversity than previously thought."

ScienceAdviser


Study uncovers lost ‘golden age’ of languages (original news release) "A new study coauthored by Yale linguist Claire Bowern suggests that tens of thousands of languages were spoken between 1,000 and 3,000 years ago."






Map of currently spoken and signed languages of the world, colored by endangerment status.


Wednesday, May 20, 2026

Specialized medical transcription speech-to-text model beats frontier AI in real time and accuracy

Good news! Impressive! Errors in translation could be deadly!

When will all and any patient doctor encounters be quickly transcribed? 

"Copenhagen-based Corti launched Symphony for Speech-to-Text, a clinical-grade recognition model that achieved a 1.4 percent word error rate on English medical terminology—versus OpenAI’s 17.7 percent, ElevenLabs’ 18.1 percent, Whisper’s 17.4 percent, and Parakeet’s 18.9 percent. The gap widens further on structured clinical entities like medication dosages: Corti hit 98.3 percent recall while the strongest generalist model managed 44.3 percent. That difference matters more now than it used to.
As healthcare shifts toward autonomous AI agents making real-time clinical decisions, transcription errors compound—if a model mishears “hyperthyroidism” as “hypothyroidism,” every downstream system operates on corrupted data. Corti also outperformed legacy incumbent Dragon Medical One in dictation accuracy and now serves over 100 million patients annually across health systems including the UK’s National Health Service. ..."

From the abstract:
"After decades of use in dictation and, more recently, ambient documentation, speech is emerging as a primary modality for interacting with technology and AI in healthcare.
Yet medical speech recognition remains difficult: systems must capture specialized terminology, resolve contextual ambiguity, and render measurements, abbreviations, and clinical shorthand precisely.
Existing solutions are typically optimized either for general-purpose transcription or narrow dictation workflows, limiting their reliability in safety-critical settings and their usefulness for broader clinical workflows.
We introduce Symphony for Speech-to-Text, a medical-grade speech recognition system for real-time streaming and batch file-based clinical use. Symphony decomposes the transcription process into specialized components for recognition, formatting, and contextual correction to optimize medical term recall while producing clinically structured text in real time and adapting across use cases. Evaluations on public benchmark and medical speech datasets show that Symphony substantially outperforms state-of-the-art systems in clinical settings while matching or exceeding them in general-domain settings, suggesting robust generalization rather than overfitting.
We release a clinical benchmark dataset to support reliable validation and further progress in medical speech recognition. Symphony is available through a production-grade API for live dictation, conversational transcription, and batch audio file processing."

Data Points: Cursor Composer undercuts competition

Sunday, May 17, 2026

Several new smart rings promise to break sign language barriers by turning hand movements into instant text

Good news!

"Researchers in South Korea have developed a new sign language translation system based on users wearing seven rings equipped with sensors. According to a new study  ... the technology can reliably recognize and translate both American and International sign language words with roughly 88% accuracy. ..."

From the abstract:
"Sign language translation systems have long aimed to bridge the communication between signers and nonsigners. However, preliminary systems rely on glove-type wearables or wired sensor arrays, which constrain hand movement, reduce comfort, and require nonpersonalized sensor positions that limit adaptability across users.
Here, we introduce a wirelessly connected, ring-type sign language translator (WRSLT) designed to overcome these limitations by enabling full finger mobility through independent sensor rings and multilink communication. The system supports static and dynamic gesture detection using selected fingers via quantitative relevance analysis and achieves robust user-independent performance without per-user calibration.
WRSLT demonstrated high recognition accuracy on large-scale datasets comprising 100 American Sign Language and 100 International Sign Language words, achieving 88.3 and 88.5% accuracy, respectively, under unseen-user conditions (i.e., test users not included in model training).
Furthermore, a custom sequential word detection framework enables sentence-level translation from continuous signing input without requiring separate training on entire sentence structures."

Seven smart rings promise to break sign language barriers by turning hand movements into instant text

Fig. 1. Overall concept and design of WRSLT.


Friday, May 01, 2026

Celebrate the 20th anniversary of Google Translate

Happy anniversary! A nice example how time flies! And fast and how far machine translation has developed in those 20 years.

Nowaday, you can scan almost any text in a foreign language you encounter with your smartphone and have it instantly translated.

"Translate now has the pronunciation tool ..."

"... 1 billion users who now translate around 1 trillion words every single month ..."

20 fun facts to celebrate Google Translate turning 20 "From its beginning as an AI experiment in 2006 to supporting about 250 languages today, Translate has come a long way in two decades. Here’s how 1 billion users use Translate to learn, speak and connect more deeply than ever before."

Wednesday, December 24, 2025

Google Translate brings real-time speech translations to any headphones for over 70 languages

Good news! The Tower of Babel resurrected!

"Google Translate’s latest update brings live speech translations, originally available only on the Pixel Buds, to any headphones you want, with support for over 70 languages. It’s rolling out today [12/12/2025] in beta and just requires a compatible Android phone with the Translate app (unlike Apple’s similar feature, which requires AirPods). ..."

"... Starting today the beta is rolling out in the Translate app on Android in the U.S., Mexico, and India, works with any pair of headphones, and supports more than 70 languages. And we’ll be bringing it to iOS and more countries in 2026. ..."

Google Translate brings real-time speech translations to any headphones | The Verge

Bringing state-of-the-art Gemini translation capabilities to Google Translate (original news release) "We’re bringing Gemini’s most powerful translation capabilities to Google Translate for text, launching a beta experience for live speech-to-speech translations with headphones, and adding new languages to the app for practice and skill building."

Credits: Last Week in AI

Saturday, April 12, 2025

All the hype about AI in one image

I thought, current large language models are multilingual up to hundreds of languages? 😊 What a disappointment!

This was my prompt today!



Wednesday, October 02, 2024

Nice example of Google translate junk. Beware!

Unfortunately, I notice too many times of those cases!!! It is also very annoying that Google translate usually provides only one translation!!!! Grrrr!

What a joke!!! Google presented a primitive word for word translation in the advanced age of machine translation and AI! Dunkelziffer (a concatenation of two German words dunkel and Ziffer) means something like unreported or under reported cases/figures!



Friday, September 06, 2024

Comment on (Perhaps) Beyond Human Translation: Harnessing Multi-Agent Collaboration for Translating Ultra-Long Literary Texts

A very interesting paper!

It appears effective machine translation of literary books is within reach! Cooperative multi-agents can do the job!

From the abstract:
"Recent advancements in machine translation (MT) have significantly enhanced translation quality across various domains. However, the translation of literary texts remains a formidable challenge due to their complex language, figurative expressions, and cultural nuances. In this work, we introduce a novel multi-agent framework based on large language models (LLMs) for literary translation, implemented as a company called TransAgents, which mirrors traditional translation publication process by leveraging the collective capabilities of multiple agents, to address the intricate demands of translating literary works. To evaluate the effectiveness of our system, we propose two innovative evaluation strategies: Monolingual Human Preference (MHP) and Bilingual LLM Preference (BLP). MHP assesses translations from the perspective of monolingual readers of the target language, while BLP uses advanced LLMs to compare translations directly with the original texts. Empirical findings indicate that despite lower d-BLEU scores, translations from TransAgents are preferred by both human evaluators and LLMs over human-written references, particularly in genres requiring domain-specific knowledge. We also highlight the strengths and limitations of TransAgents through case studies and suggests directions for future research."

[2405.11804] (Perhaps) Beyond Human Translation: Harnessing Multi-Agent Collaboration for Translating Ultra-Long Literary Texts

Thursday, August 15, 2024

Machine Translation Goes Agentic for Translating Ultra-Long Literary Texts

Good news!

Caveat: I have not yet read the paper.

"Literary works are challenging to translate. Their relative length, cultural nuances, idiomatic expressions, and expression of an author’s individual style call for skills beyond swapping words in one language for semantically equivalent words in another. Researchers built a machine translation system to address these issues.  ..."

From the abstract:
"Recent advancements in machine translation (MT) have significantly enhanced translation quality across various domains. However, the translation of literary texts remains a formidable challenge due to their complex language, figurative expressions, and cultural nuances. In this work, we introduce a novel multi-agent framework based on large language models (LLMs) for literary translation, implemented as a company called TransAgents, which mirrors traditional translation publication process by leveraging the collective capabilities of multiple agents, to address the intricate demands of translating literary works. To evaluate the effectiveness of our system, we propose two innovative evaluation strategies: Monolingual Human Preference (MHP) and Bilingual LLM Preference (BLP). MHP assesses translations from the perspective of monolingual readers of the target language, while BLP uses advanced LLMs to compare translations directly with the original texts. Empirical findings indicate that despite lower d-BLEU scores, translations from TransAgents are preferred by both human evaluators and LLMs over human-written references, particularly in genres requiring domain-specific knowledge. We also highlight the strengths and limitations of TransAgents through case studies and suggests directions for future research."






Saturday, August 10, 2024

A Love Song - Anne Murray (live)

Enjoy! 
Turn on closed captioning and you will witness how machine translation frequently fails. So much for AI hype! 😊

Thursday, June 27, 2024

Google Translate adds support for 110 languages, representing 614 million speakers

Amazing stuff! The Tower of Babel reincarnated! Anyone can talk to anyone in any language!

The Google blog post forgot to mention how many of the goal of 1,000 languages their translation service now supports!

"Google said today that it is adding support for 110 languages to its translation service. The company has used its PaLM 2 AI model to power translations. 

These languages include Afar, Cantonese, Manx, Nko, Punjabi (Shahmukhi), Tamazight (Amazigh) and Tok Pisin. The company said the newly added languages represent over 614 million speakers or roughly 8% of the world’s population. 

Google noted that these languages are in different stages of usage. While some of them have 100 million speakers, some of them don’t have any active speakers — but people are working to preserve those languages. ..."

Google Translate adds support for 110 languages, representing 614 million speakers | TechCrunch

110 new languages are coming to Google Translate (original announcement) We’re using AI to add 110 new languages to Google Translate, including Cantonese, NKo and Tamazight.




Thursday, June 13, 2024

Promoting Endangered Languages with AI - IBM Research

Food for though!

To ponder:
  1. Would the world not be a better place and more peaceful if we did not speak so many different languages? 
  2. Or will machine translation really make us all understand each other better all the time at any time? We'll find out!
  3. Will we be able to preserve some of the unwritten wisdom of minority people?

"Nearly half of the world’s 7,000 languages are expected to go extinct by 2100. Now, large language models are being trained to document threatened languages and help encourage their use in everyday life." (Source)

"Today, Nheengatu, which translates to "good language" in English, is down to about 20,000 speakers and is on UNESCO’s list of “severely endangered” languages, meaning it’s no longer being passed down to most children. And it’s not the only language with a diminishing number of speakers. Nearly half of the world’s 7,000 languages are expected to go extinct by 2100. To try and stave off their decline, the United Nations has declared 2022-2032 the International Decade of Indigenous Languages. ...
Working with students at Brazil’s University of Campinas, IBM researchers recently built a prototype for an AI-powered writing assistant in Nheengatu. It’s a language that many students grew up hearing their parents and grandparents speak, but that most never learned to read or write. ..."

Promoting Endangered Languages with AI - IBM Research From about 1600 to 1800, Nheengatu was the lingua franca of the Amazon, one of 800 Indigenous languages spoken during the Portuguese colonization of Brazil. Linguistically, Nheengatu evolved as a way for Indigenous communities to communicate with the Portuguese — and each other.

Thursday, March 28, 2024

How Japan deals with its birth rate hitting a record low. Robots instead of babies

There are also more diapers sold for adults than for babies in Japan.
For know maybe machine translation for health care & senior care workers from e.g. the Philippines are more urgent.

A book about the subject by the author interviewed in this video



Wednesday, February 07, 2024

AI breakthrough enables scientists to read whole Roman scrolls once buried by the eruption of volcano Mount Vesuvius

Very recommendable! Amazing stuff! What an exciting effort! 2024's challenge is to go from reading a few passages to entire scrolls.

"Keep scrolling" 😊

"... The goal was to take computed tomography (CT) scans of what are known as the Herculaneum scrolls as well as machine-learning-based software and put these in the hands of tech-savvy sleuths from around the world in hopes someone could read the scrolls without even touching them. ...
While about 65 feet of hot ash might seem like the worst possible outcome for papyrus, the heat carbonized the scrolls, preserving them from the natural deteriorating effects of air. ..."

AI breakthrough enables scientists to read Roman scrolls once buried by Mount Vesuvius | ZDNET ZDNET In Depth: Go inside the 20-year journey to decode the Herculaneum scrolls, carbonized by a historic volcanic eruption two thousand years ago - and unreadable until now.

Resurrect an ancient library from the ashes of a volcano. Win $100,000. Make History. The Vesuvius Challenge is a machine learning and computer vision competition that in 2023 cracked the riddle of the Herculanum Papyri & awarded over $1,000,000 in prizes. 2024's challenge is to go from reading a few passages to entire scrolls.

Herculaneum scroll with red laser lines being scanned at Institut de France by Brent Seales and his team.

Brent Seales and Seth Parker (Digital Restoration Initiative project lead) scanning a replica of the Herculaneum scroll on the University of Kentucky campus.

Vesuvius Challenge




Saturday, October 07, 2023

AI translates 5,000-year-old cuneiform tablets into English

Good news! I am sure we will learn more interesting things about our ancestors. This should be interesting reading. Probably many of these tablets contain very ordinary information, but who knows.

"Cuneiform is one of the earliest writing systems in human history. Archaeologists have traced it back to 3400 BC, a whopping 5,400 years ago. It also lasted for a pretty long time, over 3,000 years. Researchers have found thousands of texts written in cuneiform in the Sumerian and Akkadian languages -- now, they've trained a neural network that can translate these texts into English effortlessly. ..."

From the abstract:
"Cuneiform is one of the earliest writing systems in recorded human history (ca. 3,400 BCE–75 CE). Hundreds of thousands of such texts were found over the last two centuries, most of which are written in Sumerian and Akkadian. We show the high potential in assisting scholars and interested laypeople alike, by using natural language processing (NLP) methods such as convolutional neural networks (CNN), to automatically translate Akkadian from cuneiform Unicode glyphs directly to English (C2E) and from transliteration to English (T2E). We show that high-quality translations can be obtained when translating directly from cuneiform to English, as we get 36.52 and 37.47 Best Bilingual Evaluation Understudy 4 (BLEU4) scores for C2E and T2E, respectively. For C2E, our model is better than the translation memory baseline in 9.43, and for T2E, the difference is even higher and stands at 13.96. The model achieves best results in short- and medium-length sentences (c. 118 or less characters). As the number of digitized texts grows, the model can be improved by further training as part of a human-in-the-loop system which corrects the results."

AI translates 5,000-year-old cuneiform tablets into English

Wednesday, August 09, 2023

K-pop's biggest music label HYBE looks to simultaneously release songs in several languages

It is not the latest news, but it is amazing nevertheless! A little machine translations goes a long way.

"... The technology enabled HYBE ..., South Korea's largest music label, to release a track by singer MIDNATT in six languages – Korean, English, Spanish, Chinese, Japanese and Vietnamese in May. ..."

K-pop's biggest music label HYBE looks to lift language barrier with AI | Reuters

Tuesday, July 11, 2023

Translating one of the world's oldest languages Akkadian to English with neural machine translation

Good news! This should be a tremendous boost! The Rosetta Stone reinvented!

From the abstract:
"Cuneiform is one of the earliest writing systems in recorded human history (ca. 3,400 BCE–75 CE). Hundreds of thousands of such texts were found over the last two centuries, most of which are written in Sumerian and Akkadian. We show the high potential in assisting scholars and interested laypeople alike, by using natural language processing (NLP) methods such as convolutional neural networks (CNN), to automatically translate Akkadian from cuneiform Unicode glyphs directly to English (C2E) and from transliteration to English (T2E). We show that high-quality translations can be obtained when translating directly from cuneiform to English, as we get 36.52 and 37.47 Best Bilingual Evaluation Understudy 4 (BLEU4) scores for C2E and T2E, respectively. For C2E, our model is better than the translation memory baseline in 9.43, and for T2E, the difference is even higher and stands at 13.96. The model achieves best results in short- and medium-length sentences (c. 118 or less characters). As the number of digitized texts grows, the model can be improved by further training as part of a human-in-the-loop system which corrects the results."

Translating Akkadian to English with neural machine translation | PNAS Nexus | Oxford Academic (open access)

Credits (partial): Last Week in AI

Fig. 1 A schematic overview of the NMT model's pipeline.


Monday, October 24, 2022

A new AI-powered speech translation system for a primarily oral language

Good news! This is a breakthrough! Coming closer to the single language before the mythical Tower of Babel!

"Until now, AI translation has mainly focused on written languages. Yet nearly half of the world’s 7,000+ living languages are primarily oral and do not have a standard or widely used writing system. This makes it impossible to build machine translation tools using standard techniques, which require large amounts of written text in order to train an AI model. To address this challenge, we've built the first AI-powered translation system for a primarily oral language, Hokkien. Hokkien is widely spoken within the Chinese diaspora but lacks a standard written form. Our technology allows Hokkien speakers to hold conversations with English speakers. ..."

Meta’s new AI-powered speech translation system for Hokkien pioneers a new approach for an unwritten language