Showing posts with label linguistics. Show all posts
Showing posts with label linguistics. Show all posts

Saturday, July 25, 2026

Language diversity: A historical context for language loss for the past 6,000 years or more

Over human history, many languages have disappeared and continue to disappear!

Why should the evolution of human languages be so different than the biological evolution? Do we need about 7,000 existing and actively spoken languages on our planet?

Hopefully, the Tower of Babel will go away in the 21st century. The world would be a much better place if we all speak the same language (with or without machine translation)!

Is this article another way of trying to indoctrinate with the DEI (Diversity, Equity, and Inclusion) ideology?

The approach taken in this study, i.e. a kind of extrapolation from present-day hunter-gatherers, is certainly questionable if not junk! This assumption is possibly a very strong inductive bias! What if e.g. the first humans spoke only one or a few languages and that for a long time before language diversity emerged maybe with the geographic spread of humans?

Notice that this study does not blame colonialism, but the preceding multinational empire building and expansion for the loss of languages.

"Humanity’s linguistic diversity is in trouble [???]. Roughly half of the 7500 languages signed or spoken today are endangered [???], and at least four go extinct every year. But while preserving the languages that do persist is a top priority, a new study in Science suggests that we are already more than one millennium beyond humanity’s linguistic “golden age.”

Linguists have long debated when the most languages were in use: Some have argued that the total number of languages remained roughly steady from the last ice age until modern-day European colonialism, others suggested that diversity has steadily declined since the dawn of agriculture about 12,000 years ago, and a third camp theorized a high point sometime before large states became established 4000 years ago.
Since writing was not developed until around 6000 years ago, however, understanding dynamics of early language remained challenging.
So, a team of anthropologists applied data on the languages, distributions, and behaviors of present-day hunter-gatherers to estimates of the total human population size and the size and number of individual cultures over the last 12,000 years. This allowed them to create a model of the total number of languages in use at any point since the last ice age ended.

The researchers showed that linguistic diversity rose steadily for around 10,000 years, reaching a “golden age” of tens of thousands of languages around 1000 to 3000 years ago before entering a language loss free-fall that continues to this day. The results reveal that the language crisis is not the result of modern colonialism but rather the culmination of millennia of cultural exchange and homogenization."

"... The golden age was followed by a period of rapid decline in linguistic diversity that coincided with the rise of large states and multinational empires, such as the Roman Empire, the researchers found. The finding challenges a commonly held view that widespread language extinction began later, about 500 years ago, with the onset of European colonial expansion. ..."

"To the Point
  • Fewer languages before agriculture: The researchers estimate that between approximately 4,500 and 6,200 languages were spoken at the beginning of the Holocene, around 12,000 years ago - probably fewer than the roughly 7,500 languages spoken and signed today.
  • A linguistic “golden age”: As the global human population grew, language diversity also increased. The models suggest that tens of thousands of languages may have existed between 1,000 and 3,000 years ago.
  • A deep history of language loss: The decline in linguistic diversity appears to have begun with the expansion of large states and empires, long before modern European colonialism.
  • Today’s languages are survivors [???]: The languages spoken today represent a small and historically biased sample of past linguistic diversity [???], with important implications for how researchers explain global patterns in language and culture.
...
They began with ethnographic data from 171 hunter-gatherer and fisher societies whose traditional subsistence and mobility had not been profoundly transformed by contact with food-producing populations. These data allowed the researchers to estimate the likely distribution of ethnolinguistic group sizes near the beginning of the Holocene.

The team then combined these estimates with reconstructions suggesting that the global human population 12,000 years ago was between approximately 4.4 and seven million. Assuming that ethnolinguistic groups at that time generally corresponded to distinct languages, the models produced an early-Holocene estimate centred on roughly 4,500 to 6,200 languages, although broader plausible estimates ranged from around 3,300 to 7,800. ..."


From the editor's summary and abstract:
"Editor’s summary
More than four languages are lost every year, and this rate is predicted to increase. Blasi et al. investigated how language diversity has changed over the Holocene using a Bayesian modeling approach, ethnographic data, and estimates of the total human population size over time.
They assumed a maximum size of ethnolinguistic groups and that the number of different ethnolinguistic groups is related to the number of languages. Their findings suggested that language loss is not a new phenomenon, and that language diversity has been decreasing rapidly over the past one to three millennia following a peak of an order of magnitude higher diversity than we see today. ...

Abstract
Characterizing the factors that have shaped linguistic diversity is fundamental for understanding human history, culture, and cognition.
In this study, we combined statistical and social computational modeling, ethnographic data, and paleodemographic inference to model trajectories of global linguistic diversity.
Before the onset of plant and animal domestication, the number of languages was smaller than it is today (4500 to 6000 compared with 7500).
Subsequent increases in global population precipitated increased linguistic diversity.
We uncovered a linguistic “golden age” with tens of thousands of languages 3000 to 1000 years ago. Great loss of linguistic diversity did not begin with recent colonial expansion but as multinational empires first spread along with their languages, pathogens, and cultures. Thus, extinction has likely played a much greater role in shaping linguistic and cultural diversity than previously thought."

ScienceAdviser


Study uncovers lost ‘golden age’ of languages (original news release) "A new study coauthored by Yale linguist Claire Bowern suggests that tens of thousands of languages were spoken between 1,000 and 3,000 years ago."






Map of currently spoken and signed languages of the world, colored by endangerment status.


Saturday, July 11, 2026

Logic and language is separated in the brain

Amazing stuff!

"... In research ... researchers ... have shown that people can perform well on tasks that require logical reasoning even if their language abilities are severely impaired. What’s more, brain imaging shows that language-processing parts of the brain are not called on for logical reasoning. ..."

From the significance and abstract:
"Significance
Which cognitive mechanisms allow humans to reason logically, to understand whether a conclusion follows from the premises? Are they the same ones that allow the assembly of words into structured representations?
Scholars have debated for millennia whether logical reasoning is inextricably tied to natural language, or instead relies on a distinct “language of thought” (LOT).
Using fMRI in healthy adults and evaluating logical ability in individuals with severe aphasia, we find that distinct neural systems support language processing vs. logical (inductive and deductive) reasoning. These results suggest that, at least in mature brains, language processing does not underpin logical inference, perhaps due to the distinct representational format of the logical LOT.

Abstract
Humans are endowed with a powerful capacity for inductive and deductive logical thought: we easily form generalizations based on a few examples and draw conclusions from known premises.
Humans also arguably have the most sophisticated communication system in the animal kingdom: natural language allows us to express complex and structured meanings.
Some have therefore argued for a tight relationship between complex thought and language, postulating that reasoning, including logical reasoning, relies on linguistic representations.
We systematically investigated the relationship between logical reasoning and language using two complementary approaches.
First, we used noninvasive brain imaging (fMRI) to examine neural activity as healthy adults engaged in logical reasoning tasks.
And second, we behaviorally evaluated logical abilities in individuals with extensive lesions to the language brain areas and consequent severe linguistic impairment.
Our findings reveal that the language brain network is not engaged during logical reasoning, and patients with severe aphasia exhibit intact performance on logic tasks.
Instead, inductive reasoning recruits the domain-general multiple demand network implicated broadly in goal-directed behaviors,
whereas deductive reasoning draws on brain regions that are distinct from both the language and the multiple demand networks.
Together, these results indicate that linguistic representations are neither utilized nor required for inductive or deductive logical reasoning."

Separating logic and language | MIT News | Massachusetts Institute of Technology "Neuroscientists find logical reasoning does not involve language-processing parts of the brain."



A functional brain scan of a neurotypical participant in a new study shows a distinct separation between logic (green) and language (red/yellow) activations.






Sunday, May 17, 2026

China: Are the Chinese people abandoning the Chinese characters?

It will probably take several decades if not longer! Slowly, but surely? Caution: speculation!

When Chinese people type messages on their smartphone using e.g. the WeChat app they use a QWERTY keyboard to find Chinese characters. Just an observation!

If true, then this would be one of the greatest achievements in human history. Let's get rid of the Tower of Babel!

Chinese characters are beautiful calligraphy and brush artwork, but hard to learn e.g. for Westerners.

P.S. I just blogged here about one of the major mistakes of Israel since 1948. They missed the opportunity to give up the traditional Hebrew written language right from the start.


100 Basic Chinese Characters (Source)



Monday, May 11, 2026

What is perhaps the most gravest mistake of Israel since 1948?

Why does Israel not give up the Hebrew written language for the Latin/Roman script?

Israel could have made a contribution to end the global Tower of Babel! What a missed opportunity since 1948!

Set an example for the entire Middle East and other Arab countries!

Perhaps, then China and other Asian countries could be more willing to do the same! What about India?

No disrespect intended for the very old Hebrew language, impressive ancient history and traditions!


The Tower of Babel by Pieter Bruegel the Elder c. 1563 AD (Source)


Sunday, February 15, 2026

Florida Becomes Only the Fourth State to Offer English-Only Driver’s License Exams

Is this controversial or not? Until now the only official language in the US is english. 

I bet most other US states do not officially offer multilingual exams either.

According to Google search there are about only a total of 19 countries around the world having two or more official languages. 

According to Google search only three US states have officially more than one language (Alaska, Hawaii, South Dakota).

"... Recently, the Florida Highway Safety and Motor Vehicles office announced that all driver’s license exams will be administered in English only, including all knowledge exams, commercial learner’s permit, and commercial driver’s license exams. All road signs are in English, so it’s common sense for all driving exams to only be administered in English. ..."

Florida Will Now Administer English-Only Driver’s License Exams

Sunday, February 01, 2026

Mafiöse Strukturen in der Windkraft-Lobby. Wirklich!

Deutsche Sprache, schwierige Sprache! Mafiose, nicht Mafiöse Strukturen. Die Mafia ist keine Möse!

Wieder so ein ätzender schmal  Format Kurzfilm von YouTube!

Mafiöse Strukturen in der Windkraft-Lobby - YouTube

Friday, November 28, 2025

Enduring patterns in world's languages: One-third of grammatical 'universals' stand up to rigorous testing

Amazing stuff!

"... An international team ... used Grambank, the world's most comprehensive database of grammatical features, to test 191 proposed universals across more than 1,700 languages. Traditionally, linguists have attempted to circumvent the genealogical and geographic non-independence of languages by sampling widely separated languages.

However, sampling can fail to remove all dependencies, reduce statistical power and does not identify historical pathways. The Bayesian spatio-phylogenetic analyses used by the authors accounted for both the genealogical and geographic non-independence of languages—a level of statistical rigor rarely achieved in previous work. ...

The study found strong evidence for patterns involving word order (such as whether verbs precede or follow objects) and hierarchical universals (such as dependencies in which arguments are marked in grammatical agreement). The patterns predicted by the supported universals have evolved repeatedly across the world's languages, suggesting deep-rooted constraints in how humans structure communication. ..."

"To the Point
  • Linguistic universals: Of the 191 proposed linguistic universals, about one-third are statistically supported across more than 1,700 languages.
  • A wealth of data and state-of-the-art statistical methods: Using Grambank and Bayesian statistical models that control for genealogical and geographic influences, the strongest evidence emerges for patterns of word order and hierarchical agreement.
  • Evolutionary framework: The repeated evolution of these patterns suggests there are shared cognitive and communicative constraints, which narrows the search for truly universal features of human language.
..."

From the abstract:
"Human languages show astonishing variety, yet their diversity is constrained by recurring patterns. Linguists have long argued over the extent and causes of these grammatical ‘universals’.
Using Grambank—a comprehensive database of grammatical features across the world’s languages—we tested 191 proposed universals with Bayesian analyses that account for both genealogical descent and geographical proximity.
We find statistical support for about a third of the proposed linguistic universals.
The majority of these concern word order and hierarchical universals: two types that have featured prominently in earlier work.
Evolutionary analyses show that languages tend to change in ways that converge on these preferred patterns.
This suggests that, despite the vast design space of possible grammars, languages do not evolve entirely at random.
Shared cognitive and communicative pressures repeatedly push languages towards similar solutions."

Enduring patterns in world's languages: One-third of grammatical 'universals' stand up to rigorous testing

Enduring patterns in the world’s languages (original news release) "New study finds one-third of grammatical ‘universals’ stand up to rigorous testing"


The evolution of a word-order universal on the global language tree. In our analysis of the universal “With overwhelmingly greater than chance frequency, languages with normal subject–object–verb order are postpositional”, the absence or presence of the two features defines the ‘state’: state 11 (red) is the prediction made by the universal; in state 00 (black), both features are absent; in states 01 (orange) and 10 (light blue), one feature is absent and the other is present. The ancestral state reconstruction shows that in multiple language families and areas, pathways of language change repeatedly lead to the predicted outcome. 


Fig. 2: Median natural log BF and their 95% HDI from the BayesTraits analyses showing support for co-evolutionary models.


Tuesday, November 18, 2025

Aiong Taigi is an American YouTuber based in Taiwan who promotes and helps others to learn the native Taiwanese language

 Good news!

Google: "Aiong Taigi is an American YouTuber based in Taiwan known for creating content that promotes and helps others learn the Taiwanese language, or Tâi-gí. Through his videos, he shares his experience learning Taiwanese as a foreigner, discusses Taiwan's language policies, and encourages the use of the language in everyday life, as explained in sources like this Instagram post and this Spotify podcast episode. He has gained a significant following and has collaborated with other Taiwanese language content creators, notes this Reddit post. 

Focus on language preservation: Aiong Taigi's YouTube channel is dedicated to preserving and promoting the Taiwanese language."

Global Taiwan Institute: "YouTube Content Creater "Aiong Taigi" on Preserving Taiwanese and Learning the Inner Language

Taiwanese (台語), also known as Tâi-gí, is the most widely spoken native language in Taiwan. During both the Japanese colonial period and subsequent martial law era (1949-1987), Taiwanese and other indigenous languages faced severe repression. While democratization has restored the freedom to speak Taiwanese openly, new challenges have emerged. As fewer young people speak the language, the question looms: how can this once-flourishing mother tongue survive in an era of globalization?"

阿勇台語 Aiong Taigi (his YouTube channel)

Saving Tâi-gí: Taiwan’s Largest Heritage Language

Tuesday, October 28, 2025

Rapid Support Forces (RSF) or Reporters sans frontières (RSF)

One abbreviation for two very different things!

Hint: RSF translates to Reporters Without Borders. The The Rapid Support Forces (RSF) militia allegedly are committing grave atrocities in Darfur.

Monday, October 27, 2025

Saviours of Sanskrit during British colonialism in India

Recommendable!

"‘Pundits’ kept Sanskrit scholarship alive in remote settlements as British control swept across India, a major new research project will show. The largely forgotten literary figures and their works – ranging from erotic plays to legal treatises – are neglected treasures of Indian intellectual achievement, its researchers argue.

English speakers are familiar with the word ‘pundit’ but few know that it comes from the Sanskrit word paṇḍita, meaning ‘learned’. Now a Cambridge University-led project is going in search of the pundits, Brahmin scholars, who kept writing poems, plays, philosophy, theology, legal texts and other forms of literature in Sanskrit as Britain tightened its grip on India.

It has long been assumed that the expansion of British power in India from the seventeenth century steadily suffocated Sanskrit scholarship. But the experts behind an ambitious new project argue that the two centuries leading up to the establishment of the British Raj in 1858 were, in fact, a golden age of Sanskrit intellectual thought, literature, and arts. They point to the scholarly activities of hundreds of pundits dispersed across the Indian countryside in Brahmin settlements (agrahāra) and monasteries (maṭha). ..."

Saviours of Sanskrit "Indian literary genius survived British imperialism in forgotten villages, new research reveals."

Tuesday, October 21, 2025

What is a holistic approach?

A popular term of art or art for art's sake! Sounds impressive, but may indicate cluelessness or pretension!

So next time you come across someone using this term beware! 😊

Thursday, July 17, 2025

Ancient DNA solves mystery of Hungarian, Estonian, Finnish language origins

Amazing stuff! Members of two major language families were moving in opposite directions influencing each other.

"Where did Europe’s distinct Uralic family of languages — which includes Hungarian, Finnish, and Estonian — come from? New research puts their origins a lot farther east than many thought.

The analysis ... integrated genetic data on 180 newly sequenced Siberians with more than 1,000 existing samples covering many continents and about 11,000 years of human history. The results ... identify the prehistoric progenitors of two important language families, including Uralic, spoken today by more than 25 million people. ...

The study finds the ancestors of present-day Uralic speakers living about 4,500 years ago in northeastern Siberia, within an area now known as Yakutia. ...

Proto-Uralic speakers overlapped in time with the Yamnaya, the culture of horseback herders credited with transmitting Indo-European across Eurasia’s grasslands. A pair of recent papers ... zeroed in on the Yamnaya homeland, showing it was mostly likely within the current borders of Ukraine just over 5,000 years ago. 

“We can see these waves going back and forth — and interacting — as these two major language families expanded,”  ... “Just as we see Yakutia ancestry moving east to west, our genetic data show Indo-Europeans spreading west to east.” ...

Previous studies established that Finns, Estonians, and other Uralic-speaking populations today share an Eastern Eurasian genetic signature. Ancient DNA researchers ruled out the region’s best-known archaeological cultures from contributing to the Uralic expansion ..."

"... The team analyzed genomes from 180 ancient individuals from northern Eurasia, dated between 17,000 and 3,000 years ago. Their findings identify two ancestral populations that gave rise to these two language families: one from the Lena River Basin in eastern Siberia, which contributed significantly to nearly all modern Uralic-speaking populations, and another from the Baikal region in southern Siberia, associated with the genetic legacy of the Ket people.  ..."

From the abstract:
"The North Eurasian forest and forest-steppe zones have sustained millennia of sociocultural connections among northern peoples, but much of their history is poorly understood. In particular, the genomic formation of populations that speak Uralic and Yeniseian languages today is unknown.
Here, by generating genome-wide data for 180 ancient individuals spanning this region, we show that the Early-to-Mid-Holocene hunter-gatherers harboured a continuous gradient of ancestry from fully European-related in the Baltic, to fully East Asian-related in the Transbaikal.
Contemporaneous groups in Northeast Siberia were off-gradient and descended from a population that was the primary source for Native Americans, which then mixed with populations of Inland East Asia and the Amur River Basin to produce two populations whose expansion coincided with the collapse of pre-Bronze Age population structure.
Ancestry from the first population, Cis-Baikal Late Neolithic–Bronze Age (Cisbaikal_LNBA), is associated with Yeniseian-speaking groups and those that admixed with them, and 
ancestry from the second, Yakutia Late Neolithic–Bronze Age (Yakutia_LNBA), is associated with migrations of prehistoric Uralic speakers.
We show that Yakutia_LNBA first dispersed westwards from the Lena River Basin around 4,000 years ago into the Altai-Sayan region and into West Siberian communities associated with Seima-Turbino metallurgy—a suite of advanced bronze casting techniques that expanded explosively from the Altai.
The 16 Seima-Turbino period individuals were diverse in their ancestry, also harbouring DNA from Indo-Iranian-associated pastoralists and from a range of hunter-gatherer groups. Thus, both cultural transmission and migration were key to the Seima-Turbino phenomenon, which was involved in the initial spread of early Uralic-speaking communities."

Ancient DNA solves mystery of Hungarian, Finnish language origins — Harvard Gazette "Parent emerged over 4,000 years ago in Siberia, farther east than many thought, then rapidly spread west"



Map of all the sites that are sources of samples used in the study.


Sunday, June 15, 2025

On Improving large language models with concept-aware fine-tuning. Really!

The abstract of this new paper suggests that ML & AI researchers are reinventing the wheel!

The long existing ambiguity about what tokens are in natural language processing is not helpful! Breaking up words into chunks may even be counterproductive! What unit in the spectrum between a single character, a word chunk, a whole word, a sentence or even a whole paragraph etc. should be used for the training of language models? A combination of fine-grained and coarse-grained units or hierarchical-level units etc.

Caveat: I have not read the paper.

From the abstract:
"Large language models (LLMs) have become the cornerstone of modern AI. However, the existing paradigm of next-token prediction fundamentally limits their ability to form coherent, high-level concepts, making it a critical barrier to human-like understanding and reasoning.
Take the phrase "ribonucleic acid" as an example: an LLM will first decompose it into tokens, i.e., artificial text fragments ("rib", "on", ...), then learn each token sequentially, rather than grasping the phrase as a unified, coherent semantic entity. This fragmented representation hinders deeper conceptual understanding and, ultimately, the development of truly intelligent systems.
In response, we introduce Concept-Aware Fine-Tuning (CAFT), a novel multi-token training method that redefines how LLMs are fine-tuned. By enabling the learning of sequences that span multiple tokens, this method fosters stronger concept-aware learning.
Our experiments demonstrate significant improvements compared to conventional next-token finetuning methods across diverse tasks, including traditional applications like text summarization and domain-specific ones like de novo protein design.
Multi-token prediction was previously only possible in the prohibitively expensive pretraining phase; CAFT, to our knowledge, is the first to bring the multi-token setting to the post-training phase, thus effectively democratizing its benefits for the broader community of practitioners and researchers. ..."

[2506.07833] Improving large language models with concept-aware fine-tuning

Thursday, June 12, 2025

Trockenfrüchte oder Dörrobst

Schweizerdeutsch (auch Schwizerdütsch) ist manchmal köstlich und pikant!

Credits: Ein Müesli mit viel Trockenobst – das schmeckt gut. Doch es kann der Leber schaden "Weshalb man getrocknete Feigen, Rosinen und anderes Dörrobst nur in geringen Mengen essen sollte und warum frisches Obst gesünder ist."

Saturday, May 03, 2025

When will the Chinese people voluntarily adopt the Western alphabet?

Probably not in my lifetime! Food for thought! Voluntary is the operative word here!

One thing seems to be quite sure, outside of Asia most people will probably never adopt the Chinese written language/writing system.

I am not sure machine translation and or AI, now and in the future, can fully make up for the differences in written languages and its consequences.

Will machine learning & AI make the written language obsolete or fully automated? We probably find out in the next 10-20 years.

Would it not be nice, peace and community promoting if all people used the Western alphabet?

More human progress please!

How about e.g. Chinese street signs?




Saturday, April 12, 2025

There is apparently no proper English word for the German word Wortschöpfung

This is interesting! The official, common English translation for this German word is neologism. Sounds very exciting and academic! What a turn/put off! Caution: irony!

The English word wordsmith does not even come close. Until now, I thought wordsmith is the closest translation. I was wrong.

Deutschland, das Land der Denker und Dichter! (Germany, the land of thinkers and poets). Notice I reversed the expression Dichter und Denker.

Saturday, April 05, 2025

Uniquely human language capacity found in wild bonobos apes

Amazing stuff! This could be strong hint that animals are much more capable of communication than previously thought!

Wow, this field research was conducted in the long time war torn Democratic Republic of Congo!

"... A bonobo dictionary
In a first step, the researchers applied a method developed by linguists to quantify the meaning of human words. “This allowed us to create a bonobo dictionary of sorts – a complete list of bonobo calls and their meaning,” says Mélissa Berthet, a postdoctoral researcher at the Department of Evolutionary Anthropology of UZH and lead researcher of the study. “This represents an important step towards understanding the communication of other species, as it is the first time that we have determined the meaning of calls across the whole vocal repertoire of an animal.”

Compositionality is not unique to humans

After determining the meaning of single bonobo vocalizations, the researchers then moved on to investigating call combinations, using another approach borrowed from linguistics. “With our approach, we were able to quantify how the meaning of bonobo single calls and call combinations relate to each other,” ... The researchers found numerous call combinations whose meaning was related to the meaning of their single parts, a key hallmark of compositionality.  Furthermore, some of the call combinations bore a striking resemblance to the more complex nontrivial compositional structures in human language. “This suggests that the capacity to combine call types in complex ways is not as unique to humans as we once thought,” ..."

"After determining the meaning of single bonobo vocalizations, the researchers then moved on to investigating call combinations, using another approach borrowed from linguistics. “With our approach, we were able to quantify how the meaning of bonobo single calls and call combinations relate to each other,” says Simon Townsend, UZH Professor and senior author of the study. The researchers found numerous call combinations whose meaning was related to the meaning of their single parts, a key hallmark of compositionality.  Furthermore, some of the call combinations bore a striking resemblance to the more complex nontrivial compositional structures in human language. “This suggests that the capacity to combine call types in complex ways is not as unique to humans as we once thought,” ... "

From the editor's summary and abstract:
"Editor’s summary
One hallmark of human language is the combination of elements into larger meaningful structures, a pattern referred to as compositionality. Compositionality can be trivial, in which the two parts are added together to give meaning, or nontrivial, in which the meaning in one part modifies the meaning in the other. Recent research has found the presence of trivial compositionality across a number of species, but it has been argued that nontrivial compositionality is unique to humans. 
Berthet et al. used a large dataset of bonobo vocalizations in conjunction with a distributional semantics approach and found that not only did they display compositionality, but three of the four types were nontrivial. ...

Abstract
Compositionality, the capacity to combine meaningful elements into larger meaningful structures, is a hallmark of human language. Compositionality can be trivial (the combination’s meaning is the sum of the meaning of its parts) or nontrivial (one element modifies the meaning of the other element).
Recent studies have suggested that animals lack nontrivial compositionality, representing a key discontinuity with language.
In this work, using methods borrowed from distributional semantics, we investigated compositionality in wild bonobos and found that not only does each call type of their repertoire occur in at least one compositional combination, but three of these compositional combinations also exhibit nontrivial compositionality.
These findings suggest that compositionality is a prominent feature of the bonobo vocal system, revealing stronger parallels with human language than previously thought."

‘Uniquely human’ language capacity found in bonobos | Science | AAAS "In a first, researchers have seen a nonhuman animal combine different calls to make new meanings"

Bonobos Combine Calls in Similar Ways to Human Language (original news release) "Bonobos – our closest living relatives – create complex and meaningful combinations of calls resembling the word combinations of humans. This study ... challenges long-held assumptions about what makes human communication unique and suggests that key aspects of language are evolutionary ancient."




Monday, March 17, 2025

When did human language emerge? Genomic evidence may have an answer

Amazing stuff! So Homo sapiens was speechless (or in linguistic  incapacity) for about the first 100,000 years? Hard to stomach! 😀

This research seems to be a bit to speculative for my taste! I guess the crucial term here is "linguistic capacity".

"A new survey of genomic evidence suggests our unique language capacity was present at least 135,000 years ago. Subsequently, language might have entered social use 100,000 years ago.

Our species, Homo sapiens, is about 230,000 years old. Estimates of when language originated vary widely, based on different forms of evidence, from fossils to cultural artifacts. The authors of the new analysis took a different approach. ...

“Every population branching across the globe has human language, and all languages are related.” Based on what the genomics data indicate about the geographic divergence of early human populations, he adds, “I think we can say with a fair amount of certainty that the first split occurred about 135,000 years ago, so human language capacity must have been present by then, or before.” ..."

From the abstract:
"Recent genome-level studies on the divergence of early Homo sapiens, based on single nucleotide polymorphisms, suggest that the initial population division within H. sapiens from the original stem occurred approximately 135 thousand years ago.
Given that this and all subsequent divisions led to populations with full linguistic capacity, it is reasonable to assume that the potential for language must have been present at the latest by around 135 thousand years ago, before the first division occurred. Had linguistic capacity developed later, we would expect to find some modern human populations without language, or with some fundamentally different mode of communication. Neither is the case. While current evidence does not tell us exactly when language itself appeared, the genomic studies do allow a fairly accurate estimate of the time by which linguistic capacity must have been present in the modern human lineage. Based on the lower boundary of 135 thousand years ago for language, we propose that language may have triggered the widespread appearance of modern human behavior approximately 100 thousand years ago."

When did human language emerge? | MIT News | Massachusetts Institute of Technology "A new analysis suggests our language capacity existed at least 135,000 years ago, with language used widely perhaps 35,000 years after that."



Table 1. Summary of estimates of divergence times for Khoisan lineage.