Showing posts with label genomics. Show all posts
Showing posts with label genomics. Show all posts

Thursday, September 10, 2026

Google unveils giant human DNA prediction database for all 9 billion possible single-letter DNA variant

Amazing stuff! Very impressive! This should significantly speed up genomic research around the world!

"Google DeepMind released AlphaGenome Atlas, a database of precomputed predictions for the molecular effects of all 9 billion possible single-letter DNA variants in the human genome.
The dataset is 1 petabyte in size, more than 30 times larger than the AlphaFold Database, and is built by running the AlphaGenome model at scale across the genome.
It ships with a new AlphaGenome Variant Impact (AVI) score that combines AlphaGenome with AlphaMissense to rank both coding and non-coding variants by likely impact, plus a catalogue of over 2,500 recurring DNA sequence motifs.
The Atlas is free for academic use via a web portal, the AlphaGenome API, and as a skill in Google Antigravity, with commercial access coming later on Google Cloud.
External collaborators have already used the Atlas to identify a disease-causing DNM1 variant and to uncover 22 percent more non-coding genetic associations in UK Biobank data. This gives researchers a precomputed shortcut to variant-impact predictions instead of running AlphaGenome themselves for every query, though DeepMind stresses it is not validated for clinical use."

"... By precomputing AlphaGenome’s predictions at scale, we have created an easily accessible resource that vastly expands the [AlphaGenome] model's reach. Just as an atlas is a collection of maps, linking together features of the land like altitude and location, AlphaGenome Atlas charts the molecular effects of DNA variants across the genome. ...

AlphaGenome Atlas provides several powerful, interconnected resources, allowing researchers to link variants directly to the functional DNA sequences they disrupt.

  • Molecular effect predictions Atlas contains thousands of molecular effect predictions for each variant, across multiple important aspects of gene regulation, spanning hundreds of human and mouse cell types and tissues. This serves as the starting point for further resources.
  • AVI score A single number describing the impact for each genetic variant.
  • AVI feature attributions Each AVI score is also linked to distinct biological features driving it, such as the aspects of gene regulation predicted by AlphaGenome or the protein impact score from AlphaMissense.
  • DNA sequence motifs A comprehensive collection of over 2,500 recurrent DNA sequences — the "words" of the genome — and their locations.
..."

"... “This represents the first time that any researcher in the world can access a comprehensive map of the human genome and its variations by simply opening a browser,” ..."

From the abstract:
"A major challenge in genomics is deciphering the functional consequences of non-coding genetic variation.
Here we present AlphaGenome Atlas, a comprehensive resource that enables the joint interpretation and prioritization of variant effects across the entire human genome.
Using AlphaGenome, we predicted the regulatory effects across thousands of molecular phenotypes for every possible human single nucleotide variant and many observed indels. These predictions were then used to derive a unified and interpretable AlphaGenome Variant Impact (AVI) score and to map cis-regulatory motifs across the genome. AVI achieved state-of-the-art performance across diverse benchmarks with improved prioritization of deleterious noncoding variants. Application of the combined Atlas resource helped solve an epileptic encephalopathy rare disease case, increased the statistical power to detect rare non-coding variants driving population-level phenotypes, and enhanced the mechanistic interpretation of these variants.
Thus, AlphaGenome Atlas improves the prioritization and molecular interpretation of non-coding variants with genetic and clinical significance."

Data Points | The Batch | DeepLearning.AI

New Google DeepMind atlas could transform our understanding of genetic diseases "AlphaGenome Atlas predicts the effect of any single change in the fundamental structure of DNA"

AlphaGenome Atlas: A predictive map of every possible DNA letter change in the human genome "How predicting the molecular impact of every possible single-letter DNA variant in the human genome will help accelerate our understanding of biology."









Sunday, May 17, 2026

India to build a comprehensive catalogue of genetic variations and diversity of the entire population, so far uncovering 44 million new genetic variants from less then 10,000 healthy individuals

Good news! What a gigantic undertaking to the benefit of humanity!

Note the Nature journal could not resist to use the ideological term "Eurocentric"!

"GenomeIndia is a pioneering scientific project funded by the Department of Biotechnology, Ministry of Science and Technology, Government of India. The project marks a landmark collaboration of 20 academic and research institutions to drive a genomics-based health revolution for India.

The primary objective of GenomeIndia is to build a comprehensive catalogue of genetic variations that reflect the unique diversity of the Indian population. ..."

"“A genetic atlas emerging from India’s most extensive genomic sequencing exercise has revealed vast diversity in the population, with nearly 130 million genetic variants, almost a third of which have not been reported previously.

The GenomeIndia project analysed the whole genomes of 9,768 healthy people from 83 populations, uncovering 44 million variants absent from global scientific databases, including gnomAD, 1000 Genomes Project and GenomeAsia. ..."

"... The atlas captures many rare variants from the DNA of specific communities, reflecting a long history of migration, isolation, and marriage within a group (endogamy) ..."

"... Some tribal groups show genetic homozygosity more than five times higher than Ashkenazi Jewish and Finnish populations, considered global benchmarks of genetic isolation. This elevates the risk of recessive genetic diseases — conditions requiring two defective copies of the same gene — which become far more likely when both parents trace back to the same small ancestral pool.
27 of 29 tribal populations carried at least one disease-causing variant at clinically meaningful frequencies.
In one tribal group from southern India, a harmful change in the HGD gene linked to alkaptonuria — a rare metabolic disease that can cause serious damage to joints and organs — was found in 12.5% of people. It was absent from widely used reference datasets, and standard genetic tests built on existing databases likely miss it. ...

The dataset flags other such loss-of-function (LoF) variants linked to metabolic disorders — genetic alterations that inactivate or reduce a gene's functional capacity — such as in LPA and CD36, both central to lipid metabolism and cardiovascular risk. In all, it reports 15,849 high-confidence LoF variants across more than 7,000 genes, some of which heighten disease risk, while others may be neutral or even protective. ...

Among the known variants, the BCHE rs104893684 — linked to anaesthesia-related complications — stands out as far more widespread than previously thought. Detected in 29 of the 83 populations studied, it appears at frequencies above 1% in three of those groups, a level not previously recognised. Earlier evidence tied this variant to just one Indian population, prompting targeted screening advice. ..."

From the abstract:
"India, the most populous country, remains significantly underrepresented in the global genomics landscape. Previous efforts to catalog Indian genetic diversity were limited in scale, scope, and representation.
Here, we present the GenomeIndia dataset, comprising whole genome sequences of 9,768 healthy individuals from 83 populations spanning the ethnolinguistic and biogeographic spectrum of India.
We identify 129.93 million high-confidence biallelic variants, 44.03 million of which are previously unreported in global databases.
In contrast to large populations that show steady population growth and internal homogeneity, we observe low effective population sizes, significant genetic drift, and profound homozygosity in small tribal groups, likely shaped by antiquity, isolation, and endogamy.
We report multiple population-specific pharmacogenomic and deleterious variants, necessitating the integration of local genetic architecture and the inclusion of underrepresented South Asian genomes in global reference resources. Finally, we highlight the limited transferability of Eurocentric polygenic scores to Indian populations, and present an imputation panel that outperforms existing resources for both rare and common variants.
Together, our work fills a significant gap in the equity of global human genomics, and paves way for precision medicine strategies that will benefit a quarter of the world population."

India’s DNA Map Uncovers Millions of Genetic Variants - Human Progress

GenomeIndia (project website)

India’s DNA map uncovers millions of missing genetic variants "A vast study reveals deep diversity, hidden disease risks and exposes the limits of Eurocentric medicine."







Tuesday, October 28, 2025

The multiomics blueprint of the individual with the most extreme lifespan (died at age 117)

Amazing stuff! Scientists studied the genes of a woman (Maria Branyas Morera from Spain) who lived 117 years.

From the highlights and abstract:
"Highlights
• (Epi)genome, transcriptome, metabolome, proteome, and microbiome study of the oldest human
• Despite molecular hallmarks of aging, absence of major age-associated diseases
• Resilient genetic variants and low-inflammation metabolic profile reduce aging risks
• Bacteria occurrence and epigenome profile resembling younger individuals

Summary
Extreme human lifespan, exemplified by supercentenarians, presents a paradox in understanding aging: despite advanced age, they maintain relatively good health. To investigate this duality, we have performed a high-throughput multiomics study of the world’s oldest living person, interrogating her genome, transcriptome, metabolome, proteome, microbiome, and epigenome, comparing the results with larger matched cohorts.
The emerging picture highlights different pathways attributed to each process: the record-breaking advanced age is manifested by telomere attrition, abnormal B cell population, and clonal hematopoiesis, whereas absence of typical age-associated diseases is associated with rare European-population genetic variants, low inflammation levels, a rejuvenated bacteriome, and a younger epigenome. These findings provide a fresh look at human aging biology, suggesting biomarkers for healthy aging, and potential strategies to increase life expectancy. The extrapolation of our results to the general population will require larger cohorts and longitudinal prospective studies to design potential anti-aging interventions."

The multiomics blueprint of the individual with the most extreme lifespan - ScienceDirect


Maria Branyas Morera on her 117th birthday (Source)


Graphical abstract


Monday, October 20, 2025

New world record set for fastest whole human genome sequencing in less than 4 hours

Good news!

"Boston Children's Hospital, along with Broad Clinical Labs and Roche Sequencing Solutions, has demonstrated that rapid genomic sequencing and interpretation are achievable in a matter of hours. ..."

"Broad Clinical Labs, in collaboration with Roche Sequencing Solutions and Boston Children’s Hospital, today announced official recognition by Guinness World Records for achieving the fastest DNA sequencing technique to date. Leveraging Roche’s new SBX sequencing technology and a streamlined, integrated workflow, the teams completed sequencing and analysis of the whole human genome in less than 4 hours, surpassing the previous benchmark of 5 hours and 2 minutes.

The team subsequently applied this process to samples obtained from the Boston Children’s Hospital neonatal intensive care unit to demonstrate a same day workflow from blood to report. ..."

New world record set for fastest human whole genome sequencing

Broad Clinical Labs sets new GUINNESS WORLD RECORDS™ title for fastest DNA sequencing technique (original news release) "Whole human genome sequencing and analysis was completed in less than 4 hours, surpassing the previous benchmark of 5 hours and 2 minutes."



Fig. 3. Expandable nucleotide structure.




Wednesday, October 08, 2025

Genomic evolution of major malaria-transmitting mosquito species uncovered

Good news! Can we eradicate bloodsuckers like mosquitos (without great harm)?

"New research into the genetics of Anopheles funestus (An. funestus), one of the most neglected but prolific malaria-transmitting mosquitoes in Africa, has revealed how this species is evolving in response to malaria control efforts. ...

The mosquito species An. funestus is one of the most widespread in Africa. Females of the species are highly anthropophilic, meaning they are attracted to humans as a source of blood, which they need to develop their eggs. They also have a significantly longer lifespan than other malaria-transmitting mosquito species. An. funestus is also extraordinarily adaptive. For example, in some areas, it has evolved from biting indoors in the evening to biting outdoors during the day, likely in response to the use of mosquito nets. Together, these characteristics make them formidable malaria transmitters in the part of the world where malaria remains most devastating. In 2023 the World Health Organisation African Region reported 569,000 malaria-related deaths. ...

To support this, mosquito biologists across Africa together with the team at the Sanger Institute collected and sequenced the whole genomes of 656 modern An. funestus mosquito specimens that were collected from 2014 to 2018.
They also sequenced 45 historic specimens from the Natural History Museum in London and the French National Research Institute for Sustainable Development (IRD) that were collected between 1927 and 1967 to understand the evolutionary patterns and changes in the species across 16 African countries during the last century.

The team found high levels of genetic variation in An. funestus across Africa and discovered that samples originating from equatorial countries shared many genetic similarities despite covering a 4,000-kilometre range. This suggests that they likely belong to one large, interconnected population. However, some samples from this region, such as those from North Ghana and South Benin, were isolated and genetically distinct from the interconnected population. This shows some populations mix widely, while others remain separate. Such population structure has important implications for mosquito control.

By looking at the DNA of the historic samples, the team was able to highlight the fast-evolving nature of An. funestus. One key mutation linked to insecticide resistance, which is widespread among the modern populations, was already present in the mosquitoes from the 1960s. However, other mutations that make mosquitoes resistant to insecticides were absent from the historic mosquitoes, suggesting that these became beneficial for the mosquitoes only later, as different insecticides were used in subsequent decades. ..."

From the editor's summary and abstract:
"Editor’s summary
Mosquitoes serve as vectors for diseases such as dengue and malaria; however, mosquito species are unequally represented in genetics studies. Two groups collected and analyzed extensive genomic data from mosquito disease vector species ... Crawford et al. sequenced 1206 individuals of the dengue and Zika vector Aedes aegypti and found that highly invasive populations split from earlier lineages during the Atlantic slave trade. They identified genomic regions potentially underlying human-specializing behavioral adaptations. Boddé et al. examined 656 modern Anopheles funestus individuals, as well as 45 museum specimens. They found multiple instances of insecticide-resistance variants in this malaria vector, although most of these weren’t shared with museum specimens collected as recently as 1967, suggesting rapid emergence. Such results will help to inform gene drive and insecticide efforts, as well as public health initiatives. ...

Structured Abstract
INTRODUCTION
The mosquito species Anopheles funestus is a major contributor to human malaria transmission across its vast sub-Saharan African range. Vector control of the other three major malaria-transmitting species in the Gambiae Complex has benefited from a deep understanding of genetic diversity, population structure, and the emergence and spread of insecticide resistance through the whole-genome sequencing of hundreds of individuals from many African countries. We completed whole-genome sequencing of 656 modern samples collected since 2014 and 45 historic samples collected between 1927 and 1967 to create a foundational understanding of genomic diversity in An. funestus across the continent.

RATIONALE
Since large scale deployment of insecticides began in the 1950s, An. funestus has rapidly evolved resistance throughout much of its range.
However, it is an open question whether resistance alleles have evolved independently in multiple locations, whether they are shared between different populations through gene flow, or whether resistant populations have entirely replaced historically susceptible populations.
A clearer genomic view on continental population structure is crucial for implementing strategic use of insecticides, taking into account the potential emergence and spread of insecticide resistance alleles. Additionally, with the implementation of gene drive release for vector control likely in the coming years, we need to be able to predict the spread of gene drive under different release scenarios, which is only possible if detailed knowledge of population connectivity across the continent, and how it varies along the genome, is in place.

RESULTS
We found that the 17 geographic regions from which our samples originated form six population clusters with varying degrees of genome-wide differentiation.
One of these populations, the Equatorial cohort, spans more than 4000 km and comprises individuals from seven countries. In close geographic proximity to this cohort, we found two genetically distinct ecotypes that appear to have a restricted range and distinct chromosomal karyotypes.
Using a windowed principal components analysis (PCA) approach, we explored structure across the genome. We used this approach to identify segregating inversions and classify every individual into its specific inversion karyotype. We also identified genomic regions that have exceptional levels of divergence in comparison to other collinear parts of the genome.
Some of these outlier regions are clearly driven by selection for insecticide resistance, as they contain loci with excessive haplotype sharing, often centered on genes known to play a role in insecticide resistance in many insect species. We show that the Gste2 resistance allele has at least two independent origins and that, despite reports of DDT resistance emerging in the 1950s, none of the historic samples in this study carry DDT resistance alleles found in modern-day populations.

CONCLUSION
Variable structure—such as that observed in this work, with some populations readily sharing alleles across the continent, and others clearly geographically proximal but genetically distinct—is a challenge for vector control. Even if the Gambiae Complex disappeared today, malaria would still rage through Africa until An. funestus is also effectively targeted. The greater understanding of the high levels of genetic diversity and the complex population structure of An. funestus presented in this study will underpin smarter surveillance and targeted vector control."

Genomic evolution of major malaria-transmitting mosquito species uncovered "Sequencing hundreds of Anopheles funestus mosquitoes provides new insights into the evolutionary patterns of this important human malaria-transmitting species."




Fig. 1 Population structure of 656 Anopheles funestus specimens collected across Africa.


Thursday, September 11, 2025

New blood Test detects HPV-associated head, neck cancer 10 years early and with highest accuracy 4 years early

Good news! Cancer is history (soon)!

"Human papillomavirus (HPV) causes an estimated 70 percent of head and neck cancers in the U.S., making it the most common cancer caused by the virus. Yet unlike cervical cancers caused by HPV, there is no screening test for HPV-associated head and neck cancers.

In a new ... study, ... researchers show that a novel liquid biopsy tool they developed, called HPV-DeepSeek, can identify HPV-associated head and neck cancer up to 10 years before symptoms appear.  ..."

From the abstract:
"Purpose
Early detection of HPV-associated oropharyngeal cancer (HPV+OPSCC), the most common HPV cancer in the United States, could reduce disease-related morbidity and mortality, yet currently, there are no early detection tests. Circulating tumor HPV DNA (ctHPVDNA) is a sensitive and specific biomarker for HPV+OPSCC at diagnosis. It is unknown if ctHPVDNA is detectable prior to diagnosis, and thus it’s potential as an early detection test.

Methods
Plasma samples from the MassGeneralBrigham biobank collected 1.3-10.8 years prior to diagnosis from HPV+OPSCC patients (n = 28) and age- and sex-matched controls (n = 28) were blinded and run on a newly developed and validated multi-feature HPV whole genome sequencing liquid biopsy assay and a validated HPV antibody assay.

Results
ctHPVDNA results were positive in 22/28 pre-diagnostic samples from HPV+OPSCC cases (sensitivity 79%) with a maximum lead time of 7.8 years. ctHPVDNA results were negative in all controls (0/28 controls, 100% specificity). Diagnostic accuracy was highest within four years of cancer diagnosis and was higher than HPV Ab detection within the same time frame (p-value 0.004). Application of a machine learning model trained and tested on an independent cohort of 306 cases and controls increased the sensitivity of detection to 27/28 cases (overall sensitivity 96%) and the maximum lead time to 10.3 years.

Conclusions
Circulating tumor HPV DNA can be detected in the blood years prior to diagnosis with HPV+OPSCC, with high specificity, in a case-control cohort of 56 participants. ctHPVDNA detection alone, or in combination with previously identified serological biomarkers may be a feasible approach to early detection of HPV+OPSCC."

Test detects HPV-associated head, neck cancer 10 years early — Harvard Gazette "Tool identifies disease before symptoms appear"

Friday, July 25, 2025

The most complete view of the human genome yet from populations around the world

Amazing stuff!

"... The latest [human genome] datasets, published in 2 back-to-back studies in the journal Nature, present what may be the most complete overview of the human genome to date. ...

“For too long, our genetic references have excluded much of the world’s population. This work captures essential variation that helps explain why disease risk isn’t the same for everyone.”

The first paper analysed the genomes of 1,019 people from 5 continents and 28 population groups.

It focused on genomic structural variants – large sections of DNA that have been deleted, duplicated, inserted, inverted or shuffled and which introduce changes to thousands of DNA bases at a time. They contribute to genetic diversity but are also increasingly associated with diseases and cancers. ...

The team found and categorised more than 167,000 structural variants, doubling the known amount in the human pangenome. ...

“For example, 50.9% of insertions and 14.5% of deletions we found have not been reported in previous variation catalogues. It’s an important step to map blind spots in the human genome and reduce the bias that has long favoured genomes of European descent.” ...

The second study took a slightly different approach by sequencing fewer genomes at much greater detail. The researchers used several sequencing technologies to combine highly accurate medium-length DNA reads with longer, lower-accuracy ones.

This strategy allowed them to piece together the near-complete genomes of 65 individuals. They also decoded some of the most difficult to read stretches, including the highly repetitive centromeres, Y chromosomes and an intricate region associated with the immune system’s Major Histocompatibility Complex. ..."

"... This milestone builds on two foundational studies that reshaped the field of genomics.
In 2022, researchers achieved the first-ever complete sequence of a single human genome, filling in major gaps left by the original Human Genome Project.
In 2023, scientists released a draft pangenome constructed from 47 individuals—a critical step toward representing global genetic diversity.
The new study significantly expands on both efforts, closing 92% of the remaining data gaps and mapping genomic variation across ancestries with a breadth and resolution never achieved. ..."

From the abstract (1):
"Diverse sets of complete human genomes are required to construct a pangenome reference and to understand the extent of complex structural variation.
Here we sequence 65 diverse human genomes and build 130 haplotype-resolved assemblies (median continuity of 130 Mb), closing 92% of all previous assembly gaps and reaching telomere-to-telomere status for 39% of the chromosomes.
We highlight complete sequence continuity of complex loci, including the major histocompatibility complex (MHC), SMN1/SMN2, NBPF8 and AMY1/AMY2, and fully resolve 1,852 complex structural variants.
In addition, we completely assemble and validate 1,246 human centromeres. We find up to 30-fold variation in α-satellite higher-order repeat array length and characterize the pattern of mobile element insertions into α-satellite higher-order repeat arrays. Although most centromeres predict a single site of kinetochore attachment, epigenetic analysis suggests the presence of two hypomethylated regions for 7% of centromeres.
Combining our data with the draft pangenome reference significantly enhances genotyping accuracy from short-read data, enabling whole-genome inference to a median quality value of 45. Using this approach, 26,115 structural variants per individual are detected, substantially increasing the number of structural variants now amenable to downstream disease association studies."

From the abstract (2):
"Genomic structural variants (SVs) contribute substantially to genetic diversity and human diseases, yet remain under-characterized in population-scale cohorts. Here we conducted long-read sequencing in 1,019 humans to construct an intermediate-coverage resource covering 26 populations from the 1000 Genomes Project.
Integrating linear and graph genome-based analyses, we uncover over 100,000 sequence-resolved biallelic SVs and we genotype 300,000 multiallelic variable number of tandem repeats, advancing SV characterization over short-read-based population-scale surveys.
We characterize deletions, duplications, insertions and inversions in distinct populations. Long interspersed nuclear element-1 (L1) and SINE-VNTR-Alu (SVA) retrotransposition activities mediate the transduction of unique sequence stretches in 5′ or 3′, depending on source mobile element class and locus. SV breakpoint analyses point to a spectrum of homology-mediated processes contributing to SV formation and recurrent deletion events.
Our open-access resource underscores the value of long-read sequencing in advancing SV characterization and enables guiding variant prioritization in patient genomes."

Stubborn, overlooked regions of the human genome decoded

The most complete view of the human genome yet sets new standard for use in precision medicine (original news release) "New research decodes the most elusive, difficult-to-sequence regions of the genome from populations around the world, rewriting knowledge of human biology and setting a new benchmark for precision medicine."




Fig. 3: Prevalence of distinct SV [structural variation] classes in our SAGA-based data resource.


Wednesday, June 25, 2025

3D Genome Rewiring Alters Gene Expression in 15 primary human cancer types

Amazing stuff! Cancer is history (soon)! This seems to be an impressive work!

"... leveraging The Cancer Genome Atlas (TCGA) and high-throughput conformational and sequencing techniques to map the 3D cancer genome. The findings ... highlight how changes to enhancer binding influence gene expression in tumors. ...

developed HiChIP, which captures 3D structural information about the genome at specific loci and with smaller numbers of cells. It works by first crosslinking DNA-associated proteins, immunoprecipitating the crosslinked proteins, and then sequencing the bound DNA. ...

To find enhancer activity in cancer genomes, the team applied HiChip. Using histone H3 lysine 27 acetylation (H3K27ac) as a target for enhancer sequences, the researchers profiled 69 tumors from 15 different types of cancer available in TCGA to determine how enhancer binding activity changed across cancers. These samples included information from ATAC-seq, RNA-seq, and whole genome sequencing, allowing the team to combine their conformation data with that of accessible chromatin regions, gene expression, and mutations.

The researchers observed DNA loops with unique interactions of enhancer elements not previously reported. Some of these connections, for example those at the locus for the oncogene MYC, differed between cancer types; in one colon tumor, the researchers observed H34K27ac enrichment at the 5’ end of this gene, whereas in a liver tumor, they saw these marks in a 3’ regulatory region.

As they explored enhancer rewiring across 110 oncogenes, the team noticed three overall patterns:
the same enhancer rewiring occurred around a gene across cancer types,
a specific enhancer rewiring pattern only occurred in one cancer type, and
as in the example of MYC—enhancer rewiring patterns differed at the same gene in different cancers.

Next, the researchers investigated how enhancer rewiring and other DNA structural changes, such as gene duplications that result in higher gene expression, affected oncogene expression in their samples. Leveraging their HiChIP data with information from RNA-seq and whole genome sequencing, the researchers found that increased enhancer activity led to increased mRNA expression in more than 70 percent of oncogenes studied. ..."

From the abstract:
"Genome conformation underlies transcriptional regulation by distal enhancers, and genomic rearrangements in cancer can alter critical regulatory interactions.
Here we profiled the three-dimensional genome architecture and enhancer connectome of 69 tumor samples spanning 15 primary human cancer types from The Cancer Genome Atlas.
We discovered the following three archetypes of enhancer usage for over 100 oncogenes across human cancers: static, selective gain or dynamic rewiring. Integrative analyses revealed the enhancer landscape of noncancer cells in the tumor microenvironment for genes related to immune escape.
Deep whole-genome sequencing and enhancer connectome mapping provided accurate detection and validation of diverse structural variants across cancer genomes and revealed distinct enhancer rewiring consequences from noncoding point mutations, genomic inversions, translocations and focal amplifications. Extrachromosomal DNA promoted more extensive enhancer rewiring among several types of focal amplification mechanisms.
These results suggest a systematic approach to understanding genome topology in cancer etiology and therapy."

3D Genome Rewiring Alters Gene Expression in Cancer | The Scientist "Investigating how mutations in tumors alter DNA’s 3D structure—and subsequently, regulatory sequences called enhancers—can offer new treatment opportunities."



Fig. 1: HiChIP identifies high-resolution chromosome conformation in primary human cancers across multiple scales.




Fig. 4: Integration of WGS and HiChIP identifies cancer-relevant regulatory mutations and target genes.


Sunday, April 27, 2025

New strategy may enable cancer monitoring and early detection from blood tests alone

Good news!

"A new, error-corrected method for detecting cancer from blood samples is much more sensitive and accurate than prior methods and may be useful for monitoring disease status in patients following treatment ...

The method, based on whole-genome sequencing of DNA, also represents an important step toward the goal of routine blood test-based screening for early cancer detection. ...

the researchers benchmarked the cancer-detection performance of a new commercial sequencing platform from Ultima Genomics. They demonstrated that a low-cost platform such as this one enables a very high “depth” of coverage – a measure of the sequencing data quality – allowing investigators to detect extremely low concentrations of circulating tumor DNA. Adding an error-correcting method greatly improved the accuracy of the technique.

“We’re now entering an era of low-cost DNA sequencing, and in this study, we took advantage of that to apply whole-genome sequencing techniques that in the past would have been considered wildly impractical,” ...

Blood-test-based “liquid biopsy” technology for early cancer detection and monitoring of cancer burden in patients could revolutionize cancer care. However, sensitively and accurately identifying the mutational signatures of cancer, just from tiny concentrations of tumor DNA in blood samples, has involved major challenges. ... using methods based on whole-genome sequencing – not just targeted sequencing of stretches of DNA where mutations are expected. In a study published last year, they showed that they could reliably detect advanced melanoma and lung cancer from patient blood samples, even without access to sequence data from tumor samples.

In the new study, they took their approach a step further. First, they showed that the low cost of a new sequencing platform enables a depth of whole-genome sequencing that would have been prohibitively expensive with older technology. Using that platform alone, and having the known mutational patterns in patient tumors as a guide, they were able to detect tumor DNA in patient blood samples at concentrations in the part per million range. ..."

From the abstract:
"Differentiating sequencing errors from true variants is a central genomics challenge, calling for error suppression strategies that balance costs and sensitivity. For example, circulating cell-free DNA (ccfDNA) sequencing for cancer monitoring is limited by sparsity of circulating tumor DNA, abundance of genomic material in samples and preanalytical error rates.
Whole-genome sequencing (WGS) can overcome the low abundance of ccfDNA by integrating signals across the mutation landscape, but higher costs limit its wide adoption.
Here, we applied deep (~120×) lower-cost WGS (Ultima Genomics) for tumor-informed circulating tumor DNA detection within the part-per-million range.
We further leveraged lower-cost sequencing by developing duplex error-corrected WGS of ccfDNA, achieving 7.7 × 10−7 error rates, allowing us to assess disease burden in individuals with melanoma and urothelial cancer without matched tumor sequencing.
This error-corrected WGS approach will have broad applicability across genomics, allowing for accurate calling of low-abundance variants at efficient cost and enabling deeper mapping of somatic mosaicism as an emerging central aspect of aging and disease."

New strategy may enable cancer monitoring from blood tests alone | Cornell Chronicle


Whole genome error-corrected sequencing for sensitive circulating tumor DNA cancer monitoring (preprint, open access, but it was published in November 2022 and was apparently never updated since, possibly quite outdated)


Fig. 1 Ultralow ctDNA detection requires deep sequencing coverage and low error rates.





Monday, March 17, 2025

When did human language emerge? Genomic evidence may have an answer

Amazing stuff! So Homo sapiens was speechless (or in linguistic  incapacity) for about the first 100,000 years? Hard to stomach! 😀

This research seems to be a bit to speculative for my taste! I guess the crucial term here is "linguistic capacity".

"A new survey of genomic evidence suggests our unique language capacity was present at least 135,000 years ago. Subsequently, language might have entered social use 100,000 years ago.

Our species, Homo sapiens, is about 230,000 years old. Estimates of when language originated vary widely, based on different forms of evidence, from fossils to cultural artifacts. The authors of the new analysis took a different approach. ...

“Every population branching across the globe has human language, and all languages are related.” Based on what the genomics data indicate about the geographic divergence of early human populations, he adds, “I think we can say with a fair amount of certainty that the first split occurred about 135,000 years ago, so human language capacity must have been present by then, or before.” ..."

From the abstract:
"Recent genome-level studies on the divergence of early Homo sapiens, based on single nucleotide polymorphisms, suggest that the initial population division within H. sapiens from the original stem occurred approximately 135 thousand years ago.
Given that this and all subsequent divisions led to populations with full linguistic capacity, it is reasonable to assume that the potential for language must have been present at the latest by around 135 thousand years ago, before the first division occurred. Had linguistic capacity developed later, we would expect to find some modern human populations without language, or with some fundamentally different mode of communication. Neither is the case. While current evidence does not tell us exactly when language itself appeared, the genomic studies do allow a fairly accurate estimate of the time by which linguistic capacity must have been present in the modern human lineage. Based on the lower boundary of 135 thousand years ago for language, we propose that language may have triggered the widespread appearance of modern human behavior approximately 100 thousand years ago."

When did human language emerge? | MIT News | Massachusetts Institute of Technology "A new analysis suggests our language capacity existed at least 135,000 years ago, with language used widely perhaps 35,000 years after that."



Table 1. Summary of estimates of divergence times for Khoisan lineage.


Tuesday, March 04, 2025

Mollusk menagerie/family tree from whole genome sequencing of 13 species and detailed genomic analysis of 77 species

Amazing stuff!

The first author, i.e. Zeyuan Chen, of this study works at the Senckenberg Research Institute and Natural History Museum Frankfurt, in my hometown Frankfurt am Main, Germany. 

"... To create an accurate evolutionary ‘tree’ for the phylum, researchers sequenced whole genomes from 13 species, allowing them to look at genetic markers from 77 species in total, including members of every major subgroup of mollusks.

The findings revealed that monoplacophorans were one of the first offshoots of the Conchifera, a group within mollusks that includes almost all the highly recognizable types, including octopuses and their relatives as well as clams and snails.

The analyses also revealed just how genomically diverse mollusks are, with lots of rearrangements and repetitive sequences—features that may have allowed the group’s physical diversity to emerge. ..."

"... In their new study, scientists analysed the genomes of 77 mollusc species, representing all eight major living groups, including lesser-known forms like deep-sea monoplacophorans and worm-like solenogasters. Using cutting-edge genomic techniques, the team reconstructed a detailed evolutionary tree and confirmed key hypotheses about mollusc ancestry. ..."


From the editor's summary and abstract:
"Editor’s summary
Genome sequencing has allowed for a much greater understanding of how species relate to one another than did earlier morphology-based approaches. A particularly difficult phylum to study using genomic data has been mollusks, which encompass species ranging from squid to sea snails, in part due to their high levels of heterozygosity and repetitiveness. Chen et al. sequenced 13 new complete genomes from across the phylum to assemble a new phylogeny for Mollusca. They resolved several highly debated nodes and provide additional genomes for future study of this highly diverse and genomically complex phylum. ...
Abstract
Extreme morphological disparity within Mollusca has long confounded efforts to reconstruct a stable backbone phylogeny for the phylum. Familiar molluscan groups—gastropods, bivalves, and cephalopods—each represent a diverse radiation with myriad morphological, ecological, and behavioral adaptations.
The phylum further encompasses many more unfamiliar experiments in animal body-plan evolution.
In this work, we reconstructed the phylogeny for living Mollusca on the basis of metazoan BUSCO (Benchmarking Universal Single-Copy Orthologs) genes extracted from 77 (13 new) genomes, including multiple members of all eight classes with two high-quality genome assemblies for monoplacophorans. Our analyses confirm a phylogeny proposed from morphology and show widespread genomic variation. The flexibility of the molluscan genome likely explains both historic challenges with their genomes and their evolutionary success."

ScienceAdviser

Cracking the Mollusc Code: Genomes Reveal the Ancient Ancestor of Your Garden Snail (original news release) "An international team of scientists has cracked a longstanding evolutionary mystery surrounding molluscs, one of the most diverse groups of animals on Earth. The groundbreaking research, published today in Science, resolves the family tree for molluscs, bringing long-awaited clarity on their evolutionary history and resolving debates that have persisted for decades."



Nudibranchs like this Flabellina affinis are among the varied forms of mollusks.



Fig. 1. Mollusca timetree.


Monday, January 06, 2025

Why are bed bugs virtually unkillable? It might be genes (or atomic bomb blasts?)

Amazing stuff! 

They researched descendants of bed bugs hit by atomic bombs in 1945? How transferable are these results to bed bugs from other regions of the world?

"Bed bugs are notoriously difficult to remove once they’ve moved in – and they’re getting more difficult, thanks to their evolving resistance to insecticides.

A team of researchers has mapped the genomes of bed bug strains, aiming to find out why a “superstrain” has become 20,000 times more resistant to treatment. ...

In this study, researchers collected DNA from 2 bed bug sources.
One population, judged vulnerable to pesticides, had been originally collected from fields in Nagasaki, Japan, but maintained in a lab for 68 years.
The other population was collected from a hotel in Hiroshima, Japan, in 2010. ..."

From the abstract:
"Insecticide resistance in the bed bug Cimex lectularius is poorly understood due to the lack of genome sequences for resistant strains. In Japan, we identified a resistant strain of C. lectularius that exhibits a higher pyrethroid resistance ratio compared to many previously discovered strains.
We sequenced the genomes of the pyrethroid-resistant and susceptible strains using long-read sequencing, resulting in the construction of highly contiguous genomes (N50 of the resistant strain: 2.1 Mb and N50 of the susceptible strain: 1.5 Mb). Gene prediction was performed by BRAKER3, and the functional annotation was performed by the Fanflow4insects workflow.
Next, we compared their amino acid sequences to identify gene mutations, identifying 729 mutated transcripts that were specific to the resistant strain. Among them, those defined previously as resistance genes were included. Additionally, enrichment analysis implicated DNA damage response, cell cycle regulation, insulin metabolism, and lysosomes in the development of pyrethroid resistance. Genome editing of these genes can provide insights into the evolution and mechanisms of insecticide resistance. This study expanded the target genes to monitor allele distribution and frequency changes, which will likely contribute to the assessment of resistance levels. These findings highlight the potential of genome-wide approaches to understand insecticide resistance in bed bugs."

Why are bed bugs virtually unkillable? It might be genes



Figure 2. Transcripts with mutations in the resistant strain. (a) The number of mutation sites per transcript is shown by circles. (b–g) Mutation sites of candidate resistance genes are shown with different amino acids. 


Tuesday, November 26, 2024

New survey reveals thousands of new human genes

Stunning news! How is this possible!

"One of the biggest surprises to emerge when the human genome was first sequenced more than 20 years ago was how few genes it contained, with that number now hovering around 20,000. But a new systematic analysis of what some call the “dark proteome” suggests scientists have missed thousands of non-traditional genes that make smaller-than-average proteins.

Genes typically consist of a long protein-coding DNA sequence known as an open reading frame or ORF; this is preceded by a snippet of DNA that attracts the proteins needed for the gene to be read. But biologists studying everything from yeast to snakes to humans have recently unearthed a plethora of so-called non-canonical ORFs, which lack these snippets and can code for just a few amino acids. Initially dismissed as “noise,” [or artifacts?] studies have demonstrated that many do make proteins and at least a few of these proteins matter.

So, after a multidisciplinary, international consortium tracked down 7264 of these non-canonical ORFs in 2022, they teamed up with the Human Proteome Organization and PeptideAtlas  to see which ones make proteins—and they found at least 25% of those ORFs do. The newly discovered miniproteins “help provide a more complete picture of the coding portion of the genome,” ..."

"... Where does all this leave the tally of human genes? The dark proteome has clearly boosted the total, but no one knows the true number.

“My gut feeling it is probably not as high as 100,000,” Martinez says, “but 50,000 is in the realm of possibility.”"

About 43 authors affiliated with 24 different organisations contributed to this research paper. 

From the abstract:
"A major scientific drive is to characterize the protein-coding genome as it provides the primary basis for the study of human health. But the fundamental question remains: what has been missed in prior genomic analyses? Over the past decade, the translation of non-canonical open reading frames (ncORFs) has been observed across human cell types and disease states, with major implications for proteomics, genomics, and clinical science. However, the impact of ncORFs has been limited by the absence of a large-scale understanding of their contribution to the human proteome. Here, we report the collaborative efforts of stakeholders in proteomics, immunopeptidomics, Ribo-seq ORF discovery, and gene annotation, to produce a consensus landscape of protein-level evidence for ncORFs. We show that at least 25% of a set of 7,264 ncORFs give rise to translated gene products, yielding over 3,000 peptides in a pan-proteome analysis encompassing 3.8 billion mass spectra from 95,520 experiments. With these data, we developed an annotation framework for ncORFs and created public tools for researchers through GENCODE and PeptideAtlas. This work will provide a platform to advance ncORF-derived proteins in biomedical discovery and, beyond humans, diverse animals and plants where ncORFs are similarly observed."

ScienceAdvisor

‘Dark proteome’ survey reveals thousands of new human genes "Database confirms that overlooked segments of the genome code for a multitude of tiny proteins"



What a gigantic, global effort!




Saturday, November 16, 2024

Metagenomic sequencing test proves effective in diagnosing almost any kind of pathogen

Good news! A good diagnosis precedes a successful treatment!

"... It uses a powerful genomic sequencing technique, called metagenomic next-generation sequencing (mNGS).

Rather than looking for one type of pathogen at a time, mNGS analyzes all the nucleic acids, RNA and DNA, that are present in a sample. ...

The test has now been performed on thousands of patients with unexplained neurological symptoms, both at UCSF and other hospitals across the country.

In a paper ... the team demonstrated that the mNGS test correctly identified 86% of neurological infections. ...

In a companion study ... the team also used mNGS to identify pathogens in respiratory fluid that can cause pneumonia, and automated it to get results faster. ..."

From the abstract (1):
"Metagenomic next-generation sequencing (mNGS) of cerebrospinal fluid (CSF) is an agnostic method for broad-based diagnosis of central nervous system (CNS) infections. Here we analyzed the 7-year performance of clinical CSF mNGS testing of 4,828 samples from June 2016 to April 2023 performed by the University of California, San Francisco (UCSF) clinical microbiology laboratory. Overall, mNGS testing detected 797 organisms from 697 (14.4%) of 4,828 samples, consisting of 363 (45.5%) DNA viruses, 211 (26.4%) RNA viruses, 132 (16.6%) bacteria, 68 (8.5%) fungi and 23 (2.9%) parasites. We also extracted clinical and laboratory metadata from a subset of the samples (n = 1,164) from 1,053 UCSF patients. Among the 220 infectious diagnoses in this subset, 48 (21.8%) were identified by mNGS alone. The sensitivity, specificity and accuracy of mNGS testing for CNS infections were 63.1%, 99.6% and 92.9%, respectively. mNGS testing exhibited higher sensitivity (63.1%) than indirect serologic testing (28.8%) and direct detection testing from both CSF (45.9%) and non-CSF (15.0%) samples (P < 0.001 for all three comparisons). When only considering diagnoses made by CSF direct detection testing, the sensitivity of mNGS testing increased to 86%. These results justify the routine use of diagnostic mNGS testing for hospitalized patients with suspected CNS infection."

From the abstract (2):
"Tools for rapid identification of novel and/or emerging viruses are urgently needed for clinical diagnosis of unexplained infections and pandemic preparedness. Here we developed and clinically validated a largely automated metagenomic next-generation sequencing (mNGS) assay for agnostic detection of respiratory viral pathogens from upper respiratory swab and bronchoalveolar lavage samples in <24 h. The mNGS assay achieved mean limits of detection of 543 copies/mL, viral load quantification with 100% linearity, and 93.6% sensitivity, 93.8% specificity, and 93.7% accuracy compared to gold-standard clinical multiplex RT-PCR testing. Performance increased to 97.9% overall predictive agreement after discrepancy testing and clinical adjudication, which was superior to that of RT-PCR (95.0% agreement). To enable discovery of novel, sequence-divergent human viruses with pandemic potential, de novo assembly and translated nucleotide algorithms were incorporated into the automated SURPI+ computational pipeline used by the mNGS assay for pathogen detection. Using in silico analysis, we showed that after removal of all human viral sequences from the reference database, 70 (100%) of 70 representative human viral pathogens could still be identified based on homology to related animal or plant viruses. Our assay, which was granted breakthrough device designation from the US Food and Drug Administration (FDA) in August of 2023, demonstrates the feasibility of routine mNGS testing in clinical and public health laboratories, thus facilitating a robust and rapid response to the next viral pandemic."

Metagenomic sequencing test proves effective in diagnosing almost any kind of pathogen "A genomic test developed at UC San Francisco to rapidly detect almost any kind of pathogen—virus, bacteria, fungus or parasite—has proved successful after a decade of use."

1 Genomic Test Can Diagnose Nearly Any Infection (original news release) "Next-generation metagenomic sequencing test developed at UCSF proves its effectiveness in quickly diagnosing almost any kind of pathogen."

Sunday, March 17, 2024

Global survey of gut microbes uncovers 18 new bacterial species and clues to antibiotic resistance

Only 18 new bacterial species? Surprising and disappointing given the enormous effort! You almost have to believe these researchers or the involved "elite adventurers" made a mistake!

If they did not make a mistake, then this could well be a fundamental result.

"... The microbe is just one of 18 new species of Enterococcus discovered by the team, led by researchers at the Broad Institute and Massachusetts Eye and Ear. The researchers analyzed hundreds of scat, soil, and other samples that were collected from an unexplored peak in Nepal, a remote trail in Uganda, and many more places across the globe by an international team of scientists and elite adventurers.
The new species they found expand the genus diversity of known enterococcal strains by more than 25 percent and add unprecedented detail to this microbial family tree. In the new microbes’ DNA, the researchers found hundreds of novel genes that may offer clues to how enterococci are able to resist antibiotic treatment and thrive in the hospital environment.  ...
Genetic analysis of the new and existing enterococcal species allowed the team to expand the microbe’s family tree and refine its branches, or clades, yielding clues to how certain species are able to colonize particular hosts. For example, members of one clade have particularly small genomes that lack genes for amino acid biosynthesis, a hallmark of adaptation to mammalian hosts. Species in another clade carry large genomes that include genes necessary for producing the essential vitamin B12, indicating that they do not need their hosts to provide it. ..."

From the abstract:
"Enterococci are gut microbes of most land animals. Likely appearing first in the guts of arthropods as they moved onto land, they diversified over hundreds of millions of years adapting to evolving hosts and host diets. Over 60 enterococcal species are now known. Two species, Enterococcus faecalis and Enterococcus faecium, are common constituents of the human microbiome. They are also now leading causes of multidrug-resistant hospital-associated infection. The basis for host association of enterococcal species is unknown. To begin identifying traits that drive host association, we collected 886 enterococcal strains from widely diverse hosts, ecologies, and geographies. This identified 18 previously undescribed species expanding genus diversity by >25%. These species harbor diverse genes including toxins and systems for detoxification and resource acquisition. Enterococcus faecalis and E. faecium were isolated from diverse hosts highlighting their generalist properties. Most other species showed a more restricted distribution indicative of specialized host association. The expanded species diversity permitted the Enterococcus genus phylogeny to be viewed with unprecedented resolution, allowing features to be identified that distinguish its four deeply rooted clades, and the entry of genes associated with range expansion such as B-vitamin biosynthesis and flagellar motility to be mapped to the phylogeny. This work provides an unprecedentedly broad and deep view of the genus Enterococcus, including insights into its evolution, potential new threats to human health, and where substantial additional enterococcal diversity is likely to be found."

Global survey of gut microbes uncovers 18 new bacterial species and clues to antibiotic resistance | Broad Institute Researchers examined scat from hundreds of animals around the world and pinpointed genes that could fuel drug-resistant infections.


Adventure Scientist Stevie Anna Plummer holds water and scat samples collected in 2016 during an expedition to an unexplored peak in the Himalayas. 


Fig. 1. Sources for 852 Enterococcus isolates.