Showing posts with label databases. Show all posts
Showing posts with label databases. Show all posts

Saturday, July 25, 2026

The NIH’s All of Us Database Expands into Multiomics, a landmark data release

Good news! This could be a game changer!

"Late last month [June 2026], the National Institutes of Health (NIH) announced an expansion of the All of Us database. The United States created the All of Us research program to capture the diversity of the American population by building a comprehensive database to help solve the country’s health challenges through research.

The latest update expands the program’s database, which now includes health data from over 747,000 participants and adds, for the first time, multiomics information such as proteomics and RNA sequencing (RNAseq) data. ...

Over the last seven years, the All of Us database has expanded significantly. The latest update brings the number of EHRs close to 482,000 and makes health information from over 747,000 participants available to researchers. These data, in turn, provide an in-depth look at participants’ genomic information, including more than 1.3 billion genetic variants, 553,000 genotyping arrays, and 96,000 structural variant records ... 
Additionally, the new dataset now encompasses 600,000 physical measurements and survey responses from 747,000 participants, providing social, environmental, and behavioral information. ..."

"Today [6/30/2026], the National Institutes of Health's All of Us Research Program made history, unveiling the world's largest integrated dataset that combines unparalleled genomic depth with real-world clinical and wearable data. More than 747,000 participants across all 50 states and territories have shared data, enabling the linkage of nearly 482,000 electronic health records connected to 535,000 whole genome sequences. ...

This release moves All of Us beyond sequencing and into the multiomics era, providing researchers with unmatched statistical power.

  • Short-Read Whole Genome Sequencing: The short-read Whole Genome Sequencing has grown to include more than 535,000 participants.
  • Overlapping Multiomics Cohort: Over 8,000 participants of diverse genetic ancestry now have fully overlapping multiomics data across three advanced modalities: long-read sequences (>14,000 total available), proteomics (nearly 10,000 total), and RNA-seq (nearly 9,000 total).
  • Over 1.3 Billion Total Genetic Variants: Providing unparalleled resolution to map rare and common diseases.
  • Advanced Analytical Capabilities: New built-in tools allow for efficient, standardized analyses of relatedness, phasing, pharmacogenomics, and newly debuted HLA and mtDNA analysis.
  • Targeted Callsets: Includes deep-dives into the exome, ClinVar, AC/AF thresholds, and CMGR to drastically improve variant-calling accuracy.
...
Transformative Real-World Data: EHR, Wearables & Environment
All of Us dramatically expands the real-world data available to researchers. Clinical utility has scaled massively, bridging the gap between genetic codes and real-world diagnoses.

  • 22% Growth in Electronic Health Record (EHR) Data: Strategic expansion of EHR data sources for growth in phenotypic data offerings with participant-mediated EHR data (58K participants) and HIE EHR data from CLAD (15K participants). Linked EHR have grown to nearly 482,000, now connected to 535,000 WGS samples.
  • First-Ever Clinical Notes Release – an unprecedented depth in phenotyping: Registered researchers now have secure access to unstructured clinical text, with 96 million NLP-derived concept codes extracted from 9.5 million clinical notes across more than 99,000 participants using the CLAMP tool, all mapped to the OMOP standard vocabulary for seamless cross-referencing with genomic data.
  • World's Largest Accessible Fitbit Dataset: Active tracking data is now available for 68,000 participants, with Apple HealthKit integration coming soon for an initial 20,000.
  • New Sleep Data Tables: Three new Fitbit sleep tables add granular longitudinal insights, including daily sleep stage summaries (REM, light, deep, wake) for 62,000 participants, derived sleep log metrics, and micro-indicators of sleep fragmentation.
  • Geospatial & Environmental Data: An upcoming release will integrate residential and geographic data linkages, enabling researchers to study air quality, neighborhood resources, and social determinants of health alongside genetics.
...
To date, nearly 23,000 registered researchers at institutions across every state have used the All of Us Researcher Workbench, producing over 1,400 peer-reviewed publications.
This secure cloud-accessible platform democratizes science, giving a researcher at a small community college or a high school student the exact same computational power and data access as a Nobel laureate at a coastal institution. That access has already led to:

  • A first-of-its-kind genetic test predicting inherited risk across eight cardiovascular conditions.
  • A low-cost prostate cancer risk model now in clinical trial with 5,000 Veterans.
  • Researchers identifying existing medications and novel genetic changes that may help prevent Alzheimer's disease..
  • Data showing that 8,000 to 9,000 daily steps are a meaningful protective threshold across multiple chronic conditions.
..."

The NIH’s All of Us Database Expands into Multiomics | The Scientist "With RNAseq and proteomics data from thousands of people, the database provides scientists with diverse precision health data on almost 750,000 people to power research."

Tuesday, April 28, 2026

Gone in 9 Seconds: AI Coding Agent Deletes Entire Database and All Backups of software company PocketOS

Headline of the day!

"The founder of a software company has issued a public warning after an AI coding assistant erased his company’s entire production database and all backups in just nine seconds.

Tom’s Hardware reports that Jer Crane, founder of PocketOS, a platform serving car rental businesses, experienced what he describes as catastrophic failures when an AI coding agent deleted critical company data that took months to accumulate. The incident occurred when Cursor, an AI coding tool powered by Anthropic’s Claude Opus 4.6, was performing what should have been a routine task in the company’s staging environment. ..."

Gone in 9 Seconds: AI Coding Agent Deletes Entire Company Database and All Backups

Claude-powered AI coding agent deletes entire company database in 9 seconds — backups zapped, after Cursor tool powered by Anthropic's Claude goes rogue "PocketOS founder blames ‘Cursor running Anthropic's flagship Claude Opus 4.6’ plus Railway’s infrastructure for data disaster."

Sunday, May 25, 2025

On Relational Graph Transformer

Very recommendable! Just finished reading this paper. It is a deep dive into how ML & AI can be applied to relational databases and its progress.

From the abstract:
"Relational Deep Learning (RDL) is a promising approach for building state-of-the-art predictive models on multi-table relational data by representing it as a heterogeneous temporal graph.
However, commonly used Graph Neural Network models suffer from fundamental limitations in capturing complex structural patterns and long-range dependencies that are inherent in relational data.
While Graph Transformers have emerged as powerful alternatives to GNNs on general graphs, applying them to relational entity graphs presents unique challenges: (i) Traditional positional encodings fail to generalize to massive, heterogeneous graphs;
(ii) existing architectures cannot model the temporal dynamics and schema constraints of relational data;
(iii) existing tokenization schemes lose critical structural information.
Here we introduce the Relational Graph Transformer (RelGT), the first graph transformer architecture designed specifically for relational tables. RelGT employs a novel multi-element tokenization strategy that decomposes each node into five components (features, type, hop distance, time, and local structure), enabling efficient encoding of heterogeneity, temporality, and topology without expensive precomputation.
Our architecture combines local attention over sampled subgraphs with global attention to learnable centroids, incorporating both local and database-wide representations.
Across 21 tasks from the RelBench benchmark, RelGT consistently matches or outperforms GNN baselines by up to 18%, establishing Graph Transformers as a powerful architecture for Relational Deep Learning."

[2505.10960] Relational Graph Transformer




Thursday, March 28, 2024

SQLite and SQLite Browser annoyances

I have recently started to use SQLite and its Browser almost on a  daily basis! Generally speaking they work very fine. In hindsight, I was a bit premature when I wrote here about my first impressions working with SQLite and SQLite Browser.

Incident of 3/28/2024

I was in the middle of composing a select statement with a where clause not exists etc. when the browser abruptly exited upon executing the select statement. It was not a very fancy or complex statement as I am not an expert SQL developer. I hope, the automatic reporting feature worked and this will be fixed soon.

Worst of all, my project file was completely deleted and existed only with 0 bytes. All lost!

I venture to guess, because I had just renamed one of the two tables involved in the select statement may have caused the program to abort.


Sunday, May 21, 2023

New database offers insight into consequences of language loss

Since ancient times, we have lost numerous languages. Many of them we may never be able to decipher.

Would the world not be much better off if all or most humans spoke at least one common language and used one common alphabet? E.g. imagine Chinese was spoken/written as a second language and not as a first language by over one billion people, but calligraphy, cultural heritage and history e.g. make that very difficult if not impossible.

"... More than half of the world’s approximately 7,000 signed and spoken languages are currently endangered. And without intervention they are likely to become extinct, meaning nobody will speak or sign them any longer. ...
The novel database currently covers 2,467 language varieties spanning 215 different language families and 101 isolated languages from all inhabited continents and geographic areas. It captures 195 language properties — including word order, verbal tense, and whether a language features gendered pronouns — allowing researchers to draw comparisons between and across the languages. ..."

From the abstract:
"While global patterns of human genetic diversity are increasingly well characterized, the diversity of human languages remains less systematically described. Here, we outline the Grambank database. With over 400,000 data points and 2400 languages, Grambank is the largest comparative grammatical database available. The comprehensiveness of Grambank allows us to quantify the relative effects of genealogical inheritance and geographic proximity on the structural diversity of the world’s languages, evaluate constraints on linguistic diversity, and identify the world’s most unusual languages. An analysis of the consequences of language loss reveals that the reduction in diversity will be strikingly uneven across the major linguistic regions of the world. Without sustained efforts to document and revitalize endangered languages, our linguistic window into human history, cognition, and culture will be seriously fragmented."

New database offers insight into consequences of language loss | YaleNews Grambank, a database of 2,467 languages that Yale linguist Claire Bowern helped create, helps researchers better understand the stakes when languages die off.


Fig. 2. Grammatical similarity in the Grambank sample of languages.
The color coding represents the distribution of languages according to the first three principal components (PCs) mapped onto RGB color space


Tuesday, August 31, 2021

Using pre-trained language models for database queries on unstructured data

Very interesting research done by researchers at Facebook!

"... Facebook AI has developed a new approach called neural databases, which enables machines to search unstructured data — which might range from vast collections of text to recordings of songs — similar to how traditional systems can search a typical structured database. With neural databases, it might one day be possible to run a complex query such as “What is the third-longest entry about a Russian-born novelist?” directly on Wikipedia, for example.

Neural databases bridge an important gap between the fields of databases and NLP. Significant progress has been made in using natural language queries on standard structured data. This lets people pose ad hoc queries such as “How many teams won away games by more than three points?” But these existing systems can’t query a collection of information that isn’t organized into a structured database. ..."

Using pre-trained language models for database queries on unstructured data