Showing posts with label Allen Institute for Artificial Intelligence. Show all posts
Showing posts with label Allen Institute for Artificial Intelligence. Show all posts

Saturday, October 03, 2026

Open-sourcing AstaBrief, the fast report-generation model in Asta

Good news!

"This week, we’re releasing two new open artifacts for researchers and model builders: AstaBrief 8B, a model for generating fast, cited scientific reports, and Olmo-core 3, our redesigned framework for training large mixture-of-experts (MoE) models. ...

AstaBrief
 
Reviewing and synthesizing scientific literature can be time-consuming. AstaBrief 8B is an 8-billion-parameter model trained to generate a cited report from a research question and retrieved scientific literature, giving researchers a faster starting point for exploring a topic, checking sources, and refining their questions. ..."


Open-sourcing AstaBrief, the fast report-generation model in Asta | Ai2

Thursday, October 02, 2025

Asta DataVoyager: Data-driven discovery and analysis of structured experimental data for scientists

Good news! How much will this accelerate and boost scientific research?

".. Experimental logs live in spreadsheets; instrument readings arrive as CSVs; and results tables pile up across projects. Turning those structured files into answers takes time and often requires advanced programming skills to be done efficiently.

To fill the gap, we’re launching DataVoyager in Asta, our ecosystem for scientific research agents. Built to address the challenges scientists face in drilling down into structured datasets, Asta DataVoyager delivers data-driven discovery and analysis capabilities that allow you to ask questions about structured files in plain language and get clearly cited, explainable answers with copyable code, clear visuals, and a concise, well-supported summary. ...

Users upload a dataset and ask a question (e.g., “Which treatment shows the most improvement after week 6?”), along with an optional prompt to establish context so that Asta DataVoyager makes better initial choices. ...

Moreover, Asta DataVoyager allows teams to stay in full control of their data—they can delete datasets at any time from Asta’s hosted portal or secure on-premises, datacenter, and private cloud deployments. ..."

"... Asta DataVoyager is a trusted AI collaborator—one that lets researchers make queries about data in natural language and get transparent, reproducible answers they can act on. It was designed from the start to be intuitive for users, regardless of their comfort level working with dataset analysis tooling. ...

Users upload a dataset in CSV, Excel (.xlsx), JSON (.json/.jsonl), HDF5, TSV, or Parquet format and ask a question (e.g., “Which treatment arm shows the steepest improvement after week 6?”), along with an optional prompt to establish context (e.g., “use these units, measurement cadence, treatment conditions, and outcome variables”) so that Asta DataVoyager makes better initial choices.

Asta DataVoyager then outputs: 

  • A crisp answer to the user’s question, written for scientists
  • Copyable visuals that make the finding understandable at a glance
  • Copyable code that reproduces the analysis
  • A methods section that documents assumptions, detailed reasoning steps, and statistical tests conducted—so users can cite the procedure or adapt it
..."

Asta DataVoyager, fluid benchmarking, and build your own AskOlmo bot

Monday, August 04, 2025

AutoDS: A prototype engine for autonomous, open-ended scientific discovery

This is one of the latest and hottest topics in machine learning & AI! This is just one of several recent papers on this subject.

How much will ML & AI accelerate scientific research? Beyond imagination?

No more waiting for the next Albert Einstein!

From the abstract:
"The promise of autonomous scientific discovery (ASD) hinges not only on answering questions, but also on knowing which questions to ask. Most recent works in ASD explore the use of large language models (LLMs) in goal-driven settings, relying on human-specified research questions to guide hypothesis generation. However, scientific discovery may be accelerated further by allowing the AI system to drive exploration by its own criteria. The few existing approaches in open-ended ASD select hypotheses based on diversity heuristics or subjective proxies for human interestingness, but the former struggles to meaningfully navigate the typically vast hypothesis space, and the latter suffers from imprecise definitions.
This paper presents AutoDS -- a method for open-ended ASD that instead drives scientific exploration using Bayesian surprise. Here, we quantify the epistemic shift from the LLM's prior beliefs about a hypothesis to its posterior beliefs after gathering experimental results. To efficiently explore the space of nested hypotheses, our method employs a Monte Carlo tree search (MCTS) strategy with progressive widening using surprisal as the reward function.
We evaluate AutoDS in the setting of data-driven discovery across 21 real-world datasets spanning domains such as biology, economics, finance, and behavioral science.
Our results demonstrate that under a fixed budget, AutoDS substantially outperforms competitors by producing 5--29% more discoveries deemed surprising by the LLM.
Our human evaluation further finds that two-thirds of AutoDS discoveries are surprising to the domain experts, suggesting this is an important step forward towards building open-ended ASD systems."

AutoDS: A prototype engine for autonomous, open-ended scientific discovery | Ai2







Wednesday, April 09, 2025

AI2 OLMoTrace points model output back to training data

Good news! Does this new model also address hallucination? 😊

Always keep in mind: When was the training/last training update cutoff date? Or fast do these large language/multi-modal models get outdated?

"For years it’s been an open question — how much is a language model learning and synthesizing information, and how much is it just memorizing and reciting? 

Introducing OLMoTrace, a new feature in the Ai2 Playground that begins to shed some light.

OLMoTrace connects phrases or even whole sentences in the language model’s output back to verbatim matches in its training data. It does this by searching billions of documents and trillions of tokens in real time and highlighting where it finds compelling matches. ...

Through OLMoTrace, you can gain insights into why the model generates certain sequences of words. ..."

OLMoTrace points model output back to training data

Thursday, February 06, 2025

Scaling the Tülu 3 post-training recipes to surpass the performance of DeepSeek V3

Competition is good, more competition is better! Except that this latest model by AI2 is much larger (10x) than the DeepSeek model.

"Following the success of our Tülu 3 release in November, we are thrilled to announce the launch of Tülu 3 405B—The first application of fully open post-training recipes to the largest open-weight models. With this release, we demonstrate the scalability and effectiveness of our post-training recipe applied at 405B parameter scale.

As outlined below, Tülu 3 405B achieves competitive or superior performance to both Deepseek v3 and GPT-4o, while surpassing prior open-weight post-trained models of the same size including Llama 3.1 405B Instruct and Nous Hermes 3 405B on many standard benchmarks. Interestingly, we found that our Reinforcement Learning from Verifiable Rewards (RLVR) framework improved the MATH performance more significantly at a larger scale, i.e., 405B compared to 70B and 8B, similar to the findings in the DeepSeek-R1 report. Overall, our results show a consistent edge over DeepSeek V3, especially with the inclusion of safety benchmarks. ..."

Scaling the Tülu 3 post-training recipes to surpass the performance of DeepSeek V3 | Ai2

Thursday, June 06, 2024

AI2: First Real-Time System for Detecting Ships and monitoring their movements in Public Optical satellite Imagery Globally

Good news! Amazing stuff!

However, 10 meter resolution seems to be a little coarse.

"... Over the last year, our research and conservation teams have collaborated on building a highly specialized computer vision model to pinpoint vessels in data collected by the European Space Agency’s (ESA) Sentinel-2 satellite. Freely available to governments and NGOs worldwide, Sentinel-2 data empowers the fight against illegal fishing and other maritime crimes. It offers a near real-time view of ships crossing oceans at a remarkable 10-meter resolution, striking a valuable balance between capturing a vast swathe of the ocean and revealing crucial details about vessels on its surface. Since its release, vessel detections from Sentinel-2 have become one of Skylight's most popular data sources – setting a new standard in maritime surveillance that enables swift analysis and response. ..."

First Real-Time System for Detecting Ships in Public Optical Imagery Globally

Sentinel-2 optical imagery, as seen in Skylight.


Skylight's Sentinel-2 detections seen in SeaVision. Available through our API, this helps extends the impact of this data source.


AI2: Data-driven Discovery with Large Generative Models

Yes, data-driven discovery will make exponential progress with machine learning & AI! This is just the beginning!

Caveat: I did not read the entire blog post by AI2.

"... Untapped Data: Underutilized Scientific Goldmines
Many datasets in observational and experimental sciences are underutilized today, ranging in topics from computational sciences, social science, and health to climate science and astrophysics. Recognizing this potential, we aimed to harness the power of massive datasets and advancements in Large Generative Models (LGMs) to accelerate scientific discovery. Our work initiates a series of research articles that seek to achieve this goal and build systems that scientists can use to improve scientific processes and efficiency. ..."

Data-driven Discovery with Large Generative Models | by AI2 | May, 2024 | AI2 Blog



Thursday, March 07, 2024

Top Story New Paper Highlights Covert Racism in LLMs. Really!

This obsession with Afro-American racism continues in the ML & AI research community! In this case, the Allen Institute for AI. Very tiresome and annoying!

In particular and again, the focus is exclusively on black Americans. They claim "raciolinguistic stereotypes" about black Americans, as if there are not e.g. such stereotypes against southern or lower class white Americans by e.g. black Americans or Chinese and so on.

Racism is unfortunately very human and occurs among most races and groups of people.

To obsess with sanitizing large language models from human flaws is a dangerous path!

"... In a new paper, ... authors show that LLMs also exhibit covert racism, specifically in the form of dialect prejudice. ...
They extend research showing that Americans hold raciolinguistic stereotypes about speakers of African American English and find that LLMs have the same prejudice, ..."

From the abstract:
"Hundreds of millions of people now interact with language models, with uses ranging from serving as a writing aid to informing hiring decisions. Yet these language models are known to perpetuate systematic racial prejudices, making their judgments biased in problematic ways about groups like African Americans. While prior research has focused on overt racism in language models, social scientists have argued that racism with a more subtle character has developed over time. It is unknown whether this covert racism manifests in language models. Here, we demonstrate that language models embody covert racism in the form of dialect prejudice: we extend research showing that Americans hold raciolinguistic stereotypes about speakers of African American English and find that language models have the same prejudice, exhibiting covert stereotypes that are more negative than any human stereotypes about African Americans ever experimentally recorded, although closest to the ones from before the civil rights movement. By contrast, the language models' overt stereotypes about African Americans are much more positive. We demonstrate that dialect prejudice has the potential for harmful consequences by asking language models to make hypothetical decisions about people, based only on how they speak. Language models are more likely to suggest that speakers of African American English be assigned less prestigious jobs, be convicted of crimes, and be sentenced to death. Finally, we show that existing methods for alleviating racial bias in language models such as human feedback training do not mitigate the dialect prejudice, but can exacerbate the discrepancy between covert and overt stereotypes, by teaching language models to superficially conceal the racism that they maintain on a deeper level. Our findings have far-reaching implications for the fair and safe employment of language technology."

2024-03-newsletter

Friday, February 02, 2024

AI2: OLMo 7B: Open Language Model. A State-Of-The-Art, Truly Open LLM and Framework

Good news! Empowering AI research! Dive in!

"AI2 opens its framework for training and experimenting with large language models on Hugging Face and GitHub with the launch of our first Open Language Model (OLMo). The AI2 LLM framework is intentionally designed to provide access to data, training code, models, and evaluation code necessary to advance AI through open research to empower academics and researchers to study the science of language models collectively. This approach enables the AI community to access a broader range of research questions, such as understanding the specific impact of certain subsets of pretraining data on downstream performance or investigating new pretraining methods and understanding instabilities.
This effort's first batch of models includes four final variants of our language model at the 7B scale corresponding to different architectures, optimizers, and training hardware, and one model at the 1B scale, all trained on at least 2T tokens. This is the first step in a long series of planned releases, continuing with larger models, instruction-tuned models, and more variants down the line.

Each model comes with the following:
  • Full training data used for these models, including code that produces the training data, from AI2’s Dolma, and WIMBD for analyzing pretraining data.
  • Full model weights, training code, training logs, training metrics in the form of Weights & Biases logs, and inference code.
  • 500+ checkpoints per model, from every 1000 steps during the training process, available as revisions on HuggingFace.
  • Evaluation code under the umbrella of AI2’s Catwalk and Paloma.
  • Fine-tuning code and adapted models (coming soon with Open Instruct)
  • All code, weights, and intermediate checkpoints are released under the Apache 2.0 License.
 ..."

OLMo: Open Language Model. A State-Of-The-Art, Truly Open LLM and… | by AI2 | Feb, 2024 | AI2 Blog

Monday, July 31, 2023

AIAi: Top 8 machine learning trends in 2023

Recommendable! 

However, the first (General Adversarial Network (GAN)) and the penultimate  (Unsupervised machine learning) trend are actually very old hats!

Included is also a brief history of machine learning, but the history is too short and lacks a lot of events etc.

Top 8 machine learning trends in 2023

Saturday, June 25, 2022

Unified-IO, a new general purpose model from AI2

I have not had time to check it out this latest model by AllenAI, but it seems to be another significant step towards comprehensive multimodal models.

How much information gets lost due to compression and sparsification of input data?

"Unified-IO is the first neural model to perform a large and diverse set of AI tasks spanning classical computer vision, image synthesis, vision-and-language, and natural language processing (NLP). Unified-IO achieves this broad unification by homogenizing every task's input and output into a sequence of tokens drawn from a discrete and finite vocabulary. Dense inputs such as images, masks, and depth maps are converted to sequences using a universal compressor, and sparse structured inputs such as bounding boxes and human joint locations are transcribed into language, which is naturally sequential.

This approach of unifying input and output data enables us to train a single sequence-to-sequence Unified IO model to perform tasks across more than 80 diverse computer vision and NLP benchmarks. ..."

Unified-IO, a new general purpose model from AI2

Wednesday, March 09, 2022

Hierarchical Graph Representation Learning with Differentiable Pooling

Recommendable! Well written paper! 

Two things that surprised me:
  1. The researchers used the COLLAB dataset to exemplify some of the results although it has only a single level cluster structure ("This is because many collaboration graphs in COLLAB show only single-layer community structures, which can be captured well with pre-computed graph clustering algorithm"). Given that the researchers used so several other chemical and social network datasets.
  2. The researchers used only maximal two DiffPool layers for all experiments

Hierarchical Graph Representation Learning with Differentiable Pooling (NIPS not NeurIPS 2018)

Friday, February 05, 2021

Researchers find that debiasing doesn’t eliminate racism from hate speech detection models

More baloney and obsessions coming out of the highly compensated ivory tower! The ideological hygiene police at work! 

What language is toxic? The one, the researchers are using?

"Biased associations have been a challenge in the development of classifiers for detecting toxic language, hindering both fairness and accuracy. ..."

Researchers find that debiasing doesn’t eliminate racism from hate speech detection models - Reader Mode

Here is the underlying research paper: