Showing posts with label University of California Berkeley. Show all posts
Showing posts with label University of California Berkeley. Show all posts

Sunday, July 12, 2026

DiPOD: Diffusion Policy Optimization without Drifting Apart

This could be an interesting new paper by Pieter Abbeel and his team!

From the abstract:
"RL post-training has become increasingly pivotal for improving diffusion policies, but existing diffusion policy-gradient methods are often unstable and cannot achieve reliable policy improvement.
We identify the cause as the double-drift phenomenon: optimizing a variational surrogate can let the ELBO separate from the true log-likelihood, which then makes the resulting proxy policy gradient misaligned with the true policy gradient of expected return.
We propose DiPOD, a diffusion policy optimization framework that maintains tight-bound behavior throughout training by interleaving self-distillation with policy-improving gradient updates.
This leads to a simple and practical algorithm: augmenting each diffusion policy-gradient update with an on-policy ELBO regularizer.
Across diffusion language model post-training and continuous-control diffusion policies, DiPOD substantially stabilizes training and reaches higher rewards than previous methods."


[2606.13795] DiPOD: Diffusion Policy Optimization without Drifting Apart (preprint, open access)







Saturday, August 30, 2025

On Do What? Teaching Vision-Language-Action Models to Reject the Impossible

An interesting attempt!

Caveat: I did not read this paper. It is very short with only 9 pages total.

From the abstract:
"Recently, Vision-Language-Action (VLA) models have demonstrated strong performance on a range of robotic tasks. These models rely on multimodal inputs, with language instructions playing a crucial role -- not only in predicting actions, but also in robustly interpreting user intent, even when the requests are impossible to fulfill.
In this work, we investigate how VLAs can recognize, interpret, and respond to false-premise instructions: natural language commands that reference objects or conditions absent from the environment. We propose Instruct-Verify-and-Act (IVA), a unified framework that
(i) detects when an instruction cannot be executed due to a false premise,
(ii) engages in language-based clarification or correction, and 
(iii) grounds plausible alternatives in perception and action.
Towards this end, we construct a large-scale instruction tuning setup with structured language prompts and train a VLA model capable of handling both accurate and erroneous requests. Our approach leverages a contextually augmented, semi-synthetic dataset containing paired positive and false-premise instructions, enabling robust detection and natural language correction. Our experiments show that IVA improves false premise detection accuracy by 97.56% over baselines, while increasing successful responses in false-premise scenarios by 50.78%."

[2508.16292] Do What? Teaching Vision-Language-Action Models to Reject the Impossible




Thursday, February 20, 2025

Generative AI tool Evo 2 marks a milestone in biology by predicting form and function of proteins in the DNA of all domains of life except viruses

Amazing stuff!

"Imagine being able to speed up evolution – hypothetically – to learn which genes might have a harmful or beneficial effect on human health. Imagine, further, being able to rapidly generate new genetic sequences that could help cure disease or solve environmental challenges. Now, scientists have developed a generative AI tool that can predict the form and function of proteins coded in the DNA of all domains of life, identify molecules that could be useful for bioengineering and medicine, and allow labs to run dozens of other standard experiments with a virtual query – in minutes or hours instead of years (or millennia). ...

Evo 2 was trained on a dataset that includes all known living species, including humans, plants, bacteria, amoebas, and even a few extinct species. ...

In this way, Evo 2 is able to generate – to write – new genetic code that has never existed before. With Evo 2, you can enter a sequence of up to 1 million nucleotides. ...

Evo 2, on the other hand, also includes the known genomes of 15,000 or so plants and animals – the eukaryotes – which includes humans. Our dataset has now expanded from about 300 billion nucleotides to almost 9 trillion with Evo 2. In terms of safety, we have left out the genomes of viruses to prevent Evo 2 from being used to create new or more dangerous diseases. ...

If you want to design a new gene, you prompt the model with the beginning of a gene sequence of base pairs, and Evo 2 will autocomplete the gene. ...

With Evo 2, we can be more direct and steer toward mutations that have useful functions. Evo 2 also includes machine learning models that will tell you if the sequence exists in nature and predict how this new sequence will function in real life. ...

The model is actually very good at distinguishing which mutations are just random, harmless variations and which cause disease. ..."

From the abstract:
"All of life encodes information with DNA. While tools for sequencing, synthesis, and editing of genomic code have transformed biological research, intelligently composing new biological systems would also require a deep understanding of the immense complexity encoded by genomes. We introduce Evo 2, a biological foundation model trained on 9.3 trillion DNA base pairs from a highly curated genomic atlas spanning all domains of life. We train Evo 2 with 7B and 40B parameters to have an unprecedented 1 million token context window with single-nucleotide resolution. Evo 2 learns from DNA sequence alone to accurately predict
the functional impacts of genetic variation—from noncoding pathogenic mutations to clinically significant BRCA1 variants—without task-specific fine tuning.
Applying mechanistic interpretability analyses, we reveal that Evo 2 autonomously learns a breadth of biological features, including exon–intron boundaries, transcription factor binding sites, protein structural elements, and prophage genomic regions.
Beyond its predictive capabilities, Evo 2 generates mitochondrial, prokaryotic, and eukaryotic sequences at genome scale with greater naturalness and coherence than previous methods.
Guiding Evo 2 via inference-time search enables controllable generation of epigenomic structure, for which we demonstrate the first inference-time scaling results in biology. We make Evo 2 fully open, including model parameters, training code, inference code, and the OpenGenome2 dataset, to accelerate the exploration and design of biological complexity."

Generative AI tool marks a milestone in biology | Stanford Report "Trained on a dataset that includes all known living species – and a few extinct ones – Evo 2 can predict the form and function of proteins in the DNA of all domains of life and run experiments in a fraction of the time it would take a traditional lab."

AI can now model and design the genetic code for all domains of life with Evo 2 "Arc Institute develops the largest AI model for biology to date in collaboration with NVIDIA, bringing together Stanford University, UC Berkeley, and UC San Francisco researchers"





Wednesday, March 27, 2024

On Unfamiliar Finetuning Examples Control How Language Models Hallucinate

Appears to be an interesting research paper on the psychedelics of large language models. 😊

Caveat: I have not yet read this paper published by UC Berkeley and Google. The senior author, i.e. Sergey Levine, is a well known and highly cited researcher.

From the abstract: 
"Large language models (LLMs) have a tendency to generate plausible-sounding yet factually incorrect responses, especially when queried on unfamiliar concepts. In this work, we explore the underlying mechanisms that govern how finetuned LLMs hallucinate. Our investigation reveals an interesting pattern: as inputs become more unfamiliar, LLM outputs tend to default towards a ``hedged'' prediction, whose form is determined by how the unfamiliar examples in the finetuning data are supervised. Thus, by strategically modifying these examples' supervision, we can control LLM predictions for unfamiliar inputs (e.g., teach them to say ``I don't know''). Based on these principles, we develop an RL approach that more reliably mitigates hallucinations for long-form generation tasks, by tackling the challenges presented by reward model hallucinations. We validate our findings with a series of controlled experiments in multiple-choice QA on MMLU, as well as long-form biography and book/movie plot generation tasks."

[2403.05612] Unfamiliar Finetuning Examples Control How Language Models Hallucinate

Friday, January 01, 2021

New framework can train a robotic arm on 6 grasping tasks in less than an hour

Good news! I have not had time yet to read this new research paper by Pieter Abbeel and his collaborators, but Pieter Abbeel is a well known and highly cited researcher in the field.

"... Building on these advances, we present a Framework for Efficient Robotic Manipulation (FERM) that utilizes data augmentation and unsupervised learning to achieve extremely sample-efficient training of robotic manipulation policies with sparse rewards. We show that, given only 10 demonstrations, a single robotic arm can learn sparse-reward manipulation policies from pixels, such as reaching, picking, moving, pulling a large object, flipping a switch, and opening a drawer in just 15-50 minutes of real-world training time. ..."

New framework can train a robotic arm on 6 grasping tasks in less than an hour | VentureBeat

Here is the link to the respective research paper:



Tuesday, July 21, 2020

Notes on CURL: Contrastive Unsupervised Representations for Reinforcement Learning

Very interesting state of the art research in reinforcement learning (RL) by University of California Berkeley with Pieter Abbeel.

However, when you look at table 2 on page 8 (or see below) of this paper you will be astonished to find out that of the 26 Atari games (How old is Atari? 1984-2013) that were benchmarked against some of the best state of the art RL models, human performance still wins 22 games. At least, super human performance is achieved in every of those 4 games where RL wins. However, human performance still beats the best RL by a wide margin in 18 of the 26 games. This is almost a bit shocking! 

[2004.04136] CURL: Contrastive Unsupervised Representations for Reinforcement Learning