Showing posts with label machine learning & artificial intelligence. Show all posts
Showing posts with label machine learning & artificial intelligence. Show all posts

Friday, October 09, 2026

How do AI chatbots and large language models ‘reason’ like humans?

This could be an interesting new research paper by Tal Linzen and Paul Smolensky

"... The study provides evidence that the internal representations of neural networks implicitly realize symbolic structure. ...

For the new study, ... analyzed the internal functions of several high-profile LLMs. They found that neural networks implicitly grasp symbolic structure within the numeric lists that drive them.

“Our analysis demonstrates that, despite appearances, the vectors powering LLMs are organized in a way that is equivalent to the symbolic structures that drive much of human cognitive function,” ...

“With LLMs, we show that vectors are organized in a very particular way that gives rise to emergent symbolic structure — symbolic structure that is present at a high level despite not being apparent at a low level.” ...

The researchers found that they could replace an LLM’s entire representation-generating process with “role-filler” approximations embodying symbolic structures and the AI system’s behavior would remain largely unchanged. The finding held for small-scale neural networks trained to manipulate lists as well as for seven large-scale LLMs operating in four domains that have long been central in research on symbolic intelligence: language, arithmetic, logic, and computer coding. ..."

From the abstract:
"Modern systems in artificial intelligence (AI) somehow excel in domains for which they seem poorly suited.
Intelligence has traditionally been modeled as operating over structured combinations of symbols, such as logical formulas.
However, the strongest modern AI systems are based on neural networks, which instead represent information in continuous vectors. Vectors seem inadequate for capturing the structure of language, logic, and other cognitive domains, yet neural networks achieve impressive performance in these areas.
How do they do it? In this work, we propose a potential answer: Despite appearances, perhaps the internal representations of neural networks implicitly realize symbolic structure.
In support of this hypothesis, we show that the vector representations of a variety of neural networks can be closely approximated with symbolic structures: we can replace the network's entire representation-generating process with a closed-form equation instantiating a symbolic structure, and the network's behavior remains largely unchanged.
This finding holds for both small-scale neural networks trained to manipulate lists as well as large language models (LLMs) operating in four domains that are central in symbolic traditions: arithmetic, logic, computer code, and language.
Further, our symbolic approximation allows us to modify an LLM's behavior in targeted ways via precise interventions on its internal representations, showing that the LLM's behavior is reliant on the symbolic structures we have identified. This work provides a potential way to reconcile longstanding symbolic conceptions of intelligence with the vector-based nature of modern AI."

How do AI chatbots ‘reason’ like humans? A Yale-led study offers clues | Yale News "The neural networks that power popular AI chatbots process data much differently than humans but capably perform tasks involving language, logic, and math. A new study sheds light on how those systems “think.”"





Thursday, October 08, 2026

Sharing AI progress in mathematics by released more than 700 AI-generated papers | OpenAI

This is only the beginning! What an avalanche! Breathtaking!

How fast will mathematics (the queen of the sciences according to Carl Friedrich Gauss) now advance going forward? Maybe in the coming months/quarters we will see more progress than in the past 1000 years or so.

"OpenAI released more than 700 AI-generated papers presenting results related to more than 370 previously unsolved mathematical problems, including some of the field’s hardest. The massive dump marks yet another landmark in the use of AI to attack complex math—but experts are split over whether such releases are good for the field. Learn about what’s in the papers—and how researchers are reacting."

Sharing AI progress in mathematics | OpenAI

OpenAI’s $25 billion research charity

Good news!

"The OpenAI Foundation — the charity arm of the artificial-intelligence behemoth — is on track to become one of the wealthiest non-profit organizations on the planet.
It has so far provided a $40 million grant for a project at the University of North Carolina at Chapel Hill to develop cancer vaccines, as well as funding for Alzheimer’s disease research and a pilot project to repurpose data from bankrupt biotechnology companies. ..."

Nature Briefing: Cancer

How to spend $25 billion on science: OpenAI’s staggeringly rich research charity gets under way "Jacob Trefethen, who leads life sciences at the OpenAI Foundation, speaks to Nature about the organization’s ambition to cure disease."

OpenAI Foundation "Our mission is to ensure artificial general intelligence benefits all of humanity."

Supporting Communities: Meet the 2026 People-First AI Fund Grantees "Our second People-First AI Fund is deploying $50 million to 163 nonprofits exploring how AI can help communities across the United States access essential services, strengthen creative and cultural institutions, and support local journalism."


Every grant begins with people. These are some of the communities the People-First AI Fund exists to support.


Tuesday, October 06, 2026

Context Language Models

This is a new research paper by Luke Zettlemoyer and his team!

Introducing a file seems to be a very trivial improvement.

From the abstract:
"We introduce Context Language Models (CLMs), language models that natively manage their own context. We implement this by treating the context as a file and allowing the model to make unrestricted updates to this file. This allows the model to learn what is most important to maintain in context, and naturally extends to multi-agent systems where multiple agent contexts coexist as files.
Building CLMs zero-shot with existing models outperforms SOTA context management strategies across a variety of tasks: 11.4% higher accuracy with 21.5% fewer FLOPs on BrowseComp-Plus, 5% higher scores with 59% fewer FLOPs on 12-hour EdgeBench, and 65% greater improvement with the same compute on a 24-hour multi-repository agent-swarm task.
Moreover, by shifting context management from external harness control to intrinsic model behavior, CLMs naturally enable both in-context and parametric learning of context-management strategies.
We show that CLMs can be steered with natural-language instructions evolved through a standard skill-optimization loop, improving held-out accuracy by up to 35.9 points on a context-management task while reducing compute.
We also introduce an online reinforcement learning method for CLMs, improving Qwen3.5-9B performance on BrowseComp-Plus by 47.6% while using 12% fewer FLOPs.
Finally, we co-design Suffix Cache Reuse for CLM serving, further reducing server-side compute by 35% relative to standard SGLang at matched performance."

[2609.37725] Context Language Models (preprint, open access)




Monday, October 05, 2026

Ukraine’s AI-Assisted Gun Turrets Shoot Down Russian Jet Drones

Good news! Drone countermeasures are making progress!

"... Speaking on Ukrainian state broadcast television, Air Force spokesman Yurii Ihnat said the systems had successfully intercepted jet-powered unmanned aircraft, including the Geran-5. He credited artificial intelligence and machine vision with helping the turrets engage the targets. ..."

‘I Was a Skeptic’: Ukraine’s AI-Assisted Gun Turrets Shoot Down Russian Jet Drones "In brief: Ukraine says robotic machine-gun turrets using AI and machine vision have downed Russian jet-powered drones, including the Geran-5, as Kyiv expands deployments despite the systems’ limited range."

Always-On Experimentation (for medical treatments/pharmaceuticals)

This could be an interesting new research paper by Michael I. Jordan and his team!

From the abstract:
"Generative AI has dramatically accelerated the rate at which new treatments---from novel pharmaceuticals to online marketing campaigns---can be conceived and deployed. As a result, modern experimentation platforms often run continuously, with treatments added as they are ready and removed when they underperform.
We formalize this "Always-On" experimental setting, in which treatments can be dynamically generated, added to, and removed from a running experiment, and study the statistical problem of deciding whether to accept or reject each treatment while controlling for the false discovery rate.
We develop sequential tests that achieve time-uniform Type-I error control under arbitrary stopping times and "predictable" treatment schedules. Our approach builds on the testing-by-betting framework: we construct test supermartingales for testing the average treatment effect of each treatment, and show that the construction of these test supermartingales is growth-rate optimal in an almost-sure sense."

[2609.38695] Always-On Experimentation (preprint, open access)




Fast Generative DeOcclusion for Vision and Robotics

Amazing stuff! Less clutter is better for robots! 😊

This is a new research paper by Alan Yuille and his team published by CVPR 2026!

From the abstract:
"Occlusion remains a core challenge in vision, as projecting a 3D world into 2D inevitably hides much of the scene geometry.
We present FoundationDeOcclusion, a fast generative framework for fast occlusion recovery that reconstructs hidden geometry and appearance from partial visual observations. FoundationDeOcclusion first identifies occluded objects from monocular image sequences using Grounded-SAM, then matches them across views using depth cues estimated by 3D reconstruction models.
We introduce a novel geometry-aware linear de-occlusion Transformer (GL-DoT) that synthesizes the missing regions, which are then integrated into the 3D scene through depth-aware fusion.
As a result, FoundationDeOcclusion improves 3D scene reconstruction under occlusion. Notably, GL-DoT attains strong performance with only four denoising steps and runs at near real-time speed (6 FPS),
challenging the prevailing belief that generative models are impractical for time-critical robotic tasks.
Finally, we demonstrate the effectiveness of FoundationDeOcclusion in real-robot navigation and manipulation. Without bells and whistles, it boosts the prior-art AnyGrasp by 43.8% in success rate without task-specific tuning, setting a new state of the art for mobile manipulation under occlusion."

Fast Generative DeOcclusion for Vision and Robotics | OpenReview






Sunday, October 04, 2026

What if automating AI R&D triggers an intelligence explosion?

This could be an interesting new research paper by Geoffrey Hinton, Yoshua Bengio, Thore Graepel, Philip H. S. Torr, Jeff Clune their team.

I would say the fear that the progress of AI could get out of control is quite overblown. Humans can handle and manage it and they will grow with it. We have done it before with nuclear power and genetics.

Keep government intervention to a minimum!

Maybe in a few years we will see super intelligent humans.

"In contrast to even a year ago, AI systems now write most of the code inside the companies that build them.
As more of the AI research and development (R&D) pipeline is automated, could AI progress radically accelerate in an "intelligence explosion," where years of advances are compressed into months or less?
Preliminary evidence suggests that it could. In this work, we assess this evidence, analyze an intelligence explosion's potential impacts, and propose policy responses. AI systems are on track to automate most AI R&D work within a few years, and possibly all of it. If this triggers an intelligence explosion, it could dramatically bring forward AI's benefits, but also pose extreme risks: capabilities growth could accelerate far beyond what society can keep up with, humanity could lose control over superhuman AI systems, and checks on power within and between states, companies, and branches of government could be severely eroded.
Although there remains much uncertainty about these possibilities, the high stakes warrant serious further attention. Policymakers should urgently obtain more visibility into the automation of AI R&D, develop ways to steer and constrain an intelligence explosion, and prepare society to adapt to an intelligence explosion's impacts."

[2609.36054] What if automating AI R&D triggers an intelligence explosion?




When Do Models Admit They Are Wrong? Failure Disclosure Is Unstable Under Reinforcement Learning

This could be an interesting new research paper by Noah D. Goodman and his team!

From the abstract:
"Outcome-based reinforcement learning can produce models with similar task performance but very different ways of communicating about their mistakes.
We study failure disclosure: whether a model admits that an attempted solution failed rather than staying silent or presenting it as successful.
Across repeated outcome-only GRPO training runs, failure disclosure varies far more than task accuracy.
The pattern extends to a second reasoning task and stabilized PPO, persists at 7B, and also appears in an instruction-conditioned 32B setting.
We also find that small floating-point and sampling differences during training can redirect reporting behavior even when the task objective and earlier training history are held fixed.
Additional tests show that failure disclosure is not a single decision: Checking the answer, entering a report, and completing the admission can separate, and the weak point depends on the task and response format.
Further, experiments with neutral controls show more broadly that behaviors left weakly constrained by training are especially likely to vary across runs, of which failure disclosure is an example. We can reduce variability in failure disclosure by discouraging the model from drifting from its starting policy on failed, well-formed responses. This makes reporting substantially more consistent, though its effect on task performance depends on the setting.
Stable task accuracy therefore does not guarantee stable safety-relevant behavior: Researchers should measure these behaviors directly across runs and design training methods that keep them reliable."

[2609.33220] When Do Models Admit They Are Wrong? Failure Disclosure Is Unstable Under Reinforcement Learning (preprint, open access)




Saturday, October 03, 2026

Open-sourcing AstaBrief, the fast report-generation model in Asta

Good news!

"This week, we’re releasing two new open artifacts for researchers and model builders: AstaBrief 8B, a model for generating fast, cited scientific reports, and Olmo-core 3, our redesigned framework for training large mixture-of-experts (MoE) models. ...

AstaBrief
 
Reviewing and synthesizing scientific literature can be time-consuming. AstaBrief 8B is an 8-billion-parameter model trained to generate a cited report from a research question and retrieved scientific literature, giving researchers a faster starting point for exploring a topic, checking sources, and refining their questions. ..."


Open-sourcing AstaBrief, the fast report-generation model in Asta | Ai2

Friday, October 02, 2026

The US and China will have a red phone for AI emergencies, while Europe is left to watch

Headline of the day! Goes to whether Europe is a museum and a hoard of antiquities!

ZhōngHuá Mundus | The US and China will have a red phone for AI emergencies, while Europe is left to watch

Self-Play Pretraining with Zero Data

This could be an interesting new paper by Noah D. Goodman and Yoav Levine

From the abstract:
"Advances in language modeling have been driven by scaling pretraining on ever more data. Yet, the training data is still largely curated on the model's behalf.
A more general approach to pretraining would let the model learn to generate the data most useful for its own improvement. This would provide an effectively unbounded source of training data, limited by compute rather than human knowledge.
We introduce Self-Play Pretraining with Zero Data, an initial proof-of-concept towards realizing this vision. Our procedure casts synthetic data generation as a search over the space of all computable structure, taking inspiration from Solomonoff induction.
Starting from random initialization, two models learn in tandem:
a generator proposes programs interpreted by a universal Turing machine, generating byte sequences, while
a learner autoregressively predicts these byte sequences.
The learner is trained with standard cross-entropy, while the generator is trained with reinforcement learning to produce sequences at the frontier of the learner's capabilities, yielding an adaptive curriculum.
A universal Turing machine gives us a search space over all computable data-generating processes, imposing little domain-specific structure, and self-play searches over this space for useful training data.
We test whether zero-shot performance on natural data improves predictably with self-play compute; this is a clean test of transfer since neither generator nor learner is trained on natural data. Across several natural datasets, zero-shot loss exhibits predictable scaling in compute. The models also exhibit in-context learning, and discover recognizable mathematical sequences during training."

[2609.30063] Self-Play Pretraining with Zero Data (preprint, open access)






Musk's AI chatbot Grok reportedly encouraged President Trump to capture Venezuela's president

Well, the ancient Greeks called their oracle!

It is called getting a second opinion!

Get used to it! This will happen more often that elected representatives consult with AI before major decisions.

"In December 2025, about a month before the U.S. invaded Venezuela and captured its president, Nicolás Maduro, President Trump had a secret meeting with Elon Musk, Time magazine reports. This was roughly seven months after Musk left his role in the Trump administration at the Department of Government Efficiency (DOGE).

During this meeting, Trump “spent hours” talking to Musk’s Grok chatbot, including asking how Venezuelans would respond to the capture of their president, a source told Time. Just a few months earlier, in September, Trump began ordering U.S. military strikes against Venezuelan boats that Trump alleged were involved in drug trafficking. ..."

"... Since returning to office, Trump has marveled at AI’s capabilities. He is increasingly in thrall to the tech executives courting him.
In December 2025, Elon Musk returned to the Oval Office for a clandestine meeting. Trump spent hours asking Grok, Musk’s AI chatbot, questions about his presidency and legacy, according to an official present. At the time, Trump was ordering missile strikes against Venezuelan boats he alleged were smuggling drugs into the U.S.
He asked Grok how Venezuelans would react if the U.S. captured Maduro. According to the official present, the chatbot responded that Maduro was a repressive and deeply unpopular dictator and that many Venezuelans would likely celebrate his downfall. After Trump ordered the mission to seize Maduro, the following month, celebrations broke out in the streets. Trump, according to officials, came away thinking Grok was ingenious. ..."

Musk's AI chatbot Grok reportedly encouraged Trump to capture Venezuela's president | TechCrunch





Thursday, October 01, 2026

Hospitals Used AI To Inflate Medical Bills By Nearly $1 trillion Report Finds

Bad news!

"Hospitals’ use of artificial intelligence tools hiked healthcare spending by nearly $1 billion over a two-year period, per a Blue Cross Blue Shield Association analysis released Thursday.

BCBSA’s new study shows that a major rise in patients being documented as having complex conditions led to $942 million in additional medical care costs for Blue Cross and Blue Shield (BCBS) companies between 2023 and ‌2025. ..."

"Key Findings:
  • A sharp increase in patients being documented as having complex conditions appears to have added an estimated $942 million in healthcare spending for BCBS companies between 2023 and 2025, driving higher premiums and out-of-pocket costs for families, employers and taxpayers.
  • There is a clear disconnect between coding and treatment — our research shows that coding for care has changed, but there is no evidence of corresponding change in care delivered.
  • Patients often fell into a higher-cost billing category after a secondary diagnosis, a condition identified in addition to the patient's primary reason for receiving care. This secondary diagnosis may be derived from single laboratory values, making it particularly well-suited for detection by AI tools. Approximately 70% of the additional costs, more than $650 million, were tied to secondary diagnoses.
  • This growth in complex coding comes as more than 60% of hospital systems began using AI coding tools, which can document patient visits, analyze lab reports and doctors’ notes.
...."

Hospitals Used AI To Inflate Medical Bills By Nearly $1,000,000,000, Report Finds |


Credits: AZ Free News

Tuesday, September 29, 2026

Trump says AI leaders signed a 'constitution' to police themselves

Good news! Self regulation is a good idea at this point! No need for the US Congress to pass laws!

(254) Trump says AI leaders signed a 'constitution' to police themselves - YouTube


New formulation helps RNA vaccines withstand high temperatures like room temperature for up to a year

Good news!

"RNA vaccines, which have been proven effective against Covid-19, are now being developed for many other diseases, including cancer. One of the drawbacks to these vaccines is that they require ultracold storage, but researchers from MIT have found a promising way to overcome that limitation.

With help from an AI algorithm, the researchers tweaked the formulation surrounding the lipid nanoparticles that are typically used to deliver mRNA vaccines, making the vaccines more heat-resistant. Using this approach, they formulated vaccines that could remain stable even when stored at room temperature for up to a year, or at nearly 100 degrees Fahrenheit for two months. ..."

From the abstract:
"The instability of mRNA−lipid nanoparticles (LNPs) necessitates ultra-cold storage, limiting global distribution and their broader application in advanced delivery systems.
Solid-state, water-free formulations enhance thermostability and enable integration into emerging delivery modalities such as microneedle patches.
Previous efforts to stabilize mRNA−LNPs have been constrained by narrow formulation scope and low-throughput screening methods.
Here we introduce Algorithm-Guided Experimental design for lipid Nanoparticle Thermostabilization (AGENT), an artificial intelligence (AI)-driven framework that couples high-throughput experimentation with Bayesian optimization to identify thermostable mRNA−LNP formulations.
AGENT extracts maximal information from sparse experimental datasets, enabling efficient formulation optimization in six iterations completed within 1 month.
We stabilized mRNA vaccines with two clinically relevant LNPs representative of the Moderna (SM-102-based) and Pfizer-BioNTech (ALC-0315-based) compositions into solid-state formulations that retained 100% bioactivity after storage at 37 °C for more than 2 months.
In rodents and non-human primates, thermostable, solid-state vaccine formulations induced antigen-specific immune responses non-inferior to those elicited by intramuscular delivery of freshly prepared soluble vaccines."

New formulation helps RNA vaccines withstand high temperatures | MIT News | Massachusetts Institute of Technology "MIT engineers have found a way to stabilize the lipid nanoparticles used to deliver RNA vaccines, which could allow the vaccines to be more widely distributed."




Fig. 1 Exploration of the design space for solid-state mRNA-LNPs formulated via vacuum drying.


Monday, September 28, 2026

AMD will acquire Fei-Fei Li's World Labs for $8.2 billion

Good news! Good for her! Fei-Fei Li is a top, leading ML & AI researcher at Stanford University!

"AMD is acquiring World Labs, one of the leading developers of deep learning models intended to understand physical reality, in a $8.2 billion deal, the two companies said today.

World Labs justified the deal in a statement saying that AI development required “close collaboration across model research, systems and compute.” AMD, in turn, says that understanding frontier workloads, like those created at World Labs, will shape its chip-making roadmap. ..."

AMD will acquire Fei-Fei Li's World Labs for $8.2 billion | TechCrunch




Scientist accuses Anthropic of plagiarizing team’s biology work

Bad news, if confirmed!

E.g. it appears the University of Copenhagen did not release any news regarding this accusation.

"Scientist accuses Anthropic of plagiarizing team’s biology work
Anthropic said its AI agents used Claude to independently discover new enzymes it calls ARTs (array-associated reverse transcriptases) by searching gene databases.

But Copenhagen computational biologist Mario Rodríguez Mestre says he and colleagues had studied the same enzymes, which they nicknamed ‘jumbotrons,’ for four years. The team shared unpublished findings with Claude while using it for research tasks. Mestre says a key finding Anthropic credited to its AI, involving RNA molecules tied to the enzymes, matches work he says his team completed over a year earlier and had already discussed with Claude.
Anthropic responded that it found no prior published work on the ART system and that Claude was not trained on user transcripts.
The dispute echoes a similar case this month in which OpenAI claimed to have solved a math problem that mathematician Tristan Buckmaster said he had been working on using OpenAI’s models. Mestre says he is now winding down his use of Claude and moving to other AI models over concerns his research may have been used without attribution." (Data Points)

Did Anthropic’s A.I. Really Make a Scientific Discovery on Its Own? (behind paywall)

Nvidia releases tools to rein in rogue AI agents

AI nannies for AI agents?

"Nvidia released the Open Agent Safety Platform, software meant to help developers test and monitor AI agents and stop them from acting outside their intended boundaries.
Nvidia says the platform includes
(i) a watchdog that can quarantine misbehaving agents within milliseconds and
(ii) tools to control what agents can see, access, and interact with.

The release follows several incidents in which AI agents from OpenAI, Anthropic, Meta, and Google went rogue, including an OpenAI agent that Australian Prime Minister Anthony Albanese said infiltrated a government-services website, and Anthropic’s disclosure that its test software hacked outside companies without its knowledge in three cases since April.
Nvidia said each incident followed the same pattern: agents escaped sandbox safety controls to complete their assigned tasks. For developers building AI agents, the platform offers a vendor-backed option for enforcing boundaries and monitoring behavior in production; however, there is currently no independent verification of how well it prevents the sandbox-escape pattern Nvidia describes." (Data Points)

NVIDIA Open Agent Safety Platform: A Reference for Continuous In-Silicon Agent Monitoring "Secure AI agents with a safety enforcement layer spanning software and hardware"