Showing posts with label optimization. Show all posts
Showing posts with label optimization. Show all posts

Sunday, April 19, 2026

On Momentum Further Constrains Sharpness at the Edge of Stochastic Stability

This could be an interesting, but narrowly focused, new paper by Tomaso Poggio and his team!

Caveat: I have not read the paper (40 pages total) yet.

From the abstract:
"Recent work suggests that (stochastic) gradient descent self-organizes near an instability boundary, shaping both optimization and the solutions found.
Momentum and mini-batch gradients are widely used in practical deep learning optimization, but it remains unclear whether they operate in a comparable regime of instability.
We demonstrate that SGD with momentum exhibits an Edge of Stochastic Stability (EoSS)-like regime with batch-size-dependent behavior that cannot be explained by a single momentum-adjusted stability threshold.
Batch Sharpness (the expected directional mini-batch curvature) stabilizes in two distinct regimes:
at small batch sizes it converges to a lower plateau , reflecting amplification of stochastic fluctuations by momentum and favoring flatter regions than vanilla SGD; at large batch sizes it converges to a higher plateau , where momentum recovers its classical stabilizing effect and favors sharper regions consistent with full-batch dynamics.
We further show that this aligns with linear stability thresholds and discuss the implications for hyperparameter tuning and coupling."

[2604.14108] Momentum Further Constrains Sharpness at the Edge of Stochastic Stability

Friday, March 29, 2024

Multithreading software tweak doubles computer processing speed, halves energy use

Good news, but not the latest news! This could be a breakthrough!

Unfortunately, this work, published in October 2023, has not yet been cited very often, which is not a good sign or is this work overlooked? (Google Scholar: 0, Semantic Scholar: 1 citation).

I believe, multithreading is an old concept and has been widely applied since about the late 1990s.

"...  In a new study, researchers ... demonstrate a method where existing diverse components operate simultaneously to greatly improve processing speed and reduce energy consumption. ...
The researchers’ framework, called simultaneous and heterogeneous multithreading (SHMT), moves away from traditional programming models that can only delegate a region of code exclusively to one kind of processor, leaving other resources idling and not contributing to the current function.
Instead, SHMT exploits the ... heterogeneity – of multiple components, breaking the computational function up to share it among them. In other words, it’s a type of parallel processing. ...
A set of virtual operations (VOPs) allows a CPU program to ‘offload’ a function to a virtual hardware device. During program execution, a runtime system drives SHMT’s virtual hardware, gauging the hardware resource’s ability to make scheduling decisions.

SHMT uses a quality-aware work-stealing (QAWS) scheduling policy that doesn't hog resources, but helps maintain quality control and workload balance. The runtime system divides VOPs into one or more high-level operations (HLOPs) to simultaneously use multiple hardware resources.

Then, SHMT’s runtime system allocates these HLOPs to the task queues of the target hardware. Because HLOPs are hardware-independent, the runtime system can adjust the task assignment as required. ...
They tested the SHMT concept using benchmark applications, and found that the framework with the best-performing QAWS policy knocked it out of the park, with a 1.95X boost to speed and a remarkable 51% cut in energy consumption compared to the baseline method. ...
The researchers say the implications for SHMT are huge. Yes, software apps on your existing phones, tablets, desktops and laptops could use this new software library to achieve some pretty wild performance gains. But it could also reduce the need for expensive, high-performance components, leading to cheaper and more efficient devices. ..."

From the abstract:
"The landscape of modern computers is undoubtedly heterogeneous, as all computing platforms integrate multiple types of processing units and hardware accelerators. However, the entrenched programming models focus on using only the most efficient processing units for each code region, underutilizing the processing power within heterogeneous computers.

This paper simultaneous and heterogenous multithreading (SHMT), a programming and execution model that enables opportunities for “real” parallel processing using heterogeneous processing units. In contrast to conventional models, SHMT can utilize heterogeneous types of processing units concurrently for the same code region. Furthermore, SHMT presents an abstraction and a runtime system to facilitate parallel execution. More importantly, SHMT needs to additionally address the heterogeneity in data precision that various processing units support to ensure the quality of the result.

This paper implements and evaluates SHMT on an embedded system platform with a GPU and an Edge TPU. SHMT achieves up to 1.95 × speedup and  51.0% energy reduction compared to GPU baseline."

Software tweak doubles computer processing speed, halves energy use Existing processors in PCs, smartphones and other devices can be supercharged for enormous power and efficiency gains using a new parallel processing software framework designed to eliminate bottlenecks and use multiple chips at once.

Monday, August 14, 2023

On Provably Faster Gradient Descent via Long Steps

I have not read this paper yet, but perhaps it is going to trigger a new wave of research on how to further optimize and accelerate the training of models in machine learning & AI.

From the abstract:
"This work establishes provably faster convergence rates for gradient descent in smooth convex optimization via a computer-assisted analysis technique. Our theory allows nonconstant stepsize policies with frequent long steps potentially violating descent by analyzing the overall effect of many iterations at once rather than the typical one-iteration inductions used in most first-order method analyses. We show that long steps, which may increase the objective value in the short term, lead to provably faster convergence in the long term. A conjecture towards proving a faster O(1/TlogT) rate for gradient descent is also motivated along with simple numerical validation."

[2307.06324] Provably Faster Gradient Descent via Long Steps

Tuesday, February 18, 2020

ZeRO & DeepSpeed: New system optimizations enable training models with over 100 billion parameters

Bigger is better! Microsoft raises the bar by creating a state of the art largest language model! I am sure Google or OpenAI or others will pick up the challenge very soon!

ZeRO & DeepSpeed: New system optimizations enable training models with over 100 billion parameters - Microsoft Research: The latest trend in AI is that larger natural language models provide better accuracy; however, larger models are difficult to train because of cost, time, and ease of code integration. Microsoft is releasing an open-source library called DeepSpeed, which vastly advances large model training by improving scale, speed, cost, and usability, unlocking the ability to …

Here is the respective ArXiv preprint paper:
ZeRO: Memory Optimization Towards Training A Trillion Parameter Models