bandwidth
Architecting memory and storage in the AI era
With AI inference now driving enterprise workloads, organizations must rethink infrastructure for speed, efficiency, scalability, and performance per watt to unlock AI's real-world potential. The era of AI inference has arrived. Imagine a healthcare system analyzing millions of data points in real time to accelerate life-saving medical research, or an intelligent assistant instantly resolving thousands of complex customer needs at once. These real-world breakthroughs rely on advanced infrastructure acting as the engine of continuous intelligence, powering real-time services while also supporting an increasingly intelligent edge of IoT and consumer devices. However, in this inference-driven landscape, every delay, bottleneck, or wasted watt directly affects human outcomes and operating costs. This shift changes what infrastructure must deliver.
Apple's new M6 Mac mini and M5 Ultra Mac Studio: 10 details you might have missed
Gear Computers Apple's new M6 Mac mini and M5 Ultra Mac Studio: 10 details you might have missed The new Apple silicon desktops boast 2nm chips, quad-die architecture, and massive AI multipliers. Here is the fine print before you upgrade. More information Adding us as a Preferred Source in Google by using this link indicates that you would like to see more of our content in Google News results. We may earn revenue from the products available on this page and participate in affiliate programs. By signing up, you confirm you are 16+, will receive newsletters and promotional content and agree to our Terms of Use and acknowledge the data practices in our Privacy Policy .
Apple refreshes Mac mini and Mac Studio with new M6 and M5 Ultra chips
Apple hasn't forgotten about the Mac mini, after all. The company's popular diminutive desktop is getting refreshed today with a brand new M6 chip, while the Mac Studio is getting bumped up to the M5 Max and new M5 Ultra chips. Curiously, the Mac mini didn't get an upgrade to the M5 chip last year, but you'll be able to upgrade the newer model to an M5 Pro if you need more power. Pricing, unfortunately, has gotten worse since Apple's latest hikes: The Mac mini now starts at 899 with the M6 chip ( 100 more than before, with a meager 256GB of storage), while the Mac Studio starts at 2,499 with M5 Max. The Mac Studio with an M5 Ultra will run you at least 5,499. Over the last few years, AI aficionados have been buying up Mac minis en masse to run local models and AI assistants.
Lost in Transmission When and Why LLMs Fail to Reason Globally
Despite their many successes, transformer-based large language models (LLMs) continue to struggle with tasks that require complex reasoning over large parts of their input. We argue that these failures arise due to capacity limits on the accurate flow of information within LLMs. To formalize this issue, we introduce the bounded attention prefix oracle (BAPO) model, a new computational framework that models bandwidth constraints on attention heads, the mechanism for internal communication in LLMs. We show that several important reasoning problems like graph reachability require high communication bandwidth for BAPOs to solve; we call these problems BAPO-hard. Our experiments corroborate our theoretical predictions: GPT-4o, Claude, and Gemini succeed on BAPO-easy tasks and fail even on relatively small BAPO-hard tasks. BAPOs also reveal another benefit of chain of thought (CoT): we prove that breaking down a task using CoT can turn any BAPO-hard problem into a BAPO-easy one. Our results offer principled explanations for key LLM failures and suggest directions for architectures and inference methods that mitigate bandwidth limits.
Grids Often Outperform Implicit Neural Representation at Compressing Dense Signals
Implicit Neural Representations (INRs) have recently shown impressive results, but their fundamental capacity, implicit biases, and scaling behavior remain poorly understood. We investigate the performance of diverse INRs across a suite of 2D and 3D real and synthetic signals with varying effective bandwidth, as well as both overfitting and generalization tasks including tomography, super-resolution, and denoising. By stratifying performance according to model size as well as signal type and bandwidth, our results shed light on how different INR and grid representations allocate their capacity. We find that, for most tasks and signals, a simple regularized grid with interpolation trains faster and to higher quality than any INR with the same number of parameters. We also find limited settings-namely fitting binary signals such as shape contours-where INRs outperform grids, to guide future development and use of INRs towards the most advantageous applications.
Cost-Efficient LLMTraining with Lifetime-Aware Tensor Offloading via GPUDirect Storage
We present the design and implementation of a new lifetime-aware tensor offloading framework for GPU memory expansion using low-cost PCIe-based solid-state drives (SSDs). Our framework, TERAIO, is developed explicitly for large language model (LLM) training with multiple GPUs and multiple SSDs. Its design is driven by our observation that the active tensors take only a small fraction (1.7% on average) of allocated GPU memory in each LLM training iteration, the inactive tensors are usually large and will not be used for a long period of time, creating ample opportunities for offloading/prefetching tensors to/from slow SSDs without stalling the GPU training process. TERAIO accurately estimates the lifetime (active period of time in GPU memory) of each tensor with the profiling of the first few iterations in the training process. With the tensor lifetime analysis, TERAIO will generate an optimized tensor offloading/prefetching plan and integrate it into the compiled LLM program via PyTorch. TERAIO has a runtime tensor migration engine to execute the offloading/prefetching plan via GPUDirect storage, which allows direct tensor migration between GPUs and SSDs for alleviating the CPU bottleneck and maximizing the SSD bandwidth utilization. In comparison with state-of-the-art studies such as ZeRO-Offload and ZeRO-Infinity, we show that TERAIO improves the training performance of various LLMs by 1.47 on average, and achieves 80.7% of the ideal performance assuming unlimited GPU memory.
SD-KDE: Score-Debiased Kernel Density Estimation
We propose a method for density estimation that leverages an estimated score function to debias kernel density estimation (SD-KDE). In our approach, each data point is adjusted by taking a single step along the score function with a specific choice of step size, followed by standard KDE with a modified bandwidth. The step size and modified bandwidth are chosen to remove the leading order bias in the KDE, improving the asymptotic convergence rate. Our experiments on synthetic tasks in 1D, 2D and on MNIST, demonstrate that our proposed SD-KDE method significantly reduces the mean integrated squared error compared to the standard Silverman KDE, even with noisy estimates in the score function. These results underscore the potential of integrating score-based corrections into nonparametric density estimation.
Grids Often Outperform Implicit Neural Representation at Compressing Dense Signals
Implicit Neural Representations (INRs) have recently shown impressive results, but their fundamental capacity, implicit biases, and scaling behavior remain poorly understood. We investigate the performance of diverse INRs across a suite of 2D and 3D real and synthetic signals with varying effective bandwidth, as well as both overfitting and generalization tasks including tomography, super-resolution, and denoising. By stratifying performance according to model size as well as signal type and bandwidth, our results shed light on how different INR and grid representations allocate their capacity. We find that, for most tasks and signals, a simple regularized grid with interpolation trains faster and to higher quality than any INR with the same number of parameters. We also find limited settings-namely fitting binary signals such as shape contours-where INRs outperform grids, to guide future development and use of INRs towards the most advantageous applications.
Ugreen Maxidok review: The Thunderbolt 5 dock built for serious desks
When you purchase through links in our articles, we may earn a small commission. The Ugreen Maxidok combines Thunderbolt 5, DisplayPort 2.1, 2.5 Gigabit Ethernet, and an M.2 slot for SSDs up to 8TB. I put this premium docking station through its paces to see if it really delivers in everyday use. The Ugreen Maxidok 17-in-1 Thunderbolt 5 docking station is currently one of the most technically comprehensive Thunderbolt 5 docks on the market. It delivers the full bandwidth of 120Gbps, supplies the laptop with up to 140 watts, and combines this with 17 ports as well as an M.2 slot for an internal SSD upgrade.
A Unified Framework for Data-Free One-Step Sampling via Wasserstein Gradient Flows
We develop a unified theoretical framework for data-free one-step sampling from unnormalized target distributions based on Wasserstein gradient flows. For a broad class of standard f-divergence objectives, we show that the induced velocity field admits the universal form $\mathbf{V}(x)=w(r(x))\,β(x)$, where $β(x)=\nabla \log (p(x)/q(x))$ is shared across objectives and $w$ is determined solely by the choice of divergence. This decomposition shows that standard f-divergence drifts share the same asymptotic target distribution $p$ and differ primarily in how they redistribute transient repair effort across under-covered regions. To formalize this distinction, we derive a one-step regional-response theory for a soft under-coverage functional and obtain a compression--elasticity identity that links divergence choice to the geometry of mass transport into under-covered regions. We further extend the framework beyond the f-divergence family to the Log-Variance (LV) divergence, analyze how the reference distribution alters the resulting drift structure, and motivate a practical LV-inspired surrogate for data-free training. Based on this theory, we instantiate the framework with a KDE-based implementation and describe a complementary normalizing-flow route, enabling one-step inference after training. Experiments on multimodal Gaussian-mixture benchmarks are consistent with the theoretical predictions and demonstrate effective one-step sampling on these targets.