Government
Compressing Large Language Models using Low Rank and Low Precision Decomposition
Saha, Rajarshi, Sagan, Naomi, Srivastava, Varun, Goldsmith, Andrea J., Pilanci, Mert
The prohibitive sizes of Large Language Models (LLMs) today make it difficult to deploy them on memory-constrained edge devices. This work introduces $\rm CALDERA$ -- a new post-training LLM compression algorithm that harnesses the inherent low-rank structure of a weight matrix $\mathbf{W}$ by approximating it via a low-rank, low-precision decomposition as $\mathbf{W} \approx \mathbf{Q} + \mathbf{L}\mathbf{R}$. Here, $\mathbf{L}$ and $\mathbf{R}$ are low rank factors, and the entries of $\mathbf{Q}$, $\mathbf{L}$ and $\mathbf{R}$ are quantized. The model is compressed by substituting each layer with its $\mathbf{Q} + \mathbf{L}\mathbf{R}$ decomposition, and the zero-shot performance of the compressed model is evaluated. Additionally, $\mathbf{L}$ and $\mathbf{R}$ are readily amenable to low-rank adaptation, consequently enhancing the zero-shot performance. $\rm CALDERA$ obtains this decomposition by formulating it as an optimization problem $\min_{\mathbf{Q},\mathbf{L},\mathbf{R}}\lVert(\mathbf{Q} + \mathbf{L}\mathbf{R} - \mathbf{W})\mathbf{X}^\top\rVert_{\rm F}^2$, where $\mathbf{X}$ is the calibration data, and $\mathbf{Q}, \mathbf{L}, \mathbf{R}$ are constrained to be representable using low-precision formats. Theoretical upper bounds on the approximation error of $\rm CALDERA$ are established using a rank-constrained regression framework, and the tradeoff between compression ratio and model performance is studied by analyzing the impact of target rank and quantization bit budget. Results illustrate that compressing LlaMa-$2$ $7$B/$70$B and LlaMa-$3$ $8$B models obtained using $\rm CALDERA$ outperforms existing post-training LLM compression techniques in the regime of less than $2.5$ bits per parameter. The implementation is available at: \href{https://github.com/pilancilab/caldera}{https://github.com/pilancilab/caldera}.
A Causal Framework for Evaluating Deferring Systems
Palomba, Filippo, Pugnana, Andrea, Alvarez, José Manuel, Ruggieri, Salvatore
Deferring systems extend supervised Machine Learning (ML) models with the possibility to defer predictions to human experts. However, evaluating the impact of a deferring strategy on system accuracy is still an overlooked area. This paper fills this gap by evaluating deferring systems through a causal lens. We link the potential outcomes framework for causal inference with deferring systems. This allows us to identify the causal impact of the deferring strategy on predictive accuracy. We distinguish two scenarios. In the first one, we can access both the human and the ML model predictions for the deferred instances. In such a case, we can identify the individual causal effects for deferred instances and aggregates of them. In the second scenario, only human predictions are available for the deferred instances. In this case, we can resort to regression discontinuity design to estimate a local causal effect. We empirically evaluate our approach on synthetic and real datasets for seven deferring systems from the literature.
A Mallows-like Criterion for Anomaly Detection with Random Forest Implementation
Zhao, Gaoxiang, Wang, Lu, Wang, Xiaoqiang
The effectiveness of anomaly signal detection can be significantly undermined by the inherent uncertainty of relying on one specified model. Under the framework of model average methods, this paper proposes a novel criterion to select the weights on aggregation of multiple models, wherein the focal loss function accounts for the classification of extremely imbalanced data. This strategy is further integrated into Random Forest algorithm by replacing the conventional voting method. We have evaluated the proposed method on benchmark datasets across various domains, including network intrusion. The findings indicate that our proposed method not only surpasses the model averaging with typical loss functions but also outstrips common anomaly detection algorithms in terms of accuracy and robustness.
On the Role of Attention Masks and LayerNorm in Transformers
Wu, Xinyi, Ajorlou, Amir, Wang, Yifei, Jegelka, Stefanie, Jadbabaie, Ali
Self-attention is the key mechanism of transformers, which are the essential building blocks of modern foundation models. Recent studies have shown that pure self-attention suffers from an increasing degree of rank collapse as depth increases, limiting model expressivity and further utilization of model depth. The existing literature on rank collapse, however, has mostly overlooked other critical components in transformers that may alleviate the rank collapse issue. In this paper, we provide a general analysis of rank collapse under self-attention, taking into account the effects of attention masks and layer normalization (LayerNorm). In particular, we find that although pure masked attention still suffers from exponential collapse to a rank one subspace, local masked attention can provably slow down the collapse rate. In the case of self-attention with LayerNorm, we first show that for certain classes of value matrices, collapse to a rank one subspace still happens exponentially. However, through construction of nontrivial counterexamples, we then establish that with proper choice of value matrices, a general class of sequences may not converge to a rank one subspace, and the self-attention dynamics with LayerNorm can simultaneously possess a rich set of equilibria with any possible rank between one and full. Our result refutes the previous hypothesis that LayerNorm plays no role in the rank collapse of self-attention and suggests that self-attention with LayerNorm constitutes a much more expressive, versatile nonlinear dynamical system than what was originally thought.
Facing Global Outrage, Netanyahu Calls Civilian Deaths in Rafah Strike 'Tragic Accident'
Hamas, in a statement, described the Israeli strike on Rafah as "a horrific war crime" and demanded the "immediate and urgent implementation" of the World Court's decision. The group did not refer to the Israeli military's assertions that two Hamas officials had been killed in the strike. The Israeli military said it had taken a number of steps before the strike to reduce the risk of harm to civilians, including conducting aerial surveillance and using munitions characterized as precise. "Based on these measures, it was assessed that there would be no expected harm to uninvolved civilians," it said. But an Israeli official, speaking on the condition of anonymity to discuss a sensitive matter, said on Monday that an initial investigation by the military had concluded that the strike, or shrapnel from it, may have unexpectedly ignited a flammable substance at the camp.
Musk's Neuralink seeks to enroll three patients in brain implant study
Neuralink, Elon Musk's brain-chip company, aims to enroll three patients to evaluate its brain implant device in a study expected to take several years to complete, according to details on the U.S. government's clinical trials database. The company had sought to enroll 10 patients when it applied to U.S. regulators to begin clinical trials, Reuters reported last year. Neuralink is testing its implant designed to give paralyzed patients the ability to use digital devices by thinking alone, a prospect that could help people with spinal cord injuries.
Can you bequeath your Steam account? Maybe, but there's a catch
We've all got to die sometime. But whatever you think awaits us after death, it's unlikely to involve a suped-up gaming PC, a fiber connection, and tons of digital video games. Last week, a Steam support representative said that in addition to not being able to transfer your account to another person, you also can't leave it to your beneficiaries. But this policy against the inheriting of PC games might be in violation of a relatively recent United States law. A ResetEra poster named delete12345 asked a Steam support representative if they could transfer the ownership of their Steam account after they died through their will.
Anduril Is Building Out the Pentagon's Dream of Deadly Drone Swarms
When Palmer Luckey cofounded the defense startup Anduril in 2017, three years after selling his virtual reality startup Oculus to Facebook, the idea of a twentysomething from the tech industry challenging the giant contractors that build fighter jets, tanks, and warships for the US military seemed somewhat far-fetched. Seven years on, Luckey is showing that Anduril can not only compete with those contractors--it can win. Last month, Anduril was one of two companies, along with the established defense contractor General Atomics, chosen to prototype a new kind of autonomous fighter jet called the Collaborative Combat Aircraft, or CCA, for the US Air Force and Navy. Anduril was chosen ahead of a pack of what Beltway lingo dubs "defense primes"--Boeing, Lockheed Martin, and Northrup Grummond. "Anduril is proving that with the right team and business model, a seven-year-old company can go toe-to-toe with players that have been around for 70," Luckey wrote on social media platform X shortly after the contract was announced.
ChatGPT as the Marketplace of Ideas: Should Truth-Seeking Be the Goal of AI Content Governance?
As one of the most enduring metaphors within legal discourse, the marketplace of ideas has wielded considerable influence over the jurisprudential landscape for decades. A century after the inception of this theory, ChatGPT emerged as a revolutionary technological advancement in the twenty-first century. This research finds that ChatGPT effectively manifests the marketplace metaphor. It not only instantiates the promises envisaged by generations of legal scholars but also lays bare the perils discerned through sustained academic critique. Specifically, the workings of ChatGPT and the marketplace of ideas theory exhibit at least four common features: arena, means, objectives, and flaws. These shared attributes are sufficient to render ChatGPT historically the most qualified engine for actualizing the marketplace of ideas theory. The comparison of the marketplace theory and ChatGPT merely marks a starting point. A more meaningful undertaking entails reevaluating and reframing both internal and external AI policies by referring to the accumulated experience, insights, and suggestions researchers have raised to fix the marketplace theory. Here, a pivotal issue is: should truth-seeking be set as the goal of AI content governance? Given the unattainability of the absolute truth-seeking goal, I argue against adopting zero-risk policies. Instead, a more judicious approach would be to embrace a knowledge-based alternative wherein large language models (LLMs) are trained to generate competing and divergent viewpoints based on sufficient justifications. This research also argues that so-called AI content risks are not created by AI companies but are inherent in the entire information ecosystem. Thus, the burden of managing these risks should be distributed among different social actors, rather than being solely shouldered by chatbot companies.
Spanish and LLM Benchmarks: is MMLU Lost in Translation?
Plaza, Irene, Melero, Nina, del Pozo, Cristina, Conde, Javier, Reviriego, Pedro, Mayor-Rocher, Marina, Grandury, María
The evaluation of Large Language Models (LLMs) is a key element in their continuous improvement process and many benchmarks have been developed to assess the performance of LLMs in different tasks and topics. As LLMs become adopted worldwide, evaluating them in languages other than English is increasingly important. However, most LLM benchmarks are simply translated using an automated tool and then run in the target language. This means that the results depend not only on the LLM performance in that language but also on the quality of the translation. In this paper, we consider the case of the well-known Massive Multitask Language Understanding (MMLU) benchmark. Selected categories of the benchmark are translated into Spanish using Azure Translator and ChatGPT4 and run on ChatGPT4. Next, the results are processed to identify the test items that produce different answers in Spanish and English. Those are then analyzed manually to understand if the automatic translation caused the change. The results show that a significant fraction of the failing items can be attributed to mistakes in the translation of the benchmark. These results make a strong case for improving benchmarks in languages other than English by at least revising the translations of the items and preferably by adapting the tests to the target language by experts.