Government
Distill or Annotate? Cost-Efficient Fine-Tuning of Compact Models
Kang, Junmo, Xu, Wei, Ritter, Alan
Fine-tuning large models is highly effective, however, inference can be expensive and produces carbon emissions. Knowledge distillation has been shown to be a practical solution to reduce inference costs, but the distillation process itself requires significant computational resources. Rather than buying or renting GPUs to fine-tune, then distill a large model, an NLP practitioner might instead choose to allocate the available budget to hire annotators and manually label additional fine-tuning data. In this paper, we investigate how to most efficiently use a fixed budget to build a compact model. Through extensive experiments on six diverse tasks, we show that distilling from T5-XXL (11B) to T5-Small (60M) is almost always a cost-efficient strategy compared to annotating more data to directly train a compact model (T5-Small). We further investigate how the optimal budget allocated towards computation varies across scenarios. We will make our code, datasets, annotation cost estimates, and baseline models available as a benchmark to support further work on cost-efficient training of compact models.
Chaos Theory and Adversarial Robustness
Neural networks, being susceptible to adversarial attacks, should face a strict level of scrutiny before being deployed in critical or adversarial applications. This paper uses ideas from Chaos Theory to explain, analyze, and quantify the degree to which neural networks are susceptible to or robust against adversarial attacks. To this end, we present a new metric, the "susceptibility ratio," given by $\hat \Psi(h, \theta)$, which captures how greatly a model's output will be changed by perturbations to a given input. Our results show that susceptibility to attack grows significantly with the depth of the model, which has safety implications for the design of neural networks for production environments. We provide experimental evidence of the relationship between $\hat \Psi$ and the post-attack accuracy of classification models, as well as a discussion of its application to tasks lacking hard decision boundaries. We also demonstrate how to quickly and easily approximate the certified robustness radii for extremely large models, which until now has been computationally infeasible to calculate directly.
Language Detoxification with Attribute-Discriminative Latent Space
Kwak, Jin Myung, Kim, Minseon, Hwang, Sung Ju
Transformer-based Language Models (LMs) have achieved impressive results on natural language understanding tasks, but they can also generate toxic text such as insults, threats, and profanity, limiting their real-world applications. To overcome this issue, a few text generation approaches aim to detoxify toxic texts using additional LMs or perturbations. However, previous methods require excessive memory, computations, and time which are serious bottlenecks in their real-world application. To address such limitations, we propose an effective yet efficient method for language detoxification using an attribute-discriminative latent space. Specifically, we project the latent space of an original Transformer LM onto a discriminative latent space that well-separates texts by their attributes using a projection block and an attribute discriminator. This allows the LM to control the text generation to be non-toxic with minimal memory and computation overhead. We validate our model, Attribute-Discriminative Language Model (ADLM) on detoxified language and dialogue generation tasks, on which our method significantly outperforms baselines both in performance and efficiency.
A Detailed Study of Interpretability of Deep Neural Network based Top Taggers
Khot, Ayush, Neubauer, Mark S., Roy, Avik
Recent developments in the methods of explainable AI (XAI) allow researchers to explore the inner workings of deep neural networks (DNNs), revealing crucial information about input-output relationships and realizing how data connects with machine learning models. In this paper we explore interpretability of DNN models designed to identify jets coming from top quark decay in high energy proton-proton collisions at the Large Hadron Collider (LHC). We review a subset of existing top tagger models and explore different quantitative methods to identify which features play the most important roles in identifying the top jets. We also investigate how and why feature importance varies across different XAI metrics, how correlations among features impact their explainability, and how latent space representations encode information as well as correlate with physically meaningful quantities. Our studies uncover some major pitfalls of existing XAI methods and illustrate how they can be overcome to obtain consistent and meaningful interpretation of these models. We additionally illustrate the activity of hidden layers as Neural Activation Pattern (NAP) diagrams and demonstrate how they can be used to understand how DNNs relay information across the layers and how this understanding can help to make such models significantly simpler by allowing effective model reoptimization and hyperparameter tuning. These studies not only facilitate a methodological approach to interpreting models but also unveil new insights about what these models learn. Incorporating these observations into augmented model design, we propose the Particle Flow Interaction Network (PFIN) model and demonstrate how interpretability-inspired model augmentation can improve top tagging performance.
Drones with AI targeting system claimed to be 'better than human'
Drones being evaluated by the US military could soon be equipped with an artificial intelligence that is claimed to be better than humans at identifying targets, although the classified nature of the work makes it difficult to verify this claim. Stephen Bornstein of Athena AI, the Australian company behind the system, says the AI will assist human drone operators, who can lose concentration after hours spent looking at streaming video. "The AI will do a lot of the heavy lifting for them," says โฆ
Russian military claims it prevented Ukrainian attack on Moscow by shooting down four drones
Senior foreign affairs correspondent Greg Palkot reports 80% of all Ukrainians have a relative or friend killed as the result of the war on'America's Newsroom.' The Russian military claimed Tuesday that it shot down several Ukrainian drones that were headed for the capital city of Moscow, an attack that also prompted officials to briefly close one of the city's airports. The Russian Defense Ministry said air defenses on the outskirts of Moscow shot down four of five drones heading toward the capital, adding that the fifth was jammed by electronic means. Ukrainian authorities, who usually do not comment on attacks inside Russia's proper territory, have not claimed responsibility. Moscow Mayor Sergei Sobyanin said no one was harmed in the attack and no buildings were damaged.
'We have to flip the AI debate towards hope': Labour's techno-optimist, Darren Jones
In the same way as you upgrade your iPhone, we need to upgrade Britain." Labour MP Darren Jones believes artificial intelligence will bring an economic change on the scale of the industrial revolution, which politicians must be ready to shape. As chair of the business and trade select committee, the ambitious 36-year-old backbencher, who represents Bristol North West, has built a reputation for himself in Westminster as a tough interrogator. With speculation raging last week about the future of Thames Water, he took to the airwaves to criticise the way the heavily indebted sector has been regulated, saying he was "increasingly sick" of its failures. However, Jones is at his most animated when talking about AI. He has clashed with company bosses over their use of technology to monitor and control staff โ including at Amazon and Royal Mail. But he is an evangelist for the upsides of innovation, including the arrival of large language models (LLMs) such as the hit dialogue-based AI software ChatGPT. "It's really important that we flip this debate.
Local primordial non-Gaussianity from the large-scale clustering of photometric DESI luminous red galaxies
Rezaie, Mehdi, Ross, Ashley J., Seo, Hee-Jong, Kong, Hui, Porredon, Anna, Samushia, Lado, Chaussidon, Edmond, Krolewski, Alex, de Mattia, Arnaud, Beutler, Florian, Aguilar, Jessica Nicole, Ahlen, Steven, Alam, Shadab, Avila, Santiago, Bahr-Kalus, Benedict, Bermejo-Climent, Jose, Brooks, David, Claybaugh, Todd, Cole, Shaun, Dawson, Kyle, de la Macorra, Axel, Doel, Peter, Font-Ribera, Andreu, Forero-Romero, Jaime E., Gontcho, Satya Gontcho A, Guy, Julien, Honscheid, Klaus, Kisner, Theodore, Landriau, Martin, Levi, Michael, Manera, Marc, Meisner, Aaron, Miquel, Ramon, Mueller, Eva-Maria, Myers, Adam, Newman, Jeffrey A., Nie, Jundan, Palanque-Delabrouille, Nathalie, Percival, Will, Poppett, Claire, Rossi, Graziano, Sanchez, Eusebio, Schubnell, Michael, Tarlรฉ, Gregory, Weaver, Benjamin Alan, Yรจche, Christophe, Zhou, Zhimin, Zou, Hu
We use angular clustering of luminous red galaxies from the Dark Energy Spectroscopic Instrument (DESI) imaging surveys to constrain the local primordial non-Gaussianity parameter fNL. Our sample comprises over 12 million targets, covering 14,000 square degrees of the sky, with redshifts in the range 0.2< z < 1.35. We identify Galactic extinction, survey depth, and astronomical seeing as the primary sources of systematic error, and employ linear regression and artificial neural networks to alleviate non-cosmological excess clustering on large scales. Our methods are tested against log-normal simulations with and without fNL and systematics, showing superior performance of the neural network treatment in reducing remaining systematics. Assuming the universality relation, we find fNL $= 47^{+14(+29)}_{-11(-22)}$ at 68\%(95\%) confidence. With a more aggressive treatment, including regression against the full set of imaging maps, our maximum likelihood value shifts slightly to fNL$ \sim 50$ and the uncertainty on fNL increases due to the removal of large-scale clustering information. We apply a series of robustness tests (e.g., cuts on imaging, declination, or scales used) that show consistency in the obtained constraints. Despite extensive efforts to mitigate systematics, our measurements indicate fNL > 0 with a 99.9 percent confidence level. This outcome raises concerns as it could be attributed to unforeseen systematics, including calibration errors or uncertainties associated with low-\ell systematics in the extinction template. Alternatively, it could suggest a scale-dependent fNL model--causing significant non-Gaussianity around large-scale structure while leaving cosmic microwave background scales unaffected. Our results encourage further studies of fNL with DESI spectroscopic samples, where the inclusion of 3D clustering modes should help separate imaging systematics.
A Bibliographic Study on Artificial Intelligence Research: Global Panorama and Indian Appearance
Tiwari, Amit, Bardhan, Susmita, Kumar, Vikas
The present study identifies and assesses the bibliographic trend in Artificial Intelligence (AI) research for the years 2015-2020 using the science mapping method of bibliometric study. The required data has been collected from the Scopus database. To make the collected data analysis-ready, essential data transformation was performed manually and with the help of a tool viz. OpenRefine. For determining the trend and performing the mapping techniques, top five open access and commercial journals of AI have been chosen based on their citescore driven ranking. The work includes 6880 articles published in the specified period for analysis. The trend is based on Country-wise publications, year-wise publications, topical terms in AI, top-cited articles, prominent authors, major institutions, involvement of industries in AI and Indian appearance. The results show that compared to open access journals; commercial journals have a higher citescore and number of articles published over the years. Additionally, IEEE is the prominent publisher which publishes 84% of the top-cited publications. Further, China and the United States are the major contributors to literature in the AI domain. The study reveals that neural networks and deep learning are the major topics included in top AI research publications. Recently, not only public institutions but also private bodies are investing their resources in AI research. The study also investigates the relative position of Indian researchers in terms of AI research. Present work helps in understanding the initial development, current stand and future direction of AI.