costa
A ZeNN architecture to avoid the Gaussian trap
Carvalho, Luís, Costa, João L., Mourão, José, Oliveira, Gonçalo
We propose a new simple architecture, Zeta Neural Networks (ZeNNs), in order to overcome several shortcomings of standard multi-layer perceptrons (MLPs). Namely, in the large width limit, MLPs are non-parametric, they do not have a well-defined pointwise limit, they lose non-Gaussian attributes and become unable to perform feature learning; moreover, finite width MLPs perform poorly in learning high frequencies. The new ZeNN architecture is inspired by three simple principles from harmonic analysis: i) Enumerate the perceptons and introduce a non-learnable weight to enforce convergence; ii) Introduce a scaling (or frequency) factor; iii) Choose activation functions that lead to near orthogonal systems. We will show that these ideas allow us to fix the referred shortcomings of MLPs. In fact, in the infinite width limit, ZeNNs converge pointwise, they exhibit a rich asymptotic structure beyond Gaussianity, and perform feature learning. Moreover, when appropriate activation functions are chosen, (finite width) ZeNNs excel at learning high-frequency features of functions with low dimensional domains.
Keyformer: KV Cache Reduction through Key Tokens Selection for Efficient Generative Inference
Adnan, Muhammad, Arunkumar, Akhil, Jain, Gaurav, Nair, Prashant J., Soloveychik, Ilya, Kamath, Purushotham
Transformers have emerged as the underpinning architecture for Large Language Models (LLMs). In generative language models, the inference process involves two primary phases: prompt processing and token generation. Token generation, which constitutes the majority of the computational workload, primarily entails vector-matrix multiplications and interactions with the Key-Value (KV) Cache. This phase is constrained by memory bandwidth due to the overhead of transferring weights and KV cache values from the memory system to the computing units. This memory bottleneck becomes particularly pronounced in applications that require long-context and extensive text generation, both of which are increasingly crucial for LLMs. This paper introduces "Keyformer", an innovative inference-time approach, to mitigate the challenges associated with KV cache size and memory bandwidth utilization. Keyformer leverages the observation that approximately 90% of the attention weight in generative inference focuses on a specific subset of tokens, referred to as "key" tokens. Keyformer retains only the key tokens in the KV cache by identifying these crucial tokens using a novel score function. This approach effectively reduces both the KV cache size and memory bandwidth usage without compromising model accuracy. We evaluate Keyformer's performance across three foundational models: GPT-J, Cerebras-GPT, and MPT, which employ various positional embedding algorithms. Our assessment encompasses a variety of tasks, with a particular emphasis on summarization and conversation tasks involving extended contexts. Keyformer's reduction of KV cache reduces inference latency by 2.1x and improves token generation throughput by 2.4x, while preserving the model's accuracy.
A novel corrective-source term approach to modeling unknown physics in aluminum extraction process
Robinson, Haakon, Lundby, Erlend, Rasheed, Adil, Gravdahl, Jan Tommy
With the ever-increasing availability of data, there has been an explosion of interest in applying modern machine learning methods to fields such as modeling and control. However, despite the flexibility and surprising accuracy of such black-box models, it remains difficult to trust them. Recent efforts to combine the two approaches aim to develop flexible models that nonetheless generalize well; a paradigm we call Hybrid Analysis and modeling (HAM). In this work we investigate the Corrective Source Term Approach (CoSTA), which uses a data-driven model to correct a misspecified physics-based model. This enables us to develop models that make accurate predictions even when the underlying physics of the problem is not well understood. We apply CoSTA to model the Hall-H\'eroult process in an aluminum electrolysis cell. We demonstrate that the method improves both accuracy and predictive stability, yielding an overall more trustworthy model.
'AI Bumblebees:' These AI Robots Act Like Bees to Pollinate Tomato Plants
Is AI taking over the jobs of bumblebees? Bumblebees are typically used to pollinate plants in glasshouses all over the world. However, they are prohibited in Australia, so pollination must be done manually. Hence, prominent Australian fresh produce company Costa Group is deploying AI to implement robotic pollination in one of its tomato glasses, thanks to its partnership with Israeli firm Arugga AI Farming. The AI-powered robot is named "Polly" and will pollinate truss tomato plants in Costa's tomato glasshouse facilities in Guyra, New South Wales.
Supervised Machine Learning for Effective Missile Launch Based on Beyond Visual Range Air Combat Simulations
Dantas, Joao P. A., Costa, Andre N., Medeiros, Felipe L. L., Geraldo, Diego, Maximo, Marcos R. O. A., Yoneyama, Takashi
This work compares supervised machine learning methods using reliable data from constructive simulations to estimate the most effective moment for launching missiles during air combat. We employed resampling techniques to improve the predictive model, analyzing accuracy, precision, recall, and f1-score. Indeed, we could identify the remarkable performance of the models based on decision trees and the significant sensitivity of other algorithms to resampling techniques. The models with the best f1-score brought values of 0.379 and 0.465 without and with the resampling technique, respectively, which is an increase of 22.69%. Thus, if desirable, resampling techniques can improve the model's recall and f1-score with a slight decline in accuracy and precision. Therefore, through data obtained through constructive simulations, it is possible to develop decision support tools based on machine learning models, which may improve the flight quality in BVR air combat, increasing the effectiveness of offensive missions to hit a particular target.
Text characterization based on recurrence networks
Souza, Bárbara C. e, Silva, Filipi N., de Arruda, Henrique F., da Silva, Giovana D., Costa, Luciano da F., Amancio, Diego R.
Several complex systems are characterized by presenting intricate characteristics taking place at several scales of time and space. These multiscale characterizations are used in various applications, including better understanding diseases, characterizing transportation systems, and comparison between cities, among others. In particular, texts are also characterized by a hierarchical structure that can be approached by using multi-scale concepts and methods. The multiscale properties of texts constitute a subject worth further investigation. In addition, more effective approaches to text characterization and analysis can be obtained by emphasizing words with potentially more informational content. The present work aims at developing these possibilities while focusing on mesoscopic representations of networks. More specifically, we adopt an extension to the mesoscopic approach to represent text narratives, in which only the recurrent relationships among tagged parts of speech (subject, verb and direct object) are considered to establish connections among sequential pieces of text (e.g., paragraphs). The characterization of the texts was then achieved by considering scale-dependent complementary methods: accessibility, symmetry and recurrence signatures. In order to evaluate the potential of these concepts and methods, we approached the problem of distinguishing between literary genres (fiction and non-fiction). A set of 300 books organized into the two genres was considered and were compared by using the aforementioned approaches. All the methods were capable of differentiating to some extent between the two genres. The accessibility and symmetry reflected the narrative asymmetries, while the recurrence signature provided a more direct indication about the non-sequential semantic connections taking place along the narrative.
Artificial Intelligence Robot Writes Valentines Cards For Shy Lovers
Known as the PEOPLE'S PAPER, Euro Weekly News is the leading English language newspaper in Spain. Covering the Costa del Sol, Costa Blanca, Almeria, Axarquia, Mallorca and beyond, EWN supports and inspires the individuals, neighbourhoods, and communities we serve, by delivering news with a social conscience. Whether it's local news in Spain, UK news or international stories, we are proud to be the voice for the expat communities who now call Spain home. With around half a million print readers a week and over 1.5 million web views per month, EWN has the biggest readership of any English language newspaper in Spain. The paper prints over 150 news stories a week with many hundreds more on the web – no one else even comes close.
The Newest Weapon Against Covid-19: AI That Speed-Reads Faxes
Alison Stribling has learned a lot about infectious disease since she transferred onto Covid-19 response at the health department in Contra Costa County near San Francisco. One of her discoveries: How vital fax machines are to US pandemic response. Across the country, labs and health providers report new Covid-19 cases to local health departments. At Contra Costa Health Services, officials use the data to start contact tracing or send extra help in certain cases, such as at a care home or to an infected health care worker. On a typical day in Contra Costa, only around half of those reports arrive electronically; the rest, as many as hundreds, flow in via the fax line, creating a Sisyphean reading list.
Investor Sues Company Over Artificial Intelligence Advice - My TechDecisions
Decision makers are becoming more wary about certain uses of AI, including trusting solutions to make money for them. When things go wrong and losses are up, it's tough to know who's responsible – the machine, or the person behind the machine. Bloomberg reports on a recent example of this, where an investor is suing for damages after losing money based on the decisions made by a money management solution. According to Bloomberg, back in 2017, investor Samathur Li Kin-kan bought into a money management AI solution that another investor, Raffaele Costa, planned on using to mange the money made by his company, Tyndaris. "The idea of a fully automated money manager inspired Li instantly," driving him to invest in the solution to grow his own money – $2.5 billion- $250 million worth.