Goto

Collaborating Authors

 Dedham


Confidence Matters: Revisiting Intrinsic Self-Correction Capabilities of Large Language Models

arXiv.org Artificial Intelligence

The recent success of Large Language Models (LLMs) has catalyzed an increasing interest in their self-correction capabilities. This paper presents a comprehensive investigation into the intrinsic self-correction of LLMs, attempting to address the ongoing debate about its feasibility. Our research has identified an important latent factor - the "confidence" of LLMs - during the self-correction process. Overlooking this factor may cause the models to over-criticize themselves, resulting in unreliable conclusions regarding the efficacy of self-correction. We have experimentally observed that LLMs possess the capability to understand the "confidence" in their own responses. It motivates us to develop an "If-or-Else" (IoE) prompting framework, designed to guide LLMs in assessing their own "confidence", facilitating intrinsic self-corrections. We conduct extensive experiments and demonstrate that our IoE-based Prompt can achieve a consistent improvement regarding the accuracy of self-corrected responses over the initial answers. Our study not only sheds light on the underlying factors affecting self-correction in LLMs, but also introduces a practical framework that utilizes the IoE prompting principle to efficiently improve self-correction capabilities with "confidence". The code is available at https://github.com/MBZUAI-CLeaR/IoE-Prompting.git.


Statistical-Computational Trade-offs in Tensor PCA and Related Problems via Communication Complexity

arXiv.org Machine Learning

Tensor PCA is a stylized statistical inference problem introduced by Montanari and Richard to study the computational difficulty of estimating an unknown parameter from higher-order moment tensors. Unlike its matrix counterpart, Tensor PCA exhibits a statistical-computational gap, i.e., a sample size regime where the problem is information-theoretically solvable but conjectured to be computationally hard. This paper derives computational lower bounds on the run-time of memory bounded algorithms for Tensor PCA using communication complexity. These lower bounds specify a trade-off among the number of passes through the data sample, the sample size, and the memory required by any algorithm that successfully solves Tensor PCA. While the lower bounds do not rule out polynomial-time algorithms, they do imply that many commonly-used algorithms, such as gradient descent and power method, must have a higher iteration count when the sample size is not large enough. Similar lower bounds are obtained for Non-Gaussian Component Analysis, a family of statistical estimation problems in which low-order moment tensors carry no information about the unknown parameter. Finally, stronger lower bounds are obtained for an asymmetric variant of Tensor PCA and related statistical estimation problems. These results explain why many estimators for these problems use a memory state that is significantly larger than the effective dimensionality of the parameter of interest.


Research and Education Towards Smart and Sustainable World

arXiv.org Artificial Intelligence

We propose a vision for directing research and education in the ICT field. Our Smart and Sustainable World vision targets at prosperity for the people and the planet through better awareness and control of both human-made and natural environment. The needs of the society, individuals, and industries are fulfilled with intelligent systems that sense their environment, make proactive decisions on actions advancing their goals, and perform the actions on the environment. We emphasize artificial intelligence, feedback loops, human acceptance and control, intelligent use of basic resources, performance parameters, mission-oriented interdisciplinary research, and a holistic systems view complementing the conventional analytical reductive view as a research paradigm especially for complex problems. To serve a broad audience, we explain these concepts and list the essential literature. We suggest planning research and education by specifying, in a step-wise manner, scenarios, performance criteria, system models, research problems and education content, resulting in common goals and a coherent project portfolio as well as education curricula. Research and education produce feedback to support evolutionary development and encourage creativity in research. Finally, we propose concrete actions for realizing this approach.


mmFall: Fall Detection using 4D MmWave Radar and Variational Recurrent Autoencoder

arXiv.org Machine Learning

In this paper we propose mmFall - a novel fall detection system, which comprises of (i) the emerging millimeter-wave (mmWave) radar sensor to collect the human body's point cloud along with the body centroid, and (ii) a variational recurrent autoencoder (VRAE) to compute the anomaly level of the body motion based on the acquired point cloud. A fall is claimed to have occurred when the spike in anomaly level and the drop in centroid height occur simultaneously. The mmWave radar sensor provides several advantages, such as privacycompliance and high-sensitivity to motion, over the traditional sensing modalities. However, (i) randomness in radar point cloud data and (ii) difficulties in fall collection/labeling in the traditional supervised fall detection approaches are the two main challenges. To overcome the randomness in radar data, the proposed VRAE uses variational inference, a probabilistic approach rather than the traditional deterministic approach, to infer the posterior probability of the body's latent motion state at each frame, followed by a recurrent neural network (RNN) to learn the temporal features of the motion over multiple frames. Moreover, to circumvent the difficulties in fall data collection/labeling, the VRAE is built upon an autoencoder architecture in a semi-supervised approach, and trained on only normal activities of daily living (ADL) such that in the inference stage the VRAE will generate a spike in the anomaly level once an abnormal motion, such as fall, occurs. During the experiment, we implemented the VRAE along with two other baselines, and tested on the dataset collected in an apartment. The receiver operating characteristic (ROC) curve indicates that our proposed model outperforms the other two baselines, and achieves 98% detection out of 50 falls at the expense of just 2 false alarms.


A generalized multivariate Student-t mixture model for Bayesian classification and clustering of radar waveforms

arXiv.org Machine Learning

In this paper, a generalized multivariate Student-t mixture model is developed for classification and clustering of Low Probability of Intercept radar waveforms. A Low Probability of Intercept radar signal is characterized by a pulse compression waveform which is either frequency-modulated or phase-modulated. The proposed model can classify and cluster different modulation types such as linear frequency modulation, non linear frequency modulation, polyphase Barker, polyphase P1, P2, P3, P4, Frank and Zadoff codes. The classification method focuses on the introduction of a new prior distribution for the model hyper-parameters that gives us the possibility to handle sensitivity of mixture models to initialization and to allow a less restrictive modeling of data. Inference is processed through a Variational Bayes method and a Bayesian treatment is adopted for model learning, supervised classification and clustering. Moreover, the novel prior distribution is not a well-known probability distribution and both deterministic and stochastic methods are employed to estimate its expectations. Some numerical experiments show that the proposed method is less sensitive to initialization and provides more accurate results than the previous state of the art mixture models.


Contextual Deep Learning Makes Artificial Intelligence More Real

#artificialintelligence

According to Tech Spot, the concept of having a machine capable of reacting in an intelligent way has been until very recently a matter of science fiction. However, this concept is certainly very compelling and scientists were working on transform this into reality. We are now on the verge of creating this new reality. The general public, however, is not yet informed of what concepts such as neural networks, artificial intelligence and deep learning represent. Much of the current efforts in the field of deep learning technology are related from the simplest level to the very rapid recognition and classification of objects.


RAGE Frameworks' Artificial Intelligence Solution Automates Financial Data Processing for Major Financial Institution

#artificialintelligence

DEDHAM, MA--(Marketwired - Dec 7, 2016) - RAGE Frameworks, a provider of artificial intelligence (AI) for the Enterprise, today announced that a leading diversified investment and financial services company has deployed RAGE LiveSpread to automate the extraction, interpretation and processing of financial statements and other documents needed for credit analysis. RAGE LiveSpread is a contextual, traceable machine learning solution built on the RAGE-AI platform. Dealing with variations in form (electronic files, pdfs, and paper statements), format, language, accounting standards across countries, and data delivery methods makes financial statement automation a major challenge. Currently, this is a manual, error-prone, and non-scalable process at every major bank around the world. At this leading investment and financial services company, RAGE's solution completely automates the ingestion, extraction, interpretation of investment reports, bank statements and Income Tax returns.


Leading Financial Services Firm Uses RAGE Artificial Intelligence Solution to Generate Signals for Alpha

#artificialintelligence

DEDHAM, MA--(Marketwired - Sep 7, 2016) - Rage Frameworks, a provider of knowledge-based automation technology and services, today announced that a leading multinational financial services firm has selected its Artificial Intelligence platform (RAGE AI) to drive improved results for its investment customers by using artificial intelligence to discover signals captured in a wide variety of data sources with Rage's innovative deep learning capabilities. RAGE AI significantly extends the frontier of deep learning and machine intelligence technology as it incorporates proprietary linguistics-based machine learning innovations to understand market developments in the context of individual companies and interpret those signals as a human would. After demonstrating via historical back-testing that the Rage AI platform repeatedly delivered returns in excess of what the firm's quantitative team was able to produce, Rage's solution was integrated in order to drive significant lift in the returns generated for the firm's clients. In fact, Rage has repeatedly shown that its deep background in computational linguistics and Natural Language Understanding can systematically discover Alpha by forming assessments of a company's financial projections that effectively predict future performance for businesses such as Wal-Mart (attached), where Rage AI predicted an upward trend in stock price months in advance. The RAGE AI platform does this by continuously interpreting unstructured content from over 100,000 sources and translating it into valuable intelligence.


Leading Financial Services Firm Uses RAGE Artificial Intelligence Solution to Generate Signals for Alpha

#artificialintelligence

DEDHAM, MA–(Marketwired – Sep 7, 2016) – Rage Frameworks, a provider of knowledge-based automation technology and services, today announced that a leading multinational financial services firm has selected its Artificial Intelligence platform (RAGE AI) to drive improved results for its investment customers by using artificial intelligence to discover signals captured in a wide variety of data sources with Rage's innovative deep learning capabilities. RAGE AI significantly extends the frontier of deep learning and machine intelligence technology as it incorporates proprietary linguistics-based machine learning innovations to understand market developments in the context of individual companies and interpret those signals as a human would. After demonstrating via historical back-testing that the Rage AI platform repeatedly delivered returns in excess of what the firm's quantitative team was able to produce, Rage's solution was integrated in order to drive significant lift in the returns generated for the firm's clients. In fact, Rage has repeatedly shown that its deep background in computational linguistics and Natural Language Understanding can systematically discover Alpha by forming assessments of a company's financial projections that effectively predict future performance for businesses such as Wal-Mart (attached), where Rage AI predicted an upward trend in stock price months in advance. The RAGE AI platform does this by continuously interpreting unstructured content from over 100,000 sources and translating it into valuable intelligence.


Rage Frameworks Expands Its Artificial Intelligence Platform

#artificialintelligence

RAGE AI significantly extends the frontier of deep learning and machine intelligence technology from "natural language processing" to "natural language understanding." RAGE AI incorporates deep linguistic parsing and proprietary innovations to understand meaning in context, which makes its solutions completely transparent, auditable and flexible. The platform facilitates unsupervised to supervised learning and contains several innovations to support automated knowledge acquisition including pragmatic knowledge. RAGE AI is not a black box and does not rely on statistical patterns present in training data. Introduced in 2011, RAGE AI is an integral part of the broader RAGE Enterprise platform, a provider of all process orchestration and automation capabilities.