dudley
LanguaShrink: Reducing Token Overhead with Psycholinguistics
Liang, Xuechen, Tao, Meiling, Xia, Yinghui, Shi, Tianyu, Wang, Jun, Yang, JingSong
As large language models (LLMs) improve their capabilities in handling complex tasks, the issues of computational cost and efficiency due to long prompts are becoming increasingly prominent. To accelerate model inference and reduce costs, we propose an innovative prompt compression framework called LanguaShrink. Inspired by the observation that LLM performance depends on the density and position of key information in the input prompts, LanguaShrink leverages psycholinguistic principles and the Ebbinghaus memory curve to achieve task-agnostic prompt compression. This effectively reduces prompt length while preserving essential information. We referred to the training method of OpenChat.The framework introduces part-of-speech priority compression and data distillation techniques, using smaller models to learn compression targets and employing a KL-regularized reinforcement learning strategy for training.\cite{wang2023openchat} Additionally, we adopt a chunk-based compression algorithm to achieve adjustable compression rates. We evaluate our method on multiple datasets, including LongBench, ZeroScrolls, Arxiv Articles, and a newly constructed novel test set. Experimental results show that LanguaShrink maintains semantic similarity while achieving up to 26 times compression. Compared to existing prompt compression methods, LanguaShrink improves end-to-end latency by 1.43 times.
Agnostic Active Learning of Single Index Models with Linear Sample Complexity
Gajjar, Aarshvi, Tai, Wai Ming, Xu, Xingyu, Hegde, Chinmay, Li, Yi, Musco, Christopher
We study active learning methods for single index models of the form $F({\mathbf x}) = f(\langle {\mathbf w}, {\mathbf x}\rangle)$, where $f:\mathbb{R} \to \mathbb{R}$ and ${\mathbf x,\mathbf w} \in \mathbb{R}^d$. In addition to their theoretical interest as simple examples of non-linear neural networks, single index models have received significant recent attention due to applications in scientific machine learning like surrogate modeling for partial differential equations (PDEs). Such applications require sample-efficient active learning methods that are robust to adversarial noise. I.e., that work even in the challenging agnostic learning setting. We provide two main results on agnostic active learning of single index models. First, when $f$ is known and Lipschitz, we show that $\tilde{O}(d)$ samples collected via {statistical leverage score sampling} are sufficient to learn a near-optimal single index model. Leverage score sampling is simple to implement, efficient, and already widely used for actively learning linear models. Our result requires no assumptions on the data distribution, is optimal up to log factors, and improves quadratically on a recent ${O}(d^{2})$ bound of \cite{gajjar2023active}. Second, we show that $\tilde{O}(d)$ samples suffice even in the more difficult setting when $f$ is \emph{unknown}. Our results leverage tools from high dimensional probability, including Dudley's inequality and dual Sudakov minoration, as well as a novel, distribution-aware discretization of the class of Lipschitz functions.
On Generalization Bounds for Deep Compound Gaussian Neural Networks
Lyons, Carter, Raj, Raghu G., Cheney, Margaret
Algorithm unfolding or unrolling is the technique of constructing a deep neural network (DNN) from an iterative algorithm. Unrolled DNNs often provide better interpretability and superior empirical performance over standard DNNs in signal estimation tasks. An important theoretical question, which has only recently received attention, is the development of generalization error bounds for unrolled DNNs. These bounds deliver theoretical and practical insights into the performance of a DNN on empirical datasets that are distinct from, but sampled from, the probability density generating the DNN training data. In this paper, we develop novel generalization error bounds for a class of unrolled DNNs that are informed by a compound Gaussian prior. These compound Gaussian networks have been shown to outperform comparative standard and unfolded deep neural networks in compressive sensing and tomographic imaging problems. The generalization error bound is formulated by bounding the Rademacher complexity of the class of compound Gaussian network estimates with Dudley's integral. Under realistic conditions, we show that, at worst, the generalization error scales $\mathcal{O}(n\sqrt{\ln(n)})$ in the signal dimension and $\mathcal{O}(($Network Size$)^{3/2})$ in network size.
On the Vapnik-Chervonenkis dimension of products of intervals in $\mathbb{R}^d$
Gรณmez, Alirio Gรณmez, Kaufmann, Pedro L.
We study combinatorial complexity of certain classes of products of intervals in $\mathbb{R}^d$, from the point of view of Vapnik-Chervonenkis geometry. As a consequence of the obtained results, we conclude that the Vapnik-Chervonenkis dimension of the set of balls in $\ell_\infty^d$ -- which denotes $\R^d$ equipped with the sup norm -- equals $\lfloor (3d+1)/2\rfloor$.
Does GPT-2 know your phone number?
Yet, OpenAI's GPT-2 language model does know how to reach a certain Peter W-- (name redacted for privacy). When prompted with a short snippet of Internet text, the model accurately generates Peter's contact information, including his work address, email, phone, and fax: In our recent paper, we evaluate how large language models memorize and regurgitate such rare snippets of their training data. We focus on GPT-2 and find that at least 0.1% of its text generations (a very conservative estimate) contain long verbatim strings that are "copy-pasted" from a document in its training set. Such memorization would be an obvious issue for language models that are trained on private data, e.g., on users' emails, as the model might inadvertently output a user's sensitive conversations. Regular readers of the BAIR blog may be familiar with the issue of data memorization in language models.
Generalization bounds for deep thresholding networks
Behboodi, Arash, Rauhut, Holger, Schnoor, Ekkehard
We consider compressive sensing in the scenario where the sparsity basis (dictionary) is not known in advance, but needs to be learned from examples. Motivated by the well-known iterative soft thresholding algorithm for the reconstruction, we define deep networks parametrized by the dictionary, which we call deep thresholding networks. Based on training samples, we aim at learning the optimal sparsifying dictionary and thereby the optimal network that reconstructs signals from their low-dimensional linear measurements. The dictionary learning is performed via minimizing the empirical risk. We derive generalization bounds by analyzing the Rademacher complexity of hypothesis classes consisting of such deep networks. We obtain estimates of the sample complexity that depend only linearly on the dimensions and on the depth.
Bounding the expectation of the supremum of empirical processes indexed by H\"older classes
We obtain upper bounds on the expectation of the supremum of empirical processes indexed by H\"older classes of any smoothness and for any distribution supported on a bounded set. Another way to see it is from the point of view of integral probability metrics (IPM), a class of metrics on the space of probability measures: our rates quantify how quickly the empirical measure obtained from $n$ independent samples from a probability measure $P$ approaches $P$ with respect to the IPM indexed by H\"older classes. As an extremal case we recover the known rates for the Wassertein-1 distance.
Mount Sinai spinout leverages machine learning in fertility tool
A novel approach designed to enable women to accurately measure and monitor key fertility hormones through daily urine samples is being piloted as an at-home, consumer diagnostic tool. OOVA, a Mount Sinai Health System spinout, is partnering with nutritional supplement vendor Thorne Research to make OOVA's fertility monitoring tool available to more than 3 million potential customers and 35,000 clinicians through bundled sales with Thorne's supplements. The company plans to launch in the third quarter of 2019. "[OOVA] is empowering patients to take control over their own fertility. This product combines technological innovation and human behavior to meet an unmet demand in the market," says Alan Copperman, MD, director of the Division of Reproductive Endocrinology and Infertility, and vice chair of the Department of Obstetrics, Gynecology and Reproductive Science at Mount Sinai.
Mount Sinai Hospital to Explore Blockchain Applications - CoinDesk
The New York-based medical school founded by Mount Sinai Hospital has launched a new research center focused on blockchain applications in healthcare. On Tuesday, The Icahn School of Medicine said the Center for Biomedical Blockchain Research would be created inside the school's Institute for Next Generation Healthcare, which researches the application of artificial intelligence, robotics, genomic sequencing, sensors and wearable devices in medicine, New York-based news organization Crain's reported. The center's staff will conduct academic research on blockchain in medicine, as well as create their own prototype networks. The possible use cases include drug development and preventing the sale of counterfeit drugs, clinical trials and a better research reproducibility, Healthcare IT News wrote. The new center will be run by Joel Dudley, executive vice president of Precision Health at Mount Sinai and a former senior data scientist at Pivotal Software, which researches the use of artificial intelligence in biology.
Artificial Intelligence in Cardiology
Dr. Dudley is supported by the following grants from the National Institutes of Health: National Institute of Diabetes and Digestive and Kidney Diseases grant R01DK098242; National Cancer Institute grant U54CA189201; Illuminating the Druggable Genome; Knowledge Management Center sponsored by the National Institutes of Health Common Fund; National Cancer Institute grant U54-CA189201-02; and the National Center for Advancing Translational Sciences and Clinical and Translational Science Award UL1TR000067. Dr. Shameer has received consulting fees or honoraria from McKinsey, Google, LEK Consulting, Parthenon-EY, Philips Healthcare, and Kencore Health. Dr. Dudley has received consulting fees or honoraria from Janssen Pharmaceuticals, GlaxoSmithKline, AstraZeneca, and Hoffman-La Roche; is a scientific advisor to LAM Therapeutics, NuMedii, and Ayasdi; and holds equity in NuMedii, Ayasdi, and Ontomics. Dr. Ashley is founder of Personalis Inc. and Deepcell Inc; and is an advisor to Genome Medical and SequenceBio. All other authors have reported that they have no relationships relevant to the contents of this paper to disclose.