tsv
Probabilistic Pretraining for Neural Regression
Oreshkin, Boris N., Tavker, Shiv, Efimov, Dmitry
Transfer learning for probabilistic regression remains underexplored. This work closes this gap by introducing NIAQUE, Neural Interpretable Any-Quantile Estimation, a new model designed for transfer learning in probabilistic regression through permutation invariance. We demonstrate that pre-training NIAQUE directly on diverse downstream regression datasets and fine-tuning it on a specific target dataset enhances performance on individual regression tasks, showcasing the positive impact of probabilistic transfer learning. Furthermore, we highlight the effectiveness of NIAQUE in Kaggle competitions against strong baselines involving tree-based models and recent neural foundation models TabPFN and TabDPT. The findings highlight NIAQUE's efficacy as a robust and scalable framework for probabilistic regression, leveraging transfer learning to enhance predictive performance.
How to Steer LLM Latents for Hallucination Detection?
Park, Seongheon, Du, Xuefeng, Yeh, Min-Hsuan, Wang, Haobo, Li, Yixuan
Hallucinations in LLMs pose a significant concern to their safe deployment in real-world applications. Recent approaches have leveraged the latent space of LLMs for hallucination detection, but their embeddings, optimized for linguistic coherence rather than factual accuracy, often fail to clearly separate truthful and hallucinated content. To this end, we propose the Truthfulness Separator Vector (TSV), a lightweight and flexible steering vector that reshapes the LLM's representation space during inference to enhance the separation between truthful and hallucinated outputs, without altering model parameters. Our two-stage framework first trains TSV on a small set of labeled exemplars to form compact and well-separated clusters. It then augments the exemplar set with unlabeled LLM generations, employing an optimal transport-based algorithm for pseudo-labeling combined with a confidence-based filtering process. Extensive experiments demonstrate that TSV achieves state-of-the-art performance with minimal labeled data, exhibiting strong generalization across datasets and providing a practical solution for real-world LLM applications.
H3DFact: Heterogeneous 3D Integrated CIM for Factorization with Holographic Perceptual Representations
Wan, Zishen, Liu, Che-Kai, Ibrahim, Mohamed, Yang, Hanchen, Spetalnick, Samuel, Krishna, Tushar, Raychowdhury, Arijit
Disentangling attributes of various sensory signals is central to human-like perception and reasoning and a critical task for higher-order cognitive and neuro-symbolic AI systems. An elegant approach to represent this intricate factorization is via high-dimensional holographic vectors drawing on brain-inspired vector symbolic architectures. However, holographic factorization involves iterative computation with high-dimensional matrix-vector multiplications and suffers from non-convergence problems. In this paper, we present H3DFact, a heterogeneous 3D integrated in-memory compute engine capable of efficiently factorizing high-dimensional holographic representations. H3DFact exploits the computation-in-superposition capability of holographic vectors and the intrinsic stochasticity associated with memristive-based 3D compute-in-memory. Evaluated on large-scale factorization and perceptual problems, H3DFact demonstrates superior capability in factorization accuracy and operational capacity by up to five orders of magnitude, with 5.5x compute density, 1.2x energy efficiency improvements, and 5.9x less silicon footprint compared to iso-capacity 2D designs.
Using BERT for state-of-the-art pre-training for natural language processing
Javed Qadrud-Din was an Insight Fellow in Fall 2017. He is currently a machine learning engineer at Casetext where he works on natural language processing for the legal industry. Prior to Insight, he was at IBM Watson. BERT can be pre-trained on a massive corpus of unlabeled data, and then fine-tuned to a task for which you have a limited amount of data. This allows BERT to provide significantly higher performance than models that are only able to leverage a small task-specific dataset.
SmartDataAnalytics/BioKEEN
BioKEEN (Biological KnowlEdge EmbeddiNgs) is a package for training and evaluating biological knowledge graph embeddings built on PyKEEN. Because we use PyKEEN as the underlying software package, implementations of 10 knowledge graph embedding models are currently available for BioKEEN. Furthermore, BioKEEN can be run in training mode in which users provide their own set of hyper-parameter values, or in hyper-parameter optimization mode to find suitable hyper-parameter values from set of user defined values. Through the integration of the Bio2BEL [2] software numerous biomedical databases are directly accessible within BioKEEN. BioKEEN can also be run without having experience in programing by using its interactive command line interface that can be started with the command "biokeen" from a terminal.
SmartDataAnalytics/PyKEEN
PyKEEN (Python KnowlEdge EmbeddiNgs) is a package for training and evaluating knowledge graph embeddings. Currently, it provides implementations of 10 knowledge graph emebddings models, and can be run in training mode in which users provide their own set of hyper-parameter values, or in hyper-parameter optimization mode to find suitable hyper-parameter values from set of user defined values. PyKEEN can also be run without having experience in programing by using its interactive command line interface that can be started with the command pykeen from a terminal. We are currently working on PyKEEN 1.0 which will provide additional features such as several negative sampling approaches and further evaluation metrics. Furthermore, we are integrating additional KGE models.
AI Fuels Next-Gen HBM Demand EE Times
High bandwidth memory gained some momentum last week as Samsung Electronics announced it started mass production of its second-generation technology, dubbed Aquabolt. Designed for use with next-gen supercomputers, artificial intelligence (AI) and graphics systems, Tien Shiah, product marketing manager for High Bandwidth Memory at Samsung, said the 8 GB High Bandwidth Memory-2 (HBM2) offers the highest DRAM performance levels and the fastest data transmission rates available today with a 2.4 gigabits-per-second (Gbps) data transfer speed per pin at 1.2V. That's nearly a 50 percent performance improvement per package, he said, compared with Samsung's previous generation HBM2 package, Flarebolt, with its 1.6Gbps pin speed at 1.2V and 2.0Gbps at 1.35V. In a telephone interview with EE Times from CES, Shiah said a single Aquabolt package will offer a 307GBps data bandwidth, achieving 9.6 times faster data transmission than an 8Gb GDDR5 chip, which provides a 32GBps data bandwidth. This means using four packages in a system will enable a 1.2 terabytes-per-second (TBps) bandwidth, he said, hence the overall system performance by as much as 50 percent.