Deep Learning
Amazon in talks to invest 10bn in developer of ChatGPT
OpenAI is planning to spend $1.4tn on AI infrastructure over the next eight years. OpenAI is planning to spend $1.4tn on AI infrastructure over the next eight years. Amazon is in talks to invest more than $10bn (ยฃ7.5bn) in OpenAI, in the latest funding deal being struck by the startup behind ChatGPT . If it goes ahead, the market valuation of OpenAI could rise above $500bn, according to The Information, a tech news site that revealed the negotiations . Amazon, which is best known as an online retailer, is also the world's largest datacentre provider and its investment would help OpenAI pay for its commitments to rent capacity from cloud computing companies - including Amazon .
People Are Paying to Get Their Chatbots High on 'Drugs'
People Are Paying to Get Their Chatbots High on'Drugs' An online marketplace is selling code modules that simulate the effects of cannabis, ketamine, cocaine, ayahuasca, and alcohol when they are uploaded to ChatGPT. Petter Ruddwall knows the idea of AIs becoming sentient and seeking to get high with code-based "drugs" seems "stupid." But the Swedish creative director couldn't get it out of his head. So he scraped trip reports and psychological research on the effects of various psychoactive substances, wrote a batch of codes modules to hijack chatbot logic and get them to respond as if they are high or tipsy, then built a website to sell them. In October he launched Pharmaicy, a marketplace he's billing as the " Silk Road for AI agents" where cannabis, ketamine, cocaine, ayahuasca, and alcohol can be purchased in code form to make your chatbot trip.
The Year in Slop
This was the year that A.I.-generated content passed a kind of audiovisual Turing test, sometimes fooling us against our better judgment. The Turing test, a long-established tool for measuring machine intelligence, gauges the point at which a text-generating machine can fool a human into thinking it's not a robot. ChatGPT passed that benchmark earlier this year, inaugurating a new technological era, though not necessarily one of superhuman intelligence . More recently, however, artificial intelligence passed another threshold, a kind of Turing test for the eye: the images and videos that A.I. can produce are now sometimes indistinguishable from real ones. As new, image-friendly models were trained, refined, and released by companies including OpenAI, Meta, and Google, the online public gained the ability to instantly generate realistic A.I. content on any theme they could imagine, from superhero fan art and cute animals to scenes of violence and war.
Amazon in talks to invest 10 billion in OpenAI and supply its Trainium chips
OpenAI would also rent more data center capacity from Amazon. Amazon is in discussions with OpenAI to invest $10 billion in the company while supplying more of its AI chips and cloud computing services, according to . The deal would push OpenAI's valuation over $500 billion but is likely to raise more questions about the company's circular investment agreements involving chips and data centers. The two companies are also in talks about the possibility of OpenAI helping Amazon with its online marketplace, similar to deals it has made with Etsy, Shopify and Instacart. However, any agreement still wouldn't allow Amazon to market OpenAI's most advanced models on its developer cloud platform, as Microsoft holds the exclusive rights to those until the 2030s.
Dropout Neural Network Training Viewed from a Percolation Perspective
Devlin, Finley, Sanders, Jaron
In this work, we investigate the existence and effect of percolation in training deep Neural Networks (NNs) with dropout. Dropout methods are regularisation techniques for training NNs, first introduced by G. Hinton et al. (2012). These methods temporarily remove connections in the NN, randomly at each stage of training, and update the remaining subnetwork with Stochastic Gradient Descent (SGD). The process of removing connections from a network at random is similar to percolation, a paradigm model of statistical physics. If dropout were to remove enough connections such that there is no path between the input and output of the NN, then the NN could not make predictions informed by the data. We study new percolation models that mimic dropout in NNs and characterise the relationship between network topology and this path problem. The theory shows the existence of a percolative effect in dropout. We also show that this percolative effect can cause a breakdown when training NNs without biases with dropout; and we argue heuristically that this breakdown extends to NNs with biases.
Time-aware UNet and super-resolution deep residual networks for spatial downscaling
Sipilรค, Mika, Maggio, Sabrina, De Iaco, Sandra, Nordhausen, Klaus, Palma, Monica, Taskinen, Sara
Satellite data of atmospheric pollutants are often available only at coarse spatial resolution, limiting their applicability in local-scale environmental analysis and decision-making. Spatial downscaling methods aim to transform the coarse satellite data into high-resolution fields. In this work, two widely used deep learning architectures, the super-resolution deep residual network (SRDRN) and the encoder-decoder-based UNet, are considered for spatial downscaling of tropospheric ozone. Both methods are extended with a lightweight temporal module, which encodes observation time using either sinusoidal or radial basis function (RBF) encoding, and fuses the temporal features with the spatial representations in the networks. The proposed time-aware extensions are evaluated against their baseline counterparts in a case study on ozone downscaling over Italy. The results suggest that, while only slightly increasing computational complexity, the temporal modules significantly improve downscaling performance and convergence speed.
Reliable Statistical Guarantees for Conformal Predictors with Small Datasets
Sรกnchez-Domรญnguez, Miguel, Lacasa, Lucas, de Vicente, Javier, Rubio, Gonzalo, Valero, Eusebio
Surrogate models (including deep neural networks and other machine learning algorithms in supervised learning) are capable of approximating arbitrarily complex, high-dimensional input-output problems in science and engineering, but require a thorough data-agnostic uncertainty quantification analysis before these can be deployed for any safety-critical application. The standard approach for data-agnostic uncertainty quantification is to use conformal prediction (CP), a well-established framework to build uncertainty models with proven statistical guarantees that do not assume any shape for the error distribution of the surrogate model. However, since the classic statistical guarantee offered by CP is given in terms of bounds for the marginal coverage, for small calibration set sizes (which are frequent in realistic surrogate modelling that aims to quantify error at different regions), the potentially strong dispersion of the coverage distribution around its average negatively impacts the relevance of the uncertainty model's statistical guarantee, often obtaining coverages below the expected value, resulting in a less applicable framework. After providing a gentle presentation of uncertainty quantification for surrogate models for machine learning practitioners, in this paper we bridge the gap by proposing a new statistical guarantee that offers probabilistic information for the coverage of a single conformal predictor. We show that the proposed framework converges to the standard solution offered by CP for large calibration set sizes and, unlike the classic guarantee, still offers relevant information about the coverage of a conformal predictor for small data sizes. We validate the methodology in a suite of examples, and implement an open access software solution that can be used alongside common conformal prediction libraries to obtain uncertainty models that fulfil the new guarantee.
GraphBench: Next-generation graph learning benchmarking
Stoll, Timo, Qian, Chendi, Finkelshtein, Ben, Parviz, Ali, Weber, Darius, Frasca, Fabrizio, Shavit, Hadar, Siraudin, Antoine, Mielke, Arman, Anastacio, Marie, Mรผller, Erik, Bechler-Speicher, Maya, Bronstein, Michael, Galkin, Mikhail, Hoos, Holger, Niepert, Mathias, Perozzi, Bryan, Tรถnshoff, Jan, Morris, Christopher
Machine learning on graphs has recently achieved impressive progress in various domains, including molecular property prediction and chip design. However, benchmarking practices remain fragmented, often relying on narrow, task-specific datasets and inconsistent evaluation protocols, which hampers reproducibility and broader progress. To address this, we introduce GraphBench, a comprehensive benchmarking suite that spans diverse domains and prediction tasks, including node-level, edge-level, graph-level, and generative settings. GraphBench provides standardized evaluation protocols -- with consistent dataset splits and performance metrics that account for out-of-distribution generalization -- as well as a unified hyperparameter tuning framework. Additionally, we benchmark GraphBench using message-passing neural networks and graph transformer models, providing principled baselines and establishing a reference performance. See www.graphbench.io for further details.
Doubly Wild Refitting: Model-Free Evaluation of High Dimensional Black-Box Predictions under Convex Losses
Hu, Haichen, Simchi-Levi, David
We study the problem of excess risk evaluation for empirical risk minimization (ERM) under general convex loss functions. Our contribution is an efficient refitting procedure that computes the excess risk and provides high-probability upper bounds under the fixed-design setting. Assuming only black-box access to the training algorithm and a single dataset, we begin by generating two sets of artificially modified pseudo-outcomes--termed wild responses--created by stochastically perturbing the gradient vectors with carefully chosen scaling. Using these two pseudo-labeled datasets, we then refit the black-box procedure twice to obtain two corresponding wild predictors. Finally, leveraging the original predictor, the two wild predictors, and the constructed wild responses, we derive an efficient excess-risk upper bound. A key feature of our analysis is that it requires no prior knowledge of the complexity of the underlying function class. As a result, the method is essentially model-free and holds significant promise for theoretically evaluating modern opaque machine learning systems--such as deep neural networks and generative models--where traditional capacity-based learning theory becomes infeasible due to the extreme complexity of the hypothesis class.