Goto

Collaborating Authors

 Deep Learning


Diplomatic duties for Tim Cook after stepping down as Apple CEO

The Guardian

John Ternus ascends the throne - but Cook will stay on to manage tech giant's foreign policy as executive chair Tim Cook becomes Apple's elder statesman Apple announced late on Monday that Tim Cook will step down as CEO but will not leave the iPhone maker. Head of hardware engineering John Ternus will succeed him on 1 September. "I love Apple with all of my being," Cook said in a press release announcing his succession. Cook, 65, who succeeded Apple co-founder Steve Jobs, has been CEO since 2011. With a reputation for operational and supply chain management, he has overseen the global expansion of the company and its steady series of new, updated devices, though he never attained the same visionary status as Jobs.


Bing is the anti-AI search engine you should be using

PCWorld

PCWorld argues that Bing serves as a superior alternative to AI-heavy search engines by prioritizing human-authored content over automated summaries. AI search engines like Google's AI Mode often hide original sources and provide misleading information, with traffic to publishers dropping significantly.


Japanet expands its VC fund after bets on Anthropic and xAI pay off

The Japan Times

Japanet is expanding its venture capital fund with Pegasus Tech Ventures, after early investments in firms like SpaceX, OpenAI, Anthropic and xAI showed strong growth. Japanese home shopping company Japanet is expanding its venture capital fund with San Jose-based Pegasus Tech Ventures, following the success of early bets in SpaceX, OpenAI, Anthropic and xAI. The Nagasaki-based retailer known for infomercials targeting seniors in aging Japan will allocate $200 million to the fund, up from an initial $50 million in 2021, following significant growth" in investments so far, the companies said in a statement. The fund, of which Pegasus is general partner, will focus on areas such as generative AI, robotics and space technology. Its Japan portfolio includes startup Aillis, which seeks to use artificial intelligence to analyze medical scans. Asian companies have struggled to win stakes in promising startups in Silicon Valley, hampered by a lack of personal connections and reputation for slow decision-making. Pegasus also manages startup investments on behalf of Toyota Motor-affiliate Aisin, Japanese chemical maker Denka, Taiwan's Asustek Computer and Acer and Indonesia's pharma company Kalbe Farma. Everybody wants a piece of the Silicon Valley AI action," Pegasus Chief Executive Officer Anis Uzzaman said on a video call.


Woman says chatbot pushed her son to suicide and these 'guardrails' are crucial

Los Angeles Times

Things to Do in L.A. Tap to enable a layout that focuses on the article. Woman says chatbot pushed her son to suicide and these'guardrails' are crucial ChatGPT is among companion chatbots that minors use that one legislator said Monday could be "extremely dangerous." This is read by an automated voice. Please report any issues or inconsistencies here . As the mother of a teen boy who killed himself after using a chatbot, Maria Raine said she was dealing with constant grief.


FUSE: Ensembling Verifiers with Zero Labeled Data

arXiv.org Machine Learning

Verification of model outputs is rapidly emerging as a key primitive for both training and real-world deployment of large language models (LLMs). In practice, this often involves using imperfect LLM judges and reward models since ground truth acquisition can be time-consuming and expensive. We introduce Fully Unsupervised Score Ensembling (FUSE), a method for improving verification quality by ensembling verifiers without access to ground truth correctness labels. The key idea behind FUSE is to control conditional dependencies between verifiers in a manner that improves the unsupervised performance of a class of spectral algorithms from the ensembling literature. Despite requiring zero ground truth labels, FUSE typically matches or improves upon semi-supervised alternatives in test-time scaling experiments with diverse sets of generator models, verifiers, and benchmarks. In particular, we validate our method on both conventional academic benchmarks such as GPQA Diamond and on frontier, unsaturated benchmarks such as Humanity's Last Exam and IMO Shortlist questions.


mlr3torch: A Deep Learning Framework in R based on mlr3 and torch

arXiv.org Machine Learning

Deep learning (DL) has become a cornerstone of modern machine learning (ML) praxis. We introduce the R package mlr3torch, which is an extensible DL framework for the mlr3 ecosystem. It is built upon the torch package, and simplifies the definition, training, and evaluation of neural networks for both tabular data and generic tensors (e.g., images) for classification and regression. The package implements predefined architectures, and torch models can easily be converted to mlr3 learners. It also allows users to define neural networks as graphs. This representation is based on the graph language defined in mlr3pipelines and allows users to define the entire modeling workflow, including preprocessing, data augmentation, and network architecture, in a single graph. Through its integration into the mlr3 ecosystem, the package allows for convenient resampling, benchmarking, preprocessing, and more. We explain the package's design and features and show how to customize and extend it to new problems. Furthermore, we demonstrate the package's capabilities using three use cases, namely hyperparameter tuning, fine-tuning, and defining architectures for multimodal data. Finally, we present some runtime benchmarks.


Differentially Private Conformal Prediction

arXiv.org Machine Learning

Conformal prediction (CP) has attracted broad attention as a simple and flexible framework for uncertainty quantification through prediction sets. In this work, we study how to deploy CP under differential privacy (DP) in a statistically efficient manner. We first introduce differential CP, a non-splitting conformal procedure that avoids the efficiency loss caused by data splitting and serves as a bridge between oracle CP and private conformal inference. By exploiting the stability properties of DP mechanisms, differential CP establishes a direct connection to oracle CP and inherits corresponding validity behavior. Building on this idea, we develop Differentially Private Conformal Prediction (DPCP), a fully private procedure that combines DP model training with a private quantile mechanism for calibration. We establish the end-to-end privacy guarantee of DPCP and investigate its coverage properties under additional regularity conditions. We further study the efficiency of both differential CP and DPCP under empirical risk minimization and general regression models, showing that DPCP can produce tighter prediction sets than existing private split conformal approaches under the same privacy budget. Numerical experiments on synthetic and real datasets demonstrate the practical effectiveness of the proposed methods.


Towards E-Value Based Stopping Rules for Bayesian Deep Ensembles

arXiv.org Machine Learning

Bayesian Deep Ensembles (BDEs) represent a powerful approach for uncertainty quantification in deep learning, combining the robustness of Deep Ensembles (DEs) with flexible multi-chain MCMC. While DEs are affordable in most deep learning settings, (long) sampling of Bayesian neural networks can be prohibitively costly. Yet, adding sampling after optimizing the DEs has been shown to yield significant improvements. This leaves a critical practical question: How long should the sequential sampling process continue to yield significant improvements over the initial optimized DE baseline? To tackle this question, we propose a stopping rule based on E-values. We formulate the ensemble construction as a sequential anytime-valid hypothesis test, providing a principled way to decide whether or not to reject the null hypothesis that MCMC offers no improvement over a strong baseline, to early stop the sampling. Empirically, we study this approach for diverse settings. Our results demonstrate the efficacy of our approach and reveal that only a fraction of the full-chain budget is often required.


This prompt trick forces AI to stop flattering you and think harder

PCWorld

When you purchase through links in our articles, we may earn a small commission. Worried your AI chatbot is just yessing you? Here's a prompt that will make it challenge its own assumptions. I wish I had nickel for every time ChatGPT, Claude, or Gemini told me I'd hit the nail on the head, stumbled onto a genius idea, or otherwise patted me on the back for a half-formed idea or ill-conceived plan. Flattery and premature congratulations are common foibles of generative AI chatbots, with some models more susceptible to being "yes-bots" than others.


LinkedIn's new Crosscheck feature lets premium subscribers test competing AI models for free

Engadget

LinkedIn's new Crosscheck feature lets premium subscribers test competing AI models for free The feature is a blind taste test for AI models from Anthropic, Google, OpenAI and other companies. You can now use LinkedIn to test out some of the latest AI models from OpenAI, Anthropic, Google, Microsoft and other companies without having to worry about token limits or paying for an extra subscription. The professional network is experimenting with a new feature that allows people to test AI platforms' latest offerings within LinkedIn. It's called Crosscheck, and it's rolling out now to anyone with a LinkedIn Premium subscription in the United States. The feature is meant to be a kind of blind taste test for AI models, according to the company's Chief Product Officer Hari Srinivasan.