Large Language Model
LeanReasoner: Boosting Complex Logical Reasoning with Lean
Jiang, Dongwei, Fonseca, Marcio, Cohen, Shay B.
Large language models (LLMs) often struggle with complex logical reasoning due to logical inconsistencies and the inherent difficulty of such reasoning. We use Lean, a theorem proving framework, to address these challenges. By formalizing logical reasoning problems into theorems within Lean, we can solve them by proving or disproving the corresponding theorems. This method reduces the risk of logical inconsistencies with the help of Lean's symbolic solver. It also enhances our ability to treat complex reasoning tasks by using Lean's extensive library of theorem proofs. Our method achieves state-of-the-art performance on the FOLIO dataset and achieves performance near this level on ProofWriter. Notably, these results were accomplished by fine-tuning on fewer than 100 in-domain samples for each dataset.
A Large Language Model Enhanced Sequential Recommender for Joint Video and Comment Recommendation
Zheng, Bowen, Lin, Zihan, Liu, Enze, Yang, Chen, Bai, Enyang, Ling, Cheng, Zhao, Wayne Xin, Wen, Ji-Rong
In online video platforms, reading or writing comments on interesting videos has become an essential part of the video watching experience. However, existing video recommender systems mainly model users' interaction behaviors with videos, lacking consideration of comments in user behavior modeling. In this paper, we propose a novel recommendation approach called LSVCR by leveraging user interaction histories with both videos and comments, so as to jointly conduct personalized video and comment recommendation. Specifically, our approach consists of two key components, namely sequential recommendation (SR) model and supplemental large language model (LLM) recommender. The SR model serves as the primary recommendation backbone (retained in deployment) of our approach, allowing for efficient user preference modeling. Meanwhile, we leverage the LLM recommender as a supplemental component (discarded in deployment) to better capture underlying user preferences from heterogeneous interaction behaviors. In order to integrate the merits of the SR model and the supplemental LLM recommender, we design a twostage training paradigm. The first stage is personalized preference alignment, which aims to align the preference representations from both components, thereby enhancing the semantics of the SR model. The second stage is recommendation-oriented fine-tuning, in which the alignment-enhanced SR model is fine-tuned according to specific objectives. Extensive experiments in both video and comment recommendation tasks demonstrate the effectiveness of LSVCR. Additionally, online A/B testing on the KuaiShou platform verifies the actual benefits brought by our approach. In particular, we achieve a significant overall gain of 4.13% in comment watch time.
Apple's MM1 AI Model Shows a Sleeping Giant Is Waking Up
While the tech industry went gaga for generative artificial intelligence, one giant has held back: Apple. The company has yet to introduce so much as an AI-generated emoji, and according to a New York Times report today and earlier reporting from Bloomberg, it is in preliminary talks with Google about adding the search company's Gemini AI model to iPhones. Yet a research paper quietly posted online last Friday by Apple engineers suggests that the company is making significant new investments into AI that are already bearing fruit. It details the development of a new generative AI model called MM1 capable of working with text and images. The researchers show it answering questions about photos and displaying the kind of general knowledge skills shown by chatbots like ChatGPT.
Microsoft hires DeepMind cofounder to lead its new consumer AI division
Microsoft now has a lone leader overseeing consumer AI for the first time. Suleyman will try to push the consumer-facing Copilot assistant into the future, preparing for what may be a long battle with Google for artificial intelligence supremacy among Silicon Valley's Big Five companies. Suleyman's official title will be executive vice president and CEO of a new division called Microsoft AI, reporting directly to CEO Satya Nadella. Joining him will be fellow Inflection AI cofounder Karén Simonyan, who takes the title of chief scientist. "Messy" could be one way to describe Microsoft's Copilot rollout.
DeepMind and Liverpool FC develop AI to advise on football tactics
An artificial intelligence model can predict the outcome of corner kicks in football matches and help coaches design tactics that increase or decrease the probability of a player taking a shot at goal. Petar Veličković at Google DeepMind and his colleagues developed the tool, called TacticAI, as part of a three-year research collaboration with Liverpool Football Club. Corner kicks are awarded when the ball goes out of play over the goal line, and can be a good scoring opportunity for the attacking team. Because of this, football coaches develop detailed plans for various scenarios, which players learn ahead of games. TacticAI was trained on data from 7176 corner kicks in England's 2020 to 2021 Premier League season, including each player's position over time and their height and weight.
Nvidia: what's so good about the tech firm's new AI superchip?
The chipmaker Nvidia has extended its lead in artificial intelligence with the unveiling of a new "superchip", a quantum computing service, and a new suite of tools to help develop the ultimate sci-fi dream: general purpose humanoid robotics. Here we look at what the company is doing and what it might mean. The main announcement of the company's annual develop conference on Monday was the "Blackwell" series of AI chips, used to power the fantastically expensive datacentres that train frontier AI models such as the latest generations of GPT, Claude and Gemini. One, the Blackwell B200, is a fairly straightforward upgrade over the company's pre-existing H100 AI chip. Training a massive AI model, the size of GPT-4, would currently take about 8,000 H100 chips, and 15 megawatts of power, Nvidia said – enough to power about 30,000 typical British homes.
The Morning After: NVIDIA says its Blackwell GPUs are the world's most powerful chips
NVIDIA's H100 chips are used by nearly every AI company in the world to train large language models hooked into services like ChatGPT. It's been great for business. Now, the company is ready to make those chips look terrible, announcing a next-generation platform called Blackwell. Named for David Harold Blackwell, a mathematician who specialized in game theory and statistics, NVIDIA claims Blackwell is the world's most powerful chip, reaching speeds of 20 petaflops compared to just 4 petaflops the H100 provided. Yeah, throw it in the trash.
The Lifelike Illusions of A.I.
In January, 1999, the Washington Post reported that the National Security Agency had issued a memo on its intranet with the subject "Furby Alert." According to the Post, the memo decreed that employees were prohibited from bringing to work any recording devices, including "toys, such as'Furbys,' with built-in recorders that repeat the audio with synthesized sound." That holiday season, the Furby, an animatronic toy resembling a small owl, had been a retail sensation; nearly two million were sold by year's end. They were now banned from N.S.A. headquarters. A worry, according to one source for the Post, was that the toy might "start talking classified." Tiger Electronics, the makers of the Furby, was perplexed.
Nvidia's Blackwell AI 'superchip' is the most powerful yet
Nvidia has unveiled a "superchip" for training artificial intelligence models, the most powerful it has ever produced. The US computing firm, which has recently rocketed in value to become the world's third-largest company, has not yet revealed the cost of its new chips, but observers expect a high price tag that will make them accessible to only a few organisations. The chips were announced by Nvidia CEO Jensen Huang at a press conference in San Jose, California on 18 March. He showed off the company's new Blackwell B200 graphics processing units (GPUs), each of which has 208 billion transistors – the tiny switches at the heart of modern computing devices – compared to the 80 billion transistors of Nvidia's current-generation Hopper chips. He also revealed the GB200 Grace Blackwell Superchip, which combines two of the B200 chips.
NVIDIA's GPUs powered the AI revolution. Its new Blackwell chips are up to 30 times faster
In less than two years, NVIDIA's H100 chips, which are used by nearly every AI company in the world to train large language models that power services like ChatGPT, made it one of the world's most valuable companies. On Monday, NVIDIA announced a next-generation platform called Blackwell, whose chips are between seven and 30 times faster than the H100 and use 25 times less power. "Blackwell GPUs are the engine to power this new Industrial Revolution," said NVIDIA CEO Jensen Huang at the company's annual GTC event in San Jose attended by thousands of developers, and which some compared to a Taylor Swift concert. "Generative AI is the defining technology of our time. Working with the most dynamic companies in the world, we will realize the promise of AI for every industry," Huang added in a press release.