Generative AI
FinAI Data Assistant: LLM-based Financial Database Query Processing with the OpenAI Function Calling API
Kim, Juhyeong, Kim, Yejin, Lee, Youngbin, Byun, Hyunwoo
We present FinAI Data Assistant, a practical approach for natural-language querying over financial databases that combines large language models (LLMs) with the OpenAI Function Calling API. Rather than synthesizing complete SQL via text-to-SQL, our system routes user requests to a small library of vetted, parameterized queries, trading generative flexibility for reliability, low latency, and cost efficiency. We empirically study three questions: (RQ1) whether LLMs alone can reliably recall or extrapolate time-dependent financial data without external retrieval; (RQ2) how well LLMs map company names to stock ticker symbols; and (RQ3) whether function calling outperforms text-to-SQL for end-to-end database query processing. Across controlled experiments on prices and fundamentals, LLM-only predictions exhibit non-negligible error and show look-ahead bias primarily for stock prices relative to model knowledge cutoffs. Ticker-mapping accuracy is near-perfect for NASDAQ-100 constituents and high for S\&P~500 firms. Finally, FinAI Data Assistant achieves lower latency and cost and higher reliability than a text-to-SQL baseline on our task suite. We discuss design trade-offs, limitations, and avenues for deployment.
SpecExit: Accelerating Large Reasoning Model via Speculative Exit
Yang, Rubing, Bai, Huajun, Liu, Song, Yu, Guanghua, Fan, Runzhi, Dang, Yanbin, Zhang, Jiejing, Liu, Kai, Zhu, Jianchen, Chen, Peng
Despite their strong performance on reasoning tasks, large reasoning models (LRMs) often suffer from overthinking, producing unnecessarily long outputs and incurring high end-to-end latency, a significant limitation to their real-world deployment. To address overthinking, early-exit mechanisms have been proposed to terminate reasoning before typical completion, showing that this approach can effectively shorten generation length with minimal impact on accuracy. However, their reliance on probing mechanisms introduces a detection overhead that limits their end-to-end latency gains and compromises their generalizability across diverse problems. Inspired by the use of hidden states in speculative decoding, we propose SpecExit, a novel framework that predicts both future tokens and an early-exit signal directly from a lightweight draft model without probing overhead. Our method offers significant improvements, reducing average generation length by 66% and achieving a 2.5x speedup in end-to-end latency compared to the speculative decoding baseline, without compromising accuracy. Our method leverages the inherent signals from hidden states to provide effective early-exit signals, suggesting broader use of hidden states for efficient reasoning. Large reasoning models (LRMs) such as OpenAI-o1 (OpenAI, 2024), DeepSeek-R1 (DeepSeek-AI et al., 2025) and Qwen (Qwen et al., 2025) have recently achieved state-of-the-art performance in complex tasks.
OpenAI launches AI browser Atlas in latest challenge to Google
OpenAI has unveiled ChatGPT Atlas, a long-anticipated artificial intelligence-powered web browser built around its popular chatbot, in a direct challenge to Google Chrome's dominance. OpenAI on Tuesday unveiled ChatGPT Atlas, a long-anticipated artificial intelligence-powered web browser built around its popular chatbot, in a direct challenge to Google Chrome's dominance. The launch marks OpenAI's latest move to capitalize on 800 million weekly active ChatGPT users, as it expands into more aspects of users' online lives by collecting data about consumers' browser behavior. It could accelerate a broader shift toward AI-driven search, as users increasingly turn to conversational tools that synthesize information instead of relying on traditional keyword-based results from Google -- intensifying competition between OpenAI and Google. Shares of Alphabet, which owns the Chrome browser, were down 1.8% in afternoon trading.
OpenAI's Atlas Browser Takes Direct Aim at Google Chrome
OpenAI's Atlas Browser Takes Direct Aim at Google Chrome The new ChatGPT-powered web browser is OpenAI's boldest play yet to reinvent how people use the web. OpenAI announced on Tuesday it's rolling out a new internet browser called Atlas that integrates directly with ChatGPT . Atlas includes features like a sidebar window people can use to ask ChatGPT questions about the web pages they visit. "We think that AI represents a rare, once a decade opportunity to rethink what a browser can be about," OpenAI CEO Sam Altman said during a livestream announcing Atlas. "Tabs were great, but we haven't seen a lot of browser innovation since then."
ChatGPT Atlas: OpenAI launches web browser centered around its chatbot
OpenAI's CEO, Sam Altman, testifies on Capitol Hill in Washington DC on 8 May. OpenAI's CEO, Sam Altman, testifies on Capitol Hill in Washington DC on 8 May. Company's AI-powered browser built around marquee bot is designed to provide more personalized web experience OpenAI on Tuesday launched an AI-powered web browser built around its marquee chatbot. The browser is designed to provide a more personalized web experience and includes a ChatGPT sidebar that enables users to asks questions about or engage with various aspects of each website they visit, as demonstrated in a video posted alongside the announcement. Atlas is now available globally on Apple's Mac operating system and will soon be made available on Windows, iOS and Android, according to OpenAI's announcement.
Forget SEO. Welcome to the World of Generative Engine Optimization
This holiday season, more shoppers are expected to use chatbots to figure out what to buy. This holiday season, rather than searching on Google, more Americans will likely be turning to large language models to find gifts, deals, and sales. Retailers could see up to a 520 percent increase in traffic from chatbots and AI search engines this year compared to 2024, according to a recent shopping report from Adobe . OpenAI is already moving to capitalize on the trend: Last week, the ChatGPT maker announced a major partnership with Walmart that will allow users to buy goods directly within the chat window. As people start relying on chatbots to discover new products, retailers are having to rethink their approach to online marketing.
Meta Poaches Key Google AI Researcher
Upon its release earlier this month, OpenAI's Sora 2 model took the Internet by storm, thanks to its ability to generate realistic videos from just a text prompt. But Sora is about more than just capturing eyeballs with viral content. "On the surface, Sora, for example, does not look like it is AGI-relevant," OpenAI CEO Sam Altman said on a podcast earlier this month. "But I would bet that if we can build really great world models, that will be much more important to AGI than people think." Altman was speaking to a growing belief inside the AI industry at large: that if you can simulate the world with enough accuracy, you could drop AI agents into those simulations. There, they could learn more skills than they currently can from just text, photos, and videos--because they could interact with a simulated world. That form of training could be highly efficient, in part because simulated time can be accelerated, and because many simulations can be run in parallel.
Salesforce's CEO backtracks after saying Trump should send troops into San Francisco
Salesforce's CEO backtracks after saying Trump should send troops into San Francisco In tech this week: The CEO of the city's largest private employer apologizes, Amazon Web Services' outage and OpenAI's Sora makes waves What I'm watching this week: South Park's caricature of Peter Thiel and his obsession with the antichrist . Read our reporting on the show's inspiration: Thiel's bizarre off-the-record lectures on the subject. And now, let's get into things. The co-founder and CEO of Salesforce, said last week that Donald Trump should make good on his threats to send the US national guard into San Francisco, despite resistance from local leaders. Even Marc Benioff's own public relations manager was aghast at his remarks, according to the New York Times .
Bryan Cranston thanks OpenAI for cracking down on Sora 2 deepfakes
Bryan Cranston pictured speaking at a Sag-Aftra strike rally in 2023 in New York. The Breaking Bad actor went to the union with concerns after users of OpenAI's generative video platform Sora 2 were able to generate his likeness without his consent. Bryan Cranston pictured speaking at a Sag-Aftra strike rally in 2023 in New York. The Breaking Bad actor went to the union with concerns after users of OpenAI's generative video platform Sora 2 were able to generate his likeness without his consent. Users of generative AI video app were able to recreate the Breaking Bad actor's likeness without his consent, which OpenAI called'unintentional' Bryan Cranston has said he is "grateful" to OpenAI for cracking down on deepfakes of himself on the company's generative AI video platform Sora 2, after users were able to generate his voice and likeness without his consent.
VERINA: Benchmarking Verifiable Code Generation
Ye, Zhe, Yan, Zhengxu, He, Jingxuan, Kasriel, Timothe, Yang, Kaiyu, Song, Dawn
Large language models (LLMs) are increasingly integrated in software development, but ensuring correctness in LLM-generated code remains challenging and often requires costly manual review. Verifiable code generation -- jointly generating code, specifications, and proofs of code-specification alignment -- offers a promising path to address this limitation and further unleash LLMs' benefits in coding. Yet, there exists a significant gap in evaluation: current benchmarks often focus on only individual components rather than providing a holistic evaluation framework of all tasks. In this paper, we introduce Verina (Verifiable Code Generation Arena), a high-quality benchmark enabling a comprehensive and modular evaluation of code, specification, and proof generation as well as their compositions. Verina consists of 189 manually curated coding tasks in Lean, with detailed problem descriptions, reference implementations, formal specifications, and extensive test suites. Our extensive evaluation of state-of-the-art LLMs reveals significant challenges in verifiable code generation, especially in proof generation, underscoring the need for improving LLM-based theorem provers in verification domains. The best model, OpenAI o4-mini, achieves a 61.4\% code correctness rate, 51.0\% for specification soundness and completeness, and a mere 3.6\% proof success rate (based on one trial per task). We hope Verina will catalyze progress in verifiable code generation by providing a rigorous and comprehensive benchmark. We release our dataset on https://huggingface.co/datasets/sunblaze-ucb/verina and our evaluation code on https://github.com/sunblaze-ucb/verina.