Financial News
DBOT: Artificial Intelligence for Systematic Long-Term Investing
DBOT can value any public traded company on the basis of Damodaran's analysis, and generates a report to support its position in an attempt to mimic its analytic parent. Until recently, such capabilities of analytic twins for financial valuation were not feasible. However, with advances in large language models (LLMs) and generative artificial intelligence (GenAI), it has become possible to conduct valuations that marry numbers and reasoning to generate credible valuations that can be used for long-term investing. The implications for automation and support of various parts of the valuation exercise are profound. In this paper, we provide a method for creating a digital analytic twin, DBOT, which is designed to mimic the investment analysis of individual companies by Damodaran. Since DBOT can value every company in an index such as the S&P500, it also provide an analysis in a macro sense, for example, by valuing the S&P500 market index relative to the valuation of its individual components. From the perspective of generative AI, DBOT presents a multitude of challenges. First and foremost, LLMs must be able to reason over financial texts, charts, tables, and spreadsheets. Furthermore, DBOT requires the AI system to follow Damodaran's
Assessing Consistency and Reproducibility in the Outputs of Large Language Models: Evidence Across Diverse Finance and Accounting Tasks
Wang, Julian Junyan, Wang, Victor Xiaoqi
This study provides the first comprehensive assessment of consistency and reproducibility in Large Language Model (LLM) outputs in finance and accounting research. We evaluate how consistently LLMs produce outputs given identical inputs through extensive experimentation with 50 independent runs across five common tasks: classification, sentiment analysis, summarization, text generation, and prediction. Using three OpenAI models (GPT-3.5-turbo, GPT-4o-mini, and GPT-4o), we generate over 3.4 million outputs from diverse financial source texts and data, covering MD&As, FOMC statements, finance news articles, earnings call transcripts, and financial statements. Our findings reveal substantial but task-dependent consistency, with binary classification and sentiment analysis achieving near-perfect reproducibility, while complex tasks show greater variability. More advanced models do not consistently demonstrate better consistency and reproducibility, with task-specific patterns emerging. LLMs significantly outperform expert human annotators in consistency and maintain high agreement even where human experts significantly disagree. We further find that simple aggregation strategies across 3-5 runs dramatically improve consistency. Simulation analysis reveals that despite measurable inconsistency in LLM outputs, downstream statistical inferences remain remarkably robust. These findings address concerns about what we term "G-hacking," the selective reporting of favorable outcomes from multiple Generative AI runs, by demonstrating that such risks are relatively low for finance and accounting tasks.
SoftBank seals 6.5 billion deal for chip designer Ampere
SoftBank Group has agreed to acquire semiconductor designer Ampere Computing in a move that further broadens the Japanese investment firm's push into artificial intelligence infrastructure. SoftBank is buying Ampere in an all-cash transaction that values the Santa Clara, California-based firm at 6.5 billion, according to a statement. The deal for Ampere, whose early backers included Oracle and private equity firm Carlyle Group, adds to a wave of chip companies looking to capitalize on a spending boom in AI.
Extract, Match, and Score: An Evaluation Paradigm for Long Question-context-answer Triplets in Financial Analysis
Hu, Bo, Yuan, Han, Pandelea, Vlad, Luo, Wuqiong, Zhao, Yingzhu, Ma, Zheng
The rapid advancement of large language models (LLMs) has sparked widespread adoption across diverse applications, making robust evaluation frameworks crucial for assessing their performance. While conventional evaluation metrics remain applicable for shorter texts, their efficacy diminishes when evaluating the quality of long-form answers. This limitation is particularly critical in real-world scenarios involving extended questions, extensive context, and long-form answers, such as financial analysis or regulatory compliance. In this paper, we use a practical financial use case to illustrate applications that handle "long question-context-answer triplets". We construct a real-world financial dataset comprising long triplets and demonstrate the inadequacies of traditional metrics. To address this, we propose an effective Extract, Match, and Score (EMS) evaluation approach tailored to the complexities of long-form LLMs' outputs, providing practitioners with a reliable methodology for assessing LLMs' performance in complex real-world scenarios.
iRobot has new Roombas, but it doesn't sound confident it'll be around to sell them
Beyond declining sales -- the company reported that revenue decreased 47 percent in the US over the prior year in its fourth quarter earnings -- iRobot is also struggling to pay off its debts. The company took on a 200 million bridge loan to stay afloat while it waited for its 1.7 billion acquisition deal with Amazon to be approved, which it's still paying off. The European Commission ultimately investigated the acquisition in 2023, and rather than address its concerns, Amazon terminated the deal and paid out its 94 million termination fee. That wasn't enough to eliminate iRobot's problems, though. The company now plans to review its options and see if it can find another way to stick it out, including "refinancing the company's debt and exploring a potential sale or strategic transaction."
Will Neural Scaling Laws Activate Jevons' Paradox in AI Labor Markets? A Time-Varying Elasticity of Substitution (VES) Analysis
Narayanan, Rajesh P., Pace, R. Kelley
AI industry leaders often use the term ``Jevons' Paradox.'' We explore the significance of this term for artificial intelligence adoption through a time-varying elasticity of substitution framework. We develop a model connecting AI development to labor substitution through four key mechanisms: (1) increased effective computational capacity from both hardware and algorithmic improvements; (2) AI capabilities that rise logarithmically with computation following established neural scaling laws; (3) declining marginal computational costs leading to lower AI prices through competitive pressure; and (4) a resulting increase in the elasticity of substitution between AI and human labor over time. Our time-varying elasticity of substitution (VES) framework, incorporating the G\o rtz identity, yields analytical conditions for market transformation dynamics. This work provides a simple framework to help assess the economic reasoning behind industry claims that AI will increasingly substitute for human labor across diverse economic sectors.
Advanced Deep Learning Techniques for Analyzing Earnings Call Transcripts: Methodologies and Applications
Zakir, Umair, Daykin, Evan, Diagne, Amssatou, Faile, Jacob
This study presents a comparative analysis of deep learning methodologies such as BERT, FinBERT and ULMFiT for sentiment analysis of earnings call transcripts. The objective is to investigate how Natural Language Processing (NLP) can be leveraged to extract sentiment from large-scale financial transcripts, thereby aiding in more informed investment decisions and risk management strategies. We examine the strengths and limitations of each model in the context of financial sentiment analysis, focusing on data preprocessing requirements, computational efficiency, and model optimization. Through rigorous experimentation, we evaluate their performance using key metrics, including accuracy, precision, recall, and F1-score. Furthermore, we discuss potential enhancements to improve the effectiveness of these models in financial text analysis, providing insights into their applicability for real-world financial decision-making.
OpenAI's board 'unanimously' rejects Elon Musk's 97.4 billion takeover bid
Elon Musk launched a 97.4 billion bid to take control of OpenAI. The Wall Street Journal reported a group of investors led by Musk's xAI submitted an unsolicited offer to the company's board of directors on Monday. The group wants to buy the nonprofit that controls OpenAI's for-profit arm. When asked for comment, an OpenAI spokesperson pointed Engadget to an X post from CEO Sam Altman. "No thank you but we will buy twitter for 9.74 billion if you want," Altman wrote on the social media platform Musk owns.
The Tesla Revolt
Donald Trump may be pleased enough with Elon Musk, but even as the Tesla CEO is exercising his newfound power to essentially undo whole functions of the federal government, he still has to reassure his investors. Lately, Musk has delivered for them in one way: The value of the company's shares has skyrocketed since Trump was reelected to the presidency of the United States. But Musk had much to answer for on his recent fourth-quarter earnings call--not least that in 2024, Tesla's car sales had sunk for the first time in a decade. Profits were down sharply too. Usually, when this happens at a car company, the CEO issues a mea culpa, vows to cut costs, and hypes vehicles coming to market soon.
FinBloom: Knowledge Grounding Large Language Model with Real-time Financial Data
Sinha, Ankur, Agarwal, Chaitanya, Malo, Pekka
Large language models (LLMs) excel at generating human-like responses but often struggle with interactive tasks that require access to real-time information. This limitation poses challenges in finance, where models must access up-to-date information, such as recent news or price movements, to support decision-making. To address this, we introduce Financial Agent, a knowledge-grounding approach for LLMs to handle financial queries using real-time text and tabular data. Our contributions are threefold: First, we develop a Financial Context Dataset of over 50,000 financial queries paired with the required context. Second, we train FinBloom 7B, a custom 7 billion parameter LLM, on 14 million financial news articles from Reuters and Deutsche Presse-Agentur, alongside 12 million Securities and Exchange Commission (SEC) filings. Third, we fine-tune FinBloom 7B using the Financial Context Dataset to serve as a Financial Agent. This agent generates relevant financial context, enabling efficient real-time data retrieval to answer user queries. By reducing latency and eliminating the need for users to manually provide accurate data, our approach significantly enhances the capability of LLMs to handle dynamic financial tasks. Our proposed approach makes real-time financial decisions, algorithmic trading and other related tasks streamlined, and is valuable in contexts with high-velocity data flows.