Large Language Model
Evaluating Large language models on Understanding Korean indirect Speech acts
Koo, Youngeun, Lee, Jiwoo, Park, Dojun, Park, Seohyun, Lee, Sungeun
To accurately understand the intention of an utterance is crucial in conversational communication. As conversational artificial intelligence models are rapidly being developed and applied in various fields, it is important to evaluate the LLMs' capabilities of understanding the intentions of user's utterance. This study evaluates whether current LLMs can understand the intention of an utterance by considering the given conversational context, particularly in cases where the actual intention differs from the surface-leveled, literal intention of the sentence, i.e. indirect speech acts. Our findings reveal that Claude3-Opus outperformed the other competing models, with 71.94% in MCQ and 65% in OEQ, showing a clear advantage. In general, proprietary models exhibited relatively higher performance compared to open-source models. Nevertheless, no LLMs reached the level of human performance. Most LLMs, except for Claude3-Opus, demonstrated significantly lower performance in understanding indirect speech acts compared to direct speech acts, where the intention is explicitly revealed through the utterance. This study not only performs an overall pragmatic evaluation of each LLM's language use through the analysis of OEQ response patterns, but also emphasizes the necessity for further research to improve LLMs' understanding of indirect speech acts for more natural communication with humans.
AI and the Law: Evaluating ChatGPT's Performance in Legal Classification
The use of ChatGPT to analyze and classify evidence in criminal proceedings has been a topic of ongoing discussion. However, to the best of our knowledge, this issue has not been studied in the context of the Polish language. This study addresses this research gap by evaluating the effectiveness of ChatGPT in classifying legal cases under the Polish Penal Code. The results show excellent binary classification accuracy, with all positive and negative cases correctly categorized. In addition, a qualitative evaluation confirms that the legal basis provided for each case, along with the relevant legal content, was appropriate. The results obtained suggest that ChatGPT can effectively analyze and classify evidence while applying the appropriate legal rules. In conclusion, ChatGPT has the potential to assist interested parties in the analysis of evidence and serve as a valuable legal resource for individuals with less experience or knowledge in this area.
A Closer Look at System Prompt Robustness
Mu, Norman, Lu, Jonathan, Lavery, Michael, Wagner, David
System prompts have emerged as a critical control surface for specifying the behavior of LLMs in chat and agent settings. Developers depend on system prompts to specify important context, output format, personalities, guardrails, content policies, and safety countermeasures, all of which require models to robustly adhere to the system prompt, especially when facing conflicting or adversarial user inputs. In practice, models often forget to consider relevant guardrails or fail to resolve conflicting demands between the system and the user. In this work, we study various methods for improving system prompt robustness by creating realistic new evaluation and fine-tuning datasets based on prompts collected from from OpenAI's GPT Store and HuggingFace's HuggingChat. Our experiments assessing models with a panel of new and existing benchmarks show that performance can be considerably improved with realistic fine-tuning data, as well as inference-time interventions such as classifier-free guidance. Finally, we analyze the results of recently released reasoning models from OpenAI and DeepSeek, which show exciting but uneven improvements on the benchmarks we study. Overall, current techniques fall short of ensuring system prompt robustness and further study is warranted.
Why is prompting hard? Understanding prompts on binary sequence predictors
Wenliang, Li Kevin, Ruoss, Anian, Grau-Moya, Jordi, Hutter, Marcus, Genewein, Tim
Large language models (LLMs) can be prompted to do many tasks, but finding good prompts is not always easy, nor is understanding some performant prompts. We explore these issues by viewing prompting as conditioning a near-optimal sequence predictor (LLM) pretrained on diverse data sources. Through numerous prompt search experiments, we show that the unintuitive patterns in optimal prompts can be better understood given the pretraining distribution, which is often unavailable in practice. Moreover, even using exhaustive search, reliably identifying optimal prompts from practical neural predictors can be difficult. Further, we demonstrate that common prompting methods, such as using intuitive prompts or samples from the targeted task, are in fact suboptimal. Thus, this work takes an initial step towards understanding the difficulties in finding and understanding optimal prompts from a statistical and empirical perspective.
Hallucinations are inevitable but statistically negligible
Suzuki, Atsushi, He, Yulan, Tian, Feng, Wang, Zhongyuan
Hallucinations, a phenomenon where a language model (LM) generates nonfactual content, pose a significant challenge to the practical deployment of LMs. While many empirical methods have been proposed to mitigate hallucinations, a recent study established a computability-theoretic result showing that any LM will inevitably generate hallucinations on an infinite set of inputs, regardless of the quality and quantity of training datasets and the choice of the language model architecture and training and inference algorithms. Although the computability-theoretic result may seem pessimistic, its significance in practical viewpoints has remained unclear. In contrast, we present a positive theoretical result from a probabilistic perspective. Specifically, we prove that hallucinations can be made statistically negligible, provided that the quality and quantity of the training data are sufficient. Interestingly, our positive result coexists with the computability-theoretic result, implying that while hallucinations on an infinite set of inputs cannot be entirely eliminated, their probability can always be reduced by improving algorithms and training data. By evaluating the two seemingly contradictory results through the lens of information theory, we argue that our probability-theoretic positive result better reflects practical considerations than the computability-theoretic negative result.
DeepSeek sparks investor pessimism on SoftBank's 500 billion Stargate push
For SoftBank Group investors looking for the stock to climb back to all-time highs on a revival of the artificial intelligence boom, DeepSeek poses a major hurdle. SoftBank is steering a 500 billion fundraising for the Stargate Project to develop AI infrastructure in the U.S., a plan that is key to founder Masayoshi Son's drive to establish a leading position in the emerging field. But now DeepSeek's low-cost AI model is begging the question of whether such massive spending is even necessary. DeepSeek may spark a "near-term market correction" for SoftBank and other AI stocks, said Jung In Yun, chief executive officer at Fibonacci Asset Management Global Pte. While the AI rally should pick up again longer term, the focus will be on monetization, he said, adding that "will take years."
OpenAI rejects 97.4bn Musk bid and says company is not for sale
OpenAI on Friday rejected a 97.4bn bid from a consortium led by billionaire Elon Musk for the ChatGPT maker, saying the startup is not for sale. The unsolicited approach is Musk's latest attempt to block the startup he co-founded with CEO Sam Altman โ but later left โ from becoming a for-profit firm, as it looks to secure more capital and stay ahead in the AI race. "OpenAI is not for sale, and the board has unanimously rejected Mr Musk's latest attempt to disrupt his competition. Any potential reorganization of OpenAI will strengthen our nonprofit and its mission to ensure AGI benefits all of humanity," OpenAI said on X, quoting its chair Bret Taylor, on behalf of its board. On Tuesday, Altman told news website Axios that OpenAI was not for sale.
OpenAI's board 'unanimously' rejects Elon Musk's 97.4 billion takeover bid
Elon Musk launched a 97.4 billion bid to take control of OpenAI. The Wall Street Journal reported a group of investors led by Musk's xAI submitted an unsolicited offer to the company's board of directors on Monday. The group wants to buy the nonprofit that controls OpenAI's for-profit arm. When asked for comment, an OpenAI spokesperson pointed Engadget to an X post from CEO Sam Altman. "No thank you but we will buy twitter for 9.74 billion if you want," Altman wrote on the social media platform Musk owns.
The Guardian is the latest news organization to partner with OpenAI
The Guardian Media Group, owner of The Guardian and The Observer newspapers, is partnering with OpenAI. The deal will see reporting from The Guardian appear as a news source within ChatGPT, alongside article extracts and short summaries. In return, OpenAI will provide the Guardian Media Group with access to ChatGPT Enterprise, which the company says it will use to develop new products, features and tools. "This new partnership with OpenAI reflects the intellectual property rights and value associated with our award-winning journalism, expanding our reach and impact to new audiences and innovative platform services," said Keith Underwood, chief financial and operating officer of the Guardian Media Group. The Guardian Media Group joins a growing list of news publishers that are now working with OpenAI after an initial period of uncertainty over the company and its business model.
Arm is reportedly developing its own in-house chip
Chip designer Arm plans to unveil its own processor this year with Meta as the launch customer, The Financial Times reported. The chip would be a CPU designed for servers in data centers and would have the potential to be customized for clients. Manufacturing would be outsourced to a contract fab plant like TSMC (Taiwan Semiconductor Manufacturing Co.) and the first in-house chip could be revealed as early as this summer, according to the FT's sources. Last month, Arm parent Softbank announced the Stargate project, a partnership with OpenAI to build up to 500 billion worth of AI infrastructure. Arm, along with Microsoft and NVIDIA, is a key technology partner for the project.