Large Language Model
Can AI Understand Our Universe? Test of Fine-Tuning GPT by Astrophysical Data
Wang, Yu, Zhang, Shu-Rui, Momtaz, Aidin, Moradi, Rahim, Rastegarnia, Fatemeh, Sahakyan, Narek, Shakeri, Soroush, Li, Liang
ChatGPT has been the most talked-about concept in recent months, captivating both professionals and the general public alike, and has sparked discussions about the changes that artificial intelligence (AI) will bring to the world. As physicists and astrophysicists, we are curious about if scientific data can be correctly analyzed by large language models (LLMs) and yield accurate physics. In this article, we fine-tune the generative pre-trained transformer (GPT) model by the astronomical data from the observations of galaxies, quasars, stars, gamma-ray bursts (GRBs), and the simulations of black holes (BHs), the fine-tuned model demonstrates its capability to classify astrophysical phenomena, distinguish between two types of GRBs, deduce the redshift of quasars, and estimate BH parameters. We regard this as a successful test, marking the LLM's proven efficacy in scientific research. With the ever-growing volume of multidisciplinary data and the advancement of AI technology, we look forward to the emergence of a more fundamental and comprehensive understanding of our universe. This article also shares some interesting thoughts on data collection and AI design. Using the approach of understanding the universe - looking outward at data and inward for fundamental building blocks - as a guideline, we propose a method of series expansion for AI, suggesting ways to train and control AI that is smarter than humans.
FiP: a Fixed-Point Approach for Causal Generative Modeling
Scetbon, Meyer, Jennings, Joel, Hilmkil, Agrin, Zhang, Cheng, Ma, Chao
Modeling true world data-generating processes lies at the heart of empirical science. Structural Causal Models (SCMs) and their associated Directed Acyclic Graphs (DAGs) provide an increasingly popular answer to such problems by defining the causal generative process that transforms random noise into observations. However, learning them from observational data poses an ill-posed and NP-hard inverse problem in general. In this work, we propose a new and equivalent formalism that does not require DAGs to describe them, viewed as fixed-point problems on the causally ordered variables, and we show three important cases where they can be uniquely recovered given the topological ordering (TO). To the best of our knowledge, we obtain the weakest conditions for their recovery when TO is known. Based on this, we design a two-stage causal generative model that first infers the causal order from observations in a zero-shot manner, thus by-passing the search, and then learns the generative fixed-point SCM on the ordered variables. To infer TOs from observations, we propose to amortize the learning of TOs on generated datasets by sequentially predicting the leaves of graphs seen during training. To learn fixed-point SCMs, we design a transformer-based architecture that exploits a new attention mechanism enabling the modeling of causal structures, and show that this parameterization is consistent with our formalism. Finally, we conduct an extensive evaluation of each method individually, and show that when combined, our model outperforms various baselines on generated out-of-distribution problems.
From boom to burst, the AI bubble is only heading in one direction John Naughton
"Are we really in an AI bubble," asked a reader of last month's column about the apparently unstoppable rise of Nvidia, "and how would we know?" Good question, so I asked an AI about it and was pointed to Investopedia, which is written by humans who know about this stuff. It told me that a bubble goes through five stages โ rather as Elisabeth Kรผbler-Ross said people do with grief. For investment bubbles, the five stages are displacement, boom, euphoria, profit-taking and panic. So let's see how this maps on to our experience so far with AI.
Which colors look best on you? These tech tools claim to know.
Now, tech tools such as TikTok effects, stand-alone apps and ChatGPT are bringing the process into our own homes, letting us experiment with different methods of color analysis on the cheap. Finding your best colors, like your star sign or Myers Briggs Type, can be a welcome distraction from life's demands. But color analysis has historically excluded people with darker skin, professional analysts said, and AI is especially liable to regurgitate those biases and misconceptions. In recent years, professionals have moved beyond the traditional four-season system toward a more tailored approach, advising clients on their most flattering garments and jewelry without grouping them into rigid categories.
America's Buggy Internet Problem
Washington Post tech writer Shira Ovide joins Felix Salmon, Emily Peck, and Elizabeth Spiers to discuss what's wrong with America's internet industry, how YouTube became the media empire no one talks about, and the promise and peril of the AI toothbrush. In the Plus segment: OpenAI is using YouTube to train ChatGPT. If you enjoy this show, please consider signing up for Slate Plus. Slate Plus members get an ad-free experience across the network and an additional segment of our regular show every week. You'll also be supporting the work we do here on Slate Money.
Bill Gates reveals 3 jobs most immune to the AI takeover
While Microsoft co-founder Bill Gates remains optimistic about the social benefits of artificial intelligence, now even the billionaire mogul fears that AI could take his job. The candid quip came during a recent podcast conversation with OpenAI CEO Sam Altman, whose company is responsible for the AI-powered chatbot ChatGPT. Over the years, Gates has maintained that the three best career paths for recent graduates are those in alternative energy, health biosciences, and advancing artificial intelligence itself -- but notably'billionaire philanthropist' is not on that list. 'I could even lose my job,' Gates said on his podcast, 'Unconfuse Me with Bill Gates.' 'When the machine says to me, "Bill, go play pickleball, I've got malaria eradication. You're just a slow thinker,"' he worried, 'then it is a philosophically confusing thing.'
Business models for the simulation hypothesis
The simulation hypothesis suggests that we live in a computer simulation. That notion has attracted significant scholarly and popular interest. This article explores the simulation hypothesis from a business perspective. Due to the lack of a name for a universe consistent with the simulation hypothesis, we propose the term simuverse. We argue that if we live in a simulation, there must be a business justification. Therefore, we ask: If we live in a simuverse, what is its business model? We identify and explore business model scenarios, such as simuverse as a project, service, or platform. We also explore business model pathways and risk management issues. The article contributes to the simulation hypothesis literature and is the first to provide a business model perspective on the simulation hypothesis. The article discusses theoretical and practical implications and identifies opportunities for future research related to sustainability, digital transformation, and Artificial Intelligence (AI).
CuriousLLM: Elevating Multi-Document QA with Reasoning-Infused Knowledge Graph Prompting
In the field of Question Answering (QA), unifying large language models (LLMs) with external databases has shown great success. However, these methods often fall short in providing the advanced reasoning needed for complex QA tasks. To address these issues, we improve over a novel approach called Knowledge Graph Prompting (KGP), which combines knowledge graphs with a LLM-based agent to improve reasoning and search accuracy. Nevertheless, the original KGP framework necessitates costly fine-tuning with large datasets yet still suffers from LLM hallucination. Therefore, we propose a reasoning-infused LLM agent to enhance this framework. This agent mimics human curiosity to ask follow-up questions to more efficiently navigate the search. This simple modification significantly boosts the LLM performance in QA tasks without the high costs and latency associated with the initial KGP framework. Our ultimate goal is to further develop this approach, leading to more accurate, faster, and cost-effective solutions in the QA domain.
Towards Efficient Resume Understanding: A Multi-Granularity Multi-Modal Pre-Training Approach
Jiang, Feihu, Qin, Chuan, Zhang, Jingshuai, Yao, Kaichun, Chen, Xi, Shen, Dazhong, Zhu, Chen, Zhu, Hengshu, Xiong, Hui
In the contemporary era of widespread online recruitment, resume understanding has been widely acknowledged as a fundamental and crucial task, which aims to extract structured information from resume documents automatically. Compared to the traditional rule-based approaches, the utilization of recently proposed pre-trained document understanding models can greatly enhance the effectiveness of resume understanding. The present approaches have, however, disregarded the hierarchical relations within the structured information presented in resumes, and have difficulty parsing resumes in an efficient manner. To this end, in this paper, we propose a novel model, namely ERU, to achieve efficient resume understanding. Specifically, we first introduce a layout-aware multi-modal fusion transformer for encoding the segments in the resume with integrated textual, visual, and layout information. Then, we design three self-supervised tasks to pre-train this module via a large number of unlabeled resumes. Next, we fine-tune the model with a multi-granularity sequence labeling task to extract structured information from resumes. Finally, extensive experiments on a real-world dataset clearly demonstrate the effectiveness of ERU.
Adapting Mental Health Prediction Tasks for Cross-lingual Learning via Meta-Training and In-context Learning with Large Language Model
Lifelo, Zita, Ning, Huansheng, Dhelim, Sahraoui
Timely identification is essential for the efficient handling of mental health illnesses such as depression. However, the current research fails to adequately address the prediction of mental health conditions from social media data in low-resource African languages like Swahili. This study introduces two distinct approaches utilising model-agnostic meta-learning and leveraging large language models (LLMs) to address this gap. Experiments are conducted on three datasets translated to low-resource language and applied to four mental health tasks, which include stress, depression, depression severity and suicidal ideation prediction. we first apply a meta-learning model with self-supervision, which results in improved model initialisation for rapid adaptation and cross-lingual transfer. The results show that our meta-trained model performs significantly better than standard fine-tuning methods, outperforming the baseline fine-tuning in macro F1 score with 18\% and 0.8\% over XLM-R and mBERT. In parallel, we use LLMs' in-context learning capabilities to assess their performance accuracy across the Swahili mental health prediction tasks by analysing different cross-lingual prompting approaches. Our analysis showed that Swahili prompts performed better than cross-lingual prompts but less than English prompts. Our findings show that in-context learning can be achieved through cross-lingual transfer through carefully crafted prompt templates with examples and instructions.