Goto

Collaborating Authors

 Generative AI


Do AI Companies Actually Care About America?

The Atlantic - Technology

In early May, Sam Altman traveled to Washington to tell a story about America. Appearing before a Senate committee, Altman described how he came of age as the internet took off, how he stayed up late in his family's attic and learned to code on products that were invented in the United States--a personal computer, its silicon chips and accompanying software. That early experience with the "spirit of American innovation," Altman told the senators, put him on a path to found OpenAI, launch ChatGPT, and set off the AI boom. "I think America is just an incredible and special thing," he said, "and it will not only be the place where the AI revolution happens but all the revolutions after." Altman's written testimony, which was submitted to the Senate, added an important asterisk that he did not speak aloud that day.


The A.I.-Profits Drought and the Lessons of History

The New Yorker

In a 1987 article in the Times Book Review, Robert Solow, a Nobel-winning economist at M.I.T., commented, "You can see the computer age everywhere but in the productivity statistics." Despite massive increases in computing power and the rising popularity of personal computers, government figures showed that over-all output per worker, a key determinant of wages and living standards, had stagnated for more than a decade. The "productivity paradox," as it came to be known, persisted into the nineteen-nineties and beyond, generating a huge and inconclusive body of literature. Some economists blamed mismanagement of the new technology; others argued that computers paled in economic importance compared to older inventions such as the steam engine and electricity; still others blamed measurement errors in the data and argued that once these were corrected the paradox disappeared. Nearly forty years after Solow's article, and almost three years since OpenAI released its ChatGPT chatbot, we may be facing a new economic paradox, this one involving generative artificial intelligence.


Set Transformer Architectures and Synthetic Data Generation for Flow-Guided Nanoscale Localization

arXiv.org Artificial Intelligence

Flow-guided Localization (FGL) enables the identification of spatial regions within the human body that contain an event of diagnostic interest. FGL does that by leveraging the passive movement of energy-constrained nanodevices circulating through the bloodstream. Existing FGL solutions rely on graph models with fixed topologies or handcrafted features, which limit their adaptability to anatomical variability and hinder scalability. In this work, we explore the use of Set Transformer architectures to address these limitations. Our formulation treats nanodevices' circulation time reports as unordered sets, enabling permutation-invariant, variable-length input processing without relying on spatial priors. To improve robustness under data scarcity and class imbalance, we integrate synthetic data generation via deep generative models, including CGAN, WGAN, WGAN-GP, and CVAE. These models are trained to replicate realistic circulation time distributions conditioned on vascular region labels, and are used to augment the training data. Our results show that the Set Transformer achieves comparable classification accuracy compared to Graph Neural Networks (GNN) baselines, while simultaneously providing by-design improved generalization to anatomical variability. The findings highlight the potential of permutation-invariant models and synthetic augmentation for robust and scalable nanoscale localization.


QU-NLP at QIAS 2025 Shared Task: A Two-Phase LLM Fine-Tuning and Retrieval-Augmented Generation Approach for Islamic Inheritance Reasoning

arXiv.org Artificial Intelligence

This paper presents our approach and results for SubTask 1: Islamic Inheritance Reasoning at QIAS 2025, a shared task focused on evaluating Large Language Models (LLMs) in understanding and reasoning within Islamic inheritance knowledge. We fine-tuned the Fanar-1-9B causal language model using Low-Rank Adaptation (LoRA) and integrated it into a Retrieval-Augmented Generation (RAG) pipeline. Our system addresses the complexities of Islamic inheritance law, including comprehending inheritance scenarios, identifying eligible heirs, applying fixed-share rules, and performing precise calculations. Our system achieved an accuracy of 0.858 in the final test, outperforming other competitive models such as, GPT 4.5, LLaMA, Fanar, Mistral and ALLaM evaluated with zero-shot prompting. Our results demonstrate that QU-NLP achieves near state-of-the-art accuracy (85.8%), excelling especially on advanced reasoning (97.6%) where it outperforms Gemini 2.5 and OpenAI's o3. This highlights that domain-specific fine-tuning combined with retrieval grounding enables mid-scale Arabic LLMs to surpass frontier models in Islamic inheritance reasoning.


Benchmarking the Legal Reasoning of LLMs in Arabic Islamic Inheritance Cases

arXiv.org Artificial Intelligence

Islamic inheritance domain holds significant importance for Muslims to ensure fair distribution of shares between heirs. Manual calculation of shares under numerous scenarios is complex, time-consuming, and error-prone. Recent advancements in Large Language Models (LLMs) have sparked interest in their potential to assist with complex legal reasoning tasks. This study evaluates the reasoning capabilities of state-of-the-art LLMs to interpret and apply Islamic inheritance laws. We utilized the dataset proposed in the ArabicNLP QIAS 2025 challenge, which includes inheritance case scenarios given in Arabic and derived from Islamic legal sources. Various base and fine-tuned models, are assessed on their ability to accurately identify heirs, compute shares, and justify their reasoning in alignment with Islamic legal principles. Our analysis reveals that the proposed majority voting solution, leveraging three base models (Gemini Flash 2.5, Gemini Pro 2.5, and GPT o3), outperforms all other models that we utilized across every difficulty level. It achieves up to 92.7% accuracy and secures the third place overall in Task 1 of the Qias 2025 challenge.


Are LLM-Powered Social Media Bots Realistic?

arXiv.org Artificial Intelligence

As Large Language Models (LLMs) become more sophisticated, there is a possibility to harness LLMs to power social media bots. This work investigates the realism of generating LLM-Powered social media bot networks. Through a combination of manual effort, network science and LLMs, we create synthetic bot agent personas, their tweets and their interactions, thereby simulating social media networks. We compare the generated networks against empirical bot/human data, observing that both network and linguistic properties of LLM-Powered Bots differ from Wild Bots/Humans. This has implications towards the detection and effectiveness of LLM-Powered Bots.


MedArabiQ: Benchmarking Large Language Models on Arabic Medical Tasks

arXiv.org Artificial Intelligence

Large Language Models (LLMs) have demonstrated significant promise for various applications in healthcare. However, their effectiveness in the Arabic medical domain remains unexplored due to the lack of high-quality domain-specific datasets and benchmarks. This study introduces MedArabiQ, a new benchmark dataset consisting of seven Arabic medical tasks, covering multiple specialties and including multiple-choice questions, fill-in-the-blank questions, and patient-doctor questions and answers. We first constructed the dataset using past medical exams as well as publicly available datasets. We conducted an extensive evaluation with eight state-of-the-art open-access and proprietary high-resource LLMs, including GPT-4, Deepseek v3, and Gemini 1.5. Our findings highlight the need for the creation of new high-quality benchmarks that span different languages to ensure fair deployment and scalability of LLMs in healthcare. By establishing this benchmark and releasing the dataset, we provide a foundation for future research aimed at evaluating and enhancing the multilingual capabilities of LLMs for the equitable use of generative AI in healthcare. Data Availability In this article, we present a new benchmark dataset, MedArabiQ, designed to evaluate the performance of LLMs on Arabic medical tasks.


Deal to get ChatGPT Plus for whole of UK discussed by Open AI boss and minister

The Guardian

The boss of the firm behind ChatGPT and the UK technology secretary discussed a multibillion-pound deal to give the entire country premium access to the AI tool, the Guardian has learned. Sam Altman, a co-founder of OpenAI, talked to Peter Kyle about a potential agreement to give UK residents access to its advanced product. According to two sources with direct knowledge of the meeting, the idea was floated as part of a broader discussion in San Francisco about opportunities for collaboration between OpenAI and the UK. Those close to the discussion say Kyle never really took the idea seriously, not least because it could have cost as much as 2bn. OpenAI offers free and subscription versions of ChatGPT.


A-levels and GCSEs need overhaul to keep pace with generative AI, experts say

The Guardian

Oral assessments, more security checks and speedier marking are all on the cards as generative artificial intelligence (AI) could transform exams for the next generation of students. As the 2025 exam season drew to a close with GCSE students picking up their results on Thursday, after mostly sitting traditional pen and paper exams, AI is already changing the landscape. Exam preparation is undergoing a revolution, with students increasingly creating personal AI tutors, available around the clock to generate learning materials to suit individual needs that potentially lead to better results. "Using AI can give a student a much better understanding of a subject because they can ask those questions they wouldn't ask in class, or at odd hours, without being judged," said Dr Andrew Rogoyski of the Surrey Institute for People-Centred AI. "It really took off this summer," said Sandra Leaton Gray, a professor of education futures at University College London's Institute of Education. "So they're able to talk to it about the marking frameworks that are in use and upload those, and then they're able to do sample answers on their own. And then they're able to say to the AI: 'How would you improve the answer?' It's like having a tireless tutor."


Join Us for WIRED's "Uncanny Valley" Live

WIRED

On September 9, WIRED is partnering with KQED for Uncanny Valley's first live show of the podcast. Join us in San Francisco to see hosts Katie Drummond, Michael Calore, and Lauren Goode shed light on the people, power, and influence of Silicon Valley. With original reporting and sharp analysis, Uncanny Valley covers today's biggest stories in tech. We demystify companies like Palantir, trends like vibe coding, and figures like Sam Altman; we break down WIRED's essential coverage of DOGE and ICE; we guide listeners through breakthrough innovation like generative AI and sweeping policy changes like the Trump Administration's tariffs. We're thrilled to have the opportunity to see our listeners in person.