Generative AI
Panmodal Information Interaction
The chat interface is an essential component of many generative artificial intelligence (GenAI)-based systems. Multi-turn dialog has long shown promise as a way to engage with information systems,5 but is now going mainstream in support of complex tasks via progress in GenAI15 and in GenAI-based conversational systems such as ChatGPTa and Pi.b SearchGPT, recently trialed by OpenAI, provides highly relevant, verifiable answers in a conversational experience. Search engines can now show GenAI answers directly on result pages--minimizing user effort in examining search results but also removing human control over answer generation,9 which can have its own drawbacks (for example, fewer learning opportunities)--and let users follow up via multi-turn conversation for clarification or to seek additional information. Traditional search still has utility for some tasks, for fact finding or navigational tasks, and may be preferred by some searchers given its focus on providing information sources directly rather than synthesized answers. GenAI is also prone to hallucinate (that is, generate nonsensical or inaccurate outputs), making sole reliance on its generated answers inadvisable, although source attribution and answer verification are now creeping in to help users better assess what they can use and trust.
The Download: Google playing AI search catchup, and forming relationships with chatbots
I've been mulling over something that Will Heaven, our senior editor for AI, pointed out not too long ago: all the big players in AI seem to be moving in the same directions and converging on the same things. It's just announced it's adding new AI features from Gemini to search, and adding search features to Gemini. What strikes me more than how well they work is that they are really just about catching up with OpenAI's ChatGPT. And their belated appearance in March of the year 2025 doesn't seem like a great sign for Google. This story originally appeared in The Debrief with Mat Honan, a weekly newsletter about the biggest stories in tech from our editor in chief.
Is Google playing catchup on search with OpenAI?
Take AI Mode, which it announced March 5. It's cool. But it's pretty much a follow-along of what OpenAI was already doing. Google already had something called AI Overviews in search, but AI Mode is different and deeper.) As the company explained in a blog post, "This new Search mode expands what AI Overviews can do with more advanced reasoning, thinking and multimodal capabilities so you can get help with even your toughest questions." Rather than a brief overview with links out, the AI will dig in and offer more robust answers. You can ask followup questions too, something AI Overviews doesn't support.
Position: Model Collapse Does Not Mean What You Think
Schaeffer, Rylan, Kazdan, Joshua, Arulandu, Alvan Caleb, Koyejo, Sanmi
The proliferation of AI-generated content online has fueled concerns over \emph{model collapse}, a degradation in future generative models' performance when trained on synthetic data generated by earlier models. Industry leaders, premier research journals and popular science publications alike have prophesied catastrophic societal consequences stemming from model collapse. In this position piece, we contend this widespread narrative fundamentally misunderstands the scientific evidence. We highlight that research on model collapse actually encompasses eight distinct and at times conflicting definitions of model collapse, and argue that inconsistent terminology within and between papers has hindered building a comprehensive understanding of model collapse. To assess how significantly different interpretations of model collapse threaten future generative models, we posit what we believe are realistic conditions for studying model collapse and then conduct a rigorous assessment of the literature's methodologies through this lens. While we leave room for reasonable disagreement, our analysis of research studies, weighted by how faithfully each study matches real-world conditions, leads us to conclude that certain predicted claims of model collapse rely on assumptions and conditions that poorly match real-world conditions, and in fact several prominent collapse scenarios are readily avoidable. Altogether, this position paper argues that model collapse has been warped from a nuanced multifaceted consideration into an oversimplified threat, and that the evidence suggests specific harms more likely under society's current trajectory have received disproportionately less attention.
Generative AI for Software Architecture. Applications, Trends, Challenges, and Future Directions
Esposito, Matteo, Li, Xiaozhou, Moreschini, Sergio, Ahmad, Noman, Cerny, Tomas, Vaidhyanathan, Karthik, Lenarduzzi, Valentina, Taibi, Davide
Context: Generative Artificial Intelligence (GenAI) is transforming much of software development, yet its application in software architecture is still in its infancy, and no prior study has systematically addressed the topic. Aim: We aim to systematically synthesize the use, rationale, contexts, usability, and future challenges of GenAI in software architecture. Method: We performed a multivocal literature review (MLR), analyzing peer-reviewed and gray literature, identifying current practices, models, adoption contexts, and reported challenges, extracting themes via open coding. Results: Our review identified significant adoption of GenAI for architectural decision support and architectural reconstruction. OpenAI GPT models are predominantly applied, and there is consistent use of techniques such as few-shot prompting and retrieved-augmented generation (RAG). GenAI has been applied mostly to initial stages of the Software Development Life Cycle (SDLC), such as Requirements-to-Architecture and Architecture-to-Code. Monolithic and microservice architectures were the dominant targets. However, rigorous testing of GenAI outputs was typically missing from the studies. Among the most frequent challenges are model precision, hallucinations, ethical aspects, privacy issues, lack of architecture-specific datasets, and the absence of sound evaluation frameworks. Conclusions: GenAI shows significant potential in software design, but several challenges remain on its path to greater adoption. Research efforts should target designing general evaluation methodologies, handling ethics and precision, increasing transparency and explainability, and promoting architecture-specific datasets and benchmarks to bridge the gap between theoretical possibilities and practical use.
Are LLMs (Really) Ideological? An IRT-based Analysis and Alignment Tool for Perceived Socio-Economic Bias in LLMs
Wachter, Jasmin, Radloff, Michael, Smolej, Maja, Kinder-Kurlanda, Katharina
We introduce an Item Response Theory (IRT)-based framework to detect and quantify socioeconomic bias in large language models (LLMs) without relying on subjective human judgments. Unlike traditional methods, IRT accounts for item difficulty, improving ideological bias estimation. We fine-tune two LLM families (Meta-LLaMa 3.2-1B-Instruct and Chat- GPT 3.5) to represent distinct ideological positions and introduce a two-stage approach: (1) modeling response avoidance and (2) estimating perceived bias in answered responses. Our results show that off-the-shelf LLMs often avoid ideological engagement rather than exhibit bias, challenging prior claims of partisanship. This empirically validated framework enhances AI alignment research and promotes fairer AI governance.
How to use ChatGPT as a personal AI research assistant
ChatGPT continues to get more and more capable as features such as web search and scheduled tasks are added, and the latest new AI tool pushed out by OpenAI is an advanced searching feature called Deep Research--yes, just like the similar tool inside Google Gemini. As you can guess from the name, the tool is designed to do a thorough search on the web for information related to your query, then present a detailed report to your specifications. According to OpenAI, Deep Research "leverages reasoning to search, interpret, and analyze massive amounts of text, images, and PDFs on the internet, pivoting as needed in reaction to information it encounters". Right now, you need to be a paying ChatGPT user to access Deep Research, so you'll have to give OpenAI at least 20 per month to make use of it. To date, there's been no word on if or when the feature will make its way to users on the free ChatGPT tier.
Pareidolic Illusions of Meaning: ChatGPT, Pseudolaw and the Triumph of Form over Substance
The early 2020s has seen the rise of two strange and potentially quite impactful social phenomena, namely pseudolaw, where users rely upon pseudolegal arguments that mimic the form and ritual of legal argumentation but fundamentally distort the content of law, and generative AI/LLMs, which generate content that uses probabilistic calculations to create outputs that look like human generated text. This article argues that the juxtaposition of the two phenomena helps to reveal that they both share two fundamental traits as both elevate form and appearance over substance and content, and users of both routinely mistake the form for the substance. In drawing upon legal theory, computer science, linguistics and cognitive psychology, the article argues that both phenomena rely upon creating illusions of meaning that users mistake for the underlying primary phenomenon. I then explore four implications of this conception of both phenomena. Firstly, both rely on human tendencies of conceptual pareidolia resulting in the erroneous perception of meaningful linguistic legal patterns from nebulous inputs. Secondly, both rely upon the confidence heuristic, the human cognitive bias for treating confidence as a proxy for competence. Thirdly, both succeed when the primary concern is with the form of the output and not its content. Fourthly, both rely heavily upon the magical thinking of users and the desire for the promise of the approach to be real. The article argues that the legal context helps to reveal a solution for the problems caused by both phenomena as it is only where users possess sufficient legal and technological literacy that it becomes possible to reveal to them the illusionary nature of the phenomena.
FedGAI: Federated Style Learning with Cloud-Edge Collaboration for Generative AI in Fashion Design
Wu, Mingzhu, Jiang, Jianan, Li, Xinglin, Deng, Hanhui, Wu, Di
Collaboration can amalgamate diverse ideas, styles, and visual elements, fostering creativity and innovation among different designers. In collaborative design, sketches play a pivotal role as a means of expressing design creativity. However, designers often tend to not openly share these meticulously crafted sketches. This phenomenon of data island in the design area hinders its digital transformation under the third wave of AI. In this paper, we introduce a Federated Generative Artificial Intelligence Clothing system, namely FedGAI, employing federated learning to aid in sketch design. FedGAI is committed to establishing an ecosystem wherein designers can exchange sketch styles among themselves. Through FedGAI, designers can generate sketches that incorporate various designers' styles from their peers, drawing inspiration from collaboration without the need for data disclosure or upload. Extensive performance evaluations indicate that our FedGAI system can produce multi-styled sketches of comparable quality to human-designed ones while significantly enhancing efficiency compared to hand-drawn sketches.
Compositional Causal Reasoning Evaluation in Language Models
Maasch, Jacqueline R. M. A., Hüyük, Alihan, Xu, Xinnuo, Nori, Aditya V., Gonzalez, Javier
Causal reasoning and compositional reasoning are two core aspirations in generative AI. Measuring the extent of these behaviors requires principled evaluation methods. We explore a unified perspective that considers both behaviors simultaneously, termed compositional causal reasoning (CCR): the ability to infer how causal measures compose and, equivalently, how causal quantities propagate through graphs. We instantiate a framework for the systematic evaluation of CCR for the average treatment effect and the probability of necessity and sufficiency. As proof of concept, we demonstrate the design of CCR tasks for language models in the LLama, Phi, and GPT families. On a math word problem, our framework revealed a range of taxonomically distinct error patterns. Additionally, CCR errors increased with the complexity of causal paths for all models except o1.