Goto

Collaborating Authors

 Generative AI


Meta just scheduled a generative AI conference called LlamaCon for April 29

Engadget

Meta just announced its first-ever LlamaCon, a dev conference dedicated to generative AI. It's scheduled for April 29. The company titled the event after its family of generative AI models. Meta promises to "share the latest on our open source AI developments to help developers do what they do best: build amazing apps and products." Beyond that vague description, we don't know much.


Elon Musk's startup rolls out new Grok-3 chatbot as AI competition intensifies

The Guardian

Elon Musk's artificial intelligence startup xAI has introduced Grok-3, the latest iteration of its chatbot that integrates with X, formerly Twitter. Grok-3 debut comes at a critical moment in the AI arms race as Musk looks to compete with the Chinese AI firm DeepSeek, Microsoft-backed OpenAI and Google. Musk's bot has seen less widespread adoption than DeepSeek's namesake chatbot, which wowed the world weeks ago and caused panic in stock markets, as well as OpenAI's ChatGPT and Google's Gemini. Grok-3 is being rolled out immediately to Premium subscribers of X, the social media platform owned by Musk. The chatbot can generate texts and images without many of the common guardrails against sexually suggestive imagery, vulgarity or the reproduction of well-known people's likenesses. "Grok-3 across the board is in a league of its own," Musk said during a livestream alongside three xAI engineers late on Monday.


'Hopeless' to potentially handy: law firm puts AI to the test

BBC News

This was the second time Linklaters had run its LinksAI benchmark tests, with the original exercise taking place in October 2023. In the first run, OpenAI's GPT 2, 3 and 4 were tested alongside Google's Bard. The exam has now been expanded to include o1, from OpenAI, and Google's Gemini 2.0, which was also released at the end of 2024. It did not involve DeepSeek's R1 - the apparently low cost Chinese model which astonished the world last month - or any other non-US AI tool. The test involved posing the type of questions which would require advice from a "competent mid-level lawyer" with two years' experience.


Musk debuts Grok-3 AI chatbot to rival OpenAI, DeepSeek

The Japan Times

Elon Musk's artificial intelligence startup, xAI, showed off the updated Grok-3 model, showcasing a version of the chatbot technology that the billionaire has said is the "smartest AI on Earth." Across math, science and coding benchmarks, Grok-3 beats Alphabet's Google Gemini, DeepSeek's V3 model, Anthropic's Claude and OpenAI's GPT-4o, the company said via a live stream on Monday. Grok-3 has "more than 10 times" the computing power of its predecessor and completed pretraining in early January, Musk said in a presentation alongside three of xAI's engineers. "We're continually improving the models every day, and literally within 24 hours, you'll see improvements," Musk said.


Towards an automated workflow in materials science for combining multi-modal simulative and experimental information using data mining and large language models

arXiv.org Artificial Intelligence

To retrieve and compare scientific data of simulations and experiments in materials science, data needs to be easily accessible and machine readable to qualify and quantify various materials science phenomena. However, a majority of information is encoded within scientific documents limiting the capability of finding suitable literature as well as material properties. This manuscript showcases an automated workflow, which unravels the encoded information from scientific literature to a machine readable data structure of texts, figures, tables, equations and meta-data, using natural language processing and language as well as vision transformer models to generate a machine-readable database. The machine-readable database can be enriched with local data, as e.g. The study shows that such an automated workflow accelerates information retrieval, proximate context detection and material property extraction from multi-modal input data exemplarily shown for the research field of microstructural analyses of face-centered cubic single crystals. Ultimately, a Retrieval-Augmented Generation (RAG) based Large Language Model (LLM) enables a fast and e fficient question answering chat bot. Introduction Understanding physical processes in materials and material microstructures is of fundamental importance in facilitating their use in engineering applications. However, analyzing the increasing amount of existing scientific knowledge and extracting the relevant information for a desired research project is a challenging task. Especially, combining information from experiments, simulations and theory is of great significance as different aspects are considered at each discipline that together, ultimately, form a holistic picture [1, 2, 3, 4]. Machine learning (ML) and artificial intelligence (AI) have been recently used as advanced computational tools to accelerate the physical understanding in materials science research [3, 5, 6, 4, 7]. Recent progress in these computational methods enabled AI-assisted models with the ability to extrapolate beyond their data basis and generate novel materials science approaches, called generative AI (genAI) [8, 9]. Applying genAI leads for example to a novel design of crystalline materials [10], of molecule properties [11] and of architected materials [12].


Breaking the bonds of generative artificial intelligence by minimizing the maximum entropy

arXiv.org Artificial Intelligence

The emergence of generative artificial intelligence (GenAI), comprising large language models, text-to-image generators, and AI algorithms for medical drug and material design, had a transformative impact on society. However, despite an initial exponential growth surpassing Moore's law, progress is now plateauing, suggesting we are approaching the limits of current technology. Indeed, these models are notoriously data-hungry, prone to overfitting, and challenging to direct during the generative process, hampering their effective professional employment. To cope with these limitations, we propose a paradigm shift in GenAI by introducing an ab initio method based on the minimal maximum entropy principle. Our approach does not fit the data. Instead, it compresses information in the training set by finding a latent representation parameterized by arbitrary nonlinear functions, such as neural networks. The result is a general physics-driven model, which is data-efficient, resistant to overfitting, and flexible, permitting to control and influence the generative process. Benchmarking shows that our method outperforms variational autoencoders (VAEs) with similar neural architectures, particularly on undersampled datasets. We demonstrate the methods effectiveness in generating images, even with limited training data, and its unprecedented capability to customize the generation process a posteriori without the need of any fine-tuning or retraining.


Two Tickets are Better than One: Fair and Accurate Hiring Under Strategic LLM Manipulations

arXiv.org Artificial Intelligence

In an era of increasingly capable foundation models, job seekers are turning to generative AI tools to enhance their application materials. However, unequal access to and knowledge about generative AI tools can harm both employers and candidates by reducing the accuracy of hiring decisions and giving some candidates an unfair advantage. To address these challenges, we introduce a new variant of the strategic classification framework tailored to manipulations performed using large language models, accommodating varying levels of manipulations and stochastic outcomes. We propose a ``two-ticket'' scheme, where the hiring algorithm applies an additional manipulation to each submitted resume and considers this manipulated version together with the original submitted resume. We establish theoretical guarantees for this scheme, showing improvements for both the fairness and accuracy of hiring decisions when the true positive rate is maximized subject to a no false positives constraint. We further generalize this approach to an $n$-ticket scheme and prove that hiring outcomes converge to a fixed, group-independent decision, eliminating disparities arising from differential LLM access. Finally, we empirically validate our framework and the performance of our two-ticket scheme on real resumes using an open-source resume screening tool.


Flow-based generative models as iterative algorithms in probability space

arXiv.org Machine Learning

Generative AI (GenAI) has revolutionized data-driven modeling by enabling the synthesis of high-dimensional data across various applications, including image generation, language modeling, biomedical signal processing, and anomaly detection. Flow-based generative models provide a powerful framework for capturing complex probability distributions, offering exact likelihood estimation, efficient sampling, and deterministic transformations between distributions. These models leverage invertible mappings governed by Ordinary Differential Equations (ODEs), enabling precise density estimation and likelihood evaluation. This tutorial presents an intuitive mathematical framework for flow-based generative models, formulating them as neural network-based representations of continuous probability densities. We explore key theoretical principles, including the Wasserstein metric, gradient flows, and density evolution governed by ODEs, to establish convergence guarantees and bridge empirical advancements with theoretical insights. By providing a rigorous yet accessible treatment, we aim to equip researchers and practitioners with the necessary tools to effectively apply flow-based generative models in signal processing and machine learning.


Advancing Generative Artificial Intelligence and Large Language Models for Demand Side Management with Internet of Electric Vehicles

arXiv.org Artificial Intelligence

Generative artificial intelligence, particularly through large language models (LLMs), is poised to transform energy optimization and demand side management (DSM) within microgrids. This paper explores the integration of LLMs into energy management, emphasizing their roles in automating the optimization of DSM strategies with Internet of electric vehicles. We investigate challenges and solutions associated with DSM and explore the new opportunities presented by leveraging LLMs. Then, we propose an innovative solution that enhances LLMs with retrieval-augmented generation for automatic problem formulation, code generation, and customizing optimization. We present a case study to demonstrate the effectiveness of our proposed solution in charging scheduling and optimization for electric vehicles, highlighting our solution's significant advancements in energy efficiency and user adaptability. This work underscores the potential of LLMs for energy optimization and fosters a new era of intelligent DSM solutions.


H-CoT: Hijacking the Chain-of-Thought Safety Reasoning Mechanism to Jailbreak Large Reasoning Models, Including OpenAI o1/o3, DeepSeek-R1, and Gemini 2.0 Flash Thinking

arXiv.org Artificial Intelligence

Warning: This paper contains potentially offensive and harmful text. Large Reasoning Models (LRMs) have recently extended their powerful reasoning capabilities to safety checks--using chain-of-thought reasoning to decide whether a request should be answered. While this new approach offers a promising route for balancing model utility and safety, its robustness remains underexplored. To address this gap, we introduce Malicious-Educator, a benchmark that disguises extremely dangerous or malicious requests beneath seemingly legitimate educational prompts. Our experiments reveal severe security flaws in popular commercial-grade LRMs, including OpenAI o1/o3, DeepSeek-R1, and Gemini 2.0 Flash Thinking. For instance, although OpenAI's o1 model initially maintains a high refusal rate of about 98%, subsequent model updates significantly compromise its safety; and attackers can easily extract criminal strategies from DeepSeek-R1 and Gemini 2.0 Flash Thinking without any additional tricks. To further highlight these vulnerabilities, we propose Hijacking Chain-of-Thought (H-CoT), a universal and transferable attack method that leverages the model's own displayed intermediate reasoning to jailbreak its safety reasoning mechanism. Under H-CoT, refusal rates sharply decline--dropping from 98% to below 2%--and, in some instances, even transform initially cautious tones into ones that are willing to provide harmful content. We hope these findings underscore the urgent need for more robust safety mechanisms to preserve the benefits of advanced reasoning capabilities without compromising ethical standards.