Goto

Collaborating Authors

 Generative AI


One Example Shown, Many Concepts Known! Counterexample-Driven Conceptual Reasoning in Mathematical LLMs

arXiv.org Artificial Intelligence

Leveraging mathematical Large Language Models (LLMs) for proof generation is a fundamental topic in LLMs research. We argue that the ability of current LLMs to prove statements largely depends on whether they have encountered the relevant proof process during training. This reliance limits their deeper understanding of mathematical theorems and related concepts. Inspired by the pedagogical method of "proof by counterexamples" commonly used in human mathematics education, our work aims to enhance LLMs' ability to conduct mathematical reasoning and proof through counterexamples. Specifically, we manually create a high-quality, university-level mathematical benchmark, CounterMATH, which requires LLMs to prove mathematical statements by providing counterexamples, thereby assessing their grasp of mathematical concepts. Additionally, we develop a data engineering framework to automatically obtain training data for further model improvement. Extensive experiments and detailed analyses demonstrate that CounterMATH is challenging, indicating that LLMs, such as OpenAI o1, have insufficient counterexample-driven proof capabilities. Moreover, our exploration into model training reveals that strengthening LLMs' counterexample-driven conceptual reasoning abilities is crucial for improving their overall mathematical capabilities. We believe that our work offers new perspectives on the community of mathematical LLMs.


Generative AI-Enhanced Cooperative MEC of UAVs and Ground Stations for Unmanned Surface Vehicles

arXiv.org Artificial Intelligence

The increasing deployment of unmanned surface vehicles (USVs) require computational support and coverage in applications such as maritime search and rescue. Unmanned aerial vehicles (UAVs) can offer low-cost, flexible aerial services, and ground stations (GSs) can provide powerful supports, which can cooperate to help the USVs in complex scenarios. However, the collaboration between UAVs and GSs for USVs faces challenges of task uncertainties, USVs trajectory uncertainties, heterogeneities, and limited computational resources. To address these issues, we propose a cooperative UAV and GS based robust multi-access edge computing framework to assist USVs in completing computational tasks. Specifically, we formulate the optimization problem of joint task offloading and UAV trajectory to minimize the total execution time, which is in the form of mixed integer nonlinear programming and NP-hard to tackle. Therefore, we propose the algorithm of generative artificial intelligence-enhanced heterogeneous agent proximal policy optimization (GAI-HAPPO). The proposed algorithm integrates GAI models to enhance the actor network ability to model complex environments and extract high-level features, thereby allowing the algorithm to predict uncertainties and adapt to dynamic conditions. Additionally, GAI stabilizes the critic network, addressing the instability of multi-agent reinforcement learning approaches. Finally, extensive simulations demonstrate that the proposed algorithm outperforms the existing benchmark methods, thus highlighting the potentials in tackling intricate, cross-domain issues in the considered scenarios.


Auditing Prompt Caching in Language Model APIs

arXiv.org Artificial Intelligence

Prompt caching in large language models (LLMs) results in data-dependent timing variations: cached prompts are processed faster than non-cached prompts. These timing differences introduce the risk of side-channel timing attacks. For example, if the cache is shared across users, an attacker could identify cached prompts from fast API response times to learn information about other users' prompts. Because prompt caching may cause privacy leakage, transparency around the caching policies of API providers is important. To this end, we develop and conduct statistical audits to detect prompt caching in real-world LLM API providers. We detect global cache sharing across users in seven API providers, including OpenAI, resulting in potential privacy leakage about users' prompts. Timing variations due to prompt caching can also result in leakage of information about model architecture. Namely, we find evidence that OpenAI's embedding model is a decoder-only Transformer, which was previously not publicly known.


DeepSeek on a Trip: Inducing Targeted Visual Hallucinations via Representation Vulnerabilities

arXiv.org Artificial Intelligence

Multimodal Large Language Models (MLLMs) represent the cutting edge of AI technology, with DeepSeek models emerging as a leading open-source alternative offering competitive performance to closed-source systems. While these models demonstrate remarkable capabilities, their vision-language integration mechanisms introduce specific vulnerabilities. We implement an adapted embedding manipulation attack on DeepSeek Janus that induces targeted visual hallucinations through systematic optimization of image embeddings. Through extensive experimentation across COCO, DALL-E 3, and SVIT datasets, we achieve hallucination rates of up to 98.0% while maintaining high visual fidelity (SSIM > 0.88) of the manipulated images on open-ended questions. Our analysis demonstrates that both 1B and 7B variants of DeepSeek Janus are susceptible to these attacks, with closed-form evaluation showing consistently higher hallucination rates compared to open-ended questioning. We introduce a novel multi-prompt hallucination detection framework using LLaMA-3.1 8B Instruct for robust evaluation. The implications of these findings are particularly concerning given DeepSeek's open-source nature and widespread deployment potential. This research emphasizes the critical need for embedding-level security measures in MLLM deployment pipelines and contributes to the broader discussion of responsible AI implementation.


Generative AI and Empirical Software Engineering: A Paradigm Shift

arXiv.org Artificial Intelligence

The widespread adoption of generative AI in software engineering marks a paradigm shift, offering new opportunities to design and utilize software engineering tools while influencing both developers and the artifacts they create. Traditional empirical methods in software engineering, including quantitative, qualitative, and mixed-method approaches, are well established. However, this paradigm shift introduces novel data types and redefines many concepts in the software engineering process. The roles of developers, users, agents, and researchers increasingly overlap, blurring the distinctions between these social and technical actors within the field. This paper examines how integrating AI into software engineering challenges traditional research paradigms. It focuses on the research phenomena that we investigate, the methods and theories that we employ, the data we analyze, and the threats to validity that emerge in this new context. Through this exploration, our goal is to understand how AI adoption disrupts established software development practices that creates new opportunities for empirical software engineering research.


Enhancing Higher Education with Generative AI: A Multimodal Approach for Personalised Learning

arXiv.org Artificial Intelligence

This research explores the opportunities of Generative AI (GenAI) in the realm of higher education through the design and development of a multimodal chatbot for an undergraduate course. Leveraging the ChatGPT API for nuanced text-based interactions and Google Bard for advanced image analysis and diagram-to-code conversions, we showcase the potential of GenAI in addressing a broad spectrum of educational queries. Additionally, the chatbot presents a file-based analyser designed for educators, offering deep insights into student feedback via sentiment and emotion analysis, and summarising course evaluations with key metrics. These combinations highlight the crucial role of multimodal conversational AI in enhancing teaching and learning processes, promising significant advancements in educational adaptability, engagement, and feedback analysis. By demonstrating a practical web application, this research underlines the imperative for integrating GenAI technologies to foster more dynamic and responsive educational environments, ultimately contributing to improved educational outcomes and pedagogical strategies.


Elon Musk-led group makes 97.4bn bid for OpenAI

Al Jazeera

A consortium led by Elon Musk said it has offered 97.4bn to buy the nonprofit that controls OpenAI, months after the billionaire sued the artificial intelligence startup to block it from transitioning to a for-profit firm. Musk's bid, revealed on Monday, could ratchet up longstanding tensions with OpenAI CEO Sam Altman over the future of the startup at the heart of a boom in generative AI technology. Altman promptly posted on X: "No thank you but we will buy twitter for 9.74 billion if you want." The two are already embroiled in an ongoing lawsuit. Musk criticised a 500bn OpenAI-led project called Stargate announced with great fanfare at the White House just after United States President Donald Trump returned to office, suggesting the investors involved lacked the funding for the project.


Elon Musk-led group makes surprise bid of nearly 100bn for OpenAI

The Guardian

Elon Musk escalated his feud with OpenAI and its CEO Sam Altman on Monday. The billionaire is leading a consortium of investors that announced it had submitted a bid of 97.4bn for "all assets" of the artificial intelligence company to OpenAI's board of directors. The startup, which operates ChatGPT, has been working to restructure itself away from its original non-profit status. OpenAI also operates a for-profit subsidiary, and Musk's unsolicited offer could complicate the company's plans. The Wall Street Journal first reported the proposed bid. "If Sam Altman and the present OpenAI, Inc. Board of Directors are intent on becoming a fully for-profit corporation, it is vital that the charity be fairly compensated for what its leadership is taking away from it: control over the most transformative technology of our time," said Marc Toberoff, the attorney representing the investors.


Musk-led group makes 97.4bn bid for ChatGPT maker OpenAI

BBC News

OpenAI is widely credited with helping bring artificial intelligence tools into the mainstream and sparking huge investment in the sector. Musk and Altman co-founded the start-up in 2015 as a non-profit company, but the relationship has soured since the Tesla and X boss departed the firm in 2018. Altman is said to be restructuring the company to become a for-profit entity, stripping it of its non-profit board - a move Musk argues means the company has abandoned its founding mission of developing AI for the benefit of humanity. But OpenAI argues its transition into a for-profit firm is required to secure the money needed for developing the best artificial intelligence models. "It's time for OpenAI to return to the open-source, safety-focused force for good it once was. We will make sure that happens," Musk said in a statement.


Elon Musk wants to buy OpenAI for 97.4 billion

Engadget

Elon Musk has launched a 97.4 billion bid to take control of OpenAI. The Wall Street Journal reports a group of investors led by Musk's xAI submitted an unsolicited offer to the company's board of directors on Monday. The group wants to buy the nonprofit that controls OpenAI's for-profit arm. When asked for comment, an OpenAI spokesperson pointed Engadget to an X post from CEO Sam Altman. "No thank you but we will buy twitter for 9.74 billion if you want," Altman wrote on the social media platform Musk owns.