Goto

Collaborating Authors

 Generative AI


Circinus: Efficient Query Planner for Compound ML Serving

arXiv.org Artificial Intelligence

The rise of compound AI serving -- integrating multiple operators in a pipeline that may span edge and cloud tiers -- enables end-user applications such as autonomous driving, generative AI-powered meeting companions, and immersive gaming. Achieving high service goodput -- i.e., meeting service level objectives (SLOs) for pipeline latency, accuracy, and costs -- requires effective planning of operator placement, configuration, and resource allocation across infrastructure tiers. However, the diverse SLO requirements, varying edge capabilities, and high query volumes create an enormous planning search space, rendering current solutions fundamentally limited for real-time serving and cost-efficient deployments. This paper presents Circinus, an SLO-aware query planner for large-scale compound AI workloads. Circinus novelly decomposes multi-query planning and multi-dimensional SLO objectives while preserving global decision quality. By exploiting plan similarities within and across queries, it significantly reduces search steps. It further improves per-step efficiency with a precision-aware plan profiler that incrementally profiles and strategically applies early stopping based on imprecise estimates of plan performance. At scale, Circinus selects query-plan combinations to maximize global SLO goodput. Evaluations in real-world settings show that Circinus improves service goodput by 3.2-5.0$\times$, accelerates query planning by 4.2-5.8$\times$, achieving query response in seconds, while reducing deployment costs by 3.2-4.0$\times$ over state of the arts even in their intended single-tier deployments.


QAOA-GPT: Efficient Generation of Adaptive and Regular Quantum Approximate Optimization Algorithm Circuits

arXiv.org Artificial Intelligence

--Quantum computing has the potential to improve our ability to solve certain optimization problems that are computationally difficult for classical computers, by offering new algorithmic approaches that may provide speedups under specific conditions. In this work, we introduce QAOA-GPT, a generative framework that leverages Generative Pretrained Transformers (GPT) to directly synthesize quantum circuits for solving quadratic unconstrained binary optimization problems, and demonstrate it on the MaxCut problem on graphs. T o diversify the training circuits and ensure their quality, we have generated a synthetic dataset using the adaptive QAOA approach, a method that incrementally builds and optimizes problem-specific circuits. The experiments conducted on a curated set of graph instances demonstrate that QAOA-GPT, generates high quality quantum circuits for new problem instances unseen in the training as well as successfully parametrizes QAOA. Our results show that using QAOA-GPT to generate quantum circuits will significantly decrease both the computational overhead of classical QAOA and adaptive approaches that often use gradient evaluation to generate the circuit and the classical optimization of the circuit parameters. Our work shows that generative AI could be a promising avenue to generate compact quantum circuits in a scalable way. Quantum computing is rapidly emerging technology with significant potential across various domains, including finance [1], chemical simulations [2], material science [3], combinatorial optimization [4], and machine learning [5], among others. V ariational quantum-classical algorithms represent one of the most promising classes of quantum algorithms in different domains, showing potential for both fault-tolerant quantum computers and near-term noisy intermediate-scale quantum (NISQ) devices. The Quantum Approximate Optimization Algorithm (QAOA) [6] and many of its subsequent versions and customizations [7] belong to this class and demonstrate great potential due to their problem/application flexibility and compatibility with various quantum architectures. The original QAOA framework employs a fixed ansatz structure, which can limit expressibility and hinder performance, particularly on near-term quantum devices where circuit depth is limited. This rigid design may not capture the problem-specific features needed for efficient optimization. Such methods as ADAPT -QAOA [8] address this challenge by iteratively constructing the ansatz in a problem-informed manner. At each step, ADAPT -QAOA selects operators from a predefined pool based on their gradient with respect to the cost function, incorporating only those that contribute most significantly to improving the objective.


Reflexive Prompt Engineering: A Framework for Responsible Prompt Engineering and Interaction Design

arXiv.org Artificial Intelligence

Responsible prompt engineering has emerged as a critical framework for ensuring that generative artificial intelligence (AI) systems serve society's needs while minimizing potential harms. As generative AI applications become increasingly powerful and ubiquitous, the way we instruct and interact with them through prompts has profound implications for fairness, accountability, and transparency. This article examines how strategic prompt engineering can embed ethical and legal considerations and societal values directly into AI interactions, moving beyond mere technical optimization for functionality. This article proposes a comprehensive framework for responsible prompt engineering that encompasses five interconnected components: prompt design, system selection, system configuration, performance evaluation, and prompt management. Drawing from empirical evidence, the paper demonstrates how each component can be leveraged to promote improved societal outcomes while mitigating potential risks. The analysis reveals that effective prompt engineering requires a delicate balance between technical precision and ethical consciousness, combining the systematic rigor and focus on functionality with the nuanced understanding of social impact. Through examination of real-world and emerging practices, the article illustrates how responsible prompt engineering serves as a crucial bridge between AI development and deployment, enabling organizations to fine-tune AI outputs without modifying underlying model architectures. This approach aligns with broader "Responsibility by Design" principles, embedding ethical considerations directly into the implementation process rather than treating them as post-hoc additions. The article concludes by identifying key research directions and practical guidelines for advancing the field of responsible prompt engineering.


aiXamine: Simplified LLM Safety and Security

arXiv.org Artificial Intelligence

Evaluating Large Language Models (LLMs) for safety and security remains a complex task, often requiring users to navigate a fragmented landscape of ad hoc benchmarks, datasets, metrics, and reporting formats. To address this challenge, we present aiXamine, a comprehensive black-box evaluation platform for LLM safety and security. aiXamine integrates over 40 tests (i.e., benchmarks) organized into eight key services targeting specific dimensions of safety and security: adversarial robustness, code security, fairness and bias, hallucination, model and data privacy, out-of-distribution (OOD) robustness, over-refusal, and safety alignment. The platform aggregates the evaluation results into a single detailed report per model, providing a detailed breakdown of model performance, test examples, and rich visualizations. We used aiXamine to assess over 50 publicly available and proprietary LLMs, conducting over 2K examinations. Our findings reveal notable vulnerabilities in leading models, including susceptibility to adversarial attacks in OpenAI's GPT-4o, biased outputs in xAI's Grok-3, and privacy weaknesses in Google's Gemini 2.0. Additionally, we observe that open-source models can match or exceed proprietary models in specific services such as safety alignment, fairness and bias, and OOD robustness. Finally, we identify trade-offs between distillation strategies, model size, training methods, and architectural choices.


ChatGPT maker OpenAI wants to buy Chrome from Google

PCWorld

Google is having a bit of a moment. It's not quite an Enron- or FTX-style "abandon ship" situation, but between two separate US antitrust rulings on its core search and advertising businesses, it's a five-alarm fire. One of the possible outcomes is Google selling off the Chrome browserโ€ฆ and it looks like one possible buyer is OpenAI, maker of ChatGPT. OpenAI's head of product for ChatGPT is named Nick Turley, and he testified at the remedy phase of the Department of Justice's successful monopoly suit against Google. When asked if OpenAI would be interested in buying the Chrome browser from Google, Turley didn't mince words.


WhatsApp defends 'optional' AI tool that cannot be turned off

BBC News

When you first use Meta AI in WhatsApp, it states the chatbot "can only read messages people share with it". "Meta can't read any other messages in your personal chats, as your personal messages remain end to end encrypted," it says. Meanwhile the Information Commissioner's Office told the BBC it would "continue to monitor the adoption of Meta AI's technology and use of personal data within WhatsApp". "Personal information fuels much of AI innovation so people need to trust that organisations are using their information responsibly," it said. "Organisations who want to use people's personal details to train or use generative AI models need to comply with all their data protection obligations, and take the necessary extra steps when it comes to processing the data of children."


ChatGPT-maker wants to buy Google Chrome

BBC News

The current trial is looking at remedies to curtail Google's dominance in online search, as the recent explosion in generative AI services such as ChatGPT has expanded the market. Newer AI models search the internet to improve results and reduce hallucination, which has been a problem from developers since chatbots started to become popular. Last year, OpenAI offered to do a deal with Google which would have integrated Google search results into ChatGPT, according to Mr Turley's testimony. But he says their offer was rejected. "We have no partnership with Google today," Mr Turley said, according to Reuters. OpenAI does however have a partnership with Microsoft, which makes the Bing search engine and Edge browser.


3 Things Caiwei Chen is into right now

MIT Technology Review

I recently saw Doomers, a new play by Matthew Gasda about the aborted 2023 coup at OpenAI, here represented by a fictional company called MindMesh. The action is set almost entirely in a meeting room; the first act follows executives immediately after the firing of company CEO Seth (a stand-in for Sam Altman), and the second re-creates the board negotiations that determined his fate. It's a solid attempt to capture the zeitgeist of Silicon Valley's AI frenzy and the world's moral panic over artificial intelligence, but the rapid-fire, high-stakes exchanges mean it sometimes seems to get lost in its own verbosity. The vastness of Chinese cuisine defies easy categorization, and even in a city with no shortage of options, I often find myself cooking--not just to recapture something closer to home, but to create a home unlike one that ever existed. Recently, I've been experimenting with a Chinese take on the charcuterie board--pairing toasted steamed buns, called mantou, with furu, a fermented tofu spread that is sharp, pungent, and full of umami. I started sewing three years ago, but only in the past year have I begun making clothes from scratch.


AI floods Amazon with strange political books before Canadian election

The Japan Times

Canada has seen a boom in political books created with generative artificial intelligence, adding to concerns about how new technologies are affecting the information voters receive during the election campaign. Canadian Prime Minister Mark Carney was the subject of at least 16 books published in March and listed on Amazon, according to a review of the site on April 16. Five of those were published on a single day. In total, some 30 titles were published about Carney this year and made available on Amazon -- but most were taken down from the site after inquiries were made.


Benchmarking LLM for Code Smells Detection: OpenAI GPT-4.0 vs DeepSeek-V3

arXiv.org Artificial Intelligence

Determining the most effective Large Language Model for code smell detection presents a complex challenge. This study introduces a structured methodology and evaluation matrix to tackle this issue, leveraging a curated dataset of code samples consistently annotated with known smells. The dataset spans four prominent programming languages Java, Python, JavaScript, and C++; allowing for cross language comparison. We benchmark two state of the art LLMs, OpenAI GPT 4.0 and DeepSeek-V3, using precision, recall, and F1 score as evaluation metrics. Our analysis covers three levels of detail: overall performance, category level performance, and individual code smell type performance. Additionally, we explore cost effectiveness by comparing the token based detection approach of GPT 4.0 with the pattern-matching techniques employed by DeepSeek V3. The study also includes a cost analysis relative to traditional static analysis tools such as SonarQube. The findings offer valuable guidance for practitioners in selecting an efficient, cost effective solution for automated code smell detection