Goto

Collaborating Authors

 Generative AI


Dedicated Feedback and Edit Models Empower Inference-Time Scaling for Open-Ended General-Domain Tasks

arXiv.org Artificial Intelligence

Inference-Time Scaling has been critical to the success of recent models such as OpenAI o1 and DeepSeek R1. However, many techniques used to train models for inference-time scaling require tasks to have answers that can be verified, limiting their application to domains such as math, coding and logical reasoning. We take inspiration from how humans make first attempts, ask for detailed feedback from others and make improvements based on such feedback across a wide spectrum of open-ended endeavors. To this end, we collect data for and train dedicated Feedback and Edit Models that are capable of performing inference-time scaling for open-ended general-domain tasks. In our setup, one model generates an initial response, which are given feedback by a second model, that are then used by a third model to edit the response. We show that performance on Arena Hard, a benchmark strongly predictive of Chatbot Arena Elo can be boosted by scaling the number of initial response drafts, effective feedback and edited responses. When scaled optimally, our setup based on 70B models from the Llama 3 family can reach SoTA performance on Arena Hard at 92.7 as of 5 Mar 2025, surpassing OpenAI o1-preview-2024-09-12 with 90.4 and DeepSeek R1 with 92.3.


Prompt Programming: A Platform for Dialogue-based Computational Problem Solving with Generative AI Models

arXiv.org Artificial Intelligence

Computing students increasingly rely on generative AI tools for programming assistance, often without formal instruction or guidance. This highlights a need to teach students how to effectively interact with AI models, particularly through natural language prompts, to generate and critically evaluate code for solving computational tasks. To address this, we developed a novel platform for prompt programming that enables authentic dialogue-based interactions, supports problems involving multiple interdependent functions, and offers on-request execution of generated code. Data analysis from over 900 students in an introductory programming course revealed high engagement, with the majority of prompts occurring within multi-turn dialogues. Problems with multiple interdependent functions encouraged iterative refinement, with progression graphs highlighting several common strategies. Students were highly selective about the code they chose to test, suggesting that on-request execution of generated code promoted critical thinking. Given the growing importance of learning dialogue-based programming with AI, we provide this tool as a publicly accessible resource, accompanied by a corpus of programming problems for educational use.


Approaching the Limits to EFL Writing Enhancement with AI-generated Text and Diverse Learners

arXiv.org Artificial Intelligence

Generative artificial intelligence (AI) chatbots, such as ChatGPT, are reshaping how English as a foreign language (EFL) students write since students can compose texts by integrating their own words with AI-generated text. This study investigated how 59 Hong Kong secondary school students with varying levels of academic achievement interacted with AI-generated text to compose a feature article, exploring whether any interaction patterns benefited the overall quality of the article. Through content analysis, multiple linear regression and cluster analysis, we found the overall number of words -- whether AI- or human-generated -- is the main predictor of writing quality. However, the impact varies by students' competence to write independently, for instance, by using their own words accurately and coherently to compose a text, and to follow specific interaction patterns with AI-generated text. Therefore, although composing texts with human words and AI-generated text may become prevalent in EFL writing classrooms, without educators' careful attention to EFL writing pedagogy and AI literacy, high-achieving students stand to benefit more from using AI-generated text than low-achieving students.


Chatbots Are Cheating on Their Benchmark Tests

The Atlantic - Technology

Generative-AI companies have been selling a narrative of unprecedented, endless progress. Just last week, OpenAI introduced GPT-4.5 as its "largest and best model for chat yet." Earlier in February, Google called its latest version of Gemini "the world's best AI model." And in January, the Chinese company DeekSeek touted its R1 model as being just as powerful as OpenAI's o1 model--which Sam Altman had called "the smartest model in the world" the previous month. Yet there is growing evidence that progress is slowing down and that the LLM-powered chatbot may already be near its peak.


UK competition watchdog drops Microsoft-OpenAI probe

BBC News

Critics though say the decision is linked to the changed political environment the CMA is now operating in. The government has instructed the country's regulators to suggest ways of stimulating economic growth. In January, the government removed the then chair of the CMA, Marcus Bokkerink, because it was unhappy with his response to that call. He was replaced on an interim basis by Doug Gurr, former boss of Amazon UK. "The CMA has sat on this decision for over a year, yet within just a few weeks of a former Amazon boss being installed as chair, it has decided everything was absolutely fine all along, nothing to see here," said Foxglove co-executive director Rosa Curling. "This is a bad sign that Big Tech has successfully convinced the prime minister to defang our competition regulator and let Big Tech gobble up the current generation of cutting-edge tech โ€“ just like they did the last one," she told the BBC.


Fox News AI Newsletter: Judge denies Musk's move against OpenAI

FOX News

Gladstone A.I. co-founders and CEOs Edouard Harris and Jeremie Harris explain the major role that A.I will play in national security and warfare on'The Will Cain Show.' Elon Musk met with members of the Senate DOGE caucus at the White House. MUSK'S MOVE BLOCKED: A California judge denied Elon Musk's move to halt OpenAI's efforts to convert it into a for-profit entity, saying in a ruling that the SpaceX and Tesla CEO hadn't met "the high burden required for a preliminary injunction." 'DOWNFALLS' OF AI: A federal judge has declined to impose sanctions on an attorney who submitted a brief that contained incorrect case citations and quotes generated by artificial intelligence. DEFEND YOUR DATA: Windows has always been a favorite target for hackers, but it seems they have now figured out how to actively target Macs as well. We've seen an alarming rise in malware affecting Mac computers, stealing personal data and cryptocurrency.


UK watchdog drops competition review of Microsoft's OpenAI partnership

The Guardian

The UK's competition watchdog will not hold a formal investigation into Microsoft's partnership with the startup behind the artificial intelligence chatbot ChatGPT, stating that while the 2.9tn ( 2.3tn) tech company has "material influence" over OpenAI it does not control it. The Competition and Markets Authority (CMA) said Microsoft, OpenAI's biggest financial backer with a 13bn investment, acquired material influence over the San Francisco-based business in 2019 but did not exercise de facto control over it โ€“ and therefore did not meet the threshold for an official inquiry. The decision follows expressions of disquiet over the appointment of the former boss of Amazon UK, Doug Gurr, as the CMA's interim chair. The organisation's chief executive, Sarah Cardell, has also said the CMA does not want to create a "chilling effect" on business confidence, amid pressure from the UK government on regulators to produce pro-growth proposals. The CMA's executive director for mergers, Joel Bamford, said: "We have found that there has not been a change of control by Microsoft from material influence to de facto control over OpenAI. Because this change of control has not happened, the partnership in its current form does not qualify for review under the UK's merger control regime."


Judge denies Musk's initial bid to halt OpenAI's for-profit shift but sets trial for fall

The Guardian

A US judge on Tuesday denied Elon Musk's request for a preliminary injunction to pause OpenAI's transition to a for-profit model but agreed to hear a trial in the fall of this year, the latest turn in the high-stakes legal fight. The tech billionaire does not have "the high burden required for a preliminary injunction" to block the conversion of OpenAI, said Yvonne Gonzalez Rogers, a US district judge in Oakland, California. But Rogers wrote in the order that she wanted to resolve the lawsuit quickly given "the public interest at stake and potential for harm if a conversion contrary to law occurred". Musk and OpenAI, which he co-founded as a non-profit in 2015 but left before it took off, have been embroiled in a yearlong legal battle. The CEO of Tesla and X, formerly Twitter, accuses OpenAI of straying from its founding mission to develop artificial intelligence for the good of humanity, not corporate profit.


Court denies Elon Musk's attempt to block OpenAI's for-profit transformation

Engadget

US federal judge Yvonne Gonzalez Rogers has denied Elon Musk's request for an injunction that would have immediately stopped OpenAI's conversion into a for-profit entity. Musk filed for an injunction late last year after suing OpenAI and Microsoft and accusing them of telling investors not to fund rival AI companies, such as his own xAI. According to the Financial Times, the judge dismissed his request based on that claim of anticompetitive behavior. Gonzalez Rogers cited a previous statement by OpenAI CEO Sam Altman, saying that the company only warned certain investors who were granted access to sensitive information that their rights would be terminated if they made a non-passive investment in rival companies. The judge also reportedly rejected the request based on Musk's claim that OpenAI and Altman broke their contract with him and violated the company's founding mission of building AI "for the benefit of humanity."


A Generative Approach to High Fidelity 3D Reconstruction from Text Data

arXiv.org Artificial Intelligence

The convergence of generative artificial intelligence and advanced computer vision technologies introduces a groundbreaking approach to transforming textual descriptions into three-dimensional representations. This research proposes a fully automated pipeline that seamlessly integrates text-to-image generation, various image processing techniques, and deep learning methods for reflection removal and 3D reconstruction. By leveraging state-of-the-art generative models like Stable Diffusion, the methodology translates natural language inputs into detailed 3D models through a multi-stage workflow. The reconstruction process begins with the generation of high-quality images from textual prompts, followed by enhancement by a reinforcement learning agent and reflection removal using the Stable Delight model. Advanced image upscaling and background removal techniques are then applied to further enhance visual fidelity. These refined two-dimensional representations are subsequently transformed into volumetric 3D models using sophisticated machine learning algorithms, capturing intricate spatial relationships and geometric characteristics. This process achieves a highly structured and detailed output, ensuring that the final 3D models reflect both semantic accuracy and geometric precision. This approach addresses key challenges in generative reconstruction, such as maintaining semantic coherence, managing geometric complexity, and preserving detailed visual information. Comprehensive experimental evaluations will assess reconstruction quality, semantic accuracy, and geometric fidelity across diverse domains and varying levels of complexity. By demonstrating the potential of AI-driven 3D reconstruction techniques, this research offers significant implications for fields such as augmented reality (AR), virtual reality (VR), and digital content creation.