Generative AI
Google is reportedly paying publishers thousands of dollars to use its AI to write stories
Google has been quietly striking deals with some publishers to use new generative AI tools to publish stories, according to a report in Adweek. The deals, reportedly worth tens of thousands of dollars a year, are apparently part of the Google News Initiative (GNI), a six-year-old program that funds media literacy projects, fact-checking tools, and other resources for newsrooms. But the move into generative AI publishing tools would be a new, and likely controversial, step for the company. According to Adweek, the program is currently targeting a "handful" of smaller publishers. "The beta tools let under-resourced publishers create aggregated content more efficiently by indexing recently published reports generated by other organizations, like government agencies and neighboring news outlets, and then summarizing and publishing them as a new article," Adweek reports.
Tumblr and WordPress posts will reportedly be used for OpenAI and Midjourney training
Tumblr and WordPress are reportedly set to strike deals to sell user data to artificial intelligence companies OpenAI and Midjourney. It isn't clear which data will be included, but the report suggests Automattic may have overreached initially. An alleged internal post from Tumblr product manager Cyle Gage suggests Automattic prepared to send private or partner-related data that wasn't supposed to be included in the deal. The questionable content reportedly included private posts on public blog posts, deleted or suspended blogs, unanswered (therefore, not publicly posted) questions, private answers, posts marked explicit and content from premium partner blogs (like Apple's former music site). The internal post suggests Automattic's engineers are preparing a list of post IDs that should have been excluded.
OpenAI claims New York Times 'hacked' ChatGPT to build copyright lawsuit
OpenAI said in a filing in Manhattan federal court on Monday that the Times caused the technology to reproduce its material through "deceptive prompts that blatantly violate OpenAI's terms of use". "The allegations in the Times's complaint do not meet its famously rigorous journalistic standards," OpenAI said. "The truth, which will come out in the course of this case, is that the Times paid someone to hack OpenAI's products." OpenAI did not name the "hired gun" whom it said the Times used to manipulate its systems and did not accuse the newspaper of breaking any anti-hacking laws. Representatives for the New York Times and OpenAI did not immediately respond to requests for comment on the filing. The Times sued OpenAI and its largest financial backer, Microsoft, in December, accusing them of using millions of its articles without permission to train chatbots to provide information to users.
Alibaba Leads Record Deal to Mint 2.5 Billion China AI Firm
Alibaba Group Holding Ltd. led the largest single financing round for a Chinese artificial intelligence startup, the latest in a string of sizable investments that suggest the e-commerce firm is again deploying capital in the hunt for growth. Alibaba joins Tencent Holdings Ltd. and Silicon Valley peers like Microsoft Corp. in placing big bets on generative AI, the technology that powers ChatGPT. It led a 1 billion funding round in Moonshot AI with existing backer Monolith Management, boosting the year-old firm's valuation eight-fold to some 2.5 billion, people familiar with the deal said. They joined previous backers including food delivery giant Meituan's investment arm Long-Z and Hongshan, formerly Sequoia China, the people said, asking not to be identified discussing a private transaction. Founded in March 2023, Moonshot AI is among the better-known startups developing generative artificial intelligence in China, hoping to eventually match the likes of OpenAI and Google.
E.U. Watchdog to Examine Microsoft's Mistral AI Investment
Microsoft Corp.'s Mistral AI investment is set to be analyzed by the European Union's competition watchdog at the same time that its deep ties to OpenAI Inc come under regulatory scrutiny. Mistral announced a "strategic partnership" with Microsoft on Monday that includes making the startup's latest artificial intelligence models available to customers of Microsoft's Azure cloud. Microsoft said the investment amounted to 15 million ( 16.3 million.) Mistral develops algorithmic models similar to those from OpenAI used for chatbots and other AI services, but Mistral models are shared openly. Microsoft's investments will convert into equity as part of Mistral's next funding round.
SoftBank, Nvidia and Microsoft team up to use AI in mobile base stations
SoftBank, Nvidia, Microsoft and others said Monday that they have formed an alliance aimed at effectively using mobile base stations with the help of artificial intelligence. The members of the AI-Ran Alliance aim to work together in preventing communications congestion and promoting the use of smartphone apps using generative AI. The initiative was unveiled at the Mobile World Congress, an international trade fair for the telecommunications industry, in Spain. The group will apply AI technology so that data processing can be performed at mobile base stations rather than in the cloud, to help save power and eliminate communication delays. The alliance "has been formed with the vision to spearhead the advancement of society through AI innovations, particularly from the telecom industry," SoftBank President and CEO Junichi Miyakawa said in a statement.
On the Societal Impact of Open Foundation Models
Kapoor, Sayash, Bommasani, Rishi, Klyman, Kevin, Longpre, Shayne, Ramaswami, Ashwin, Cihon, Peter, Hopkins, Aspen, Bankston, Kevin, Biderman, Stella, Bogen, Miranda, Chowdhury, Rumman, Engler, Alex, Henderson, Peter, Jernite, Yacine, Lazar, Seth, Maffulli, Stefano, Nelson, Alondra, Pineau, Joelle, Skowron, Aviya, Song, Dawn, Storchan, Victor, Zhang, Daniel, Ho, Daniel E., Liang, Percy, Narayanan, Arvind
Foundation models are powerful technologies: how they are released publicly directly shapes their societal impact. In this position paper, we focus on open foundation models, defined here as those with broadly available model weights (e.g. Llama 2, Stable Diffusion XL). We identify five distinctive properties (e.g. greater customizability, poor monitoring) of open foundation models that lead to both their benefits and risks. Open foundation models present significant benefits, with some caveats, that span innovation, competition, the distribution of decision-making power, and transparency. To understand their risks of misuse, we design a risk assessment framework for analyzing their marginal risk. Across several misuse vectors (e.g. cyberattacks, bioweapons), we find that current research is insufficient to effectively characterize the marginal risk of open foundation models relative to pre-existing technologies. The framework helps explain why the marginal risk is low in some cases, clarifies disagreements about misuse risks by revealing that past work has focused on different subsets of the framework with different assumptions, and articulates a way forward for more constructive debate. Overall, our work helps support a more grounded assessment of the societal impact of open foundation models by outlining what research is needed to empirically validate their theoretical benefits and risks.
Generative AI and Copyright: A Dynamic Perspective
Yang, S. Alex, Zhang, Angela Huyue
The rapid advancement of generative AI is poised to disrupt the creative industry. Amidst the immense excitement for this new technology, its future development and applications in the creative industry hinge crucially upon two copyright issues: 1) the compensation to creators whose content has been used to train generative AI models (the fair use standard); and 2) the eligibility of AI-generated content for copyright protection (AI-copyrightability). While both issues have ignited heated debates among academics and practitioners, most analysis has focused on their challenges posed to existing copyright doctrines. In this paper, we aim to better understand the economic implications of these two regulatory issues and their interactions. By constructing a dynamic model with endogenous content creation and AI model development, we unravel the impacts of the fair use standard and AI-copyrightability on AI development, AI company profit, creators income, and consumer welfare, and how these impacts are influenced by various economic and operational factors. For example, while generous fair use (use data for AI training without compensating the creator) benefits all parties when abundant training data exists, it can hurt creators and consumers when such data is scarce. Similarly, stronger AI-copyrightability (AI content enjoys more copyright protection) could hinder AI development and reduce social welfare. Our analysis also highlights the complex interplay between these two copyright issues. For instance, when existing training data is scarce, generous fair use may be preferred only when AI-copyrightability is weak. Our findings underscore the need for policymakers to embrace a dynamic, context-specific approach in making regulatory decisions and provide insights for business leaders navigating the complexities of the global regulatory environment.
Large Scale Generative AI Text Applied to Sports and Music
Baughman, Aaron, Hammer, Stephen, Agarwal, Rahul, Akay, Gozde, Morales, Eduardo, Johnson, Tony, Karlinsky, Leonid, Feris, Rogerio
We address the problem of scaling up the production of media content, including commentary and personalized news stories, for large-scale sports and music events worldwide. Our approach relies on generative AI models to transform a large volume of multimodal data (e.g., videos, articles, real-time scoring feeds, statistics, and fact sheets) into coherent and fluent text. Based on this approach, we introduce, for the first time, an AI commentary system, which was deployed to produce automated narrations for highlight packages at the 2023 US Open, Wimbledon, and Masters tournaments. In the same vein, our solution was extended to create personalized content for ESPN Fantasy Football and stories about music artists for the Grammy awards. These applications were built using a common software architecture achieved a 15x speed improvement with an average Rouge-L of 82.00 and perplexity of 6.6. Our work was successfully deployed at the aforementioned events, supporting 90 million fans around the world with 8 billion page views, continuously pushing the bounds on what is possible at the intersection of sports, entertainment, and AI.
The Deeper Problem With Google's Racially Diverse Nazis
Is there a right way for Google's generative AI to create fake images of Nazis? Gemini, Google's answer to ChatGPT, was shown last week to generate an absurd range of racially and gender-diverse German soldiers styled in Wehrmacht garb. It was, understandably, ridiculed for not generating any images of Nazis who were actually white. Prodded further, it seemed to actively resist generating images of white people altogether. The company ultimately apologized for "inaccuracies in some historical image generation depictions" and paused Gemini's ability to generate images featuring people.