Media
In the age-old good vs evil story, is Artificial Intelligence cinema's new villain?
Media mogul Barry Diller urged all parties to reach a resolution by September 1 amid ongoing Hollywood strikes during a Sunday interview on'Face the Nation.' Since the publication of the Bible, good vs. evil has long been a universal theme in literature - and in Hollywood storytelling. But could the same perceived evil force, also be good? As of May 2 of this year, 11,500 Hollywood screenwriters, represented by the Writers Guild of America (WGA) have been on strike over a three-pronged fight that boils down to money, autonomy, and Artificial Intelligence (AI). The writers are asking for increased and commensurate pay, for a guaranteed number of writers per room, and for regulated use of artificial intelligence in the writing process.
Hollywood strikers accuse NBCUniversal of blocking picket area
Hollywood's striking Writers Guild of America (WGA) and SAG-AFTRA actors' union have filed a grievance with the United States's National Labor Relations Board (NLRB) against Comcast's NBCUniversal, accusing the company of blocking a picket area. The unions said on Tuesday that NBCUniversal infringed its freedom to picket and endangered its members by obstructing a public sidewalk next to the company's studio lot in California with an ongoing construction project. The WGA's complaint said NBCUniversal "forced picketers to patrol in busy streets with significant car traffic where two picketers have already been struck by a car". SAG-AFTRA said members had been forced "to picket at the unsafe crowded location, exacerbating the dire public safety situation to interfere with striking members' right to engage in the protected, concerted activity of picketing and patrolling outside the employer's premises during a lawful strike". Hollywood actors joined film and television writers on picket lines for the first time in 63 years last week as they demanded higher streaming-era pay and curbs on the use of artificial intelligence.
GREG GUTFELD: People are tired of being talked down to about their beliefs
And what a great Tuesday it is. So SAG-AFTRA, the union for actors, claims that their profession is about as dead as a critic of Hillary Clinton. It all has to do with AI replacing real, live actors, which seems redundant of course, replacing Hollywood actors with artificial intelligence is like replacing Vin Diesel with Vin Diesel. But remember, they've done worse. They once replaced humans with "Real Housewives."
Instruction-following Evaluation through Verbalizer Manipulation
Li, Shiyang, Yan, Jun, Wang, Hai, Tang, Zheng, Ren, Xiang, Srinivasan, Vijay, Jin, Hongxia
While instruction-tuned models have shown remarkable success in various natural language processing tasks, accurately evaluating their ability to follow instructions remains challenging. Existing benchmarks primarily focus on common instructions that align well with what the model learned during training. However, proficiency in responding to these instructions does not necessarily imply strong ability in instruction following. In this paper, we propose a novel instruction-following evaluation protocol called verbalizer manipulation. It instructs the model to verbalize the task label with words aligning with model priors to different extents, adopting verbalizers from highly aligned (e.g., outputting "postive" for positive sentiment), to minimally aligned (e.g., outputting "negative" for positive sentiment). Verbalizer manipulation can be seamlessly integrated with any classification benchmark to examine the model's reliance on priors and its ability to override them to accurately follow the instructions. We conduct a comprehensive evaluation of four major model families across nine datasets, employing twelve sets of verbalizers for each of them. We observe that the instruction-following abilities of models, across different families and scales, are significantly distinguished by their performance on less natural verbalizers. Even the strongest GPT-4 model struggles to perform better than random guessing on the most challenging verbalizer, emphasizing the need for continued advancements to improve their instruction-following abilities. Large language models have achieved remarkable success in zero-shot generalization for various natural language processing (NLP) tasks via instruction tuning (Wei et al., 2022a; Ouyang et al., 2022; Sanh et al., 2022; Iyer et al., 2022). Existing benchmark datasets (Wang et al., 2018; 2019; Cobbe et al., 2021; Hendrycks et al., 2021; Li et al., 2023) primarily focus on common instructions that align well with what models learned during pre-training or instructiontuning.
Polyffusion: A Diffusion Model for Polyphonic Score Generation with Internal and External Controls
Min, Lejun, Jiang, Junyan, Xia, Gus, Zhao, Jingwei
ABSTRACT We propose Polyffusion, a diffusion model that generates polyphonic music scores by regarding music as imagelike piano roll representations. The model is capable of controllable music generation with two paradigms: internal control and external control. We show that by using tive modeling [14,15], symbolic music generation still suffers internal and external controls, Polyffusion unifies a from the lack of controllability and consistency at different wide range of music creation tasks, including melody generation time scales [16]. In our study, we experiment with given accompaniment, accompaniment generation the idea of using diffusion models to approach controllable given melody, arbitrary music segment inpainting, and music symbolic music generation. Experimental results Inspired by the high-quality and controllable image show that our model significantly outperforms existing generation that diffusion models have achieved in computer Transformer and sampling-based baselines, and using vision, we devise an image-like piano roll format as pre-trained disentangled representations as external conditions the input, and used a UNet-based diffusion model to stepwise yields more effective controls.
Challenges and Applications of Large Language Models
Kaddour, Jean, Harris, Joshua, Mozes, Maximilian, Bradley, Herbie, Raileanu, Roberta, McHardy, Robert
Large Language Models (LLMs) went from non-existent to ubiquitous in the machine learning discourse within a few years. Due to the fast pace of the field, it is difficult to identify the remaining challenges and already fruitful application areas. In this paper, we aim to establish a systematic set of open problems and application successes so that ML researchers can comprehend the field's current state more quickly and become productive.
Our Model Achieves Excellent Performance on MovieLens: What Does it Mean?
Fan, Yu-chen, Ji, Yitong, Zhang, Jie, Sun, Aixin
A typical benchmark dataset for recommender system (RecSys) evaluation consists of user-item interactions generated on a platform within a time period. The interaction generation mechanism partially explains why a user interacts with (e.g.,like, purchase, rate) an item, and the context of when a particular interaction happened. In this study, we conduct a meticulous analysis on the MovieLens dataset and explain the potential impact on using the dataset for evaluating recommendation algorithms. We make a few main findings from our analysis. First, there are significant differences in user interactions at the different stages when a user interacts with the MovieLens platform. The early interactions largely define the user portrait which affect the subsequent interactions. Second, user interactions are highly affected by the candidate movies that are recommended by the platform's internal recommendation algorithm(s). Removal of interactions that happen nearer to the last few interactions of a user leads to increasing difficulty in learning user preference, thus deteriorating recommendation accuracy. Third, changing the order of user interactions makes it more difficult for sequential algorithms to capture the progressive interaction process. Based on these findings, we further discuss the discrepancy between the interaction generation mechanism that is employed by the MovieLens system and that of typical real world recommendation scenarios. In summary, models that achieve excellent recommendation accuracy on the MovieLens dataset may not demonstrate superior performance in practice for at least two kinds of differences: (i) the differences in the contexts of user-item interaction generation, and (ii) the differences in user knowledge about the item collections.
From West to East: Who can understand the music of the others better?
Papaioannou, Charilaos, Benetos, Emmanouil, Potamianos, Alexandros
Recent developments in MIR have led to several benchmark deep learning models whose embeddings can be used for a variety of downstream tasks. At the same time, the vast majority of these models have been trained on Western pop/rock music and related styles. This leads to research questions on whether these models can be used to learn representations for different music cultures and styles, or whether we can build similar music audio embedding models trained on data from different cultures or styles. To that end, we leverage transfer learning methods to derive insights about the similarities between the different music cultures to which the data belongs to. We use two Western music datasets, two traditional/folk datasets coming from eastern Mediterranean cultures, and two datasets belonging to Indian art music. Three deep audio embedding models are trained and transferred across domains, including two CNN-based and a Transformer-based architecture, to perform auto-tagging for each target domain dataset. Experimental results show that competitive performance is achieved in all domains via transfer learning, while the best source dataset varies for each music culture. The implementation and the trained models are both provided in a public repository.
Text2Layer: Layered Image Generation using Latent Diffusion Model
Zhang, Xinyang, Zhao, Wentian, Lu, Xin, Chien, Jeff
Layer compositing is one of the most popular image editing workflows among both amateurs and professionals. Motivated by the success of diffusion models, we explore layer compositing from a layered image generation perspective. Instead of generating an image, we propose to generate background, foreground, layer mask, and the composed image simultaneously. To achieve layered image generation, we train an autoencoder that is able to reconstruct layered images and train diffusion models on the latent representation. One benefit of the proposed problem is to enable better compositing workflows in addition to the high-quality image output. Another benefit is producing higher-quality layer masks compared to masks produced by a separate step of image segmentation. Experimental results show that the proposed method is able to generate high-quality layered images and initiates a benchmark for future work.
Llama 2: Open Foundation and Fine-Tuned Chat Models
Touvron, Hugo, Martin, Louis, Stone, Kevin, Albert, Peter, Almahairi, Amjad, Babaei, Yasmine, Bashlykov, Nikolay, Batra, Soumya, Bhargava, Prajjwal, Bhosale, Shruti, Bikel, Dan, Blecher, Lukas, Ferrer, Cristian Canton, Chen, Moya, Cucurull, Guillem, Esiobu, David, Fernandes, Jude, Fu, Jeremy, Fu, Wenyin, Fuller, Brian, Gao, Cynthia, Goswami, Vedanuj, Goyal, Naman, Hartshorn, Anthony, Hosseini, Saghar, Hou, Rui, Inan, Hakan, Kardas, Marcin, Kerkez, Viktor, Khabsa, Madian, Kloumann, Isabel, Korenev, Artem, Koura, Punit Singh, Lachaux, Marie-Anne, Lavril, Thibaut, Lee, Jenya, Liskovich, Diana, Lu, Yinghai, Mao, Yuning, Martinet, Xavier, Mihaylov, Todor, Mishra, Pushkar, Molybog, Igor, Nie, Yixin, Poulton, Andrew, Reizenstein, Jeremy, Rungta, Rashi, Saladi, Kalyan, Schelten, Alan, Silva, Ruan, Smith, Eric Michael, Subramanian, Ranjan, Tan, Xiaoqing Ellen, Tang, Binh, Taylor, Ross, Williams, Adina, Kuan, Jian Xiang, Xu, Puxin, Yan, Zheng, Zarov, Iliyan, Zhang, Yuchen, Fan, Angela, Kambadur, Melanie, Narang, Sharan, Rodriguez, Aurelien, Stojnic, Robert, Edunov, Sergey, Scialom, Thomas
In this work, we develop and release Llama 2, a collection of pretrained and fine-tuned large language models (LLMs) ranging in scale from 7 billion to 70 billion parameters. Our fine-tuned LLMs, called Llama 2-Chat, are optimized for dialogue use cases. Our models outperform open-source chat models on most benchmarks we tested, and based on our human evaluations for helpfulness and safety, may be a suitable substitute for closed-source models. We provide a detailed description of our approach to fine-tuning and safety improvements of Llama 2-Chat in order to enable the community to build on our work and contribute to the responsible development of LLMs.