Goto

Collaborating Authors

 Media


Shinobi is the latest video game to get the big screen treatment

Engadget

Back in the old days, there was no sure-fire indicator of box office poison more than a video game adaptation. That has changed in recent years and now all kinds of gaming mascots are getting their chance to appear in a major motion picture or, at the very least, a streaming series. They're now making a movie based on Shinobi, as reported by Deadline. For the uninitiated, Shinobi is a famous hack-and-slash game developed by Sega in which you play as a ninja. There have been plenty of sequels throughout the years, though they mostly share the same basic story.


Low-Rank Constraints for Fast Inference in Structured Models

Neural Information Processing Systems

Structured distributions, i.e. distributions over combinatorial spaces, are commonly used to learn latent probabilistic representations from observed data. However, scaling these models is bottlenecked by the high computational and memory complexity with respect to the size of the latent representations. Common models such as Hidden Markov Models (HMMs) and Probabilistic Context-Free Grammars (PCFGs) require time and space quadratic and cubic in the number of hidden states respectively. This work demonstrates a simple approach to reduce the computational and memory complexity of a large class of structured models. We show that by viewing the central inference step as a matrix-vector product and using a low-rank constraint, we can trade off model expressivity and speed via the rank. Experiments with neural parameterized structured models for language modeling, polyphonic music modeling, unsupervised grammar induction, and video modeling show that our approach matches the accuracy of standard models at large state spaces while providing practical speedups.


298 Best Prime Day Deals, Vetted By Our Amazon Experts (Oct 2024)

WIRED

Amazon's fall Prime Day sale--also known as Big Deals Days--ends tonight. It's October, yes, but it's never too early to jump on that holiday gift shopping. We've combed through the deals and found the best ones, based on our years of testing and reviewing. WIRED's picks for the best Prime Day deals only include products someone from our team has personally tested and reviewed. We track prices using several tools to avoid falling for fake discounts. There are no shoddy knockoffs or overpriced products among our recommendations, just good deals on good stuff. We've linked our reviews and buying guide throughout to help you make fully informed buying decisions. We test products year-round and handpicked these Prime Day deals. We'll update this guide regularly throughout Prime Day by adding fresh deals and removing dead deals. This is our favorite e-reader. You'll have the choice between the base Paperwhite and the Signature Edition (8/10, WIRED Recommends), which comes with 16 gigabytes ...


em The Wild Robot /em Wants You To Cry, Really

Slate

On this week's show, Dana and Stephen are joined by Supreme Friend of the Podcast (SFOP) Isaac Butler, author of The Method: How the Twentieth Century Learned to Act. The trio first explores The Wild Robot, DreamWork Animation's handcrafted, lovingly made film that's the surprise of the year. Lupita Nyong'o voices ROZ, an old-fashioned robot powered by supremely advanced A.I. who must learn about and adapt to her new wild surroundings. Then, they dissect Nobody Wants This, a new Netflix series starring Kristen Bell (who plays a sex podcaster) and Adam Brody as a hot rabbi. Although there are obvious charms, the show's "will they, won't they" rom-com beats can often feel, at best, gratingly familiar, and at worst, bizarre and unthoughtful, particularly in its portrayal of Jewish women.


The best projector for 2024

Engadget

If you're looking to upgrade your entertainment setup, finding the best projector could be the perfect solution. Whether you're into binge-watching shows, hosting outdoor movie nights or even leveling up your gaming experience, modern projectors can help you do it all. Some are fantastic for creating that full home-theater vibe, while others are so good they could even replace your TV, offering huge screen sizes, sharp image quality and built-in smart features. Many projectors are portable enough to take outside, making them great for BBQs, yard parties, or just enjoying a cozy movie night under the stars. Some are even designed for easy room-to-room transport, meaning you can switch up your viewing experience wherever you are. If you're thinking of stepping up your viewing game, we've tested some of the best projectors out there to help you find the right one for your needs. As mentioned, ultra-short-throw models have rapidly established themselves in the market due to the extra performance and convenience, and all manufacturers sell at least a couple of models. Within the ultra-short-throw category, We'll compare two price categories: under 7,000 and 3,500, with three projectors each.


Sega's ninja game Shinobi to get the movie treatment

The Japan Times

One of Sega's most popular games, Shinobi, will be made into a movie in a joint project with Universal Pictures, the Japanese gamemaker announced Wednesday, aiming to emulate the success of "The Super Mario Bros. Movie." Sega did not give a target date for the release but said it had "started the development of a film production" with the Hollywood behemoth. Shinobi was originally created for Japanese arcades in 1987 and features a ninja character who fights to stop a criminal organization that kidnaps child ninjas. It is the latest effort to cash in on a video-game adaptation craze after "The Super Mario Bros. Movie" became the second-highest grossing film of 2023, following a 2020 adaptation of Sega's "Sonic the Hedgehog." "Shinobi is one of Sega's most popular series worldwide, along with Sonic the Hedgehog," Sega said on Wednesday.


Parameter-Efficient Fine-Tuning via Selective Discrete Cosine Transform

arXiv.org Artificial Intelligence

In the era of large language models, parameter-efficient fine-tuning (PEFT) has been extensively studied. However, these approaches usually rely on the space domain, which encounters storage challenges especially when handling extensive adaptations or larger models. The frequency domain, in contrast, is more effective in compressing trainable parameters while maintaining the expressive capability. In this paper, we propose a novel Selective Discrete Cosine Transformation (sDCTFT) fine-tuning scheme to push this frontier. Its general idea is to exploit the superior energy compaction and decorrelation properties of DCT to improve both model efficiency and accuracy. Specifically, it projects the weight change from the low-rank adaptation into the discrete cosine space. Then, the weight change is partitioned over different levels of the discrete cosine spectrum, and the most critical frequency components in each partition are selected. Extensive experiments on four benchmark datasets demonstrate the superior accuracy, reduced computational cost, and lower storage requirements of the proposed method over the prior arts. For instance, when performing instruction tuning on the LLaMA3.1-8B model, sDCTFT outperforms LoRA with just 0.05M trainable parameters compared to LoRA's 38.2M, and surpasses FourierFT with 30\% less trainable parameters. The source code will be publicly available.


Fine-tuning can Help Detect Pretraining Data from Large Language Models

arXiv.org Artificial Intelligence

In the era of large language models (LLMs), detecting pretraining data has been increasingly important due to concerns about fair evaluation and ethical risks. Current methods differentiate members and non-members by designing scoring functions, like Perplexity and Min-k%. In this paper, we first explore the benefits of unseen data, which can be easily collected after the release of the LLM. We find that the perplexities of LLMs perform differently for members and non-members, after fine-tuning with a small amount of previously unseen data. In light of this, we introduce a novel and effective method termed Fine-tuned Score Deviation (FSD), which improves the performance of current scoring functions for pretraining data detection. In particular, we propose to measure the deviation distance of current scores after fine-tuning on a small amount of unseen data within the same domain. In effect, using a few unseen data can largely decrease the scores of all non-members, leading to a larger deviation distance than members. Extensive experiments demonstrate the effectiveness of our method, significantly improving the AUC score on common benchmark datasets across various models. The impressive performance of large language models (LLMs) arises from large-scale pretraining on massive datasets collected from the internet (Achiam et al., 2023; Touvron et al., 2023b). But, model developers are often reluctant to disclose detailed information about the pretraining datasets, raising significant concerns regarding fair evaluation and ethical risks. Specifically, Recent studies reveal that the pretraining corpus may inadvertently include data from evaluation benchmarks (Sainz et al., 2023; Balloccu et al., 2024), making it difficult to assess the practical capability of LLMs. Considering the vast size of the pretraining dataset and the single iteration of pretraining, it has been increasingly important and challenging to detect pretraining data, which determines whether a piece of text is part of the pretraining dataset.


MKGL: Mastery of a Three-Word Language

arXiv.org Artificial Intelligence

Large language models (LLMs) have significantly advanced performance across a spectrum of natural language processing (NLP) tasks. Yet, their application to knowledge graphs (KGs), which describe facts in the form of triplets and allow minimal hallucinations, remains an underexplored frontier. In this paper, we investigate the integration of LLMs with KGs by introducing a specialized KG Language (KGL), where a sentence precisely consists of an entity noun, a relation verb, and ends with another entity noun. Despite KGL's unfamiliar vocabulary to the LLM, we facilitate its learning through a tailored dictionary and illustrative sentences, and enhance context understanding via real-time KG context retrieval and KGL token embedding augmentation. Our results reveal that LLMs can achieve fluency in KGL, drastically reducing errors compared to conventional KG embedding methods on KG completion. Furthermore, our enhanced LLM shows exceptional competence in generating accurate three-word sentences from an initial entity and interpreting new unseen terms out of KGs.


PublicHearingBR: A Brazilian Portuguese Dataset of Public Hearing Transcripts for Summarization of Long Documents

arXiv.org Artificial Intelligence

This paper introduces PublicHearingBR, a Brazilian Portuguese dataset designed for summarizing long documents. The dataset consists of transcripts of public hearings held by the Brazilian Chamber of Deputies, paired with news articles and structured summaries containing the individuals participating in the hearing and their statements or opinions. The dataset supports the development and evaluation of long document summarization systems in Portuguese. Our contributions include the dataset, a hybrid summarization system to establish a baseline for future studies, and a discussion on evaluation metrics for summarization involving large language models, addressing the challenge of hallucination in the generated summaries. As a result of this discussion, the dataset also provides annotated data that can be used in Natural Language Inference tasks in Portuguese.