Goto

Collaborating Authors

 Media


The best films about AI – ranked!

The Guardian

Forget the more recent TV show, which ended up so frustratingly opaque as to render it pointless. The most fun version of Westworld is Michael Crichton's original movie. A robot cowboy comes to life and goes nuts in a theme park. What more could anyone need? Eleven years on, it's still hard to believe this film exists.


Reddit Is Already on the Rebound

WIRED

Social media researchers at the Network Contagion Research Institute in Princeton, New Jersey, got a rude awakening early last month. They were roused by 6:30 am phone calls from a colleague warning that Reddit had started blocking the institute's Pushshift service from updating its ongoing archive of every post on the discussion platform. That was a problem for more than just NCRI, because some of Reddit's 50,000 volunteer moderators depend on Pushshift to quickly investigate problem users, and many academics rely on the service. If it went stale, mods, as Reddit calls moderators, would have to work overtime or let more trash content accumulate. Researchers studying online communities would be forced to put projects and doctoral dissertations on ice.


Can ChatGPT discuss current events? Chatbot has clear knowledge cutoff date

FOX News

During an appearance on "The Ingraham Angle," Jimmy Failla shares his thoughts on the latest interesting development in the world of artificial intelligence. ChatGPT has been a game changer for artificial intelligence, catapulting earlier this year to the fastest-growing web platform ever as millions of people across the world rushed to communicate with a system that can mimic human conversation. The system, however, is unable to respond to current events questions due to having a knowledge cutoff date of September 2021. When Fox News Digital, for example, attempted to ask ChatGPT questions about current events, such as if the Titan submersible implosion could have been prevented or what charges Hunter Biden was hit with this month, the chatbot responded that it does not have knowledge of current events after September 2021. "As an AI language model, I have a knowledge cutoff date because my training data only goes up until September 2021," ChatGPT responded when asked why it does not possess language beyond September 2021.


Diff-Foley: Synchronized Video-to-Audio Synthesis with Latent Diffusion Models

arXiv.org Artificial Intelligence

The Video-to-Audio (V2A) model has recently gained attention for its practical application in generating audio directly from silent videos, particularly in video/film production. However, previous methods in V2A have limited generation quality in terms of temporal synchronization and audio-visual relevance. We present Diff-Foley, a synchronized Video-to-Audio synthesis method with a latent diffusion model (LDM) that generates high-quality audio with improved synchronization and audio-visual relevance. We adopt contrastive audio-visual pretraining (CAVP) to learn more temporally and semantically aligned features, then train an LDM with CAVP-aligned visual features on spectrogram latent space. The CAVP-aligned features enable LDM to capture the subtler audio-visual correlation via a cross-attention module. We further significantly improve sample quality with `double guidance'. Diff-Foley achieves state-of-the-art V2A performance on current large scale V2A dataset. Furthermore, we demonstrate Diff-Foley practical applicability and generalization capabilities via downstream finetuning. Project Page: see https://diff-foley.github.io/


LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

arXiv.org Artificial Intelligence

Instruction tuning unlocks the superior capability of Large Language Models (LLM) to interact with humans. Furthermore, recent instruction-following datasets include images as visual inputs, collecting responses for image-based instructions. However, visual instruction-tuned models cannot comprehend textual details within images well. This work enhances the current visual instruction tuning pipeline with text-rich images (e.g., movie posters, book covers, etc.). Specifically, we first use publicly available OCR tools to collect results on 422K text-rich images from the LAION dataset. Moreover, we prompt text-only GPT-4 with recognized texts and image captions to generate 16K conversations, each containing question-answer pairs for text-rich images. By combining our collected data with previous multi-modal instruction-following data, our model, LLaVAR, substantially improves the LLaVA model's capability on text-based VQA datasets (up to 20% accuracy improvement) while achieving an accuracy of 91.42% on ScienceQA. The GPT-4-based instruction-following evaluation also demonstrates the improvement of our model on both natural images and text-rich images. Through qualitative analysis, LLaVAR shows promising interaction (e.g., reasoning, writing, and elaboration) skills with humans based on the latest real-world online content that combines text and images. We make our code/data/models publicly available at https://llavar.github.io/.


Predicting Music Hierarchies with a Graph-Based Neural Decoder

arXiv.org Artificial Intelligence

This paper describes a data-driven framework to parse musical sequences into dependency trees, which are hierarchical structures used in music cognition research and music analysis. The parsing involves two steps. First, the input sequence is passed through a transformer encoder to enrich it with contextual information. Then, a classifier filters the graph of all possible dependency arcs to produce the dependency tree. One major benefit of this system is that it can be easily integrated into modern deep-learning pipelines. Moreover, since it does not rely on any particular symbolic grammar, it can consider multiple musical features simultaneously, make use of sequential context information, and produce partial results for noisy inputs. We test our approach on two datasets of musical trees -- time-span trees of monophonic note sequences and harmonic trees of jazz chord sequences -- and show that our approach outperforms previous methods.


The Drunkard's Odometry: Estimating Camera Motion in Deforming Scenes

arXiv.org Artificial Intelligence

Estimating camera motion in deformable scenes poses a complex and open research challenge. Most existing non-rigid structure from motion techniques assume to observe also static scene parts besides deforming scene parts in order to establish an anchoring reference. However, this assumption does not hold true in certain relevant application cases such as endoscopies. Deformable odometry and SLAM pipelines, which tackle the most challenging scenario of exploratory trajectories, suffer from a lack of robustness and proper quantitative evaluation methodologies. To tackle this issue with a common benchmark, we introduce the Drunkard's Dataset, a challenging collection of synthetic data targeting visual navigation and reconstruction in deformable environments. This dataset is the first large set of exploratory camera trajectories with ground truth inside 3D scenes where every surface exhibits non-rigid deformations over time. Simulations in realistic 3D buildings lets us obtain a vast amount of data and ground truth labels, including camera poses, RGB images and depth, optical flow and normal maps at high resolution and quality. We further present a novel deformable odometry method, dubbed the Drunkard's Odometry, which decomposes optical flow estimates into rigid-body camera motion and non-rigid scene deformations. In order to validate our data, our work contains an evaluation of several baselines as well as a novel tracking error metric which does not require ground truth data. Dataset and code: https://davidrecasens.github.io/TheDrunkard'sOdometry/


Bounded (O(1)) Regret Recommendation Learning via Synthetic Controls Oracle

arXiv.org Artificial Intelligence

In online exploration systems where users with fixed preferences repeatedly arrive, it has recently been shown that O(1), i.e., bounded regret, can be achieved when the system is modeled as a linear contextual bandit. This result may be of interest for recommender systems, where the popularity of their items is often short-lived, as the exploration itself may be completed quickly before potential long-run non-stationarities come into play. However, in practice, exact knowledge of the linear model is difficult to justify. Furthermore, potential existence of unobservable covariates, uneven user arrival rates, interpretation of the necessary rank condition, and users opting out of private data tracking all need to be addressed for practical recommender system applications. In this work, we conduct a theoretical study to address all these issues while still achieving bounded regret. Aside from proof techniques, the key differentiating assumption we make here is the presence of effective Synthetic Control Methods (SCM), which are shown to be a practical relaxation of the exact linear model knowledge assumption. We verify our theoretical bounded regret result using a minimal simulation experiment.


Researchers reconstruct 3D environments from eye reflections

Engadget

Researchers at the University of Maryland have turned eye reflections into (somewhat discernible) 3D scenes. The work builds on Neural Radiance Fields (NeRF), an AI technology that can reconstruct environments from 2D photos. Although the eye-reflection approach has a long way to go before it spawns any practical applications, the study (first reported by Tech Xplore) provides a fascinating glimpse into a technology that could eventually reveal an environment from a series of simple portrait photos. The team used subtle reflections of light captured in human eyes (using consecutive images shot from a single sensor) to try to discern the person's immediate environment. They began with several high-resolution images from a fixed camera position, capturing a moving individual looking toward the camera.


Windows 11 tips and tricks you didn't know you needed until now

FOX News

Windows 11 has a lot of features you may not know about. CyberGuy shows you how to customize your computer. We recently got a question from Wayne from Burgettstown, Pennsylvania. "I would like tips on using Windows 11. CLICK TO GET KURT'S FREE CYBERGUY NEWSLETTER WITH SECURITY ALERTS, QUICK TIPS, TECH REVIEWS AND EASY HOW-TO'S TO MAKE YOU SMARTER It's been about 20 months since Windows 11 was released, and its capabilities are pretty impressive.