time and location
PuzzleGPT: Emulating Human Puzzle-Solving Ability for Time and Location Prediction
Ayyubi, Hammad, Feng, Xuande, Liu, Junzhang, Lin, Xudong, Wang, Zhecan, Chang, Shih-Fu
The task of predicting time and location from images is challenging and requires complex human-like puzzle-solving ability over different clues. In this work, we formalize this ability into core skills and implement them using different modules in an expert pipeline called PuzzleGPT. PuzzleGPT consists of a perceiver to identify visual clues, a reasoner to deduce prediction candidates, a combiner to combinatorially combine information from different clues, a web retriever to get external knowledge if the task can't be solved locally, and a noise filter for robustness. This results in a zero-shot, interpretable, and robust approach that records state-of-the-art performance on two datasets -- TARA and WikiTilo. PuzzleGPT outperforms large VLMs such as BLIP-2, InstructBLIP, LLaVA, and even GPT-4V, as well as automatically generated reasoning pipelines like VisProg, by at least 32% and 38%, respectively. It even rivals or surpasses finetuned models.
Formalization of Operational Domain and Operational Design Domain for Automated Vehicles
Specifying an Operational Design Domain (ODD) is crucial for safeguarding automated vehicle systems against conditions that exceed their capabilities. Yet, prior definitions of ODD have relied on ambiguous and unclear terms, resulting in numerous misunderstandings and misconceptions. This paper introduces a formal approach to clearly define the Operational Domain (OD) and ODD for automated vehicles. Furthermore, the absence of essential terms, such as the OD, has resulted in the creation of numerous terms that have made things more complicated and confusing. This level of complexity is unacceptable when it comes to developing safety-critical systems, where any uncertainty can lead to significant risks. This study addresses these deficiencies by providing a precise mathematical model of OD and clarifying its relationship with other terms. Also, by formalizing these terms, this work establishes a foundation for developing further concepts such as ODD specification and ODD monitoring, which are explained in this paper.
Can Vision-Language Models be a Good Guesser? Exploring VLMs for Times and Location Reasoning
Zhang, Gengyuan, Zhang, Yurui, Zhang, Kerui, Tresp, Volker
Vision-Language Models (VLMs) are expected to be capable of reasoning with commonsense knowledge as human beings. One example is that humans can reason where and when an image is taken based on their knowledge. This makes us wonder if, based on visual cues, Vision-Language Models that are pre-trained with large-scale image-text resources can achieve and even outperform human's capability in reasoning times and location. To address this question, we propose a two-stage \recognition\space and \reasoning\space probing task, applied to discriminative and generative VLMs to uncover whether VLMs can recognize times and location-relevant features and further reason about it. To facilitate the investigation, we introduce WikiTiLo, a well-curated image dataset compromising images with rich socio-cultural cues. In the extensive experimental studies, we find that although VLMs can effectively retain relevant features in visual encoders, they still fail to make perfect reasoning. We will release our dataset and codes to facilitate future studies.
Machine Learning
Please click on Timetables on the right hand side of this page for time and location of the practicals. Practicals will use Torch, a powerful programming framework for deep learning that is very popular at Google and Facebook research. Please click on Timetables on the right hand side of this page for time and location of the classes. The exercises appear below and are due Thursdays at 1pm on the specified week.
Spatio-temporal Sequence Prediction with Point Processes and Self-organizing Decision Trees
Karaahmetoglu, Oguzhan, Kozat, Suleyman S.
We investigate spatio-temporal prediction and introduce a novel prediction algorithm. Our approach is based on the point processes, which we use to model the event arrivals in both space and time. Although we specifically use the Hawkes process, other processes can be readily used as provided remarks in the paper. Moreover, we partition the given spatial region into subregions by an adaptive decision tree and model each subregion with individual and interacting point processes. With individual point processes for each subregion, we estimate the time and location of the events using the past event times and locations. Furthermore, thanks to the nonstationary and self-exciting point generation mechanism in the Hawkes process and the adaptive partitioning of the space, we model the data as nonstationary in both time and space. Finally, we provide a gradient based joint optimization algorithm for the adaptive tree parameter and the point process parameters. With the joint optimization, our algorithm can infer the source statistics and adaptive partitioning of the region. We also provide a training algorithm for the online setup, where we update the model parameters with newly arrived points. We provide experimental results on both simulated data and real-life data where we compare our approach with the standard approaches and demonstrate significant performance improvements thanks to the adaptive spatial partitioning mechanism and the joint optimization procedure.
Some Use Cases for "Automating AutoML"
In my last post I discussed why we at Auger believe that AI will eat software. Enterprises will move beyond just solving their biggest problems with painstakingly built predictive models that teams of data scientists spend months on. Instead every enterprise application can build predictive models wherever they have access to data. Such predictive models can replace hand-coded rule of thumb "business rules" (sort orders, if then else statements, switch-case statements, complex menus that users must navigate) and enable truly optimal decisions. The enabler for this transition is truly accurate AutoML: better than human data scientists.
Artificial Intelligence Isn't Just About Cutting Costs. It's Also About Growth.
What if this is the wrong question? When it comes to automating customer conversations with chatbots and AI-driven virtual agents, the real value of that automation may well be in generating revenue and growth, not in cutting headcount. Already, there are a few notable examples of chatbots generating significant revenue. For example, the casual dining chain TGI Fridays decided to reach out to millennials with an irreverent, brand-appropriate chatbot available through Facebook Messenger and Twitter . As Fridays' Chief Experience Officer Sherif Mityas explained, "This is an opportunity to think differently about how we engage with guests who talk to us." It has worked spectacularly well.
A Hierarchical Distance-dependent Bayesian Model for Event Coreference Resolution
Yang, Bishan, Cardie, Claire, Frazier, Peter
We present a novel hierarchical distance-dependent Bayesian model for event coreference resolution. While existing generative models for event coreference resolution are completely unsupervised, our model allows for the incorporation of pairwise distances between event mentions -- information that is widely used in supervised coreference models to guide the generative clustering processing for better event clustering both within and across documents. We model the distances between event mentions using a feature-rich learnable distance function and encode them as Bayesian priors for nonparametric clustering. Experiments on the ECB+ corpus show that our model outperforms state-of-the-art methods for both within- and cross-document event coreference resolution.
HVAC-Aware Occupancy Scheduling (Extended Abstract)
Lim, Boon-Ping (NICTA and Australian National University)
My research focuses on developing innovative ways to control Heating, Ventilation, and Air Conditioning (HVAC) and schedule occupancy flows in smart buildings to reduce our ecological footprint (and energy bills). We look at the potential for integrating building operations with room booking and meeting scheduling. Specifically, we improve on the effectiveness of energy-aware room-booking and occupancy scheduling approaches, by allowing the scheduling decisions to rely on an explicit model of the building's occupancy-based HVAC control. From computational standpoint, this is a challenging topic as HVAC models are inherently non-linear non-convex, and occupancy scheduling models additionally introduce discrete variables capturing the time slot and location at which each activity is scheduled. The mechanism needs to tradeoff minimizing energy cost against addressing occupancy thermal comfort and control feasibility in a highly dynamic and uncertain system.
Building a Timeline Network for Evacuation in Earthquake Disaster
Nguyen, The Minh (The University of Electro-Communications) | Kawamura, Takahiro (The University of Electro-Communications) | Tahara, Yasuyuki (The University of Electro-Communications) | Ohsuga, Akihiko (The University of Electro-Communications)
In this paper, we propose an approach that automatically extract users’ activities in sentences retrieved from Twitter. We then design a timeline action networkbased on Web Ontology Language (OWL). By using the proposed activity extraction approach, we can automatically collect data for the action network. Finally, we propose a novel action-based collaborative filtering, which predicts missing activity data, in order to complement this timeline network. Moreover, with a combination of collaborative filtering and natural language processing (NLP), our method can deal with minority actions such as successful actions. Based on evaluation of tweets which related to the massive Tohoku earthquake,we indicated that our timeline action network can provide useful action patterns in real-time. Not only earthquake disaster, our research can also be applied to other disasters and business models, such as typhoon,travel, marketing, etc.