ambiance
Learning Street View Representations with Spatiotemporal Contrast
Li, Yong, Huang, Yingjing, Mai, Gengchen, Zhang, Fan
Street view imagery is extensively utilized in representation learning for urban visual environments, supporting various sustainable development tasks such as environmental perception and socio-economic assessment. However, it is challenging for existing image representations to specifically encode the dynamic urban environment (such as pedestrians, vehicles, and vegetation), the built environment (including buildings, roads, and urban infrastructure), and the environmental ambiance (such as the cultural and socioeconomic atmosphere) depicted in street view imagery to address downstream tasks related to the city. In this work, we propose an innovative self-supervised learning framework that leverages temporal and spatial attributes of street view imagery to learn image representations of the dynamic urban environment for diverse downstream tasks. By employing street view images captured at the same location over time and spatially nearby views at the same time, we construct contrastive learning tasks designed to learn the temporal-invariant characteristics of the built environment and the spatial-invariant neighborhood ambiance. Our approach significantly outperforms traditional supervised and unsupervised methods in tasks such as visual place recognition, socioeconomic estimation, and human-environment perception. Moreover, we demonstrate the varying behaviors of image representations learned through different contrastive learning objectives across various downstream tasks. This study systematically discusses representation learning strategies for urban studies based on street view images, providing a benchmark that enhances the applicability of visual data in urban science. The code is available at https://github.com/yonglleee/UrbanSTCL.
Make Compound Sentences Simple to Analyze: Learning to Split Sentences for Aspect-based Sentiment Analysis
Seo, Yongsik, Song, Sungwon, Heo, Ryang, Kim, Jieyong, Lee, Dongha
In the domain of Aspect-Based Sentiment Analysis (ABSA), generative methods have shown promising results and achieved substantial advancements. However, despite these advancements, the tasks of extracting sentiment quadruplets, which capture the nuanced sentiment expressions within a sentence, remain significant challenges. In particular, compound sentences can potentially contain multiple quadruplets, making the extraction task increasingly difficult as sentence complexity grows. To address this issue, we are focusing on simplifying sentence structures to facilitate the easier recognition of these elements and crafting a model that integrates seamlessly with various ABSA tasks. In this paper, we propose Aspect Term Oriented Sentence Splitter (ATOSS), which simplifies compound sentence into simpler and clearer forms, thereby clarifying their structure and intent. As a plug-and-play module, this approach retains the parameters of the ABSA model while making it easier to identify essential intent within input sentences. Extensive experimental results show that utilizing ATOSS outperforms existing methods in both ASQP and ACOS tasks, which are the primary tasks for extracting sentiment quadruplets.
Read the Room: Adapting a Robot's Voice to Ambient and Social Contexts
Tuttosi, Paige, Hughson, Emma, Matsufuji, Akihiro, Lim, Angelica
How should a robot speak in a formal, quiet and dark, or a bright, lively and noisy environment? By designing robots to speak in a more social and ambient-appropriate manner we can improve perceived awareness and intelligence for these agents. We describe a process and results toward selecting robot voice styles for perceived social appropriateness and ambiance awareness. Understanding how humans adapt their voices in different acoustic settings can be challenging due to difficulties in voice capture in the wild. Our approach includes 3 steps: (a) Collecting and validating voice data interactions in virtual Zoom ambiances, (b) Exploration and clustering human vocal utterances to identify primary voice styles, and (c) Testing robot voice styles in recreated ambiances using projections, lighting and sound. We focus on food service scenarios as a proof-of-concept setting. We provide results using the Pepper robot's voice with different styles, towards robots that speak in a contextually appropriate and adaptive manner. Our results with N=120 participants provide evidence that the choice of voice style in different ambiances impacted a robot's perceived intelligence in several factors including: social appropriateness, comfort, awareness, human-likeness and competency.
This Philips Hue Smart Bulb kit is $20 Off
Smart bulbs are great starting points when creating the connected home of your dreams. Because they connect to Wi-Fi, you can adjust them from an app, and control them via voice commands with your smart speakers. They also offer an awesome level of granular control, whether that means setting them on timers, or controlling their precise brightness levels and color profiles. For the smart home-curious, Philips has put its entry-level smart bulb kit on sale. The Hue White A19 60W Equivalent Smart Bulb Starter Kit is $44.99 on Amazon now--a significant price drop from its average offering of around $65.
Yelp reveals AI that 'looks' through pictures and can work out exactly what you are eating
Yelp has always relied on its users to share their dining experience with reviews. But now, it is experimenting with an AI to help it out. The San Francisco firm has developed software that uses image analysis to detect colour, texture and shape of objects in pictures in order to gather more information about a restaurant. As we live in an era where taking pictures before eating is part of the meal, the San Francisco firm has shifted its focus to images. Yelp has developed software that uses image analysis to detect colour, texture and shape of objects in pictures in order to gather more information about a restaurant.