Goto

Collaborating Authors

 Personal Assistant Systems


A Survey on Federated Learning in Human Sensing

arXiv.org Artificial Intelligence

Human Sensing, a field that leverages technology to monitor human activities, psycho-physiological states, and interactions with the environment, enhances our understanding of human behavior and drives the development of advanced services that improve overall quality of life. However, its reliance on detailed and often privacy-sensitive data as the basis for its machine learning (ML) models raises significant legal and ethical concerns. The recently proposed ML approach of Federated Learning (FL) promises to alleviate many of these concerns, as it is able to create accurate ML models without sending raw user data to a central server. While FL has demonstrated its usefulness across a variety of areas, such as text prediction and cyber security, its benefits in Human Sensing are under-explored, given the particular challenges in this domain. This survey conducts a comprehensive analysis of the current state-of-the-art studies on FL in Human Sensing, and proposes a taxonomy and an eight-dimensional assessment for FL approaches. Through the eight-dimensional assessment, we then evaluate whether the surveyed studies consider a specific FL-in-Human-Sensing challenge or not. Finally, based on the overall analysis, we discuss open challenges and highlight five research aspects related to FL in Human Sensing that require urgent research attention. Our work provides a comprehensive corpus of FL studies and aims to assist FL practitioners in developing and evaluating solutions that effectively address the real-world complexities of Human Sensing.


InterFormer: Towards Effective Heterogeneous Interaction Learning for Click-Through Rate Prediction

arXiv.org Artificial Intelligence

Click-through rate (CTR) prediction, which predicts the probability of a user clicking an ad, is a fundamental task in recommender systems. The emergence of heterogeneous information, such as user profile and behavior sequences, depicts user interests from different aspects. A mutually beneficial integration of heterogeneous information is the cornerstone towards the success of CTR prediction. However, most of the existing methods suffer from two fundamental limitations, including (1) insufficient inter-mode interaction due to the unidirectional information flow between modes, and (2) aggressive information aggregation caused by early summarization, resulting in excessive information loss. To address the above limitations, we propose a novel module named InterFormer to learn heterogeneous information interaction in an interleaving style. To achieve better interaction learning, InterFormer enables bidirectional information flow for mutually beneficial learning across different modes. To avoid aggressive information aggregation, we retain complete information in each data mode and use a separate bridging arch for effective information selection and summarization. Our proposed InterFormer achieves state-of-the-art performance on three public datasets and a large-scale industrial dataset.


KGIF: Optimizing Relation-Aware Recommendations with Knowledge Graph Information Fusion

arXiv.org Artificial Intelligence

While deep-learning-enabled recommender systems demonstrate strong performance benchmarks, many struggle to adapt effectively in real-world environments due to limited use of user-item relationship data and insufficient transparency in recommendation generation. Traditional collaborative filtering approaches fail to integrate multifaceted item attributes, and although Factorization Machines account for item-specific details, they overlook broader relational patterns. Collaborative knowledge graph-based models have progressed by embedding user-item interactions with item-attribute relationships, offering a holistic perspective on interconnected entities. However, these models frequently aggregate attribute and interaction data in an implicit manner, leaving valuable relational nuances underutilized. This study introduces the Knowledge Graph Attention Network with Information Fusion (KGIF), a specialized framework designed to merge entity and relation embeddings explicitly through a tailored self-attention mechanism. The KGIF framework integrates reparameterization via dynamic projection vectors, enabling embeddings to adaptively represent intricate relationships within knowledge graphs. This explicit fusion enhances the interplay between user-item interactions and item-attribute relationships, providing a nuanced balance between user-centric and item-centric representations. An attentive propagation mechanism further optimizes knowledge graph embeddings, capturing multi-layered interaction patterns. The contributions of this work include an innovative method for explicit information fusion, improved robustness for sparse knowledge graphs, and the ability to generate explainable recommendations through interpretable path visualization.


RecKG: Knowledge Graph for Recommender Systems

arXiv.org Artificial Intelligence

Knowledge graphs have proven successful in integrating heterogeneous data across various domains. However, there remains a noticeable dearth of research on their seamless integration among heterogeneous recommender systems, despite knowledge graph-based recommender systems garnering extensive research attention. This study aims to fill this gap by proposing RecKG, a standardized knowledge graph for recommender systems. RecKG ensures the consistent representation of entities across different datasets, accommodating diverse attribute types for effective data integration. Through a meticulous examination of various recommender system datasets, we select attributes for RecKG, ensuring standardized formatting through consistent naming conventions. By these characteristics, RecKG can seamlessly integrate heterogeneous data sources, enabling the discovery of additional semantic information within the integrated knowledge graph. We apply RecKG to standardize real-world datasets, subsequently developing an application for RecKG using a graph database. Finally, we validate RecKG's achievement in interoperability through a qualitative evaluation between RecKG and other studies.


Want to See How Streaming Services Will Change in 2025? Check Your Phone.

Slate

Sign up for the Slatest to get the most insightful analysis, criticism, and advice out there, delivered to your inbox daily. Your favorite streaming app may soon no longer be just a "streaming app" per se but an overall entertainment app that functions as a cross-medium platform of its own--and a necessity on both your smartphone and smart TV. On Monday, Peacock, the NBCUniversal service that catapulted off a successful 2024 to become one of the most popular streamers in the world, rolled out a bunch of new goodies in its latest mobile update, part of the testing phase for a gradual expansion of platform features. Subscribers will now have access to a significant array of pilot programs: new engines for discovery and for personalized recommendations, bonus clips for beloved shows, and even NBC-themed "mini-games" that will be updated daily to adapt to the latest broadcasts in the sports and reality-TV schedules. "After presenting the Paris Olympics on Peacock and testing interactive features like'Choose Your Reality' for shows like Real Housewives, we've learned that fans want to have options in their viewing experience and dive deeper into their favorite content," John Jelley, senior vice president of product and user experience at Peacock, told me. He added that the platform is "piloting new features" to allow users to interact with sports and TV shows "so they can indulge their obsessions in a way that's fun, super simple, and all in one place."


Gemini AI smarts are coming to Google Home to make the Assistant a better conversationalist

Engadget

During CES 2025, I had a chance to check out a demo of the way Google is integrating Gemini capabilities into its smart home platform via devices like the Nest Audio, Nest Hub and Nest Cameras. The main takeaway is that the conversations you have with the Google Assistant will feel more natural. Personally, I'd appreciate being able to ask questions as they pop in my head, without having to formulate some Assistant-friendly sentence before speaking -- what I saw makes me feel like my wish could come true. To kick things off, you'll still say "Hey Google," but for follow-up questions you can skip the prompt and the Assistant will be able to hold on to the thread of your conversation. During the demonstration, held in a simulated (and very posh) kitchen, the Google representative asked things like what to cook with ingredients he had on hand (chicken and spinach).


Gemini AI is coming to Google TV devices in 2025, making them easier to talk to

Engadget

This week at CES, Google presented an early look at new software and hardware upgrades coming to Google TV devices. The new features include the integration of Gemini, Google's AI model, to the Google Assistant, as well as a new ambient experience. New smart TVs with Google TV will also gain far-field mics and proximity sensors to support the new software perks. If you've used a Google TV or Google streaming device, you may have already used the "hey Google" prompt to search for shows to watch. With the addition of Gemini, those "conversations" should now feel more natural.


Halliday promises its smart wayfarers have a 'proactive' AI assistant inside

Engadget

Smart glasses are traditionally long on promise, short on delivery, especially at these sorts of consumer electronics shindigs. There's always a steady stream of companies promising we're on the cusp of having our very own Gary-from-Veep attached to our faces before fading away. The weight of promises Halliday has laid upon the table is a sign of braggadocio, but it'll take a while before we know if it's deserved or not. Halliday has turned up at CES 2025 in Las Vegas with a pair of eponymous smart glasses filled to the brim with technology. There's a waveguide display in the right eyecup that will project the equivalent of a 3.5-inch screen into the wearer's view.


Designing Telepresence Robots to Support Place Attachment

arXiv.org Artificial Intelligence

People feel attached to places that are meaningful to them, which psychological research calls "place attachment." Place attachment is associated with self-identity, self-continuity, and psychological well-being. Even small cues, including videos, images, sounds, and scents, can facilitate feelings of connection and belonging to a place. Telepresence robots that allow people to see, hear, and interact with a remote place have the potential to establish and maintain a connection with places and support place attachment. In this paper, we explore the design space of robotic telepresence to promote place attachment, including how users might be guided in a remote place and whether they experience the environment individually or with others. We prototyped a telepresence robot that allows one or more remote users to visit a place and be guided by a local human guide or a conversational agent. Participants were 38 university alumni who visited their alma mater via the telepresence robot. Our findings uncovered four distinct user personas in the remote experience and highlighted the need for social participation to enhance place attachment. We generated design implications for future telepresence robot design to support people's connections with places of personal significance.


Personalized Fashion Recommendation with Image Attributes and Aesthetics Assessment

arXiv.org Artificial Intelligence

Personalized fashion recommendation is a difficult task because 1) the decisions are highly correlated with users' aesthetic appetite, which previous work frequently overlooks, and 2) many new items are constantly rolling out that cause strict cold-start problems in the popular identity (ID)-based recommendation methods. These new items are critical to recommend because of trend-driven consumerism. In this work, we aim to provide more accurate personalized fashion recommendations and solve the cold-start problem by converting available information, especially images, into two attribute graphs focusing on optimized image utilization and noise-reducing user modeling. Compared with previous methods that separate image and text as two components, the proposed method combines image and text information to create a richer attributes graph. Capitalizing on the advancement of large language and vision models, we experiment with extracting fine-grained attributes efficiently and as desired using two different prompts. Preliminary experiments on the IQON3000 dataset have shown that the proposed method achieves competitive accuracy compared with baselines.