Oceania
Alphabet's Wing shows off a larger delivery drone with a bigger payload capacity
Alphabet-owned Wing has been trying to make drone delivery an actual thing, but the relatively minuscule payload capacity of modern delivery aircraft has been a serious obstacle. The company just unveiled a new drone that's a step in the right direction. The new model can handle payloads of up to five pounds, which is twice as much as Wing's previous flagship drone. It can also travel up to 65 MPH, which is pretty darned fast. The onboard battery allows for a 12 mile round trip, which is in line with previous metrics, so that translates to an under six-minute delivery time. That certainly beats pizza delivery.
Efficient Reinforcemen Learning via Decoupling Exploration and Utilization
Yang, Jingpu, Zhao, Qirui, Wang, Helin, Huang, Yuxiao, Song, Zirui, Fang, Miao
Deep neural network(DNN) generalization is limited by the over-reliance of current offline reinforcement learning techniques on conservative processing of existing datasets. This method frequently results in algorithms that settle for suboptimal solutions that only adjust to a certain dataset. Similarly, in online reinforcement learning, the previously imposed punitive pessimism also deprives the model of its exploratory potential. Our research proposes a novel framework, Optimistic and Pessimistic Actor Reinforcement Learning (OPARL). OPARL employs a unique dual-actor approach: an optimistic actor dedicated to exploration and a pessimistic actor focused on utilization, thereby effectively differentiating between exploration and utilization strategies. This unique combination in reinforcement learning methods fosters a more balanced and efficient approach. It enables the optimization of policies that focus on actions yielding high rewards through pessimistic utilization strategies, while also ensuring extensive state coverage via optimistic exploration. Experiments and theoretical study demonstrates OPARL improves agents' capacities for application and exploration. In the most tasks of DMControl benchmark and Mujoco environment, OPARL performed better than state-of-the-art methods. Our code has released on https://github.com/yydsok/OPARL
Trade-off Between Dependence and Complexity for Nonparametric Learning -- an Empirical Process Approach
Deb, Nabarun, Mukherjee, Debarghya
Empirical process theory for i.i.d. observations has emerged as a ubiquitous tool for understanding the generalization properties of various statistical problems. However, in many applications where the data exhibit temporal dependencies (e.g., in finance, medical imaging, weather forecasting etc.), the corresponding empirical processes are much less understood. Motivated by this observation, we present a general bound on the expected supremum of empirical processes under standard $\beta/\rho$-mixing assumptions. Unlike most prior work, our results cover both the long and the short-range regimes of dependence. Our main result shows that a non-trivial trade-off between the complexity of the underlying function class and the dependence among the observations characterizes the learning rate in a large class of nonparametric problems. This trade-off reveals a new phenomenon, namely that even under long-range dependence, it is possible to attain the same rates as in the i.i.d. setting, provided the underlying function class is complex enough. We demonstrate the practical implications of our findings by analyzing various statistical estimators in both fixed and growing dimensions. Our main examples include a comprehensive case study of generalization error bounds in nonparametric regression over smoothness classes in fixed as well as growing dimension using neural nets, shape-restricted multivariate convex regression, estimating the optimal transport (Wasserstein) distance between two probability distributions, and classification under the Mammen-Tsybakov margin condition -- all under appropriate mixing assumptions. In the process, we also develop bounds on $L_r$ ($1\le r\le 2$)-localized empirical processes with dependent observations, which we then leverage to get faster rates for (a) tuning-free adaptation, and (b) set-structured learning problems.
HCVP: Leveraging Hierarchical Contrastive Visual Prompt for Domain Generalization
Zhou, Guanglin, Han, Zhongyi, Chen, Shiming, Huang, Biwei, Zhu, Liming, Liu, Tongliang, Yao, Lina, Zhang, Kun
Domain Generalization (DG) endeavors to create machine learning models that excel in unseen scenarios by learning invariant features. In DG, the prevalent practice of constraining models to a fixed structure or uniform parameterization to encapsulate invariant features can inadvertently blend specific aspects. Such an approach struggles with nuanced differentiation of inter-domain variations and may exhibit bias towards certain domains, hindering the precise learning of domain-invariant features. Recognizing this, we introduce a novel method designed to supplement the model with domain-level and task-specific characteristics. This approach aims to guide the model in more effectively separating invariant features from specific characteristics, thereby boosting the generalization. Building on the emerging trend of visual prompts in the DG paradigm, our work introduces the novel \textbf{H}ierarchical \textbf{C}ontrastive \textbf{V}isual \textbf{P}rompt (HCVP) methodology. This represents a significant advancement in the field, setting itself apart with a unique generative approach to prompts, alongside an explicit model structure and specialized loss functions. Differing from traditional visual prompts that are often shared across entire datasets, HCVP utilizes a hierarchical prompt generation network enhanced by prompt contrastive learning. These generative prompts are instance-dependent, catering to the unique characteristics inherent to different domains and tasks. Additionally, we devise a prompt modulation network that serves as a bridge, effectively incorporating the generated visual prompts into the vision transformer backbone. Experiments conducted on five DG datasets demonstrate the effectiveness of HCVP, outperforming both established DG algorithms and adaptation protocols.
Characterizing Online Eating Disorder Communities with Large Language Models
Chu, Minh Duc, Karnati, Aryan, He, Zihao, Lerman, Kristina
The rise in eating disorders, a dangerous mental health condition with high mortality and morbidity, has been linked to the proliferation of idealized body images on social media. However, the link between social media and eating disorders is far more complex. We argue that social media platforms create a feedback loop that amplifies the growth of content and communities that promote eating disorders like anorexia and bulimia. Specifically, social media platforms make it easy for vulnerable individuals to find and connect to like-minded others, while group dynamic processes encourage them to stay engaged within communities that promote and glorify harmful behaviors linked to eating disorders. We characterize this dynamic empirically through a combination of network and language analysis. We describe a novel framework that leverages large language models to analyze the discourse within online communities and probe their attitudes on topics related to eating disorders to identify potentially harmful content. Our work emphasizes the need for better social media moderation to disrupt harmful feedback loops and protect vulnerable individuals.
Automatic 3D Multi-modal Ultrasound Segmentation of Human Placenta using Fusion Strategies and Deep Learning
Singh, Sonit, Stevenson, Gordon, Mein, Brendan, Welsh, Alec, Sowmya, Arcot
The placenta has roles in fetal growth and development, oxygenation and nutrition, synthesising vital substances for pregnancy maintenance, including estrogen, progesterone, cytokines, and growth factors, and acting as a barrier against pathogens and drugs. Placental dysfunction is a leading cause of perinatal morbidity and mortality, including fetal growth restriction (FGR), pre-eclampsia, and stillbirth [1]. The in vivo assessment of placenta across gestation is critical to understand placental structure, function, and development and to identify strategies to optimise pregnancy outcome [2]. The primary modality for placental evaluation is two-dimensional (2D) ultrasound (US), which is non-invasive, inexpensive and more easily acceptable and accessible than other imaging modalities such as X-ray or Magnetic Resonance Imaging (MRI). It may be used to characterize location, shape, and volume of the placenta along with its interface with the endometrium and myometrium. Three-dimensional Power Doppler (PD) ultrasound permits direct visualisation of multi-directional placental vascularity, allowing assessment of both the uteroplacental and fetoplacental circulations, providing dynamic assessment of blood flow for imaging of abnormalities of the placenta. In three-dimensional (3D) ultrasound, a process called semantic segmentation could be used to separate the placenta for qualitative and quantitative analysis. Placenta segmentation is challenging because of its geometry, position, and appearance as the shape and location of placentas vary greatly across subjects [3] and fetal position can lead to shadowing artefacts. Determination of the placental boundary in relation to the uterine tissue is also challenging due to similar appearances [4], irregularity of boundary and changing size and shape with gestation, posing problems for segmentation [5]. Figure 1 shows 3D ultrasound providing placental visualisation in axial, coronal, and sagittal Springer Nature 2021 L
eipy: An Open-Source Python Package for Multi-modal Data Integration using Heterogeneous Ensembles
Bennett, Jamie J. R., Li, Yan Chak, Pandey, Gaurav
In this paper, we introduce eipy--an open-source Python package for developing effective, multi-modal heterogeneous ensembles for classification. eipy simultaneously provides both a rigorous, and user-friendly framework for comparing and selecting the best-performing multi-modal data integration and predictive modeling methods by systematically evaluating their performance using nested cross-validation. The package is designed to leverage scikit-learn-like estimators as components to build multi-modal predictive models. An up-to-date user guide, including API reference and tutorials, for eipy is maintained at https://eipy.readthedocs.io . The main repository for this project can be found on GitHub at https://github.com/GauravPandeyLab/eipy .
Large Language Models Are Neurosymbolic Reasoners
Fang, Meng, Deng, Shilong, Zhang, Yudi, Shi, Zijing, Chen, Ling, Pechenizkiy, Mykola, Wang, Jun
A wide range of real-world applications is characterized by their symbolic nature, necessitating a strong capability for symbolic reasoning. This paper investigates the potential application of Large Language Models (LLMs) as symbolic reasoners. We focus on text-based games, significant benchmarks for agents with natural language capabilities, particularly in symbolic tasks like math, map reading, sorting, and applying common sense in text-based worlds. To facilitate these agents, we propose an LLM agent designed to tackle symbolic challenges and achieve in-game objectives. We begin by initializing the LLM agent and informing it of its role. The agent then receives observations and a set of valid actions from the text-based games, along with a specific symbolic module. With these inputs, the LLM agent chooses an action and interacts with the game environments. Our experimental results demonstrate that our method significantly enhances the capability of LLMs as automated agents for symbolic reasoning, and our LLM agent is effective in text-based games involving symbolic tasks, achieving an average performance of 88% across all tasks.
Event-Based Visual Odometry on Non-Holonomic Ground Vehicles
Xu, Wanting, Zhang, Si'ao, Cui, Li, Peng, Xin, Kneip, Laurent
Despite the promise of superior performance under challenging conditions, event-based motion estimation remains a hard problem owing to the difficulty of extracting and tracking stable features from event streams. In order to robustify the estimation, it is generally believed that fusion with other sensors is a requirement. In this work, we demonstrate reliable, purely event-based visual odometry on planar ground vehicles by employing the constrained non-holonomic motion model of Ackermann steering platforms. We extend single feature n-linearities for regular frame-based cameras to the case of quasi time-continuous event-tracks, and achieve a polynomial form via variable degree Taylor expansions. Robust averaging over multiple event tracks is simply achieved via histogram voting. As demonstrated on both simulated and real data, our algorithm achieves accurate and robust estimates of the vehicle's instantaneous rotational velocity, and thus results that are comparable to the delta rotations obtained by frame-based sensors under normal conditions. We furthermore significantly outperform the more traditional alternatives in challenging illumination scenarios. The code is available at \url{https://github.com/gowanting/NHEVO}.
Risk-Aware Accelerated Wireless Federated Learning with Heterogeneous Clients
Ads, Mohamed, ElSawy, Hesham, Hassanein, Hossam S.
Wireless Federated Learning (FL) is an emerging distributed machine learning paradigm, particularly gaining momentum in domains with confidential and private data on mobile clients. However, the location-dependent performance, in terms of transmission rates and susceptibility to transmission errors, poses major challenges for wireless FL's convergence speed and accuracy. The challenge is more acute for hostile environments without a metric that authenticates the data quality and security profile of the clients. In this context, this paper proposes a novel risk-aware accelerated FL framework that accounts for the clients heterogeneity in the amount of possessed data, transmission rates, transmission errors, and trustworthiness. Classifying clients according to their location-dependent performance and trustworthiness profiles, we propose a dynamic risk-aware global model aggregation scheme that allows clients to participate in descending order of their transmission rates and an ascending trustworthiness constraint. In particular, the transmission rate is the dominant participation criterion for initial rounds to accelerate the convergence speed. Our model then progressively relaxes the transmission rate restriction to explore more training data at cell-edge clients. The aggregation rounds incorporate a debiasing factor that accounts for transmission errors. Risk-awareness is enabled by a validation set, where the base station eliminates non-trustworthy clients at the fine-tuning stage. The proposed scheme is benchmarked against a conservative scheme (i.e., only allowing trustworthy devices) and an aggressive scheme (i.e., oblivious to the trust metric). The numerical results highlight the superiority of the proposed scheme in terms of accuracy and convergence speed when compared to both benchmarks.