Rule-Based Reasoning
Towards Generating Controllable and Solvable Geometry Problem by Leveraging Symbolic Deduction Engine
Jiang, Zhuoxuan, Zhang, Tianyang, Peng, Peiyan, Chen, Jing, Xun, Yinong, Zhang, Haotian, Li, Lichi, Li, Yong, Zhang, Shaohua
Generating high-quality geometry problems is both an important and challenging task in education. Compared to math word problems, geometry problems further emphasize multi-modal formats and the translation between informal and formal languages. In this paper, we introduce a novel task for geometry problem generation and propose a new pipeline method: the Symbolic Deduction Engine-based Geometry Problem Generation framework (SDE-GPG). The framework leverages a symbolic deduction engine and contains four main steps: (1) searching a predefined mapping table from knowledge points to extended definitions, (2) sampling extended definitions and performing symbolic deduction, (3) filtering out unqualified problems, and (4) generating textual problems and diagrams. Specifically, our method supports to avoid inherent biases in translating natural language into formal language by designing the mapping table, and guarantees to control the generated problems in terms of knowledge points and difficulties by an elaborate checking function. With obtained formal problems, they are translated to natural language and the accompanying diagrams are automatically drew by rule-based methods. We conduct experiments using real-world combinations of knowledge points from two public datasets. The results demonstrate that the SDE-GPG can effectively generate readable, solvable and controllable geometry problems.
Token and Span Classification for Entity Recognition in French Historical Encyclopedias
Moncla, Ludovic, Zeghidi, Hรฉdi
Named Entity Recognition (NER) in historical texts presents unique challenges due to non-standardized language, archaic orthography, and nested or overlapping entities. This study benchmarks a diverse set of NER approaches, ranging from classical Conditional Random Fields (CRFs) and spaCy-based models to transformer-based architectures such as CamemBERT and sequence-labeling models like Flair. Experiments are conducted on the GeoEDdA dataset, a richly annotated corpus derived from 18th-century French encyclopedias. We propose framing NER as both token-level and span-level classification to accommodate complex nested entity structures typical of historical documents. Additionally, we evaluate the emerging potential of few-shot prompting with generative language models for low-resource scenarios. Our results demonstrate that while transformer-based models achieve state-of-the-art performance, especially on nested entities, generative models offer promising alternatives when labeled data are scarce. The study highlights ongoing challenges in historical NER and suggests avenues for hybrid approaches combining symbolic and neural methods to better capture the intricacies of early modern French text.
Surrogate Interpretable Graph for Random Decision Forests
Dubey, Akshat, Anลพel, Aleksandar, Hattab, Georges
The field of health informatics has been profoundly influenced by the development of random forest models, which have led to significant advances in the interpretability of feature interactions. These models are characterized by their robustness to overfitting and parallelization, making them particularly useful in this domain. However, the increasing number of features and estimators in random forests can prevent domain experts from accurately interpreting global feature interactions, thereby compromising trust and regulatory compliance. A method called the surrogate interpretability graph has been developed to address this issue. It uses graphs and mixed-integer linear programming to analyze and visualize feature interactions. This improves their interpretability by visualizing the feature usage per decision-feature-interaction table and the most dominant hierarchical decision feature interactions for predictions. The implementation of a surrogate interpretable graph enhances global interpretability, which is critical for such a high-stakes domain.
From Turbulence to Tranquility: AI-Driven Low-Altitude Network
Tekbฤฑyฤฑk, Kรผrลat, Raouf, Amir Hossein Fahim, Gรผvenรง, ฤฐsmail, Chen, Mingzhe, Kurt, Gรผneล Karabulut, Lesage-Landry, Antoine
Abstract--The Low Altitude Economy (LAE) network, with its transformative capabilities, is a candidate to become one of the major technological developments of the next decade for air mobility. However, the expected unprecedented density, mobility, and heterogeneity pose challenges and require new approaches, as it renders traditional rule-based approaches inadequate. T o address these challenges, this study introduces artificial intelligence (AI)-based approaches and validation frameworks for transitioning AI-enabled technologies from simulation-based studies to practical and deployable systems. First, AI-based spectrum sensing and coexistence utilizing the distributed nature of LAE nodes is introduced. Then, joint resource allocation and trajectory optimization driven by reinforcement learning is discussed. Bridging the gap between simulation and deployment through experimental platforms such as Aerial Experiments and Research Platform for Advanced Wireless (AERPA W), which are critical for validating models under realistic and non-stationary airspace conditions, is also addressed. The study concludes by highlighting open issues and outlining a forward-looking roadmap for the development of efficient, interoperable, and scalable AI-driven LAE ecosystems. The Low Altitude Economy (LAE) network is poised to become one of the defining technological trends of the next decade. Encompassing the use of the airspace below 3000 metres for economic, social, and operational activities, LAE covers various applications: urban air mobility (e.g., air taxis, emergency medical deliveries), precision agriculture, environmental sensing, surveillance, and logistics, as illustrated ixn Figure 1. M. Chen is with the Department of Electrical and Computer Engineering and Frost Institute for Data Science and Computing, University of Miami, Coral Gables, FL, 33146, USA (email: mingzhe.chen@miami.edu). This work is supported by the NSERC award ALLRP 579869-22 in Canada and the NSF awards CNS-2332834 and CNS-2332835 in the United States.
Exploring Domain Wall Pinning in Ferroelectrics via Automated High Throughput AFM
Barakati, Kamyar, Liu, Yu, Funakubo, Hiroshi, Kalinin, Sergei V.
Domain-wall dynamics in ferroelectric materials are strongly position-dependent since each polar interface is locked into a unique local microstructure. This necessitates spatially resolved studies of the wall-pinning using scanning-probe microscopy techniques. The pinning centers and preexisting domain walls are usually sparse within image plane, precluding the use of dense hyperspectral imaging modes and requiring time-consuming human experimentation. Here, a large area epitaxial PbTiO$_3$ film on cubic KTaO$_3$ were investigated to quantify the electric field driven dynamics of the polar-strain domain structures using ML-controlled automated Piezoresponse Force Microscopy. Analysis of 1500 switching events reveals that domain wall displacement depends not only on field parameters but also on the local ferroelectric-ferroelastic configuration. For example, twin boundaries in polydomains regions like a$_1^-$/$c^+$ $\parallel$ a$_2^-$/$c^-$ stay pinned up to a certain level of bias magnitude and change only marginally as the bias increases from 20V to 30V, whereas single variant boundaries like a$_2^+$/$c^+$ $\parallel$ a$_2^-$/$c^-$ stack are already activated at 20V. These statistics on the possible ferroelectric and ferroelastic wall orientations, together with the automated, high-throughput AFM workflow, can be distilled into a predictive map that links domain configurations to pulse parameters. This microstructure-specific rule set forms the foundation for designing ferroelectric memories.
VisualSphinx: Large-Scale Synthetic Vision Logic Puzzles for RL
Feng, Yichen, Xu, Zhangchen, Jiang, Fengqing, Li, Yuetai, Ramasubramanian, Bhaskar, Niu, Luyao, Lin, Bill Yuchen, Poovendran, Radha
Vision language models (VLMs) are expected to perform effective multimodal reasoning and make logically coherent decisions, which is critical to tasks such as diagram understanding and spatial problem solving. However, current VLM reasoning lacks large-scale and well-structured training datasets. To bridge this gap, we propose VisualSphinx, a first-of-its-kind large-scale synthetic visual logical reasoning training data. To tackle the challenge of image synthesis with grounding answers, we propose a rule-to-image synthesis pipeline, which extracts and expands puzzle rules from seed questions and generates the code of grounding synthesis image synthesis for puzzle sample assembly. Experiments demonstrate that VLM trained using GRPO on VisualSphinx benefit from logical coherence and readability of our dataset and exhibit improved performance on logical reasoning tasks. The enhanced reasoning capabilities developed from VisualSphinx also benefit other reasoning tasks such as algebraic reasoning, arithmetic reasoning and geometry reasoning.
Scene-Adaptive Motion Planning with Explicit Mixture of Experts and Interaction-Oriented Optimization
Zhu, Hongbiao, Ma, Liulong, Wu, Xian, Deng, Xin, Liang, Xiaoyao
Abstract--Despite over a decade of development, autonomous driving trajectory planning in complex urban environments continues to encounter significant challenges. These challenges include the difficulty in accommodating the multi-modal nature of trajectories, the limitations of the single expert model in managing diverse scenarios, and insufficient consideration of environmental interactions. T o address these issues, this paper introduces the EMoE-Planner, which incorporates three innovative approaches. Firstly, the Explicit MoE (Mixture of Experts) dynamically selects specialized experts based on scenario-specific information through a shared scene router . Secondly, the planner utilizes scene-specific queries to provide multi-modal priors, directing the model's focus towards relevant target areas. Lastly, it enhances the prediction model and loss calculation by considering the interactions between the ego vehicle and other agents, thereby significantly boosting planning performance. Comparative experiments were conducted on the Nuplan dataset against the state-of-the-art methods. The simulation results demonstrate that our model consistently outperforms SOT A models across nearly all test scenarios. Our model is the first pure learning model to achieve performance surpassing rule-based algorithms in almost all Nuplan closed-loop simulations. UTONOMOUS driving trajectory planning has evolved over decades, with rule-based methods [1]-[3] providing fundamental safety assurances via predefined logic and heuristics. However, in complex urban settings, three significant limitations become apparent: (1) The manual construction of rules struggles to accommodate dynamic interactions and abrupt changes in road topology, resulting in unaddressed long-tail scenarios; (2) Rigid trajectory generation fails to mimic the adaptive behaviors of human drivers, such as dynamically adjusting following distances; (3) An exponential increase in maintenance costs arises from the "combinatorial explosion" of accumulating rules. Conversely, data-driven approaches, including imitation learning [4]-[6], address edge cases like extreme weather and complex traffic, capturing human-like driving behaviors from expert data. Reinforcement learning [7], [8] enables dynamic optimization through advanced reward mechanisms. These systems offer lower costs and faster iterations compared to rule-based alternatives.
Nakatani urges closer defense tie-ups amid erosion of rules-based order
Defense Minister Gen Nakatani called Saturday for closer defense cooperation among like-minded partners in the Indo-Pacific region in order to strengthen the global rules-based order and -- in an implicit criticism of China -- act as a counter to countries seeking to erode the status quo. The Japanese defense chief used a speech before scores of his counterparts and military brass in Singapore at the Shangri-La Dialogue, Asia's leading security conference, to push for closer cooperation and coordination, "while ensuring openness, inclusiveness and transparency, with an aim of restoring a rules-based international order in the Indo-Pacific region, strengthening accountability and promoting the international public good." Nakatani said the need to unite on defense cooperation was clear, pointing to Russia's invasion of Ukraine -- a violation of the U.N. charter -- and Beijing's moves in the disputed South China Sea, including its decision to openly ignore a 2016 international arbitral tribunal ruling that dismissed the country's claim to most of the strategic waterway.
GeNRe: A French Gender-Neutral Rewriting System Using Collective Nouns
Doyen, Enzo, Todirascu, Amalia
A significant portion of the textual data used in the field of Natural Language Processing (NLP) exhibits gender biases, particularly due to the use of masculine generics (masculine words that are supposed to refer to mixed groups of men and women), which can perpetuate and amplify stereotypes. Gender rewriting, an NLP task that involves automatically detecting and replacing gendered forms with neutral or opposite forms (e.g., from masculine to feminine), can be employed to mitigate these biases. While such systems have been developed in a number of languages (English, Arabic, Portuguese, German, French), automatic use of gender neutralization techniques (as opposed to inclusive or gender-switching techniques) has only been studied for English. This paper presents GeNRe, the very first French gender-neutral rewriting system using collective nouns, which are gender-fixed in French. We introduce a rule-based system (RBS) tailored for the French language alongside two fine-tuned language models trained on data generated by our RBS. We also explore the use of instruct-based models to enhance the performance of our other systems and find that Claude 3 Opus combined with our dictionary achieves results close to our RBS. Through this contribution, we hope to promote the advancement of gender bias mitigation techniques in NLP for French.
Can Modern NLP Systems Reliably Annotate Chest Radiography Exams? A Pre-Purchase Evaluation and Comparative Study of Solutions from AWS, Google, Azure, John Snow Labs, and Open-Source Models on an Independent Pediatric Dataset
Hegde, Shruti, Ninan, Mabon Manoj, Dillman, Jonathan R., Hayatghaibi, Shireen, Babcock, Lynn, Somasundaram, Elanchezhian
A Pre - Purchase Evaluation and Comparative Study of Solutions from A WS, Google, Azure, John Snow Labs, and Open - Source Models on an Independent Pediatric Dataset Shruti Hegde MS, Mabon Manoj Ninan BS, Jonathan R. Dillman MD, MSc, Shireen Hayatghaibi PhD, Lynn Babcock MD, Elanchezhian Somasundaram PhD Abstract Purpose: General purpose clinical natural language processing tools are increasingly used for the automatic labeling of clinical reports to support various clinical, research and quality improvement applications. However, independent performance evaluations for specific tasks, such as labeling pediatric chest radiograph reports, remain scarce. This study aims to compare four leading commercial clinical NLP systems for entity extraction and assertion detection of clinically relevant findings in pediatric chest radiog raph reports . In addition, the study evaluates two dedicated chest radiograph report labelers, CheXpert and CheXbert, to provide a comprehensive performance comparison of the systems in extracting disease labels defined by CheXpert. Methods: A total of 95,008 pediatric chest radiograph (CXR) reports were obtained from a large academic pediatric hospital for this IRB - waived study. Clinically relevant terms were extracted using four general - purpose clinical NLP systems: Amazon Comprehend Medical (AWS), Google Healthcare NLP (GC), Azure Clinical NLP (AZ), and SparkNLP (SP) from John Snow Labs. After standardization, entities and their assertion statuses (positive, negative, uncertain) from the findings and impression sec tions were analyzed using descriptive statistics, paired t - tests, and Chi - square tests . Entities from the I mpression sections were mapped to 12 disease categories plus a No Findin gs category using a regular expression algorithm. In parallel, CheXpert and CheXbert processed the same reports to extract the same 13 categories (12 disease categories and a No Findings category) . Outputs from all six models were compared using Fleiss' Kappa across the assertion categories .