Semantic Networks
Hound: Relation-First Knowledge Graphs for Complex-System Reasoning in Security Audits
Hound introduces a relation-first graph engine that improves system-level reasoning across interrelated components in complex codebases. The agent designs flexible, analyst-defined views with compact annotations (e.g., monetary/value flows, authentication/authorization roles, call graphs, protocol invariants) and uses them to anchor exact retrieval: for any question, it loads precisely the code that matters (often across components) so it can zoom out to system structure and zoom in to the decisive lines. A second contribution is a persistent belief system: long-lived vulnerability hypotheses whose confidence is updated as evidence accrues. The agent employs coverage-versus-intuition planning and a QA finalizer to confirm or reject hypotheses. On a five-project subset of ScaBench[1], Hound improves recall and F1 over a baseline LLM analyzer (micro recall 31.2% vs. 8.3%; F1 14.2% vs. 9.8%) with a modest precision trade-off. We attribute these gains to flexible, relation-first graphs that extend model understanding beyond call/dataflow to abstract aspects, plus the hypothesis-centric loop; code and artifacts are released to support reproduction.
EcphoryRAG: Re-Imagining Knowledge-Graph RAG via Human Associative Memory
Cognitive neuroscience research indicates that humans leverage cues to activate entity-centered memory traces (engrams) for complex, multi-hop recollection. Inspired by this mechanism, we introduce EcphoryRAG, an entity-centric knowledge graph RAG framework. During indexing, EcphoryRAG extracts and stores only core entities with corresponding metadata, a lightweight approach that reduces token consumption by up to 94\% compared to other structured RAG systems. For retrieval, the system first extracts cue entities from queries, then performs a scalable multi-hop associative search across the knowledge graph. Crucially, EcphoryRAG dynamically infers implicit relations between entities to populate context, enabling deep reasoning without exhaustive pre-enumeration of relationships. Extensive evaluations on the 2WikiMultiHop, HotpotQA, and MuSiQue benchmarks demonstrate that EcphoryRAG sets a new state-of-the-art, improving the average Exact Match (EM) score from 0.392 to 0.474 over strong KG-RAG methods like HippoRAG. These results validate the efficacy of the entity-cue-multi-hop retrieval paradigm for complex question answering.
Supplementary Material of Rot-Pro: Modeling Transitivity by Projection in Knowledge Graph Embedding
In section 3.2 of the submitted paper, we use the conclusion that "the transitive relation can be represented as the union of transitive closures of of all transitive chains." S1, S2, and S3 datasets of Counties are separated by '/'. Our model is implemented in Python 3.6 using Pytorch 1.1.0. We list the best hyper-parameter setting of Rot-Pro on the above datasets in Table 2. The fully expressive of BoxE refers to that it is able to express inference patterns, which includes symmetry, anti-symmetry, inversion, composition, hierarchy, intersection, and mutual exclusion.
1 Additional Results 1 1.1 Synthetically generated partially annotated datasets 2 1.1.1 Knowledge-graph based partially annotated dataset generation
We use the two highest frequency ones which result in 776 label categories. Let us consider two datasets as shown in Figure 1. Let's say that the label Fine-grained mismatch problem: This problem occurs when a parent label ( e.g . We use the CIFAR100 [8] and MS COCO panoptic segmentation [7] datasets for this purpose. The third row has the similar thing, but for Dataset 2. Roughly Dataset 1 have 3x more data as the Dataset 2, with a total of 45k images across both.
Flavonoid Fusion: Creating a Knowledge Graph to Unveil the Interplay Between Food and Health
Dalal, Aryan Singh, Zhang, Yinglun, Doฤan, Duru, ฤฐleri, Atalay Mert, McGinty, Hande Kรผรงรผk
The focus on'food as medicine' is gaining traction in the field of health and several studies conducted in the past few years discussed this aspect of food in the literature. However, very little research has been done on representing the relationship between food and health in a standardized, machine - readable fo rmat using a semantic web that can help us leverage this knowledge effectively. To address this gap, this study aims to create a knowledge graph to link food and health through the knowledge graphs' ability to combine information from various platforms foc using on flavonoid contents of food found in the USDA's databases and cancer connections found in the literature. We looked closely at these relationships using KNARM methodology and represented them in machine - operable format. The proposed knowledge graph serves as an example for researchers, enabling them to explore the complex interplay between dietary choices and disease management. Future work for this study involves expanding the scope of the knowledge graph by capturing nuances, adding more related d ata, and performing inferences on the acquired knowledge to uncover hidden relationships.
Evaluating Embedding Frameworks for Scientific Domain
Ahmed, Nouman, Wu, Ronin, Botev, Victor
Finding an optimal word representation algorithm is particularly important in terms of domain specific data, as the same word can have different meanings and hence, different representations depending on the domain and context. While Generative AI and transformer architecture does a great job at generating contextualized embeddings for any given work, they are quite time and compute extensive, especially if we were to pre-train such a model from scratch. In this work, we focus on the scientific domain and finding the optimal word representation algorithm along with the tokenization method that could be used to represent words in the scientific domain. The goal of this research is two fold: 1) finding the optimal word representation and tokenization methods that can be used in downstream scientific domain NLP tasks, and 2) building a comprehensive evaluation suite that could be used to evaluate various word representation and tokenization algorithms (even as new ones are introduced) in the scientific domain. To this end, we build an evaluation suite consisting of several downstream tasks and relevant datasets for each task. Furthermore, we use the constructed evaluation suite to test various word representation and tokenization algorithms.