Goto

Collaborating Authors

 Technology


Into the Single Cell Multiverse: an End-to-End Dataset for Procedural Knowledge Extraction in Biomedical Texts

Neural Information Processing Systems

Here we describe the additional details of FlaMBรฉ's curation including structured guidelines for each annotation task, corpus curation, and file assembly. All manual curation in FlaMBรฉ was conducted by three annotators who have doctorate level expertise in computational biology. For named entity tagging annotations a set of structured guidelines were followed to ensure consistency. The guidelines given to reviewers are in the annotator guidelines section below. B.1 Tissue and cell type entities Generally, all terms, related synonyms, and text entities that can be mapped to an entry from the tissue, organ, body part, fluid, and cell type branches of the NCI thesaurus were labeled. Instead of a rigid vocabulary fixed on exact matches of NCIThesaurus (NCIT) terms and synonyms, annotators were encouraged to tag any word with the same meaning as an ontology term. For example, "Pancreatic ductal adenocarcinoma" describes cancer of the pancreas, which can be related back to the NCI Thesaurus, and thus was tagged as a "TISSUE". An initial set of rules was provided to each annotator. When one annotator encountered a corner case (e.g., "is neuron a tissue or cell type?") all annotators discussed, reached a consensus, then added the corner case to the set of annotation rules.



Neural Hybrid Automata Supplementary Material

Neural Information Processing Systems

A.1 Neural Hybrid Automata: Modules and Hyperparameters We provide a notation and summary table for Neural Hybrid Automata (NHA). The table serves as a quick reference for the core concepts introduced in the main text. Labels every subjtrajectory Xi with a mode z to ensure mode-conditioned decoder Fz can reconstruct it despite Neural ODE representation limitations (uniqueness of solutions given an initial condition). The only NHA hyperparameter beyond module architectural choices is m, or number of latent modes provided to the model at initialization. Performance effects of changing mhave been explored in Section 5.2 and Appendix B.2. Appendix B.2 further provides analyzes potential techniques to prune additional modes. A.2 Gradient Pathologies We provide some theoretical insights on the phenomenon of gradient pathologies with the simple example of a one-dimensional linear hybrid system with two modes and one timed jump, xt = axtt<ฯ„ bxtt>= ฯ„ t 6= ฯ„ x+t = cxtt= ฯ„ (A.1)




Appendices ALow-Rank Matrix Factorization with Non-Uniform Sampling

Neural Information Processing Systems

In this section, we demonstrate the effectiveness of low-rank matrix factorization in recovering the label relationship matrix. We first present four important facts: f1: the rank of the matrix is equivalent to the number of classes. Specifically, this also means that if ห†Zi,k = 1, then ห†Zj,k = 1. We consider a toy example (without self-loops), ห†Z = 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 A = 0 1 1 1 0 0 1 0 0 0 0 0 0 0 0 0 (14) In a standard LRMF problem, it is not possible to recover ห†Z from A since no entries are observed for the third and fourth rows. However, we can demonstrate how LRMF effectively performs in this situation. Recovery: We begin by assuming v1 is in class 1, resulting in U1,: = [1, 1, 1] and V1,: = [1,0,0]. By observing A1,4, we know that v4 is also in class 1, resulting in U4,: = [1, 1, 1]and V4,: = [1,0,0](f2). By analyzing A1,2 and A1,3, we determine that v2 and v3 do not belong to class 1.



Practical Near Neighbor Search via Group Testing: Supplementary Materials

Neural Information Processing Systems

In this section, we provide proofs for all of the theorems introduced in the main text. We begin with a simple extension of the results of [3] for the Bloom filter false positive and negative rates. Then, we prove our main claim, which is that the query time of our data structure is sublinear, given some relatively weak assumptions on the stability of the query. Theorem 1. Assuming the existence of an LSH family with collision probability s(x,y) = sim(x,y), the distance-sensitive Bloom filter solves the approximate membership query problem with p 1 exp 2m t/m+ SLH We begin with a brief explanation of the results from [3]. Recall that a distance-sensitive Bloom filter is a collection of mbit arrays. Array iis indexed using an independent LSH function li(x). To insert a point xinto the ith array, we set the bit at location li(x) to '1.' To query the filter, we calculate the mhash values of the query and return "true" when at least tof the corresponding bits are '1.' To bound p (the true positive rate) and q (the false positive rate), we bound the probability that a single array returns "true."


Practical Near Neighbor Search via Group Testing

Neural Information Processing Systems

We present a new algorithm for the approximate near neighbor problem that combines classical ideas from group testing with locality-sensitive hashing (LSH). We reduce the near neighbor search problem to a group testing problem by designating neighbors as "positives," non-neighbors as "negatives," and approximate membership queries as group tests.


Supplementary Material On Provable Benefits of Depth in Training Graph Convolutional Networks

Neural Information Processing Systems

Finally in GCNs, we from discuss [26, conditions Before processing, where ov let er-smoothing first recall and the o notation ver-fitting we might defined happen in Section simultaneously 3. Recall .