Goto

Collaborating Authors

 Statistical Learning




A Limitations and future work We believe that the

Neural Information Processing Systems

All real-world datasets analysed consist of sequence reads of the same part of the genome. This is a widespread set-up for sequence analysis but not ubiquitous. In this project, we work with edit distances between sequences, these are too expensive for large-scale analysis, but it is feasible to produce a large enough training set. We describe here the methods that are most closely related to our work. However, these are bound to a quadratic complexity w.r.t. the length of the input sequence, the best algorithm [ Experiments were also run on synthetic datasets formed by sequences randomly generated.


Neural Distance Embeddings for Biological Sequences

Neural Information Processing Systems

The vector space can then be used to study the relationship between sequences and, potentially, decode new ones (see Section 7.2). On the right, an example of the hierarchical clustering produced on the Poincarรฉ disk.


A used and training procedures

Neural Information Processing Systems

All the models are trained for 200 epochs with stochastic gradient descent with a batch size = 128, momentum = 0.9, and cosine All the hyperparameters were selected with a small grid search. From epoch 150 to epoch 185 the training error of the chunks with size 128/256 decreases below 0.5%, while for smaller chunk sizes it remains above 5%. Random chunks with sizes larger than 128/256 can fit the training set, thus having the same representational power as the whole network on the training data. For W > 128/256 the test accuracy is decaying approximately with the same law as that of independent networks with the same width (see Figure 1). This picture suggests that for CIFAR100 the size of a clone is 128/256, slightly larger than the size of the clones in CIFAR10.