Country
Appendix to " GraphMP: Graph Neural Network-based Motion Planning with Efficient Graph Search "
The overall network architecture is shown in Figure 1. This work was done when the author was with Rutgers University. The overall network architecture is shown in Figure 1. We also apply the ReLU activation after its first and second layers. Empirical evaluations show that NHE exhibits admissibility and consistency.
Cross-lingual Retrieval for Iterative Self-Supervised Training (supplementary materials) 1 Experiment details
Becauseof the file size limit, we will release the source code and pretrained checkpoints after the anonymity period. To be able to make a fair comparison,we followed the same preprocessingsteps as described in [13]. In each iteration, we mine all90 language pairs in parallel, using8 GPUs for each pair, each pair taking about15 30 hours to finish. We lightly tune the margin score threshold using validation BLEU (using threshold score between 1.04and1.07.) For all experiments, we use Transformerwith 12 layers of encoder and 12 layers of decoder with model dimension of1024 on 16 heads ( 680M parameters). 1 We trained for maximum20,000 steps using label-smoothed cross-entropy loss with 0.2 label smoothing,0.3