Goto

Collaborating Authors

 Country





TowardsVideoTextVisualQuestionAnswering: BenchmarkandBaseline

Neural Information Processing Systems

Therearealready sometext-based visualquestion answering(TextVQA) benchmarks for developing machine's ability to answer questions based on texts in imagesinrecentyears.







Non-Linguistic Supervision for Contrastive Learning of Sentence Embeddings Appendix

Neural Information Processing Systems

We provide hyper-parameters of our models in Table A.1. Table A.1: Hyper-parameters used for training our VisualCSE and AudioCSE. Vision, we use Dropout augmentation (the same strategy in SimCSE) for AudioCSE. We compare unsup-SimCSE and unsup-VisualCSE on a small scale retrieval test. As shown in Table C.1, VisualCSE generally retrieves qualitatively different sentences than SimCSE.