Goto

Collaborating Authors

 Technology





Three children killed in drone strike on mosque in central Sudan: Doctors

Al Jazeera

A drone attack on a mosque in central Sudan has killed two children and injured 13 more, according to a Sudanese doctor's association, amid a rise in similar attacks across the region. The Sudan Doctors Network said the attack was carried out at dawn on Wednesday by the Rapid Support Forces (RSF), a paramilitary group engaged in a three-year civil war with the Sudanese Armed Forces. "Targeting children inside mosques is a fully constituted crime that cannot be justified under any pretext and represents a dangerous escalation in the pattern of repeated violations against civilians," the doctors said. The Sudan Doctors Network said the RSF has previously targeted other religious buildings for attack, including a church in Khartoum and another mosque in el-Fasher, reflecting a "systematic pattern that shows clear disregard for the sanctity of life and religious sites". "The network calls on the international community, the United Nations, and human rights and humanitarian organizations to take urgent action to pressure for the end to the targeting of civilians, ensure their protection, open safe corridors for the delivery of medical and humanitarian aid, and work to document these violations and hold those responsible accountable," it said.





Understanding the Role of Momentum in Stochastic Gradient Methods

Neural Information Processing Systems

Different variants ofmomentum, including heavyball momentum, Nesterov's accelerated gradient (NAG), and quasi-hyperbolic momentum (QHM), havedemonstrated success onvarious tasks. Our results are most closely related to the work of Mandt et al.[19]who use stationaryanalysis of SGD with momentum to perform approximateBayesianinference.



SupplementaryMaterial: UnifiedVision-Language Pre-TrainingwithMixture-of-Modality-Experts

Neural Information Processing Systems

We perform finetuning with image-textcontrastiveand image-textmatching losses. During inference, VLMO is first used as a dual encoder to obtain top-k candidates, then the model is used as a fusionencoder torerankthecandidates. For the text-only pre-training data, we use English Wikipedia and BookCorpus [5]. Table 1: Ablation study of the shared self-attention module used in Multiway Transformer.