Industry
The Secret Life of a Winter Olympics Drone
You have a very important role! As a first-person-view camera drone, you soar high above the action at the Milan Cortina Games, capturing aerial footage of Olympic athletes as they fly through the snow and slide down the ice. You will zoom around at speeds of up to 75 miles per hour, capturing immersive, veritรฉ-style footage that makes these inherently exciting sports feel even more exciting. You make the luge come alive! Here the head of Olympic Broadcasting Services, @YiannisExarchos takes us through the journey of the drone at the fastest winter sport, luge.
Patch Diffusion: Faster and More Data-Efficient Training of Diffusion Models
Diffusion models are powerful, but they require a lot of time and data to train. We propose Patch Diffusion, a generic patch-wise training framework, to significantly reduce the training time costs while improving data efficiency, which thus helps democratize diffusion model training to broader users. At the core of our innovations is a new conditional score function at the patch level, where the patch location in the original image is included as additional coordinate channels, while the patch size is randomized and diversified throughout training to encode the cross-region dependency at multiple scales. Sampling with our method is as easy as in the original diffusion model.
A Limitations and Societal Impacts
Limitations One limitation of our model is its potential for data bias. This could limit the applications of the model. MLLMs could be used to create fake news articles or social media posts. Hyperparameters Number of layers 24 Hidden size 2,048 FFN inner hidden size 8,192 Attention heads 32 Dropout 0.1 Attention dropout 0.1 Activation function GeLU [1] V ocabulary size 64,007 Soft tokens V size 64 Max length 2,048 Relative position embedding xPos [2] Initialization Magneto [3] Table 1: Hyperparameters of causal language model of K The detailed instruction tuning hyperparameters are listed in Table 3. The models are trained on web-scale multimodal corpora.