Goto

Collaborating Authors

 Technology









Towards Global Optimal Visual In-Context Learning Prompt Selection

Neural Information Processing Systems

Visual In-Context Learning (VICL) is a prevailing way to transfer visual foundation models to new tasks by leveraging contextual information contained in in-context examples to enhance learning and prediction of query samples.


Supplementary Material for LEPARD: Learning Explicit Part Discovery for 3D Articulated Shape Reconstruction

Neural Information Processing Systems

In this section, we provide detailed derivation for the kinematics proposed in the main paper. The numbers in () indicate the dimension of output features. S is a shape matrix that we set to the identity matrix I in LEP ARD since we use one-to-one mapping for the local deformation estimation. Finally, we obtain a pseudo ground-truth object silhouette G by thresholding the minimum feature distance to the center of the clusters. In Figure 1, we provide the architecture of the encoder-decoder model proposed in the main paper.