branch structure
HiCo: Hierarchical Controllable Diffusion Model for Layout-to-image Generation
The task of layout-to-image generation involves synthesizing images based on the captions of objects and their spatial positions. Existing methods still struggle in complex layout generation, where common bad cases include object missing, inconsistent lighting, conflicting view angles, etc. To effectively address these issues, we propose a \textbf{Hi}erarchical \textbf{Co}ntrollable (HiCo) diffusion model for layout-to-image generation, featuring object seperable conditioning branch structure. Our key insight is to achieve spatial disentanglement through hierarchical modeling of layouts. We use a multi branch structure to represent hierarchy and aggregate them in fusion module. To evaluate the performance of multi-objective controllable layout generation in natural scenes, we introduce the HiCo-7K benchmark, derived from the GRIT-20M dataset and manually cleaned.
HiCo: Hierarchical Controllable Diffusion Model for Layout-to-image Generation
The task of layout-to-image generation involves synthesizing images based on the captions of objects and their spatial positions. Existing methods still struggle in complex layout generation, where common bad cases include object missing, inconsistent lighting, conflicting view angles, etc. To effectively address these issues, we propose a \textbf{Hi}erarchical \textbf{Co}ntrollable (HiCo) diffusion model for layout-to-image generation, featuring object seperable conditioning branch structure. Our key insight is to achieve spatial disentanglement through hierarchical modeling of layouts. We use a multi branch structure to represent hierarchy and aggregate them in fusion module.
Revisiting Temporal Modeling for CLIP-based Image-to-Video Knowledge Transferring
Liu, Ruyang, Huang, Jingjia, Li, Ge, Feng, Jiashi, Wu, Xinglong, Li, Thomas H.
However, it is hard to get a pretrained model as powerful as CLIP in the video Image-text pretrained models, e.g., CLIP, have shown domain due to the unaffordable demands on computation resources impressive general multi-modal knowledge learned from and the difficulty of collecting video-text data pairs large-scale image-text data pairs, thus attracting increasing as large and diverse as image-text data. Instead of directly attention for their potential to improve visual representation pursuing video-text pretrained models [17, 27], a potential learning in the video domain. In this paper, based alternative solution that benefits video downstream tasks is on the CLIP model, we revisit temporal modeling in the to transfer the knowledge in image-text pretrained models context of image-to-video knowledge transferring, which is to the video domain, which has attracted increasing attention the key point for extending image-text pretrained models to in recent years [12, 13, 26, 29, 30, 41]. the video domain. We find that current temporal modeling Extending pretrained 2D image models to the video domain mechanisms are tailored to either high-level semanticdominant is a widely-studied topic in deep learning [4, 7], and tasks (e.g., retrieval) or low-level visual patterndominant the key point lies in empowering 2D models with the capability tasks (e.g., recognition), and fail to work on the of modeling temporal dependency between video two cases simultaneously. The key difficulty lies in modeling frames while taking advantages of knowledge in the pretrained temporal dependency while taking advantage of both highlevel models. In this paper, based on CLIP [32], we revisit and low-level knowledge in CLIP model. To tackle temporal modeling in the context of image-to-video knowledge this problem, we present Spatial-Temporal Auxiliary Network transferring, and present Spatial-Temporal Auxiliary (STAN) - a simple and effective temporal modeling Network (STAN) - a new temporal modeling method that mechanism extending CLIP model to diverse video tasks. is easy and effective for extending image-text pretrained Specifically, to realize both low-level and high-level knowledge model to diverse downstream video tasks.
Image analysis and AI tech used to study branches
Researchers from Osaka University have managed the reconstruction of plant branch structures, focusing on points of interest like branch structures under leaves. This has been achieved using advanced image analysis together with artificial intelligence technology. This represents one of the first applications of this type of technology to botany. The stud is of particular importance to fruit-bearing trees. By gaining insights into the growth of branches and leaves of individual trees, those whose livelihoods are based on the cultivation of fruit trees can learn about new methods for managing trees. This can help with maximizing fruit growth and with helping to maintain and protect the trees.
3D reconstruction of hidden branch structures made by using image analysis and AI tech
Three-dimensional (3D) reconstruction from multiple images obtained from different viewpoints has been actively examined. However, it was difficult to reconstruct the structure of objects which have hidden portions, such as plants with branch structures hidden under their leaves. By combining the original image-to-image translation approach in a Bayesian deep learning framework and 3D reconstruction, a group of researchers led by Fumio Okura estimated the existence probability of branches that are hidden under leaves in images obtained. Using these estimated branch positions, they achieved 3D reconstruction of plant structure, i.e., accurate reconstruction of branch structures, including those hidden under leaves. Specifically, they converted images of leafy plants to images showing branch existence probability, thereby achieving 3D reconstruction. The results of this study will be presented at the EEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2018) to be held from June 18 through June 22, 2018.