Image Captioning: An Eye for Blind

#artificialintelligence 

The objective of this study is to design a variational autoencoder to model images as well as associated labels or captions. In this model, the dataset used was extracted from the Flickr8k dataset which consisted of 8,000 images, each paired with five different captions and provided clean descriptions of the salient entities and events. Here, our dataset inherits the same properties but only consisted of 1,000 images. The dataset was prepared by Anuj Garg and can be found on the Kaggle. This helped us in the reduction of the size of the dataset from 1GB to nearly 133Mb.

Duplicate Docs Excel Report

Title
None found

Similar Docs  Excel Report  more

TitleSimilaritySource
None found