A Deep Learning-Based Approach for Measuring the Domain Similarity of Persian Texts
Keshavarz, Hossein, Seifi, Shohreh Tabatabayi, Izadi, Mohammad
–arXiv.org Artificial Intelligence
In this paper, we propose a novel approach for measuring the degree of similarity between categories of two pieces of Persian text, which were published as descriptions of two separate advertisements. We built an appropriate dataset for this work using a dataset which consists of advertisements posted on an e-commerce website. We generated a significant number of paired texts from this dataset and assigned each pair a score from 0 to 3, which demonstrates the degree of similarity between the domains of the pair. In this work, we represent words with word embedding vectors derived from word2vec. Then deep neural network models are used to represent texts. Eventually, we employ concatenation of absolute difference and bit-wise multiplication and a fully-connected neural network to produce a probability distribution vector for the score of the pairs. Through a supervised learning approach, we trained our model on a GPU, and our best model achieved an F1 score of 0.9865.
arXiv.org Artificial Intelligence
Sep-24-2019
- Country:
- North America
- United States
- New York (0.04)
- Nevada (0.04)
- Maryland > Baltimore (0.04)
- Washington > King County
- Seattle (0.04)
- Texas > Travis County
- Austin (0.04)
- Georgia > Fulton County
- Atlanta (0.04)
- Colorado > Denver County
- Denver (0.04)
- California
- Santa Clara County > Stanford (0.04)
- San Diego County > San Diego (0.04)
- Canada
- Quebec > Montreal (0.04)
- British Columbia > Metro Vancouver Regional District
- Vancouver (0.04)
- United States
- Europe
- Asia > Middle East
- Qatar > Ad-Dawhah
- Doha (0.04)
- Israel > Haifa District
- Haifa (0.04)
- Iran > Razavi Khorasan Province
- Mashhad (0.04)
- Qatar > Ad-Dawhah
- North America
- Genre:
- Research Report > New Finding (0.46)
- Industry:
- Information Technology > Services > e-Commerce Services (0.54)
- Technology: