Learning from String Sequences

May-10-2024–arXiv.org Artificial Intelligence

The Universal Similarity Metric (USM) has been demonstrated to give practically useful measures of "similarity" between sequence data. Here we have used the USM as an alternative distance metric in a K-Nearest Neighbours (K-NN) learner to allow effective pattern recognition of variable length sequence data. We compare this USM approach with the commonly used string-to-word vector approach. Our experiments have used two data sets of divergent domains: (1) spam email filtering and (2) protein subcellular localisation. Our results with this data reveal that the USM based K-NN learner (1) gives predictions with higher classification accuracy than those output by techniques that use the string to word vector approach, and (2) can be used to generate reliable probability forecasts.

algorithm, learner, probability forecast, (15 more...)

arXiv.org Artificial Intelligence

May-10-2024

arXiv.org PDF

Add feedback

Country:
- North America > United States
  - New York (0.04)
  - Maryland > Baltimore (0.04)
  - Wisconsin > Dane County
    - Madison (0.04)
  - California > San Francisco County
    - San Francisco (0.05)

Genre:
- Research Report (0.70)

Industry:
- Health & Medicine > Pharmaceuticals & Biotechnology (0.69)

Technology:
- Information Technology > Artificial Intelligence
  - Representation & Reasoning (1.00)
  - Machine Learning
    - Performance Analysis > Accuracy (0.70)
    - Statistical Learning > Nearest Neighbor Methods (0.49)
    - Learning Graphical Models > Directed Networks
      - Bayesian Learning (0.48)

Duplicate Docs Excel Report

Title
None found

Similar Docs Excel Report more

Title	Similarity	Source
None found