Recursive Attribute Factoring

Cohn, David, Verma, Deepak, Pfleger, Karl

Dec-31-2007–Neural Information Processing Systems

Clustering, or factoring of a document collection attempts to "explain" each observed documentin terms of one or a small number of inferred prototypes. Prior work demonstrated that when links exist between documents in the corpus (as is the case with a collection of web pages or scientific papers), building a joint model of document contents and connections produces a better model than that built from contents or connections alone. Many problems arise when trying to apply these joint models to corpus at the scale of the World Wide Web, however; one of these is that the sheer overhead of representing a feature space on the order of billions of dimensions becomes impractical. Weaddress this problem with a simple representational shift inspired by probabilistic relationalmodels: instead of representing document linkage in terms of the identities of linking documents, we represent it by the explicit and inferred attributes ofthe linking documents. Several surprising results come with this shift: in addition to being computationally more tractable, the new model produces factors thatmore cleanly decompose the document collection. We discuss several variations on this model and show how some can be seen as exact generalizations of the PageRank algorithm.

artificial intelligence, machine learning, natural language, (18 more...)

Neural Information Processing Systems

Dec-31-2007

Conferences PDF

Add feedback

Country:
- North America > United States > California (0.28)

Technology:
- Information Technology
  - Information Management > Search (1.00)
  - Artificial Intelligence
    - Natural Language (1.00)
    - Machine Learning (1.00)

Duplicate Docs Excel Report

Title
Recursive Attribute Factoring
Recursive Attribute Factoring

Similar Docs Excel Report more

Title	Similarity	Source
None found