Pull out all the stops: Textual analysis via punctuation sequences
Darmon, Alexandra N. M., Bazzi, Marya, Howison, Sam D., Porter, Mason A.
–arXiv.org Artificial Intelligence
Whether enjoying the lucid prose of a favorite author or slogging through some other writer's cumbersome, heavy-set prattle (full of parentheses, em dashes, compound adjectives, and Oxford commas), readers will notice stylistic signatures not only in word choice and grammar, but also in punctuation itself. Indeed, visual sequences of punctuation from different authors produce marvelously different (and visually striking) sequences. Punctuation is a largely overlooked stylistic feature in "stylometry", the quantitative analysis of written text. In this paper, we examine punctuation sequences in a corpus of literary documents and ask the following questions: Are the properties of such sequences a distinctive feature of different authors? Is it possible to distinguish literary genres based on their punctuation sequences? Do the punctuation styles of authors evolve over time? Are we on to something interesting in trying to do stylometry without words, or are we full of sound and fury (signifying nothing)?
arXiv.org Artificial Intelligence
Jan-16-2020
- Country:
- Asia > Middle East
- Israel (0.04)
- Europe
- France (0.04)
- United Kingdom > England
- Cambridgeshire > Cambridge (0.04)
- East Sussex > Brighton (0.04)
- Greater London > London (0.04)
- Oxfordshire > Oxford (0.27)
- North America > United States
- California
- Alameda County > Berkeley (0.04)
- Los Angeles County > Los Angeles (0.28)
- Santa Clara County
- Massachusetts > Middlesex County
- Reading (0.04)
- New Jersey > Hudson County
- Hoboken (0.04)
- New York
- Bronx County > New York City (0.04)
- Kings County > New York City (0.04)
- New York County > New York City (0.14)
- Queens County > New York City (0.04)
- Richmond County > New York City (0.04)
- California
- South America
- Brazil (0.04)
- Venezuela > Monagas State
- Maturin (0.04)
- Asia > Middle East
- Genre:
- Research Report > New Finding (0.67)
- Industry:
- Government > Regional Government (0.45)
- Media > Music (0.45)
- Technology: