AlphaZip: Neural Network-Enhanced Lossless Text Compression
Narashiman, Swathi Shree, Chandrachoodan, Nitin
–arXiv.org Artificial Intelligence
Data compression, especially in the realm of text, is the art of reducing the size of information without sacrificing its integrity. In an era where vast amounts of data are constantly exchanged, efficient text compression has become more critical than ever, enabling faster communication, reduced storage costs, and enhanced performance in bandwidthlimited environments. Text in digital form is stored using different character sets (ASCII, Unicode, etc.) that are encoded into binary using encoding schemes like ASCII encoding, UTF-8, UTF-16 etc. Text files can be stored in plain text format devoid of any formatting metadata or could be stored in Rich Text, HTML or XML formats where structuring and formatting add an overhead to the storage space. Compressing text files involves identifying the redundancies in these representations and encoding them to capture recurring pattern information. This is usually achieved through lossless compression algorithms. Though there have been many purely information theoretic frameworks for compressing text in a lossless manner, Neural Network based architectures outperform most of these techniques, as they can encode additional information about the relationships between different components of text[2]. Generative pre-trained transformers (GPTs) are the state-of-the-art in text generation and prediction [3].
arXiv.org Artificial Intelligence
Sep-23-2024
- Country:
- North America > United States
- California > Santa Clara County > Palo Alto (0.04)
- Europe > Sweden
- Vaestra Goetaland > Gothenburg (0.04)
- North America > United States
- Genre:
- Research Report > New Finding (0.69)
- Technology: