The Power of Hard Attention Transformers on Data Sequences: A formal language theoretic perspective

May-27-2025, 12:50:59 GMT–Neural Information Processing Systems

Formal language theory has recently been successfully employed to unravel the power of transformer encoders. This setting is primarily applicable in Natural Language Processing (NLP), as a token embedding function (where a bounded number of tokens is admitted) is first applied before feeding the input to the transformer. In this paper, we initiate the study of the expressive power of transformer encoders on sequences of data (i.e. Our results indicate an increase in expressive power of hard attention transformers over data sequences, in stark contrast to the case of strings. In particular, we prove that Unique Hard Attention Transformers (UHAT) over inputs as data sequences no longer lie within the circuit complexity class AC0 (even without positional encodings), unlike the case of string inputs, but are still within the complexity class TC0 (even with positional encodings).

data sequence, formal language theoretic perspective, hard attention transformer, (3 more...)

Neural Information Processing Systems

May-27-2025, 12:50:59 GMT

Conferences Web Page

Add feedback

Technology:
- Information Technology > Artificial Intelligence
  - Natural Language (1.00)
  - Representation & Reasoning > Logic & Formal Reasoning (0.64)