Picture for Anej Svete

Anej Svete

Disentangling the Expressivity of RoPE

Add code
Aug 12, 2026
Viaarxiv icon

Bridging the Gap Between Latent and Explicit Reasoning with Looped Transformers

Add code
Jun 30, 2026
Viaarxiv icon

Efficiently Representing Algorithms With Chain-of-Thought Transformers

Add code
Jun 18, 2026
Viaarxiv icon

Causally Evaluating the Learnability of Formal Language Tasks

Add code
Jun 08, 2026
Viaarxiv icon

Understanding the Parameter Space Geometry of Transformers Encoding Boolean Functions

Add code
Jun 07, 2026
Viaarxiv icon

Revisiting Padded Transformer Expressivity: Which Architectural Choices Matter and Which Don't

Add code
May 28, 2026
Viaarxiv icon

Olmo Hybrid: From Theory to Practice and Back

Add code
Apr 07, 2026
Viaarxiv icon

Context-Free Recognition with Transformers

Add code
Jan 05, 2026
Viaarxiv icon

Probability Distributions Computed by Hard-Attention Transformers

Add code
Oct 31, 2025
Figure 1 for Probability Distributions Computed by Hard-Attention Transformers
Figure 2 for Probability Distributions Computed by Hard-Attention Transformers
Viaarxiv icon

Information Locality as an Inductive Bias for Neural Language Models

Add code
Jun 05, 2025
Figure 1 for Information Locality as an Inductive Bias for Neural Language Models
Figure 2 for Information Locality as an Inductive Bias for Neural Language Models
Figure 3 for Information Locality as an Inductive Bias for Neural Language Models
Figure 4 for Information Locality as an Inductive Bias for Neural Language Models
Viaarxiv icon