Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Hasan Can Biyik

When Semantic Overlap Is Not Enough: Cross-Lingual Euphemism Transfer Between Turkish and English

Feb 18, 2026

Hasan Can Biyik, Libby Barak, Jing Peng, Anna Feldman

Abstract:Euphemisms substitute socially sensitive expressions, often softening or reframing meaning, and their reliance on cultural and pragmatic context complicates modeling across languages. In this study, we investigate how cross-lingual equivalence influences transfer in multilingual euphemism detection. We categorize Potentially Euphemistic Terms (PETs) in Turkish and English into Overlapping (OPETs) and Non-Overlapping (NOPETs) subsets based on their functional, pragmatic, and semantic alignment. Our findings reveal a transfer asymmetry: semantic overlap is insufficient to guarantee positive transfer, particularly in low-resource Turkish-to-English direction, where performance can degrade even for overlapping euphemisms, and in some cases, improve under NOPET-based training. Differences in label distribution help explain these counterintuitive results. Category-level analysis suggests that transfer may be influenced by domain-specific alignment, though evidence is limited by sparsity.

* Proceedings of the Second Workshop on Natural Language Processing for Turkic Languages. EACL 2026. Rabat, Morocco March 29, 2026

Via

Access Paper or Ask Questions

Turkish Delights: a Dataset on Turkish Euphemisms

Jul 17, 2024

Hasan Can Biyik, Patrick Lee, Anna Feldman

Figure 1 for Turkish Delights: a Dataset on Turkish Euphemisms

Figure 2 for Turkish Delights: a Dataset on Turkish Euphemisms

Figure 3 for Turkish Delights: a Dataset on Turkish Euphemisms

Figure 4 for Turkish Delights: a Dataset on Turkish Euphemisms

Abstract:Euphemisms are a form of figurative language relatively understudied in natural language processing. This research extends the current computational work on potentially euphemistic terms (PETs) to Turkish. We introduce the Turkish PET dataset, the first available of its kind in the field. By creating a list of euphemisms in Turkish, collecting example contexts, and annotating them, we provide both euphemistic and non-euphemistic examples of PETs in Turkish. We describe the dataset and methodologies, and also experiment with transformer-based models on Turkish euphemism detection by using our dataset for binary classification. We compare performances across models using F1, accuracy, and precision as evaluation metrics.

* In Proceedings of The First SIGTURK workshop co-located with ACL 2024: https://sigturk.github.io/workshop/

Via

Access Paper or Ask Questions