Alert button
Picture for Felipe González-Pizarro

Felipe González-Pizarro

Alert button

Diversity-Aware Coherence Loss for Improving Neural Topic Models

May 26, 2023
Raymond Li, Felipe González-Pizarro, Linzi Xing, Gabriel Murray, Giuseppe Carenini

Figure 1 for Diversity-Aware Coherence Loss for Improving Neural Topic Models
Figure 2 for Diversity-Aware Coherence Loss for Improving Neural Topic Models
Figure 3 for Diversity-Aware Coherence Loss for Improving Neural Topic Models
Figure 4 for Diversity-Aware Coherence Loss for Improving Neural Topic Models

The standard approach for neural topic modeling uses a variational autoencoder (VAE) framework that jointly minimizes the KL divergence between the estimated posterior and prior, in addition to the reconstruction loss. Since neural topic models are trained by recreating individual input documents, they do not explicitly capture the coherence between topic words on the corpus level. In this work, we propose a novel diversity-aware coherence loss that encourages the model to learn corpus-level coherence scores while maintaining a high diversity between topics. Experimental results on multiple datasets show that our method significantly improves the performance of neural topic models without requiring any pretraining or additional parameters.

* Minor Fixes, 11 pages, Camera-Ready for ACL 2023 (Short Paper) 
Viaarxiv icon

Regional Differences in Information Privacy Concerns After the Facebook-Cambridge Analytica Data Scandal

Feb 16, 2022
Felipe González-Pizarro, Andrea Figueroa, Claudia López, Cecilia Aragon

Figure 1 for Regional Differences in Information Privacy Concerns After the Facebook-Cambridge Analytica Data Scandal
Figure 2 for Regional Differences in Information Privacy Concerns After the Facebook-Cambridge Analytica Data Scandal
Figure 3 for Regional Differences in Information Privacy Concerns After the Facebook-Cambridge Analytica Data Scandal
Figure 4 for Regional Differences in Information Privacy Concerns After the Facebook-Cambridge Analytica Data Scandal

While there is increasing global attention to data privacy, most of their current theoretical understanding is based on research conducted in a few countries. Prior work argues that people's cultural backgrounds might shape their privacy concerns; thus, we could expect people from different world regions to conceptualize them in diverse ways. We collected and analyzed a large-scale dataset of tweets about the #CambridgeAnalytica scandal in Spanish and English to start exploring this hypothesis. We employed word embeddings and qualitative analysis to identify which information privacy concerns are present and characterize language and regional differences in emphasis on these concerns. Our results suggest that related concepts, such as regulations, can be added to current information privacy frameworks. We also observe a greater emphasis on data collection in English than in Spanish. Additionally, data from North America exhibits a narrower focus on awareness compared to other regions under study. Our results call for more diverse sources of data and nuanced analysis of data privacy concerns around the globe.

* Computer Supported Cooperative Work (CSCW), 2022  
Viaarxiv icon