Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Artūras Nakvosas

Elastic Weight Consolidation for Full-Parameter Continual Pre-Training of Gemma2

May 09, 2025

Vytenis Šliogeris, Povilas Daniušis, Artūras Nakvosas

Abstract:This technical report describes an experiment on autoregressive pre-training of Gemma2 2 billion parameter large language model (LLM) with 10\% on the Lithuanian language component of CulturaX from the point of view of continual learning. We apply elastic weight consolidation (EWC) to the full set of the model's parameters and investigate language understanding benchmarks, consisting of Arc, Belebele, Gsm8K, Hellaswag, MMLU, TruthfulQA, and Winogrande sets (both in English and Lithuanian versions), and perplexity benchmarks. We empirically demonstrate that EWC regularisation allows us not only to mitigate catastrophic forgetting effects but also that it is potentially beneficial for learning of the new task with LLMs.

* 8 pages, 4 figures

Via

Access Paper or Ask Questions

Open Llama2 Model for the Lithuanian Language

Aug 23, 2024

Artūras Nakvosas, Povilas Daniušis, Vytas Mulevičius

Figure 1 for Open Llama2 Model for the Lithuanian Language

Figure 2 for Open Llama2 Model for the Lithuanian Language

Figure 3 for Open Llama2 Model for the Lithuanian Language

Figure 4 for Open Llama2 Model for the Lithuanian Language

Abstract:In this paper, we propose and describe the first open Llama2 large language models (LLMs) for the Lithuanian language, including an accompanying question/answer (Q/A) dataset and translations of popular LLM benchmarks. We provide a brief review of open regional LLMs and detailed information on the proposed LLMs and their training process. We also conduct an empirical evaluation, comparing the perplexities of the proposed LLMs with those of other modern open LLMs. In addition, benchmarking the proposed LLMs against language understanding tasks reveals that high-quality pretraining datasets may be essential for achieving models that perform efficiently on these benchmarks. The full realisations of the described LLMs are available in the accompanying open repository~\url{https://huggingface.co/neurotechnology}.

* 12 pages, 8 figures, 5 tables

Via

Access Paper or Ask Questions