Improving Fake News Detection in Low-Resource Languages through Retrieval-Augmented Generation And Parameter-Efficient Fine-Tuning
Abstract
Full Text:
PDFReferences
L. W. Abbott and D. Snidal, “Engaging the Public and the Private in Global Sustainability Governance,” International Affairs, vol. 88, no. 3, 2012, pp. 543–564.
F. Menczer and T. Hills, "Information overload helps fake news spread, and social media knows it," Scientific American, vol. 321, no. 6, pp. 54–61, 2019.
A. Bondielli and F. Marcelloni, "A survey on fake news and rumour detection techniques," Inf. Sci., vol. 497, pp. 38–55, 2019.
K. Shu, A. Sliva, S. Wang, J. Tang, and H. Liu, "Fake news detection on social media: A data mining perspective," ACM SIGKDD Explor. Newsl., vol. 19, no. 1, pp. 22–36, 2017.
Ethnologue, “Bengali,” 2026. [Online]. Available: https://www.ethnologue.com/language/ben/
F. Alam, A. Hasan, T. Alam, A. Khan, J. Tajrin, N. Khan, and S. A. Chowdhury, “A Review of Bangla Natural Language Processing Tasks and the Utility of Transformer Models,” arXiv preprint arXiv:2107.03844, 2021.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, "BERT: Pre-training of deep bidirectional transformers for language understanding," in Proc. NAACL-HLT, 2019, pp. 4171–4186.
A. Conneau et al., "Unsupervised cross-lingual representation learning at scale," in Proc. ACL, 2020, pp. 8440–8451.
Qwen Team, "Qwen3 technical report," arXiv:2505.09388, 2025.
P. Lewis et al., "Retrieval-augmented generation for knowledge-intensive NLP tasks," in Proc. NeurIPS, vol. 33, 2020, pp. 9459–9474.
E. J. Hu et al., "LoRA: Low-rank adaptation of large language models," in Proc. ICLR, 2022.
W. Y. Wang, “ ‘Liar, Liar Pants on Fire’: A New Benchmark Dataset for Fake News Detection,” in Proc. 55th Annual Meeting of the Association for Computational Linguistics, Vol. 2: Short Papers, Vancouver, Canada, 2017, pp. 422–426.
K. Shu, D. Mahudeswaran, S. Wang, D. Lee, and H. Liu, "FakeNewsNet: A data repository with news content, social context and spatial information," Big Data, vol. 8, no. 3, pp. 171–188, 2020.
T. Thorne et al., "FEVER: A large-scale dataset for fact extraction and verification," in Proc. NAACL-HLT, 2018, pp. 809–819.
Z. Jiang et al., "How can we know what language models know?," Trans. Assoc. Comput. Linguist., vol. 8, pp. 423–438, 2020.
K. Nakamura, S. Levy, and W. Y. Wang, "r/Fakeddit: A new multimodal benchmark dataset for fine-grained fake news detection," in Proc. LREC, 2020, pp. 6149–6157.
M. Z. H. George, N. Hossain, M. R. Bhuiyan, A. K. M. Masum, and S. Abujar, “Bangla Fake News Detection Based On Multichannel Combined CNN-LSTM,” in 2021 12th International Conference on Computing Communication and Networking Technologies (ICCCNT), 2021, pp. 1–5.
Hossain et al., “BanFakeNews: A Dataset for Detecting Fake News in Bangla,” in Proc. Twelfth Language Resources and Evaluation Conference, Marseille, France, 2020, pp. 2862–2871.
T. Brown et al., "Language models are few-shot learners," in Proc. NeurIPS, vol. 33, 2020, pp. 1877–1901.
V. Karpukhin et al., "Dense passage retrieval for open-domain question answering," in Proc. EMNLP, 2020, pp. 6769–6781.
F. Shi et al., "Language models are multilingual chain-of-thought reasoners," in Proc. ICLR, 2023.
N. Reimers and I. Gurevych, "Sentence-BERT: Sentence embeddings using siamese BERT-networks," in Proc. EMNLP-IJCNLP, 2019, pp. 3982–3992.
M. McCloskey and N. J. Cohen, "Catastrophic interference in connectionist networks," in Psychol. Learn. Motiv., vol. 24. Academic Press, 1989, pp. 109–165.
S. Mangrulkar et al., "PEFT: State-of-the-art parameter-efficient fine-tuning methods," GitHub, 2022. [Online]. Available: https://github.com/huggingface/peft
T. Dettmers, A. Pagnoni, A. Holtzman, and L. Zettlemoyer, "QLoRA: Efficient finetuning of quantized LLMs," in Proc. NeurIPS, vol. 36, 2023.
Unsloth AI, "Unsloth: 2x faster LLM fine-tuning," GitHub, 2024. [Online]. Available: https://github.com/unslothai/unsloth
J. Ainslie et al., "GQA: Training generalized multi-query transformer models from multi-head checkpoints," in Proc. EMNLP, 2023, pp. 4895–4901.
J. Su et al., "RoFormer: Enhanced transformer with rotary position embedding," Neurocomputing, vol. 568, p. 127063, 2024.
G. Gerganov, "GGML: Tensor library for machine learning," GitHub, 2023. [Online]. Available: https://github.com/ggerganov/ggml
Ollama, "Run large language models locally," 2024. [Online]. Available: https://ollama.com
Pinecone Systems Inc., "Pinecone: The vector database for ML applications," 2024. [Online]. Available: https://www.pinecone.io
T. Wolf et al., "Transformers: State-of-the-art natural language processing," in Proc. EMNLP: System Demonstrations, 2020, pp. 38–45.
Refbacks
- There are currently no refbacks.
Abava Кибербезопасность Monetec 2026 СНЭ
ISSN: 2307-8162