magazinelogo

Translation and Foreign Language Learning

ISSN Online: 3069-0315 ISSN Print: 3070-3077 CODEN:
Frequency: monthly Email: tfll@hillpublish.com
Total View: 336240 Downloads: 78633 Citations: 8 (From Dimensions)
ArticleInterdisciplinary Studies of Translation http://dx.doi.org/10.26855/tfll.2025.12.027

Corpus-based LDA Topic Modeling for Interdisciplinary Language-data Research: A Case Study of News Article Analysis

Xudi Zhao, Feijing Han, Tiedong Yang*

School of Foreign Languages, Inner Mongolia University of Technology, Hohhot 010051, Inner Mongolia Autonomous Region, China.

*Corresponding author: Tiedong Yang

This study is supported by the Basic Research Fund for Universities Directly under the Inner Mongolia Autonomous Region Government: Research on the Cultivation Model of AI-Empowered College Students' Foreign Language Competence in Telling China's Stories from the Perspective of Mediationism (JY20250009); the 2025 General Education and Teaching Reform Project of Inner Mongolia University of Technology: Development and Practice of an Interdisciplinary "Foreign Language + Regional Culture" Curriculum Module (2025246); and the General Project of the Research Base for Forging a Strong Sense of Community for the Chinese Nation at Inner Mongolia University of Technology: Research on Northern Frontier Narratives in Western Travelogues and the Integration of Northern Frontier Culture into Foreign Language Teaching from the Perspective of Forging a Strong Sense of Community for the Chinese Nation (NGDZLJDY2504).
Published: December 31, 2025

Abstract

This paper uses Latent Dirichlet Allocation (LDA) topic modeling on a corpus of 967 Word documents. It looks at how data-driven computational methods can fit into humanities research through a “Language + Data” framework. That framework means mixing linguistic analysis with computational data techniques so we can close the gap between language studies and data science. Using corpus linguistics and Natural Language Processing (NLP) techniques the study finds five semantically coherent topics inside a mixed corpus of 967 documents. Those documents   include ISBN-related academic research papers on economics stock market news and a selection of general news articles. The LDA model got a coherence score of 0.72 and it identified five distinct topics. Most of the documents – 89.5% – clustered around financial content. Another 8.5% focused on academic research and the rest were split between general news topics.

Keyword

Corpus linguistics; LDA topic modeling; interdisciplinary research; language data; NLP

References

Baker, M. (2000). Towards a methodology for investigating the style of a literary translator. Target, 12(2), 241-266.

Baker, M. (2011). Corpus linguistics and translation studies: Implications and applications. In M. Baker (Ed.), Critical concepts in linguistics: Translation studies (Vol. 2, pp. 73-98). Routledge.

Blei, D. M., Ng, A. Y., & Jordan, M. I. (2003). Latent Dirichlet allocation. Journal of Machine Learning Research, 3, 993-1022.

Chen, D. H., Zhou, Z., & Yang, J. (2018). Financial news mining based on LDA and sentiment analysis. Data Analysis and Knowledge Discovery, 2(10), 92-100.

Griffiths, T. L., & Steyvers, M. (2004). Finding scientific topics. Proceedings of the National Academy of Sciences, 101(Suppl 1), 5228-5235.

Hu, K. (2012). Corpus translation studies: Connotations and implications. Journal of Foreign Languages, 35(5), 59-70.

Hu, K., & Li, X. (2017). Corpus-based research on translation and China's image: Connotations and implications. Foreign Languages Research, (4), 70-75.

Hu, K., & Sheng, D. (2020). Corpus-based literary translation criticism: Connotations, implications and future. Technology Enhanced Foreign Language Education, (5), 19-24.

Hu, K., & Yang, F. (2019). Corpus-based literary studies: Connotations and implications. Journal of Zhejiang University (Humanities and Social Sciences), 49(5), 5-18.

Huang, L., & Shi, X. (2018). A corpus-based comparison of the translation styles of two Chinese versions of To the Lighthouse. Journal of PLA University of Foreign Languages, 41(2), 11-19.

Kim, K. H., & Zhu, Y. (Eds.). (2023). Researching translation in the age of technology and global conflict: Selected works of Mona Baker. Routledge.

Liu, Z. (2010). Corpus-based study of translator's style and translation strategies: A case of reporting verbs in translations of Dream of the Red Chamber. Journal of PLA University of Foreign Languages, 33(4), 87-92.

Liu, H., Huang, Y., Zhao, S., et al. (2017). Research on topic model optimization for Chinese short texts. Computer Engineering & Science, 39(3), 550-557.

Mimno, D., Wallach, H. M., Talley, E., et al. (2011). Optimizing semantic coherence in topic models. In Proceedings of EMNLP 2011 (pp. 262-272). Association for Computational Linguistics.

Newman, D., Lau, J. H., Grieser, K., & Baldwin, T. (2010). Automatic evaluation of topic coherence. In Proceedings of EMNLP 2010 (pp. 100-109). Association for Computational Linguistics.

Röder, M., Both, A., & Hinneburg, A. (2015). Exploring the space of topic coherence measures. In Proceedings of the Eighth ACM International Conference on Web Search and Data Mining (pp. 399-408). Association for Computing Machinery.

Standardization Administration of China. (2021). General technical specifications for corpus. Standardization Administration of China.

Zhao, W. X., Jiang, J., Weng, J., et al. (2010). Comparing Twitter and traditional media using topic models. In Proceedings of the 33rd International ACM SIGIR Conference on Research and Development in Information Retrieval (pp. 338-345). Association for Computing Machinery.

Copyright

© 2025 by the author(s).
This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution-NonCommercial-NoDerivatives (CC BY-NC-ND) license, which permits non-commercial use, distribution, and reproduction in any medium, provided the original work is properly cited and is not modified or adapted.
https://creativecommons.org/licenses/by-nc-nd/4.0/

How to cite this paper

Corpus-based LDA Topic Modeling for Interdisciplinary Language-data Research: A Case Study of News Article Analysis

How to cite this paper: Xudi Zhao, Feijing Han, Tiedong Yang. (2025). Corpus-based LDA Topic Modeling for Interdisciplinary Language-data Research: A Case Study of News Article Analysis. Translation and Foreign Language Learning1(5), 906-910.

DOI: http://dx.doi.org/10.26855/tfll.2025.12.027