Enhancing Information Access: Summarization Techniques For Tamil Language

Srimathi R, Premlatha K R, Abirami A M, Lohitha K, Pratika Lakshmi L G

Department of Information Technology

Thiagarajar College of Engineering, Madurai- 625015

Summary

Digital content in regional languages is increasing, and it is now more important than ever to provide summarization methods that are appropriate for low-resource languages, such as Tamil. This paper discusses an extractive method of summarization for Tamil text using word frequency and TextRank-based graph algorithms to select the most salient and relevant sentences. The method is modified slightly for Tamil's linguistic features with a custom stopword list and language-specific tokenization. Experimental evaluation revealed that the extractive method, as proposed, preserved factual representation while improving readability. The method achieved promising ROUGE scores, indicating strong alignment with human generated summaries. This system offers a dependable and scalable solution for summarizing Tamil text for different applications in the real world, such as in educational tools, organizing content, and information retrieval.

  1. Introduction

In today's digital age, the exponential growth of information has made it increasingly difficult for users to process and understand long pieces of text efficiently. This challenge becomes more significant when dealing with content in low-resource regional languages, where tools for automated processing are limited. Text summarization, a subfield of Natural Language Processing (NLP), addresses this issue by condensing large documents into concise summaries while preserving the essential meaning. It plays a vital role in enhancing information accessibility across multiple domains including education, healthcare, media, and public governance. This paper presents an extractive Tamil text summarization system that leverages word frequency and Text Rank-based graph algorithms. The proposed method has been carefully adapted to suit Tamil's unique linguistic characteristics, incorporating a custom stop word list and language-specific tokenization strategies.

Through experimental evaluation, the system has demonstrated its effectiveness in retainingfactual content while improving readability, making it suitable for real-world applications such as educational tools, news summarization, and digital governance. By developing a summarization system specifically tailored to Tamil, this work not only contributes to the field of computational linguistics but also supports the larger mission of digital inclusion and

regional language processing. The system addresses the growing demand for scalable, accurate summarization tools that can serve the needs of Tamil-speaking communities in an increasingly digital world.

Literature Survey

Text summarization has advanced from rule-based to deep learning methods, but low- resource languages like Tamil still face challenges due to limited datasets and tools. This survey highlights key contributions to Tamil summarization, spanning extractive, abstractive, and hybrid approaches.

Priyadharshan and Sumathipala [1] proposed a Tamil text summarization system for online sports news using NLP techniques like POS tagging and TF-IDF, combined with a Restricted Boltzmann Machine (RBM) for feature refinement. This hybrid approach showed the potential of combining linguistic processing with neural models for low-resource languages.Atul Kumar, Dr. Vinodani Katiyar et al. [2] conducted a comparative study of extractive and abstractive summarization techniques. They emphasized that while extractive models are lightweight and easier to implement, abstractive methods offer more fluent summaries, though they require substantial data and computational resources.

Syed Sabir Mohamed and Shanmugasundaram Hariharan [3] developed a graph-based summarization model for Tamil newspapers. Sentences were treated as graph nodes, and their similarities as edges, resulting in improved summary coherence and relevance compared to non-graph approaches. Prabhudas Janjanam and C.H. Pradeep Reddy [4] provided a historical overview of summarization techniques, highlighting the shift from statistical to neural methods. They stressed the importance of semantic coherence and suggested adapting modern models for low-resource languages like Tamil.Rahul, Surabhi Adhikari, and Monika [5] analyzed summarization techniques using various datasets, categorizing them into extractive, abstractive, and query-based types.

Banu, P. Karthika, C. Sudarmani, and T.V. Geetha [6] proposed a summarizer that used syntactic parsing to extract Subject-Object-Predicate (SOP) triples and construct semantic graphs. An SVM classifier selected the most relevant triples, enhancing summary accuracy and logical flow.Sarika M. [7] presented a comparative analysis of Tamil and English news summarization, focusing on linguistic and structural challenges. The study emphasized the unique morphological features of Tamil and the need for language-specific models and preprocessing techniques to achieve better summarization quality.

Proposed Solution

The proposed solution is an extractive Tamil text summarization system designed for a low-resource, morphologically complex language. Due to the lack of pretrained models and Tamil-specific NLP tools, the system uses efficient techniques like word frequency analysis and the TextRank graph-based algorithm. The goal is to produce accurate and readable summaries that preserve the original text's meaning. A dedicated Tamil preprocessing module handles tokenization, stop word removal, and script normalization. This is crucial because of Tamil’s agglutinative structure and free word order. A custom stopword list ensures that only meaningful words are used for scoring sentence importance.

A diagram of a diagram

AI-generated content may be incorrect.Fig 1. Proposed solution

For summarization, the system represents sentences as nodes in a graph, with edges based on sentence similarity (TextRank). Sentences are ranked by centrality, and TF-IDF scores are used to further identify key content. The most important sentences are then selected to form a coherent summary.A post-processing phase corrects punctuation, spacing, and script formatting to enhance readability. The summarized output is presented through a user-friendly interface that supports Tamil input (typing, pasting, uploading) and offers features like copying, downloading, or sharing results.

This lightweight yet effective system supports educational, professional, and content curation use cases—improving accessibility to Tamil content while promoting digital inclusion in low- resource language contexts.

Results And Discussion:

The Tamil Text Summarization system was evaluated using two extractive techniques: frequency-based summarization and the TextRank algorithm. Both were implemented in Python using NLTK, with Tamil-specific preprocessing for tokenization, stopword removal, and normalization. The aim was to test how well lightweight methods could perform without relying on large datasets or pretrained models.The frequency-based method scored sentences based on word frequencies. It performed reasonably well for short texts with repeated key terms but lacked coherence and often favored high-frequency but less meaningful content. This approach was less effective for longer documents with diverse vocabulary.

TextRank, using a graph-based model where sentences were ranked by centrality, produced more coherent and contextually relevant summaries. It better captured sentence relationships and delivered balanced outputs, although it still lacked deep semantic understanding.


A screenshot of a computer  AI-generated content may be incorrect.A screenshot of a computer  AI-generated content may be incorrect.

Fig 2. User input interface of the Tamil Text Summarization system

Preprocessing played a critical role in both methods. Custom Tamil stopwords and adapted tokenization significantly improved summary quality by cleaning and normalizing the input. These steps helped overcome the absence of Tamil-specific pretrained models.

As there were no standard reference summaries in Tamil, manual evaluation was used. Human feedback confirmed TextRank’s superiority in producing more informative and readable summaries. The system was efficient, processing inputs in under two seconds, and proved suitable for real-world texts like news, educational content, and official documents. However, for longer inputs, performance declined slightly, indicating future scope for chunk-wise

processing.Overall, the results show that with proper preprocessing and smart algorithm choices, extractive summarization for Tamil is both feasible and impactful.

While traditional methods performed well, the study also highlighted their limitations in capturing deeper semantics—pointing toward future improvements through hybrid or abstractive models.

Conclusion

The Tamil Text Summarization system developed in this project addresses the

need for automated summarization tools for low-resource languages like Tamil. Using extractive methods such as word frequency analysis and the TextRank algorithm, the system operates without relying on large, annotated corpora or deep learning models. Custom preprocessing—including Tamil-specific tokenization, stop word removal, and normalization—enabled effective handling of the language's morphological complexity.The summarizer generates concise, readable summaries that retain semantic accuracy, as confirmed by human evaluation. While TextRank offered better coherence, frequency-based methods provided a simpler alternative. This project highlights that meaningful summarization is possible through thoughtful algorithm design, even in resource-constrained settings. Beyond its technical contributions, the system supports improved content accessibility in Tamil for students, researchers, and the public. Its applications span news, education, and information retrieval. Future work includes integrating abstractive transformer models, handling multi- document inputs, and extending support to other Indian languages—laying a foundation for inclusive and scalable multilingual NLP tools.

Acknowledgement: This research work is supported by MUTHIRAI - A Global Research Center for Tamil and AI, Thiagarajar College of Engineering, Madurai, India.

References

  1. T. Priyadharshan and S. Sumathipala, "Text summarization for Tamil online sports news using NLP," in Proc. 3rd Int. Conf. on Information Technology Research (ICITR), Dec. 2018,

  2. pp. 1–5. IEEE.

  3. A. Kumar, "A Comparative Analysis Of Different Text Summarizers," IJRAR-International Journal of Research and Analytical Reviews (IJRAR), vol. 5, no. 4, pp. 610–613, 2018.

  4. S. S. Mohamed and S. Hariharan, "An investigation on graphical approach for Tamil text summary generation," in Proc. Int. Conf. on Intelligent Computing and Control (I2C2), Jun. 2017, pp. 1–5. IEEE.

  5. P. Janjanam and C. P. Reddy, "Text summarization: an essential study," in Proc. Int. Conf. on Computational Intelligence in Data Science (ICCIDS), Feb. 2019, pp. 1–6. IEEE.

  6. S. Adhikari, "NLP based machine learning approaches for text summarization," in Proc. 4th Int. Conf. on Computing Methodologies and Communication (ICCMC), Mar. 2020, pp. 535–538. IEEE.

  7. M. Banu, C. Karthika, P. Sudarmani, and T. V. Geetha, "Tamil document summarization using semantic graph method," in Proc. Int. Conf. on Computational Intelligence and Multimedia Applications (ICCIMA), vol. 2, Dec. 2007, pp. 128–134. IEEE.

  8. S. S. Mohamed and S. Hariharan, "An investigation on graphical approach for Tamil text summary generation," in Proc. Int. Conf. on Intelligent Computing and Control (I2C2), Jun. 2017, pp.


Author
கட்டுரையாளர்

Srimathi R, Premlatha K R, Abirami A M, Lohitha K, Pratika Lakshmi L G

Department of Information Technology

Thiagarajar College of Engineering, Madurai- 625015