Digital repository of Slovenian research organisations

Show document
A+ | A- | Help | SLO | ENG

Title:Comparative analysis of machine translation for Hindi-Dogri text using rule-based, statistical, and neural approaches
Authors:ID Kumar, Joginder (Author)
ID Rakhra, Manik (Author)
ID Dubey, Preeti (Author)
ID Prashar, Deepak (Author)
ID Mršić, Leo (Author)
ID Khan, Arfat Ahmad (Author)
ID Kadry, Seifedine (Author)
ID Kim, Jungeun (Author)
Files:.pdf PDF - Presentation file, download (9,80 MB)
MD5: 822BBF2E1AFE80267A1BA823DE2ACD0C
 
URL URL - Source URL, visit https://doi.org/10.7717/peerj-cs.3218
 
Language:English
Typology:1.01 - Original Scientific Article
Organization:Logo RUDOLFOVO - Rudolfovo - Science and Technology Centre Novo Mesto
Abstract:Machine translation has made significant progress in several Indian languages; however, some, known as computationally low-resourced languages, have seen very little work in this field. The Dogri language, which is listed in the 8th Schedule of the Indian Constitution, is one such language. The authors have developed a machine translation system for the Hindi-Dogri language pair in the fixed news domain using three approaches: rule-based machine translation (developed using linguistic rules), statistical machine translation (built using the Moses toolkit), and neural machine translation (developed using neural networks). A comparison of all three approaches is presented in this article. The article also discusses various research challenges identified in each approach used for machine translation. A corpus of approximately 0.1 million sentences in the news domain was used to train the corpus-based statistical machine translation (SMT) and neural machine translation (NMT) models. The authors also addressed whether NMT produces results equivalent to or better than those of SMT and rule-based machine translation (RBMT). To ensure a comprehensive evaluation, the outputs of all systems were evaluated using two approaches: manual evaluation by language experts and automatic evaluation using standard metrics—Bilingual Evaluation Understudy (BLEU), TER (Translation Edit Rate), METEOR (Metric for Evaluation of Translation with Explicit Ordering), and WER (Word Error Rate). Although RBMT achieved the highest overall scores in both automatic and manual evaluations, expert analysis revealed that translations produced by NMT and SMT exhibited less ambiguity. The study concludes that the performance of SMT and NMT systems are likely to improve further with the availability of larger bilingual parallel corpora.
Keywords:Machine translation, Hindi-Dogri language pair, Low-resourced languages, Neural machine translation (NMT), Statistical machine translation (SMT), Rule-based machine translation (RBMT), test mining, algorithms, artificial intelligence, natural language and speech
Publication status:Published
Publication version:Version of Record
Publication date:15.10.2025
Year of publishing:2025
Number of pages:23 str.
PID:20.500.12556/DiRROS-24570 New window
UDC:004.9
ISSN on article:2376-5992
DOI:10.7717/peerj-cs.3218 New window
COBISS.SI-ID:260288515 New window
Copyright:Copyright 2025 Kumar et al
Note:Nasl. z nasl. zaslona; Soavtorji: Manik Rakhra, Preeti Dubey, Deepak Prashar, Leo Mrsic, Arfat Ahmad Khan, Seifedine Kadry and Jungeun Kim; Opis vira dne 6. 12. 2025;
Publication date in DiRROS:01.09.2026
Views:51
Downloads:22
Metadata:XML DC-XML DC-RDF
:
Copy citation
  
Share:Bookmark and Share


Hover the mouse pointer over a document title to show the abstract or click on the title to get all document metadata.

Record is a part of a journal

Title:PeerJ computer science
Publisher:PeerJ
ISSN:2376-5992
COBISS.SI-ID:21549320 New window

Licences

License:CC BY 4.0, Creative Commons Attribution 4.0 International
Link:http://creativecommons.org/licenses/by/4.0/
Description:This is the standard Creative Commons license that gives others maximum freedom to do what they want with the work as long as they credit the author.

Secondary language

Language:Slovenian
Keywords:strojno prevajanje, jezikovni par hindijščina–dogri, jeziki z malo viri, nevronsko strojno prevajanje, rudarjenje besedil, umetna inteligenca


Back