Digitalni repozitorij raziskovalnih organizacij Slovenije

Izpis gradiva
A+ | A- | Pomoč | SLO | ENG

Naslov:Comparative analysis of machine translation for Hindi-Dogri text using rule-based, statistical, and neural approaches
Avtorji:ID Kumar, Joginder (Avtor)
ID Rakhra, Manik (Avtor)
ID Dubey, Preeti (Avtor)
ID Prashar, Deepak (Avtor)
ID Mršić, Leo (Avtor)
ID Khan, Arfat Ahmad (Avtor)
ID Kadry, Seifedine (Avtor)
ID Kim, Jungeun (Avtor)
Datoteke:.pdf PDF - Predstavitvena datoteka, prenos (9,80 MB)
MD5: 822BBF2E1AFE80267A1BA823DE2ACD0C
 
URL URL - Izvorni URL, za dostop obiščite https://doi.org/10.7717/peerj-cs.3218
 
Jezik:Angleški jezik
Tipologija:1.01 - Izvirni znanstveni članek
Organizacija:Logo RUDOLFOVO - Rudolfovo – Znanstveno in tehnološko središče Novo mesto
Povzetek:Machine translation has made significant progress in several Indian languages; however, some, known as computationally low-resourced languages, have seen very little work in this field. The Dogri language, which is listed in the 8th Schedule of the Indian Constitution, is one such language. The authors have developed a machine translation system for the Hindi-Dogri language pair in the fixed news domain using three approaches: rule-based machine translation (developed using linguistic rules), statistical machine translation (built using the Moses toolkit), and neural machine translation (developed using neural networks). A comparison of all three approaches is presented in this article. The article also discusses various research challenges identified in each approach used for machine translation. A corpus of approximately 0.1 million sentences in the news domain was used to train the corpus-based statistical machine translation (SMT) and neural machine translation (NMT) models. The authors also addressed whether NMT produces results equivalent to or better than those of SMT and rule-based machine translation (RBMT). To ensure a comprehensive evaluation, the outputs of all systems were evaluated using two approaches: manual evaluation by language experts and automatic evaluation using standard metrics—Bilingual Evaluation Understudy (BLEU), TER (Translation Edit Rate), METEOR (Metric for Evaluation of Translation with Explicit Ordering), and WER (Word Error Rate). Although RBMT achieved the highest overall scores in both automatic and manual evaluations, expert analysis revealed that translations produced by NMT and SMT exhibited less ambiguity. The study concludes that the performance of SMT and NMT systems are likely to improve further with the availability of larger bilingual parallel corpora.
Ključne besede:Machine translation, Hindi-Dogri language pair, Low-resourced languages, Neural machine translation (NMT), Statistical machine translation (SMT), Rule-based machine translation (RBMT), test mining, algorithms, artificial intelligence, natural language and speech
Status publikacije:Objavljeno
Verzija publikacije:Objavljena publikacija
Datum objave:15.10.2025
Leto izida:2025
Št. strani:23 str.
PID:20.500.12556/DiRROS-24570 Novo okno
UDK:004.9
ISSN pri članku:2376-5992
DOI:10.7717/peerj-cs.3218 Novo okno
COBISS.SI-ID:260288515 Novo okno
Avtorske pravice:Copyright 2025 Kumar et al
Opomba:Nasl. z nasl. zaslona; Soavtorji: Manik Rakhra, Preeti Dubey, Deepak Prashar, Leo Mrsic, Arfat Ahmad Khan, Seifedine Kadry and Jungeun Kim; Opis vira dne 6. 12. 2025;
Datum objave v DiRROS:01.09.2026
Število ogledov:47
Število prenosov:22
Metapodatki:XML DC-XML DC-RDF
:
Kopiraj citat
  
Objavi na:Bookmark and Share


Postavite miškin kazalec na naslov za izpis povzetka. Klik na naslov izpiše podrobnosti ali sproži prenos.

Gradivo je del revije

Naslov:PeerJ computer science
Založnik:PeerJ
ISSN:2376-5992
COBISS.SI-ID:21549320 Novo okno

Licence

Licenca:CC BY 4.0, Creative Commons Priznanje avtorstva 4.0 Mednarodna
Povezava:http://creativecommons.org/licenses/by/4.0/deed.sl
Opis:To je standardna licenca Creative Commons, ki daje uporabnikom največ možnosti za nadaljnjo uporabo dela, pri čemer morajo navesti avtorja.

Sekundarni jezik

Jezik:Slovenski jezik
Ključne besede:strojno prevajanje, jezikovni par hindijščina–dogri, jeziki z malo viri, nevronsko strojno prevajanje, rudarjenje besedil, umetna inteligenca


Nazaj