Computer Science > Computation and Language

arXiv:2408.06537 (cs)

[Submitted on 13 Aug 2024 (v1), last revised 21 Aug 2024 (this version, v4)]

Title:Introducing the NewsPaLM MBR and QE Dataset: LLM-Generated High-Quality Parallel Data Outperforms Traditional Web-Crawled Data

Authors:Mara Finkelstein, David Vilar, Markus Freitag

Abstract:Recent research in neural machine translation (NMT) has shown that training on high-quality machine-generated data can outperform training on human-generated data. This work accompanies the first-ever release of a LLM-generated, MBR-decoded and QE-reranked dataset with both sentence-level and multi-sentence examples. We perform extensive experiments to demonstrate the quality of our dataset in terms of its downstream impact on NMT model performance. We find that training from scratch on our (machine-generated) dataset outperforms training on the (web-crawled) WMT'23 training dataset (which is 300 times larger), and also outperforms training on the top-quality subset of the WMT'23 training dataset. We also find that performing self-distillation by finetuning the LLM which generated this dataset outperforms the LLM's strong few-shot baseline. These findings corroborate the quality of our dataset, and demonstrate the value of high-quality machine-generated data in improving performance of NMT models.

Subjects:	Computation and Language (cs.CL)
Cite as:	arXiv:2408.06537 [cs.CL]
	(or arXiv:2408.06537v4 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2408.06537

Submission history

From: Mara Finkelstein [view email]
[v1] Tue, 13 Aug 2024 00:06:56 UTC (701 KB)
[v2] Wed, 14 Aug 2024 18:38:11 UTC (701 KB)
[v3] Fri, 16 Aug 2024 18:23:41 UTC (701 KB)
[v4] Wed, 21 Aug 2024 04:03:06 UTC (865 KB)

Computer Science > Computation and Language

Title:Introducing the NewsPaLM MBR and QE Dataset: LLM-Generated High-Quality Parallel Data Outperforms Traditional Web-Crawled Data

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:Introducing the NewsPaLM MBR and QE Dataset: LLM-Generated High-Quality Parallel Data Outperforms Traditional Web-Crawled Data

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators