Which evaluation metric is standardly used to evaluate machine translation and natural...
Nopal Securities technical mcq question, verified with a worked answer. Free to practise - no sign-up.
Which evaluation metric is standardly used to evaluate machine translation and natural language generation models?
Show answer & explanation
Answer: A. BLEU score
Bilingual Evaluation Understudy (BLEU) score measures n-gram overlap between machine-generated and reference texts.
Step-by-step Derivation:
Step 1: BLEU is the standard benchmark metric for text generation and machine translation.