AIExplainer

What is a BLEURT?

A metric used to evaluate the quality of text generated by language models

Stands for: Bilingual Evaluation Underground Test

BLEURT is a way to measure how well a language model can generate human-like text. It compares the generated text to a reference text and gives a score based on how similar they are.

BLEURT is like a teacher grading a student's essay. The teacher compares the student's essay to a perfect essay and gives a grade based on how well the student's essay matches the perfect one.

A company developing a chatbot might use BLEURT to evaluate the chatbot's responses to user queries and adjust the model to generate more natural-sounding text.

BLEURT is used by researchers and developers to evaluate and improve the performance of language models, such as chatbots and language translation systems.

BLEURT is not a measure of the truth or accuracy of the generated text, but rather how well it mimics human language.

BLEURT was developed by researchers at Google in 2020 as a more robust alternative to existing metrics for evaluating language models.

ROUGE BLEU METEOR

Three products for different needs — explore what’s relevant to you.