Evaluating Large Language Models for Translating Humour in English and Urdu
Publisher : PJPCR
Author(s)
Leila M.
Abstract
Humor poses a unique challenge for translation since it relies on cultural knowledge, wordplay and societal norms. This study evaluates the performance of six different LLMs (Deepseek R1, Qwen 3, Llama 3.1, GPT 5.2, Gemini 2.5 Flash and Claude Sonnet 4.5) to translate jokes between English and Urdu. Through mixed-methods analysis including quantitative Likert scale ratings and qualitative thematic coding, we find that closed-source models significantly outperform open-weight models across humor, fluency, and accuracy metrics. Models struggle particularly with the English→Urdu direction, exhibiting mechanical errors such as grammatical mistakes and incoherency. In the Urdu→English direction, models achieve higher accuracy but fail to preserve cultural context. Results indicate that while LLMs show promise for under-resourced language translation, current limitations in capturing linguistic subtlety and cultural nuance present significant obstacles to humor translation.