CultureConverse: A Multilingual Multi-turn Simulation Harness for Culturally Grounded Assistance in East and Southeast Asia
- Bryan Chen Zhengyu Tan ,
- Weihua Zheng ,
- Thong T. Doan ,
- Ngoc Bich Doan ,
- Jia Wang Peh ,
- Xiaoyuan Yi ,
- Jing Yao ,
- Xing Xie ,
- Nancy F. Chen ,
- Zhengyuan Liu ,
- Jinyeong Bak ,
- Wafi Shamdi ,
- Soo-Kai Chie ,
- Liew Yu Siong ,
- Aina Azyyati Binti Mohamad Rezal ,
- Lew Yan Yan Vanessa ,
- Huarui Wu ,
- Dylan Raharja ,
- Nadya Yuki Wangsajaya ,
- Akane Fukushige ,
- Kazushi Kato ,
- K. Inoue ,
- Tatsuya Kawahara ,
- J. Seo ,
- Dongjun Kim ,
- Seungyoon Lee ,
- Z. Pang ,
- Rui Yang Tan ,
- Charibeth Cheng ,
- M. R. J. Estuar ,
- J. Montalan ,
- Duc Minh Pham ,
- Roy Ka-Wei Lee
arXiv
Current cultural evaluations for large language models (LLMs) often reduce culture to single-turn factual recall via MCQs, failing to capture a common use case: users seeking practical help over multiple turns in culturally grounded scenarios. We introduce CultureConverse, a scalable, multilingual simulation and evaluation harness for culturally grounded assistant dialogue that covers 10 East and Southeast Asian regions, 58 subgroup identities, and 7 domains. Each simulated and evaluated episode produces a scored interaction where the assistant assists the user and infers cultural constraints from partial information. The resulting CultureConverse-DS dataset contains 14,610 benchmark (evaluation) episodes and 274,295 oracle-guided (gold-mode) dialogues. In our benchmark evaluation of 18 models, GPT-5 mini achieves the highest assistance quality. Human annotation experiments suggest that our evaluation framework is a sufficient proxy for human judgment. Performance gains from fine-tuning on 27,860 high-quality CultureConverse-DS samples improve in-domain assistance and transfer out-of-domain to cultural MCQ and safety classification benchmarks. We release the harness, both splits, and judge prompts to support interactive evaluation of cultural competency.