Is Your Language Model Ready for Monetization Decisions?
- Jialu Gao
2026 ACL 2026 |
Organized by aclanthology.org
Large language models (LLMs) are increas
ingly deployed in monetization-driven systems
such as search engines, advertising platforms,
and e-commerce services, where decision mak
ing is shaped by complex interactions among
user intent, advertiser objectives, and plat
form constraints. Despite rapid progress, exist
ing benchmarks primarily focus on shopping
centric scenarios and user-facing data, captur
ing only a limited subset of real-world mone
tization pipelines and overlooking intermedi
ate decision stages and robustness considera
tions. In this work, we introduce MonBench,
a high-quality multi-task benchmark designed
to evaluate LLMs in realistic monetization con
texts. The benchmark is constructed from large
scale production data collected from multiple
search engines, including both intermediate
candidate pools and user-visible outcomes, bet
ter reflecting the distributional characteristics
of real monetization systems. MonBench cov
ers key capability dimensions such as intent un
derstanding, commercial matching, and user be
havior modeling, and adopts a unified multiple
choice formulation to enable systematic com
parison across models. We further propose a
comprehensive evaluation protocol that mea
sures both performance and robustness. We
evaluate a diverse set of state-of-the-art LLMs
and conduct detailed task-level analyses. Our
results reveal monetization-specific behaviors,
including gaps between relevance optimization
and broader decision-making capabilities, as
well as differences in robustness across model
families. These findings provide new insights
into the strengths and limitations of current
LLMs and highlight the need for richer domain
specific supervision in monetization-oriented
applications.