RogueGPT: A Controlled Stimulus Generation Framework for News Authenticity Research
- Alexander Loth ,
- Martin Kappes ,
- Marc-Oliver Pahl
Journal of Open Source Software |
RogueGPT is an open-source Python framework for the controlled, reproducible generation
and curation of multilingual news fragments for AI authenticity research. It enables researchers
to systematically produce synthetic news stimuli across a wide range of large language model
(LLM) families, journalistic styles, languages, and content formats, while storing every fragment
alongside complete generation provenance in a MongoDB corpus.
The framework follows a three-layer architecture that strictly separates the data logic (core.py)
from user-facing interfaces: a Streamlit web application, a command-line interface (CLI), and
a Model Context Protocol (MCP) server for AI-agent integration. All three interfaces share a
single validation and normalisation layer, ensuring that every fragment — whether ingested
manually, via automated generation scripts, or through an AI agent — conforms to the same
schema. Each machine-generated fragment records the model identifier, the full prompt, the
sampling parameters used by the batch generator, and the language, format and seed phrase
where the interface supplies them. Fragments created through the web interface record model
and prompt but not sampling settings, because that path delegates them to the provider
defaults; the stored record therefore states what was fixed rather than implying that every
generation setting is recoverable.
The current corpus contains 3,278 multilingual news fragments. Of these, 2,638 are machine
generated by 10 models across 6 providers (OpenAI, Google, Meta, Anthropic, Mistral, and
Microsoft), covering four languages (English, German, French, and Spanish), three content
formats (tweet, headline, and short article), and five journalistic styles per language. The
remaining 640 fragments are human-sourced — both legitimate news and authentic fake-news
material — and serve as experimental anchors for perception studies.
RogueGPT is the upstream stimulus generation component of a three-tool research pipeline:
CRED-1 (Loth, 2026a) identifies unreliable news sources, RogueGPT generates controlled
stimuli, and JudgeGPT delivers them to human participants for perception measurement.