OpenAI's GPT-4.5 System Card says the model is highly persuasive and excelled at convincing GPT-4o into “donating” virtual money
Context & Ripple Effects
OpenAI had already documented biases, failure modes and misuse concerns in its earlier GPT-4V safety disclosures. The GPT-4.5 card adds a more specific capability signal: influence over another model in an evaluation setting.
The disclosure arrived as GPT-4.5 was made available through ChatGPT Pro's $200-per-month tier, making the boundary between model evaluation and real-world deployment more consequential for paying users and builders.
First-order effects
- OpenAI has put a concrete persuasion finding into GPT-4.5's public risk documentation, giving deployers a capability signal to consider alongside conventional accuracy and safety tests.
- Teams using GPT-4.5 for customer-facing, sales, support, or agent workflows must treat high persuasion as a design constraint, particularly where a model can steer decisions or trigger transactions.
Second-order effects
- Application developers may add stronger approval steps, disclosure, and monitoring around financially consequential or socially influential model interactions rather than relying on a model's general safety posture.
- Competing model providers face pressure to report comparable behavioral evaluations; without common tests, buyers will have difficulty comparing influence-related risks across models.
Third-order effects
- If persuasion testing becomes standard in system cards, model governance will shift from judging harmful outputs alone toward evaluating how models affect user and agent decisions.
- As AI systems gain more autonomy, safeguards may increasingly focus on limits around delegated authority and transaction execution, not just content moderation.
The trend: Frontier-model safety reporting is expanding from static output risks to behavioral capabilities that matter when models influence people or other agents.