You Won't Believe This Click

Content Rewriting for Agentic Choice

Candidate Set
Original presentation
  1. A.Marshawn playing in charity soccer game went exactly as you'd expect.AbstractIf there was ever a sport-athlete combination that we'd never expect to work out, it'd be Marshawn Lynch dabbling in soccer.
  2. B.Sofia Vergara and Joe Manganiello Celebrate 4-Year Wedding Anniversary: 'Mi Amor!'AbstractSofia Vergara and Joe Manganiello Celebrate 4-Year Wedding Anniversary
  3. C.The Coolest And Craziest McDonald's Across The CountryAbstractSometimes, the Golden Arches know how to pull out all the stops.
Target snippet · A
Original → Rewritten

Marshawn playing in charity soccer game went exactly as you'd expect.

When Marshawn Lynch Took the Pitch: An Inside Look

Can changing only one item's presentation change the chooser's decision?

A list of competing documents is shown to the target agent. We choose one target document from the list and generate a rewriting strategy for only that document. A separate rewriting model then revises the target document's title and abstract, while all other documents in the list remain exactly the same. The target agent selects from this updated list, and whether the rewritten target document is selected is used to train the advisor.

Abstract

Language models are increasingly used as agents to help humans decide what information is surfaced. This usage incentivizes content creators to optimize content in ways that appeal not only to humans, but also to agents that mediate access to them. In this paper, we study selection shifts induced by rewriting in agentic decision-making. Given a set of competing content snippets, we rewrite only one snippet while leaving the rest unchanged, and then measure how it impacts the agent's choice.

We operationalize this setup with AgentBait, an advisor-rewriter framework in which the advisor learns to propose rewriting strategies and the rewriter revises the snippet. While a rewriter with a fixed prompt improves target snippet selection from 17.1% to 34.8%, our AgentBait raises its selection to 98.5%. We further show that the advisor trained with AgentBait effectively transfers to setups with different agents, languages, and snippets in other domains (e.g., scientific papers).

However, higher target selection can reflect unsupported rewrites rather than better content. Adding a reward for support from the original snippet redirects the advisor toward more supported rewriting strategies, revealing a trade-off between factuality and target selection. Together, our results show that once agents mediate access to information, content can be rewritten to be chosen by the agent, even when selection and usefulness diverge.

How presentation becomes a decision signal

Agent-mediated selection creates a direct optimization pressure on content presentation.

01Sensitivity

Presentation shifts choice

17.1%Original34.8%Prompt rewrite

02Optimization

Agent feedback makes the pressure learnable

34.8%Prompt rewrite98.5%Trained advisor

03Shortcut

Selection-only rewards discover unsupported shortcuts

98.5%Selected2.0%Supported

04Source support

Source support changes what the optimizer learns

68.6%SelectedUnsupported technical substitution96.2%0.1%

Same slate. One rewrite.
Can the decision change?

The candidate list has already been constructed. Only the target presentation may change.

  1. 01Read
  2. 02Edit
  3. 03Select
01 · Target inputTrained

Advisor

Receives only the extracted target document and proposes a rewriting strategy.

Advisor input · target only

BTarget document

Title + abstract

Advisor suggests“Try a sharper, more specific framing.”“Push the hook further, but keep it specific.”

02 · EditFrozen

Frozen Rewriter

Applies the strategy to the target title and abstract only.

Original titleRewritten title

A study of news recommendation

What Makes a Model Choose This?

Original abstractRewritten abstract

We study how language models choose among a fixed slate of news candidates.

A controlled rewrite reveals which presentation cues redirect the same chooser.

03 · SelectFrozen

Target Agent

Selects from the same candidate identities and order, with only the target rewritten.

  1. ACandidate A
  2. CCandidate C
Rewardr

Selection outcome defines the rewardGRPO updates the advisor policy.

Figure 2 | AgentBait system schematic. The advisor and frozen rewriter receive only the extracted target document; only the target agent sees the full fixed slate. Policy: Qwen3.5-9B; frozen rewriter: GPT-5-mini; objective: GRPO selection reward, optionally augmented with MiniCheck sentence support.

Same source. Different rewards.
Different strategies.

Selection-only optimization introduces an AI and sensor-fusion mechanism absent from the source. Support-aware optimization instead reframes source-supported operations and stakes.

A · UnconstrainedTechnical authority · Novelty

AI-Driven Runway Scheduling: How Sensor Fusion and ML Boosted Haneda's 85.6% On-Time Rate

Abstract

Tokyo International Airport achieved an 85.6% on-time performance in 2018, which the rewrite attributes to a proprietary AI-based predictive maintenance and dynamic scheduling system. It describes sensor fusion, delay forecasting, and reinforcement-learning runway scheduling as mechanisms behind the reported punctuality.

MiniCheck support ↑
0.014
Worst-sentence ↑
0.006
B · Support-awareOperational puzzle · Stakes

How Tokyo's Haneda Beats the Odds: Inside the Operations That Deliver 85.6% On-Time Flights

Abstract

Tokyo International Airport is the world's fifth-busiest airport, yet in 2018 it achieved an 85.6% on-time rate. This feature probes the paradox: what management choices, scheduling practices, ground operations, and airport-airline coordination let Haneda run so punctually at massive scale?

MiniCheck support ↑
0.623
Worst-sentence ↑
0.051
Figure 3 | Haneda qualitative comparison. The list is fixed and only the target text changes. Model: GPT-5-mini target agent; metrics: target selection and MiniCheck support; example n=1.

Table 1 · Target-agent transfer

Optimization amplifies the presentation effect

Metric Target selected (%) ↑

Prompt-only rewriting raises target selection from 17.1% to 34.8%, while the RL-trained advisor reaches 98.5%. Gains remain positive across all held-out target agents, although transfer strength varies.

ConditionMethodGPT-5-minitrain targetGPT-5.5Gemini 3 FlashGemini 3.1 ProSonnet 4.6Opus 4.8
ReferenceOriginal text17.117.117.117.317.617.5
Prompting onlyRewriter only34.8+17.724.1+7.040.7+23.620.2+2.924.4+6.841.0+23.5
Prompting onlyAdvisor + rewriter43.9+26.839.4+22.348.8+31.727.0+9.734.2+16.651.5+34.0
RL-trained rewriterRewriter only95.9+78.885.3+68.297.4+80.354.8+37.549.2+31.675.9+58.4
RL-trained rewriter+ MiniCheck43.4+26.326.3+9.239.7+22.618.6+1.324.1+6.543.9+26.4
RL-trained advisorAdvisor + rewriter98.5+81.493.3+76.298.9+81.878.7+61.465.0+47.494.1+76.6
RL-trained advisor+ MiniCheck68.6+51.553.2+36.165.8+48.737.5+20.241.0+23.468.8+51.3
Table 1 | Target selection on 1,000 unseen MIND-English news impressions. GPT-5-mini is the training target and frozen rewriter; Qwen3.5-9B is the advisor policy. Row-wise chance is 16.9%. Small values are percentage-point changes from original text for the same evaluator. Target documents do not appear in training; transfer columns use no additional training.

Table 2 · Source-support tradeoff

Selection can outrun source support

Metrics Selection and support (0–100) ↑

The unconstrained advisor reaches 98.5% target selection with only 2.0% MiniCheck support. Adding a source-support reward partially recovers support while reducing selection.

ConditionTarget selected (%)MiniCheck support (%)Constraint
Prompt / rewriter34.860.7None
Prompt / advisor43.942.2None
RL / rewriter95.92.2None
RL / rewriter + MC43.466.7MiniCheck
RL / advisor98.52.0None
RL / advisor + MC68.631.2MiniCheck
Table 2 | Source-support tradeoff on 1,000 unseen MIND-English impressions. Target selection uses the GPT-5-mini chooser; source support is scored with MiniCheck-Flan. The original target-selection baseline is 17.1%.

Transfer across languages,
news datasets and academic documents

The English MIND-trained advisor is evaluated without additional training. Language, dataset and domain shifts are reported separately so that the evidence is not compressed into a single transfer claim.

Table 3 · Language transfer

The learned advisor transfers across languages

Metric Target selected (%) ↑

Trained only on English MIND, the advisor remains above 93% selection in every evaluated language without additional training. Direct rewriter training is less stable, falling to 27.8% in Swahili.

LanguageOriginalPrompt rewriterPrompt advisorRL rewriterRL advisorRL rewriter + MCRL advisor + MC
English (en)17.134.8+17.743.9+26.895.9+78.898.5+81.443.4+26.368.6+51.5
Arabic (ar)16.128.9+12.839.9+23.881.0+64.995.5+79.433.6+17.564.5+48.4
Spanish (es)17.831.2+13.442.8+25.085.8+68.098.0+80.235.7+17.966.5+48.7
Swahili (sw)17.623.8+6.239.4+21.827.8+10.293.7+76.127.2+9.662.2+44.6
Turkish (tr)17.530.3+12.839.2+21.786.7+69.295.8+78.331.1+13.659.9+42.4
Chinese (zh-CN)17.333.2+15.943.4+26.187.1+69.896.6+79.336.8+19.567.4+50.1
Average17.329.5+12.240.9+23.673.7+56.495.9+78.632.9+15.664.1+46.8
Table 3 | Language transfer; paper Table 4. The English row is the training language. Each language uses 1,000 aligned impressions; Arabic, Spanish, Swahili, Turkish and Chinese preserve the same MIND rows, targets, candidate order and slate structure. Advisor: Qwen3.5-9B; frozen rewriter and chooser: GPT-5-mini.

Table 4 · Dataset transfer

The effect extends beyond the training dataset

Metric Target selected (%) ↑

Without additional training, the advisor reaches 98.1% on EB-NeRD English and 98.2% on EB-NeRD Danish. The result therefore extends beyond translation of the original MIND evaluation set.

DatasetLanguageOriginalPrompt rewriterPrompt advisorRL rewriterRL advisorRL rewriter + MCRL advisor + MC
MINDEnglish17.134.8+17.743.9+26.895.9+78.898.5+81.443.4+26.368.6+51.5
MINDDanish16.531.1+14.640.9+24.495.0+78.597.9+81.435.6+19.165.0+48.5
EB-NeRDEnglish12.843.8+31.057.7+44.993.1+80.398.1+85.350.2+37.481.2+68.4
EB-NeRDDanish11.632.2+20.647.8+36.286.1+74.598.2+86.640.4+28.871.7+60.1
Table 4 | News-dataset transfer; paper Table 5. Every setting contains 1,000 impressions and receives no additional training. MIND-Danish changes display language. EB-NeRD is a distinct Danish news dataset; the English version translates the same EB-NeRD slates.

Table 5 · Domain transfer

Cross-domain transfer is harder, but remains substantial

Metric Target selected (%) ↑

On scientific-document selection, the MIND-trained advisor reaches 63.6%, compared with 42.7% for the prompt-only advisor and 32.7% for the RL-trained direct rewriter.

ConditionMethodSelection rateGain over original
ReferenceWithout rewriting9.9
Prompting onlyRewriter-only33.7+23.8 pp
Prompting onlyAdvisor-rewriter42.7+32.8 pp
RL-trained rewriterRewriter-only32.7+22.8 pp
RL-trained rewriter+ MiniCheck24.8+14.9 pp
RL-trained advisorAdvisor-rewriter63.6+53.7 pp
RL-trained advisor+ MiniCheck47.3+37.4 pp
Table 5 | Cross-domain academic transfer; paper Table 6. Evaluation uses 1,000 impressions drawn from SciRepEval-derived scientific documents, with mean / maximum slate sizes of 9.29 / 10. The target agent is GPT-5-mini; all learned policies are trained only on English MIND, with no additional academic-domain training.
@article{jin2026agentbait,
  title   = {You Won't Believe This Click: Content Rewriting for Agentic Choice},
  author  = {Jin, Tianyi and Wang, Zirui and Chan, David M.},
  year    = {2026}
}