Abstract
AgenticGen improves advertising video generation by decomposing it into strategy selection and draft generation stages supervised by online business feedback and human quality rewards.
Advertising video generation is not only a video synthesis task, but also a product-conditioned reasoning problem whose success is measured by online business metrics. Recent video foundation models can generate realistic clips from multimodal conditions, yet they do not optimize how a product should be transformed into an effective advertisement or how future generation should be improved from online business feedback. To close this loop, we propose AgenticGen, a reward-guided agentic framework that decomposes advertising video generation into two trainable reasoning stages, strategy selection and draft generation, thereby exposing optimization targets that online business feedback can supervise. AgenticGen learns a performance-based reward from accumulated online feedback and a complementary rubric-based reward aligned with human quality standards, then uses them to supervise policy optimization. DPO first moves the agentic policies toward online preferences, and GRPO further refines both stages with process and outcome rewards. Offline experiments validate the reward models and successive policy optimization. Online A/B experiments in the TikTok advertising system show that AgenticGen after DPO and GRPO improves CTR by 2.72%, CVR by 2.63%, and Advv by 9.61% over the SFT baseline.
Community
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- VA-Judger: Reward Modeling from Human Preference Feedback for Joint Video-Audio Generation (2026)
- RubricRM: Generative Reward Modeling via Dynamic Rubrics for Image Generation and Editing (2026)
- WorldReward: Reward Modeling for Camera-Conditioned World Models (2026)
- Beyond Action Imitation: Learning a Decision-Aware User Simulator for Online Advertising (2026)
- What to Edit Next: Visually Aligned Image-Editing Follow-Up Suggestions in Conversational Systems (2026)
- Crayotter: Learning Long-Horizon Video Editing Agents via Group-Relative Preference Backpropagation (2026)
- FRAMEWORKERS: A Dynamic Multi-Agent Framework for AI-Generated Video Production (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper