30.1 GB
350 files
Updated 3 months ago
Name
Size
馬鹿.svg.png71.1 kB
xet
野原新之助 - 維基百科,自由的百科全書.html632 kB
xet
螢幕擷取畫面_4-7-2026_112349_zh.wikipedia.org.jpeg333 kB
xet
螢幕擷取畫面_30-6-2026_215451_zh.wikipedia.org.jpeg572 kB
xet
螢幕擷取畫面_25-6-2026_172559_zh.wikipedia.org.jpeg40.5 kB
xet
螢幕擷取畫面_25-6-2026_10315_en.wikipedia.org.jpeg175 kB
xet
螢幕擷取畫面_25-6-2026_103046_en.wikipedia.org.jpeg27.2 kB
xet
螢幕擷取畫面_21-6-2026_193915_talesrunner.funtown.com.hk.jpeg10.9 kB
xet
螢幕擷取畫面_2-7-2026_2003_zh.wikipedia.org.jpeg1.04 MB
xet
螢幕擷取畫面_2-7-2026_194910_zh.wikipedia.org.jpeg78.7 kB
xet
螢幕擷取畫面_12-6-2026_20247_hkrcpaspas-my.sharepoint.com.jpeg243 kB
xet
第一次鴉片戰爭 - 維基百科,自由的百科全書.html1.69 MB
xet
竈門禰豆子 - 維基百科,自由的百科全書.html604 kB
xet
福克蘭戰爭 - 維基百科,自由的百科全書.html1.02 MB
xet
柯南M14海报.jpg169 kB
xet
七年戰爭 - 維基百科,自由的百科全書.html1.03 MB
xet
Đối_-_Tết_2009.jpg1.73 MB
xet
yourshadow.png346 kB
xet
youkilledkenny_resize.webp67.8 kB
xet
yellowwort.entity194 Bytes
xet
x64_SamsungBrowserSetup.exe7.57 MB
xet
what_do_you_want_what_are_you_doing_2_the_winning_side_s3.webp34.4 kB
xet
vocab.json2.78 MB
xet
vit_config.json205 Bytes
xet
unnamed.jpg100 kB
xet
unnamed (2).jpg153 kB
xet
unnamed (1).jpg157 kB
xet
tokenizer_config.json7.31 kB
xet
tokenizer.json7.03 MB
xet
tmpirr2rtgr.mp4926 kB
xet
tmpbovn_7gs.mp41.62 MB
xet
tmp8flc9z66.mp41.12 MB
xet
tianjin_japanese_ink_20260531_120501.png2.01 MB
xet
tennessee_map_575.webp29 kB
xet
stoke_mandeville_japanese_ink_20260606_040308.png4.99 MB
xet
squiddleterrors.webp28.1 kB
xet
spacerace.webp196 kB
xet
sochi_japanese_ink_20260531_121359.png4.33 MB
xet
scplogo4.webp14.5 kB
xet
sakura_and_nozaki.webp192 kB
xet
remote_body.webp22.8 kB
xet
red_herring1.webp9.27 kB
xet
pyeongchang_japanese_ink_20260531_121640.png2.46 MB
xet
preprocessor_config.json392 Bytes
xet
osaka_japanese_ink_20260531_113951.png7.58 MB
xet
oc7dpswa.mp3507 kB
xet
mqdefault.webp7.38 kB
xet
moonlandings_2.webp43.5 kB
xet
model.safetensors.index.json123 kB
xet
merges.txt1.67 MB
xet
manshaped_rock.webp7.28 kB
xet
llm_config.json663 Bytes
xet
lausanne_blueprint_20260531_114921.png10.4 MB
xet
kyoto_japanese_ink_20260531_114133.png10.4 MB
xet
kabuki.webp389 kB
xet
ju_on.webp11.5 kB
xet
infidelity_takingchances.webp18.5 kB
xet
img_8269_6.webp71 kB
xet
img_1193_3.webp25.5 kB
xet
image.jpg17 kB
xet
image - 2026-05-31T173013.077.webp31.8 kB
xet
image - 2026-05-31T173001.176.webp31.7 kB
xet
image (9).png815 kB
xet
image (9).jpg38.9 kB
xet
image (8).png15.3 kB
xet
image (8).jpg297 kB
xet
image (7).png9.7 kB
xet
image (7).jpg21.2 kB
xet
image (6).jpg165 kB
xet
image (5).jpg50.7 kB
xet
image (4).png3.48 MB
xet
image (4).jpg43.7 kB
xet
image (3).jpg14.2 kB
xet
image (25).png3.99 MB
xet
image (24).png4.12 MB
xet
image (23).png3.55 MB
xet
image (22).png3.29 MB
xet
image (21).png3.67 MB
xet
image (20).png3.98 MB
xet
image (2).jpg109 kB
xet
image (19).png4.29 MB
xet
image (18).png3.98 MB
xet
image (17).png3.96 MB
xet
image (16).png3.77 MB
xet
image (15).png3.55 MB
xet
image (14).png3.45 MB
xet
image (14).jpg66.2 kB
xet
image (13).png744 kB
xet
image (13).jpg21 kB
xet
image (12).png626 kB
xet
image (12).jpg144 kB
xet
image (11).png885 kB
xet
image (11).jpg123 kB
xet
image (10).png631 kB
xet
image (10).jpg47.1 kB
xet
image (1).jpg126 kB
xet
i2v_20260625_121329_1782382409.mp41.07 MB
xet
i2v_20260621_140946_1782043786.mp4220 kB
xet
i2v_20260606_035257_1780710777.mp4723 kB
xet
i2v_20260530_154451_1780148691.mp4755 kB
xet
README.md

BAGEL

🥯 BAGEL • Unified Model for Multimodal Understanding and Generation

BAGEL Website BAGEL Paper on arXiv Github BAGEL Demo BAGEL Discord

We present BAGEL, an open‑source multimodal foundation model with 7B active parameters (14B total) trained on large‑scale interleaved multimodal data. BAGEL outperforms the current top‑tier open‑source VLMs like Qwen2.5-VL and InternVL-2.5 on standard multimodal understanding leaderboards, and delivers text‑to‑image quality that is competitive with strong specialist generators such as SD3. Moreover, BAGEL demonstrates superior qualitative results in classical image‑editing scenarios than the leading open-source models. More importantly, it extends to free-form visual manipulation, multiview synthesis, and world navigation, capabilities that constitute "world-modeling" tasks beyond the scope of previous image-editing models.

This repository hosts the model weights for BAGEL. For installation, usage instructions, and further documentation, please visit our GitHub repository.

🧠 Method

BAGEL adopts a Mixture-of-Transformer-Experts (MoT) architecture to maximize the model’s capacity to learn from richly diverse multimodal information. Following the same principle of capacity maximization, it utilizes two separate encoders to capture pixel-level and semantic-level features of an image. The overall framework follows a Next Group of Token Prediction paradigm, where the model is trained to predict the next group of language or visual tokens as a compression target.

BAGEL scales MoT’s capacity through Pre-training, Continued Training, and Supervised Finetuning on trillions of interleaved multimodal tokens spanning language, image, video, and web data. It surpasses open models on standard understanding and generation benchmarks and demonstrates advanced in-context multimodal abilities like free-form image editing, future frame prediction, 3D manipulation, world navigation, and sequential reasoning.

🌱 Emerging Properties

As we scale up BAGEL’s pretraining with more multimodal tokens, we observe consistent performance gains across understanding, generation, and editing tasks. Different capabilities emerge at distinct training stages—multimodal understanding and generation appear early, followed by basic editing, while complex, intelligent editing emerges later. This staged progression suggests an emergent pattern, where advanced multimodal reasoning builds on well-formed foundational skills. Ablation studies further show that combining VAE and ViT features significantly improves intelligent editing, underscoring the importance of visual-semantic context in enabling complex multimodal reasoning and further supporting its role in the emergence of advanced capabilities.

📊 Benchmarks

1. Visual Understanding

Model MME ↑ MMBench ↑ MMMU ↑ MM-Vet ↑ MathVista ↑
Janus-Pro-7B - 79.2 41.0 50.0 –
Qwen2.5-VL-7B 2347 83.5 58.6 67.1 68.2
BAGEL 2388 85.0 55.3 67.2 73.1

2. Text-to-Image Generation · GenEval

Model Overall ↑
FLUX-1-dev 0.82
SD3-Medium 0.74
Janus-Pro-7B 0.80
BAGEL 0.88

3. Image Editing

Model GEdit-Bench-EN (SC) ↑ GEdit-Bench-EN (PQ) ↑ GEdit-Bench-EN (O) ↑ IntelligentBench ↑
Step1X-Edit 7.09 6.76 6.70 14.9
Gemini-2-exp. 6.73 6.61 6.32 57.6
BAGEL 7.36 6.83 6.52 44.0
BAGEL+CoT – – – 55.3

License

BAGEL is licensed under the Apache 2.0 license. It is finetuned from Qwen2.5-7B-Instruct and siglip-so400m-14-384-flash-attn2 model, and uses the FLUX.1-schnell VAE model, all under Apache 2.0.

✍️ Citation

@article{deng2025bagel,
  title   = {Emerging Properties in Unified Multimodal Pretraining},
  author  = {Deng, Chaorui and Zhu, Deyao and Li, Kunchang and Gou, Chenhui and Li, Feng and Wang, Zeyu and Zhong, Shu and Yu, Weihao and Nie, Xiaonan and Song, Ziang and Shi, Guang and Fan, Haoqi},
  journal = {arXiv preprint arXiv:2505.14683},
  year    = {2025}
}
Total size
30.1 GB
Files
350
Last updated
Jul 4
Pre-warmed CDN
US EU US EU

Contributors