UniEvo-VL: Self-Distillation Training for Multimodal Model Self-Improvement
7 hours ago
- Introduces UniEvo-VL, a self-evolving framework for multimodal models to learn from self-correction feedback during test-time compute.
- Uses on-policy self-distillation without a separate teacher; the model acts as both teacher and student with different contexts.
- Student sees vanilla question, teacher conditions on privileged critique; training minimizes divergence between denoising diffusion distributions over student sampling trajectories.
- Improves image generation capabilities while maintaining sensitivity to reflection information.
- Built on Qwen-image-2512, achieving performance gains from 0.747 to 0.808 on GenEval and from 32.97 to 35.53 on GenEval2 Soft-TIFA.
- More powerful external critics (e.g., GPT5.6-Luna) suggest a higher self-evolving ceiling for models with strong judge capabilities.
- Self-improvements may not be uniform across different tasks, as shown by mixed text-rendering outcomes.