Mage-VL
3 August 2026by microsoft
Mage-VL runs Microsoft's 4B parameter multimodal model directly in a Gradio interface, handling both image and video understanding through codec-native processing. Builders working on pipelines that need a single model to reason across static and moving visuals can use this space to test how the model handles video inputs without requiring separate preprocessing steps.
Try it on Hugging FaceFrom Hugging Face
Codec-native video & image understanding with Mage-VL 4B