Members-Only
Recent Talks & Demos are for members only
You must be an AI Tinkerers active member to view these talks and demos.
June 27, 2025
·
Seattle
Teaching LLMs to See: Training a Phi-4 × FastViTHD Vision–Language Model (VLM)
Learn to build a vision-language model by merging Phi-4 and FastViT. This talk covers the end-to-end pipeline for training and demos live inference, showing it's achievable on a small budget.
Overview
I’ll walk through how I merged Microsoft’s 2.7 B-parameter text-only Phi-4-mini-reasoning LLM with Apple’s high-speed FastViT-HD image encoder to create Friday-VLM—a finetuned Vision Language Model (VLM) that can caption, reason over, and chat about high-resolution images. I’ll cover the end-to-end pipeline (pre-training, instruction fine-tuning, and image encoding) and demo live inference.
Links
Friday-VLM: PyTorch Phi-4/FastViT VLM for efficient instruction-tuned multimodal learning.
Friday-VLM: multimodal LLM fine-tuned for image-text instruction following.
Tech stack
Compose Email
Sending...
Email preview
Loading recent emails...