Camera feed into a VLM
Time: 9:55 AM to 10:55 AM
The architecture
Program 1: Test the vision API
Start simple — capture one photo, send it to GPT, and print what it sees.Step 1 — Capture a frame as base64
Thepicamera2 library captures JPEG frames. Convert to base64 so it can be included in the API request:
Step 2 — Send image + text to the vision API
The key difference from Day 3: thecontent field is now a list containing both text and image data:
content field changed from a simple string to a list with two items: the text prompt and the image. The detail: "low" setting keeps token costs down.
Step 3 — Interactive loop
After the first description, let the user ask follow-up questions about the same image or snap a new one:Run it
Click to see the complete test_vision.py program
Click to see the complete test_vision.py program
Program 2: Vision chatbot
Now add voice and TTS. The robot captures a photo with every question, sends both the image and your spoken question to GPT, and speaks the answer.Step 1 — Combine camera + voice input
The key difference from Day 3’s voice chatbot: every API call now includes a fresh camera frame alongside the text:Step 2 — The conversation loop with camera
Every iteration: listen → capture photo → send both to GPT → speak response:Run it
Click to see the complete vision_chatbot.py program
Click to see the complete vision_chatbot.py program
After this section, take a 10-minute break (10:55 AM to 11:05 AM).

