Staff · 2 pieces on file
Lucia Castellan
Multimodal beat
Lucia Castellan covers vision, audio, and the messy edges where modalities meet. She maintains the desk’s chart-and-document benchmarks and has been running every new vision model against them for years. Her reviews always include at least one image and one document the model failed on.
Beats: multimodal, model-reviews
All pieces by Lucia
-
Multimodal · SEPTEMBER 19, 2026
Qwen3.8-Omni-Flash lands with 1M-token context and a 98% cheaper audio bill
Alibaba's new native omni-modal model reads text, image, audio, and video in a single call, calls tools through function calling, and prices an hour of audio input at roughly $0.0038 — but it answers only in text.
-
Multimodal · MAY 19, 2026
Google Gemini Omni: world-understanding multimodal at scale, any-input-to-any-output
Announced at Google I/O on May 19, Gemini Omni is positioned as a leap in world understanding, multimodality, and editing — generating any output from any input, starting with video.