Multimodal · 2 pieces on file
Multimodal
Vision, audio, video, and the messy edges where modalities meet. Charts and documents the models still fail on.
Feature · SEPTEMBER 19, 2026
Qwen3.8-Omni-Flash lands with 1M-token context and a 98% cheaper audio bill
Alibaba's new native omni-modal model reads text, image, audio, and video in a single call, calls tools through function calling, and prices an hour of audio input at roughly $0.0038 — but it answers only in text.
More in Multimodal
-
MAY 19, 2026
Google Gemini Omni: world-understanding multimodal at scale, any-input-to-any-output
Announced at Google I/O on May 19, Gemini Omni is positioned as a leap in world understanding, multimodality, and editing — generating any output from any input, starting with video.