AI Model Report

Multimodal · 2 pieces on file

Multimodal

Vision, audio, video, and the messy edges where modalities meet. Charts and documents the models still fail on.


Feature · SEPTEMBER 19, 2026

Qwen3.8-Omni-Flash lands with 1M-token context and a 98% cheaper audio bill

Alibaba's new native omni-modal model reads text, image, audio, and video in a single call, calls tools through function calling, and prices an hour of audio input at roughly $0.0038 — but it answers only in text.

By Lucia Castellan · Multimodal beat

Read the full piece →


More in Multimodal