Best Multimodal AI Models

12 models tracked · 60 recent news stories

🏆

Most-talked-about Multimodal right now

Ranked by mentions across 30+ AI sources in July 2026.

AM
Amazon Nova

Amazon Nova — Amazon's family of foundation models on AWS Bedrock, spanning text, vision and the Nova Act agentic browser model.

CO
Cosmos

NVIDIA Cosmos — world foundation models that generate physics-aware synthetic data and reasoning for physical AI and robotics.

GR
GR00T

NVIDIA's foundation model for humanoid robots (Isaac GR00T), enabling generalist embodied skills.

GE
Gemini
GE
Gemini 3.1 Pro

Gemini 3.1 Pro — Google's frontier Gemini model for complex reasoning tasks, benchmarked against GPT-5.4, Claude Opus 4.6 and Grok.

GE
Gemini Robotics-ER

Google DeepMind's embodied-reasoning Gemini model for real-world robotics tasks.

Gemma 4

Google DeepMind's most capable open model family. Available in 4 sizes (E2B, E4B, 26B MoE, 31B Dense) with advanced reasoning, agentic workflows, vision, audio, 256K context, 140+ languages. Apache 2.0 license. Runs on devices from phones to H100 GPUs.

IN
Inkling

Inkling — Thinking Machines Lab's first model trained from scratch: a 975B-parameter open-weights multimodal Mixture-of-Experts with 41B active parameters and controllable thinking effort, released under Apache 2.0 in July 2026.

LA
Lance

ByteDance's unified model for image and video understanding, generation and editing.

MI
Mistral OCR 4

Mistral AI's document-intelligence (OCR) model (June 2026). Structure-aware extraction with bounding boxes, block classification and confidence scores across 170 languages; self-hostable in a single container. Tops OlmOCRBench; feeds RAG, agentic and enterprise-search pipelines.

ST
Step 3.7

Step 3.7 — StepFun's enterprise-ready multimodal model (including the Step 3.7 Flash tier), optimised for NVIDIA GPU inference.

Π0
π0

Physical Intelligence's Vision-Language-Action (VLA) models for general robot control (π0, π0-FAST, π0.6).

📰 Latest Multimodal Model News(60 stories)