Skip to content
Not available in this workspace
OpenRouterOpenRouter
© 2026 OpenRouter, Inc

Product

  • Chat
  • Rankings
  • Benchmarks
  • Apps
  • Discover
  • Models
  • Collections
  • Providers
  • Pricing
  • Enterprise
  • Labs

Company

  • About
  • Blog
  • Careers
    Hiring
  • Privacy
  • Terms of Service
  • Support
  • Works With OR
  • Data

Developer

  • Documentation
  • API Reference
  • Developer Platform
  • Status

Connect

  • Discord
  • GitHub
  • LinkedIn
  • X
  • YouTube

Note: The free trial period for OpenClaw users of MiMo v2 Pro and Omni will end on Thursday April 2 at 12:00PM EST / 5:00PM UTC

Favicon for xiaomi

Xiaomi: MiMo-V2-Omni

xiaomi/mimo-v2-omni

MiMo-V2-Omni is a frontier omni-modal model that natively processes image, video, and audio inputs within a unified architecture. It combines strong multimodal perception with agentic capability - visual grounding, multi-step planning, tool use, and code execution - making it well-suited for complex real-world tasks that span modalities, 256K context window.

Modalities

Context

262K

Released

Mar 18, 2026

ActivityFAQExplore

Activity

Token volume and request traffic to this model over time.

Explore models like this

AI Models with Vision: Multimodal LLMs for Image UnderstandingCollectionAI Model RankingsRanking

About Xiaomi: MiMo-V2-Omni

OpenRouter makes Xiaomi: MiMo-V2-Omni available through a unified, OpenAI-compatible API using the model ID xiaomi/mimo-v2-omni.

Xiaomi: MiMo-V2-Omni accepts text, audio, images and video and returns text. It has a 262,144-token context window.

It was released on March 18, 2026.

More models from Xiaomi

  • MiMo-V2.5-Pro
  • MiMo-V2.5

Frequently asked questions

MiMo-V2-Omni is a frontier omni-modal model that natively processes image, video, and audio inputs within a unified architecture. It combines strong multimodal perception with agentic capability - visual grounding, multi-step planning, tool use, and code execution - making it well-suited for complex real-world tasks that span modalities, 256K context window.

MiMo-V2-Omni has a 262,144 token context window.

MiMo-V2-Omni accepts text, audio, images and video as input and returns text.

MiMo-V2.5-Pro and MiMo-V2.5 are other text models from Xiaomi.

MiMo-V2-Omni was released on March 18, 2026.