• VKMO AI
    VKMO AI
  • Search
  • Explore
  • AI Promos Codes
  • Prompt Library
  • AI Models
  • Submit AI Tool
Categories
  • AI API
  • AI Music Generation
  • AI Data
  • AI Writer
  • AI Image Generator
  • AI Video Generator
  • AI Logo Generator
  • AI Ecommerce
  • AI Study
  • AI Chat
  • AI Voice Generator
  • AI Anime Generator
  • AI Agent
  • AI Coding Tools
  • AI Games
SearchExploreAI Promos CodesPrompt LibraryAI ModelsSubmit AI Tool

VKMO AI is a premium AI tools directory that helps users discover the best AI products worldwide.

Categories
AI APIAI Music GenerationAI Data
Resources
Submit ToolAI NewsBlog
Hot Models
GPT-5.5
© 2024 VKMO AI, All rights reserved
Privacy PolicyTerms of Service
  1. Home
  2. AI Models
  3. SANA-WM
SANA-WM

SANA-WM

Released May 17, 2026Image-to-Video

SANA is an efficiency-oriented codebase for high-resolution image and video generation, providing complete training and inference pipelines.

Context

-

tokens

Input

-

per 1M tokens

Output

-

per 1M tokens

Downloads

-

Analysis Summary

SANA-WM Overview

SANA-WM is an open-source world model from NVIDIA Research that turns one image plus a camera trajectory into minute-scale video. Its core promise is not just longer generation, but longer generation that still respects spatial structure and camera motion.

If you searched the term because it suddenly appeared in research news, the useful answer is simple: this is a model aimed at minute-long 720p worlds with precise camera control, not another ordinary short-form video generator.

SANA-WM Features

Long-horizon worlds
Generate minute-scale scenes that stay coherent across a longer camera path.

Precise camera motion
Follow 6-DoF trajectories instead of only producing unconstrained cinematic motion.

Higher-throughput evaluation
The paper reports comparable visual quality to industrial baselines with 36x higher throughput on its benchmark.

Open research footing
The project page, paper, and code repository are already public, making the model easier to inspect than a closed demo.

How it works

Start with a still image
The model takes an initial frame as the visual anchor for the world.

Add a camera path
A 6-DoF trajectory tells the model where the virtual camera should move.

Roll out the world
Hybrid linear attention keeps the long sequence tractable while preserving scene continuity.

Refine the result
A second-stage long-video refiner improves texture, motion, and later-frame quality.

Related materials

  • HuggingFace: https://huggingface.co/Efficient-Large-Model/SANA-WM_bidirectional
  • Paper: https://arxiv.org/abs/2605.15178
  • Github: https://github.com/NVlabs/SANA

Specifications

Context-
Input-
Output-
Downloads-
ReleasedMay 17, 2026
Category
Image-to-Video

Related Models

Kimi K3 Open-Source

Kimi K3 Open-Source

OvisOCR2

OvisOCR2

Cosmos 3

Cosmos 3

Stable Audio 3.0

Stable Audio 3.0

Gemini Omni Flash

Gemini Omni Flash

Pixal3D

Pixal3D

View all models →