
AI & ML
How Sora Generates Video: The AI Technology Behind OpenAI's Text-to-Video Model
OpenAI's Sora generates videos using a diffusion transformer (DiT) that processes spacetime patches of compressed latents, enabling coherent outputs with strong 3D consistency and object permanence from text prompts. At its core, the model compresses raw video into a lower-dimensional space via an autoencoder before iteratively denoising noise patches conditioned on detailed text embeddings to simulate realistic motion and world dynamics
Read →