{"id":180181,"date":"2025-04-10T10:09:08","date_gmt":"2025-04-10T10:09:08","guid":{"rendered":"https:\/\/peraltafinancing.com\/analytics\/generating-one-minute-videos-with-test-time-training\/"},"modified":"2025-04-10T10:09:08","modified_gmt":"2025-04-10T10:09:08","slug":"generating-one-minute-videos-with-test-time-training","status":"publish","type":"post","link":"https:\/\/fivemor.com\/?p=180181","title":{"rendered":"Generating One-Minute Videos with Test-Time Training"},"content":{"rendered":"<p> <br \/>\n<\/p>\n<div id=\"article-start\">\n<p>Video generation from text has come a long way, but it still hits a wall when it comes to producing longer, multi-scene stories. While diffusion models like <a href=\"https:\/\/www.analyticsvidhya.com\/blog\/2024\/12\/openai-sora\/\" target=\"_blank\" rel=\"noreferrer noopener\">Sora<\/a>, <a href=\"https:\/\/www.analyticsvidhya.com\/blog\/2024\/12\/googles-veo-2\/\" target=\"_blank\" rel=\"noreferrer noopener\">Veo<\/a>, and Movie Gen have raised the bar in visual quality, they\u2019re typically limited to clips under 20 seconds. The real challenge? Context. Generating a one-minute, story-driven video from a paragraph of text requires models to process hundreds of thousands of tokens while maintaining narrative and visual coherence. That\u2019s where this new research from NVIDIA, Stanford, UC Berkeley, and others steps in, introducing a technique called Test-Time Training (TTT) to push past current limitations.<\/p>\n<h2 class=\"wp-block-heading\" id=\"h-what-s-the-problem-with-long-videos\">What\u2019s the Problem with Long Videos?<\/h2>\n<p>Transformers, particularly those used in video generation, rely on <a href=\"https:\/\/www.analyticsvidhya.com\/blog\/2019\/11\/comprehensive-guide-attention-mechanism-deep-learning\/\" target=\"_blank\" rel=\"noreferrer noopener\">self-attention mechanisms<\/a>. These scale poorly with sequence length due to their quadratic computational cost. Attempting to generate a full minute of high-resolution video with dynamic scenes and consistent characters means juggling over 300,000 tokens of information. That makes the model inefficient and often incoherent over long stretches.<\/p>\n<p>Some teams have tried to circumvent this by using recurrent neural networks (<a href=\"https:\/\/www.analyticsvidhya.com\/blog\/2022\/03\/a-brief-overview-of-recurrent-neural-networks-rnn\/\" target=\"_blank\" rel=\"noreferrer noopener\">RNNs<\/a>) like Mamba or DeltaNet, which offer linear-time context handling. However, these models compress context into a fixed-size hidden state, which limits expressiveness. It\u2019s like trying to squeeze an entire movie into a postcard, some details just won\u2019t fit.<\/p>\n<h2 class=\"wp-block-heading\" id=\"h-how-does-ttt-test-time-training-solve-the-issue\">How Does TTT (Test-Time Training) Solve the Issue?<\/h2>\n<p><a href=\"https:\/\/arxiv.org\/pdf\/2504.05298\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">This paper<\/a> comes from the idea of making the hidden state of RNNs more expressive by turning it into a trainable neural network itself. Specifically, the authors propose using TTT layers, essentially small, two-layer MLPs that adapt on the fly while processing input sequences. These layers are updated during inference time using a self-supervised loss, which helps them dynamically learn from the video\u2019s evolving context.<\/p>\n<p>Imagine a model that adapts mid-flight: as the video unfolds, its internal memory adjusts to better understand the characters, motions, and storyline. That\u2019s what TTT enables.<\/p>\n<h2 class=\"wp-block-heading\" id=\"h-examples-of-one-minute-videos-with-test-time-training\">Examples of One-Minute Videos with Test-Time Training<\/h2>\n<p><strong>Adding TTT Layers to a Pre-Trained Transformer<\/strong><\/p>\n<p>Adding TTT layers into a pre-trained Transformer enables it to generate one-minute videos with strong temporal consistency and motion smoothness.<\/p>\n<p><strong>Prompt: <\/strong>\u201c<em>Jerry snatches a wedge of cheese and races for his mousehole with Tom in pursuit. He slips inside just in time, leaving Tom to crash into the wall. Safe and cozy, Jerry enjoys his prize at a tiny table, happily nibbling as the scene fades to black.<\/em>\u201c<\/p>\n<p>\n<iframe src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/attn.mp4\" loading=\"lazy\" title=\"mamba\" allowfullscreen=\"\"><\/iframe>\n<\/p>\n<p><strong>Baseline Comparisons<\/strong><\/p>\n<p>TTT-MLP outperforms all other baselines in temporal consistency, motion smoothness, and overall aesthetics, as measured by human evaluation Elo scores.<\/p>\n<p><strong>Prompt:<\/strong> \u201c<em>Tom is happily eating an apple pie at the kitchen table. Jerry looks longingly wishing he had some. Jerry goes outside the front door of the house and rings the doorbell. While Tom comes to open the door, Jerry runs around the back to the kitchen. Jerry steals Tom\u2019s apple pie. Jerry runs to his mousehole carrying the pie, while Tom is chasing him. Just as Tom is about to catch Jerry, he makes it through the mouse hole and Tom slams into the wall.<\/em>\u201c<\/p>\n<p>\n<iframe src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/mamba.mp4\" loading=\"lazy\" title=\"mamba\" allowfullscreen=\"\"><\/iframe>\n<\/p>\n<p><strong>Limitations<\/strong><\/p>\n<p>The generated one-minute videos demonstrate clear potential as a proof of concept, but still contain notable artifacts.<\/p>\n<p>\n<iframe src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/04\/homeless.mp4\" loading=\"lazy\" title=\"mamba\" allowfullscreen=\"\"><\/iframe>\n<\/p>\n<h2 class=\"wp-block-heading\" id=\"h-how-does-it-work\">How Does it Work?<\/h2>\n<p>The system starts with a pre-trained Diffusion Transformer model, CogVideo-X 5B, which previously could only generate 3-second clips. The researchers inserted TTT layers into the model and trained them (along with local attention blocks) to handle longer sequences.<\/p>\n<p>To manage cost, self-attention was restricted to short, 3-second segments, while the TTT layers took charge of understanding the global narrative across these segments. The architecture also includes gating mechanisms to ensure TTT layers don\u2019t degrade performance during early training.<\/p>\n<p>They further enhance training by processing sequences bidirectionally and segmenting videos into annotated scenes. For example, a storyboard format was used to describe each 3-second segment in detail, backgrounds, character positions, camera angles, and actions.<\/p>\n<h2 class=\"wp-block-heading\" id=\"h-the-dataset-tom-amp-jerry-with-a-twist\">The Dataset: Tom &amp; Jerry with a Twist<\/h2>\n<p>To ground the research in a consistent, well-understood visual domain, the team curated a dataset from over 7 hours of classic Tom and Jerry cartoons. These were broken down into scenes and finely annotated into 3-second segments. By focusing on cartoon data, the researchers avoided the complexity of photorealism and honed in on narrative coherence and motion dynamics.<\/p>\n<p>Human annotators wrote descriptive paragraphs for each segment, ensuring the model had rich, structured input to learn from. This also allowed for multi-stage training\u2014first on 3-second clips, and progressively on longer sequences up to 63 seconds.<\/p>\n<h2 class=\"wp-block-heading\" id=\"h-performance-does-it-actually-work\">Performance: Does it Actually Work?<\/h2>\n<p>Yes, and impressively so. When benchmarked against leading baselines like Mamba 2, Gated DeltaNet, and sliding-window attention, the TTT-MLP model outperformed them by an average of 34 Elo points in a human evaluation across 100 videos.<\/p>\n<p>The evaluation considered:<\/p>\n<ul class=\"wp-block-list\">\n<li><strong>Text alignment<\/strong>: How well the video follows the prompt<\/li>\n<li><strong>Motion naturalness<\/strong>: Realism in character movement<\/li>\n<li><strong>Aesthetics<\/strong>: Lighting, color, and visual appeal<\/li>\n<li><strong>Temporal consistency<\/strong>: Visual coherence across scenes<\/li>\n<\/ul>\n<p>TTT-MLP particularly excelled in motion and scene consistency, maintaining logical continuity across dynamic actions\u2014something that other models struggled with.<\/p>\n<h2 class=\"wp-block-heading\" id=\"h-artifacts-amp-limitations\">Artifacts &amp; Limitations<\/h2>\n<p>Despite the promising results, there are still artifacts. Lighting may shift inconsistently, or motion may look floaty (e.g., cheese hovering unnaturally). These issues are likely linked to the limitations of the base model, CogVideo-X. Another bottleneck is efficiency. While TTT-MLP is significantly faster than full self-attention models (2.5x speedup), it\u2019s still slower than leaner RNN approaches like Gated DeltaNet. That said, TTT only needs to be fine-tuned\u2014not trained from scratch\u2014making it more practical for many use cases.<\/p>\n<h2 class=\"wp-block-heading\" id=\"h-what-makes-this-approach-stand-out\">What Makes This Approach Stand Out<\/h2>\n<ul class=\"wp-block-list\">\n<li><strong>Expressive Memory<\/strong>: TTT turns the hidden state of RNNs into a trainable network, making it far more expressive than a fixed-size matrix.<\/li>\n<li><strong>Adaptability<\/strong>: TTT layers learn and adjust during inference, allowing them to respond in real time to the unfolding video.<\/li>\n<li><strong>Scalability<\/strong>: With enough resources, this method scales to longer and more complex video stories.<\/li>\n<li><strong>Practical Fine-Tuning<\/strong>: Researchers fine-tune only the TTT layers and gates, which keeps training lightweight and efficient.<\/li>\n<\/ul>\n<h2 class=\"wp-block-heading\" id=\"h-future-directions\">Future Directions<\/h2>\n<p>The team points out several opportunities for expansion:<\/p>\n<ul class=\"wp-block-list\">\n<li>Optimizing the TTT kernel for faster inference<\/li>\n<li>Experimenting with larger or different backbone models<\/li>\n<li>Exploring even more complex storylines and domains<\/li>\n<li>Using Transformer-based hidden states instead of MLPs for even more expressiveness<\/li>\n<\/ul>\n<h2 class=\"wp-block-heading\" id=\"h-ttt-video-generation-vs-mocha-vs-goku-vs-omnihuman1-vs-dreamactor-m1\">TTT Video Generation vs MoCha vs Goku vs OmniHuman1 vs DreamActor-M1<\/h2>\n<p>The table given below explains the difference betweeen this model and other trending video generation models out there: <\/p>\n<div class=\"responsive-table\" style=\"overflow-x:auto; border: 1px solid #ccc; padding: 1rem;\">\n<table style=\"width:100%; border-collapse: collapse; border: 1px solid #ccc;\">\n<thead>\n<tr style=\"background-color: #f2f2f2;\">\n<th style=\"border: 1px solid #ccc; padding: 8px;\">Model<\/th>\n<th style=\"border: 1px solid #ccc; padding: 8px;\">Core Focus<\/th>\n<th style=\"border: 1px solid #ccc; padding: 8px;\">Input Type<\/th>\n<th style=\"border: 1px solid #ccc; padding: 8px;\">Key Features<\/th>\n<th style=\"border: 1px solid #ccc; padding: 8px;\">How It Differs from TTT<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td style=\"border: 1px solid #ccc; padding: 8px;\">TTT (Test-Time Training)<\/td>\n<td style=\"border: 1px solid #ccc; padding: 8px;\">Long-form video generation with dynamic adaptation<\/td>\n<td style=\"border: 1px solid #ccc; padding: 8px;\">Text storyboard<\/td>\n<td style=\"border: 1px solid #ccc; padding: 8px;\">\n          \u2013 Adapts during inference<br \/>\u2013 Handles 60+ sec videos<br \/>\u2013 Coherent multi-scene stories\n        <\/td>\n<td style=\"border: 1px solid #ccc; padding: 8px;\">\n          Designed for long videos; updates internal state during generation for narrative consistency\n        <\/td>\n<\/tr>\n<tr>\n<td style=\"border: 1px solid #ccc; padding: 8px;\">MoCha<\/td>\n<td style=\"border: 1px solid #ccc; padding: 8px;\">Talking character generation<\/td>\n<td style=\"border: 1px solid #ccc; padding: 8px;\">Text + Speech<\/td>\n<td style=\"border: 1px solid #ccc; padding: 8px;\">\n          \u2013 No keypoints or reference images<br \/>\u2013 Speech-driven full-body animation\n        <\/td>\n<td style=\"border: 1px solid #ccc; padding: 8px;\">\n          Focuses on character dialogue &amp; expressions, not full-scene narrative videos\n        <\/td>\n<\/tr>\n<tr>\n<td style=\"border: 1px solid #ccc; padding: 8px;\">Goku<\/td>\n<td style=\"border: 1px solid #ccc; padding: 8px;\">High-quality video &amp; image generation<\/td>\n<td style=\"border: 1px solid #ccc; padding: 8px;\">Text, Image<\/td>\n<td style=\"border: 1px solid #ccc; padding: 8px;\">\n          \u2013 Rectified Flow Transformers<br \/>\u2013 Multi-modal input support\n        <\/td>\n<td style=\"border: 1px solid #ccc; padding: 8px;\">\n          Optimized for quality &amp; training speed; not designed for long-form storytelling\n        <\/td>\n<\/tr>\n<tr>\n<td style=\"border: 1px solid #ccc; padding: 8px;\">OmniHuman1<\/td>\n<td style=\"border: 1px solid #ccc; padding: 8px;\">Realistic human animation<\/td>\n<td style=\"border: 1px solid #ccc; padding: 8px;\">Image + Audio + Text<\/td>\n<td style=\"border: 1px solid #ccc; padding: 8px;\">\n          \u2013 Multiple conditioning signals<br \/>\u2013 High-res avatars\n        <\/td>\n<td style=\"border: 1px solid #ccc; padding: 8px;\">\n          Creates lifelike humans; doesn\u2019t model long sequences or dynamic scene transitions\n        <\/td>\n<\/tr>\n<tr>\n<td style=\"border: 1px solid #ccc; padding: 8px;\">DreamActor-M1<\/td>\n<td style=\"border: 1px solid #ccc; padding: 8px;\">Image-to-animation (face\/body)<\/td>\n<td style=\"border: 1px solid #ccc; padding: 8px;\">Image + Driving Video<\/td>\n<td style=\"border: 1px solid #ccc; padding: 8px;\">\n          \u2013 Holistic motion imitation<br \/>\u2013 High frame consistency\n        <\/td>\n<td style=\"border: 1px solid #ccc; padding: 8px;\">\n          Animates static images; doesn\u2019t use text or handle scene-by-scene story generation\n        <\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p>Also Read: <\/p>\n<h2 class=\"wp-block-heading\" id=\"h-end-note\">End Note<\/h2>\n<p>Test-Time Training offers a fascinating new lens for tackling long-context video generation. By letting the model learn and adapt during inference, it bridges a crucial gap in storytelling, a domain where continuity, emotion, and pacing matter just as much as visual fidelity.<\/p>\n<p>Whether you\u2019re a researcher in <a href=\"https:\/\/www.analyticsvidhya.com\/blog\/2023\/04\/what-is-generative-ai\/\" target=\"_blank\" rel=\"noreferrer noopener\">generative AI<\/a>, a creative technologist, or a product leader curious about what\u2019s next for AI-generated media, this work is a signpost pointing toward the future of dynamic, coherent video synthesis from text.<\/p>\n<div class=\"border-top py-3 author-info my-4\">\n<div class=\"author-card d-flex align-items-center\">\n<div class=\"flex-shrink-0 overflow-hidden\">\n                                    <a href=\"https:\/\/www.analyticsvidhya.com\/blog\/author\/nitika-sharma\/\" class=\"text-decoration-none active-avatar\"><br \/>\n                                                                       <img decoding=\"async\" src=\"https:\/\/av-eks-lekhak.s3.amazonaws.com\/media\/lekhak-profile-images\/converted_image_A027XT6.webp\" width=\"48\" height=\"48\" alt=\"Nitika Sharma\" loading=\"lazy\" class=\"rounded-circle\"\/><\/p>\n<p>                                <\/a>\n                                <\/div>\n<\/p><\/div>\n<p>Hello, I am Nitika, a tech-savvy Content Creator and Marketer. Creativity and learning new things come naturally to me. I have expertise in creating result-driven content strategies. I am well versed in SEO Management, Keyword Operations, Web Content Writing, Communication, Content Strategy, Editing, and Writing.<\/p>\n<\/p><\/div>\n<\/p><\/div>\n<p><h4 class=\"fs-24 text-dark\">Login to continue reading and enjoy expert-curated content.<\/h4>\n<p>                        <button class=\"btn btn-primary mx-auto d-table\" data-bs-toggle=\"modal\" data-bs-target=\"#loginModal\" id=\"readMoreBtn\">Keep Reading for Free<\/button>\n                    <\/p>\n\n","protected":false},"excerpt":{"rendered":"<p>Video generation from text has come a long way, but it still hits a wall when it comes to producing longer, multi-scene stories. While diffusion models like Sora, Veo, and Movie Gen have raised the bar in visual quality, they\u2019re typically limited to clips under 20 seconds. The real challenge? Context. Generating a one-minute, story-driven [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":180182,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[12033],"tags":[31746,71346,71347,324,13073],"dealstore":[],"offerexpiration":[],"class_list":["post-180181","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-analytics","tag-generating","tag-oneminute","tag-testtime","tag-training","tag-videos"],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v26.4 - https:\/\/yoast.com\/wordpress\/plugins\/seo\/ -->\n<title>Generating One-Minute Videos with Test-Time Training - Som2ny Network<\/title>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/fivemor.com\/?p=180181\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"Generating One-Minute Videos with Test-Time Training - Som2ny Network\" \/>\n<meta property=\"og:description\" content=\"Video generation from text has come a long way, but it still hits a wall when it comes to producing longer, multi-scene stories. While diffusion models like Sora, Veo, and Movie Gen have raised the bar in visual quality, they\u2019re typically limited to clips under 20 seconds. The real challenge? Context. Generating a one-minute, story-driven [&hellip;]\" \/>\n<meta property=\"og:url\" content=\"https:\/\/fivemor.com\/?p=180181\" \/>\n<meta property=\"og:site_name\" content=\"Som2ny Network\" \/>\n<meta property=\"article:published_time\" content=\"2025-04-10T10:09:08+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/04\/Generating-One-Minute-Videos-with-Test-Time-Training-Give-me-tom-and-Jerry-image-with-Test-Time-Trai.webp\" \/>\n\t<meta property=\"og:image:width\" content=\"872\" \/>\n\t<meta property=\"og:image:height\" content=\"473\" \/>\n\t<meta property=\"og:image:type\" content=\"image\/webp\" \/>\n<meta name=\"author\" content=\"admin\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"admin\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"7 minutes\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\/\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\/\/fivemor.com\/?p=180181#article\",\"isPartOf\":{\"@id\":\"https:\/\/fivemor.com\/?p=180181\"},\"author\":{\"name\":\"admin\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371\"},\"headline\":\"Generating One-Minute Videos with Test-Time Training\",\"datePublished\":\"2025-04-10T10:09:08+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\/\/fivemor.com\/?p=180181\"},\"wordCount\":1359,\"commentCount\":0,\"publisher\":{\"@id\":\"https:\/\/fivemor.com\/#organization\"},\"image\":{\"@id\":\"https:\/\/fivemor.com\/?p=180181#primaryimage\"},\"thumbnailUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/04\/Generating-One-Minute-Videos-with-Test-Time-Training-Give-me-tom-and-Jerry-image-with-Test-Time-Trai.webp\",\"keywords\":[\"generating\",\"OneMinute\",\"TestTime\",\"Training\",\"Videos\"],\"articleSection\":[\"Analytics\"],\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"CommentAction\",\"name\":\"Comment\",\"target\":[\"https:\/\/fivemor.com\/?p=180181#respond\"]}]},{\"@type\":\"WebPage\",\"@id\":\"https:\/\/fivemor.com\/?p=180181\",\"url\":\"https:\/\/fivemor.com\/?p=180181\",\"name\":\"Generating One-Minute Videos with Test-Time Training - Som2ny Network\",\"isPartOf\":{\"@id\":\"https:\/\/fivemor.com\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\/\/fivemor.com\/?p=180181#primaryimage\"},\"image\":{\"@id\":\"https:\/\/fivemor.com\/?p=180181#primaryimage\"},\"thumbnailUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/04\/Generating-One-Minute-Videos-with-Test-Time-Training-Give-me-tom-and-Jerry-image-with-Test-Time-Trai.webp\",\"datePublished\":\"2025-04-10T10:09:08+00:00\",\"breadcrumb\":{\"@id\":\"https:\/\/fivemor.com\/?p=180181#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\/\/fivemor.com\/?p=180181\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/fivemor.com\/?p=180181#primaryimage\",\"url\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/04\/Generating-One-Minute-Videos-with-Test-Time-Training-Give-me-tom-and-Jerry-image-with-Test-Time-Trai.webp\",\"contentUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/04\/Generating-One-Minute-Videos-with-Test-Time-Training-Give-me-tom-and-Jerry-image-with-Test-Time-Trai.webp\",\"width\":872,\"height\":473},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\/\/fivemor.com\/?p=180181#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\/\/fivemor.com\/?bp_activities=1\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"Generating One-Minute Videos with Test-Time Training\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\/\/fivemor.com\/#website\",\"url\":\"https:\/\/fivemor.com\/\",\"name\":\"Som2ny Network\",\"description\":\"Daily Deals\",\"publisher\":{\"@id\":\"https:\/\/fivemor.com\/#organization\"},\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\/\/fivemor.com\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"},{\"@type\":\"Organization\",\"@id\":\"https:\/\/fivemor.com\/#organization\",\"name\":\"Som2ny Network\",\"url\":\"https:\/\/fivemor.com\/\",\"logo\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/logo\/image\/\",\"url\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png\",\"contentUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png\",\"width\":300,\"height\":86,\"caption\":\"Som2ny Network\"},\"image\":{\"@id\":\"https:\/\/fivemor.com\/#\/schema\/logo\/image\/\"}},{\"@type\":\"Person\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371\",\"name\":\"admin\",\"image\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/person\/image\/\",\"url\":\"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png\",\"contentUrl\":\"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png\",\"caption\":\"admin\"},\"sameAs\":[\"https:\/\/fivemor.com\"],\"url\":\"https:\/\/fivemor.com\/?author=1\"}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"Generating One-Minute Videos with Test-Time Training - Som2ny Network","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/fivemor.com\/?p=180181","og_locale":"en_US","og_type":"article","og_title":"Generating One-Minute Videos with Test-Time Training - Som2ny Network","og_description":"Video generation from text has come a long way, but it still hits a wall when it comes to producing longer, multi-scene stories. While diffusion models like Sora, Veo, and Movie Gen have raised the bar in visual quality, they\u2019re typically limited to clips under 20 seconds. The real challenge? Context. Generating a one-minute, story-driven [&hellip;]","og_url":"https:\/\/fivemor.com\/?p=180181","og_site_name":"Som2ny Network","article_published_time":"2025-04-10T10:09:08+00:00","og_image":[{"width":872,"height":473,"url":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/04\/Generating-One-Minute-Videos-with-Test-Time-Training-Give-me-tom-and-Jerry-image-with-Test-Time-Trai.webp","type":"image\/webp"}],"author":"admin","twitter_card":"summary_large_image","twitter_misc":{"Written by":"admin","Est. reading time":"7 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/fivemor.com\/?p=180181#article","isPartOf":{"@id":"https:\/\/fivemor.com\/?p=180181"},"author":{"name":"admin","@id":"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371"},"headline":"Generating One-Minute Videos with Test-Time Training","datePublished":"2025-04-10T10:09:08+00:00","mainEntityOfPage":{"@id":"https:\/\/fivemor.com\/?p=180181"},"wordCount":1359,"commentCount":0,"publisher":{"@id":"https:\/\/fivemor.com\/#organization"},"image":{"@id":"https:\/\/fivemor.com\/?p=180181#primaryimage"},"thumbnailUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/04\/Generating-One-Minute-Videos-with-Test-Time-Training-Give-me-tom-and-Jerry-image-with-Test-Time-Trai.webp","keywords":["generating","OneMinute","TestTime","Training","Videos"],"articleSection":["Analytics"],"inLanguage":"en-US","potentialAction":[{"@type":"CommentAction","name":"Comment","target":["https:\/\/fivemor.com\/?p=180181#respond"]}]},{"@type":"WebPage","@id":"https:\/\/fivemor.com\/?p=180181","url":"https:\/\/fivemor.com\/?p=180181","name":"Generating One-Minute Videos with Test-Time Training - Som2ny Network","isPartOf":{"@id":"https:\/\/fivemor.com\/#website"},"primaryImageOfPage":{"@id":"https:\/\/fivemor.com\/?p=180181#primaryimage"},"image":{"@id":"https:\/\/fivemor.com\/?p=180181#primaryimage"},"thumbnailUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/04\/Generating-One-Minute-Videos-with-Test-Time-Training-Give-me-tom-and-Jerry-image-with-Test-Time-Trai.webp","datePublished":"2025-04-10T10:09:08+00:00","breadcrumb":{"@id":"https:\/\/fivemor.com\/?p=180181#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/fivemor.com\/?p=180181"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/fivemor.com\/?p=180181#primaryimage","url":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/04\/Generating-One-Minute-Videos-with-Test-Time-Training-Give-me-tom-and-Jerry-image-with-Test-Time-Trai.webp","contentUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/04\/Generating-One-Minute-Videos-with-Test-Time-Training-Give-me-tom-and-Jerry-image-with-Test-Time-Trai.webp","width":872,"height":473},{"@type":"BreadcrumbList","@id":"https:\/\/fivemor.com\/?p=180181#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/fivemor.com\/?bp_activities=1"},{"@type":"ListItem","position":2,"name":"Generating One-Minute Videos with Test-Time Training"}]},{"@type":"WebSite","@id":"https:\/\/fivemor.com\/#website","url":"https:\/\/fivemor.com\/","name":"Som2ny Network","description":"Daily Deals","publisher":{"@id":"https:\/\/fivemor.com\/#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/fivemor.com\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":"Organization","@id":"https:\/\/fivemor.com\/#organization","name":"Som2ny Network","url":"https:\/\/fivemor.com\/","logo":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/fivemor.com\/#\/schema\/logo\/image\/","url":"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png","contentUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png","width":300,"height":86,"caption":"Som2ny Network"},"image":{"@id":"https:\/\/fivemor.com\/#\/schema\/logo\/image\/"}},{"@type":"Person","@id":"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371","name":"admin","image":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/fivemor.com\/#\/schema\/person\/image\/","url":"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png","contentUrl":"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png","caption":"admin"},"sameAs":["https:\/\/fivemor.com"],"url":"https:\/\/fivemor.com\/?author=1"}]}},"_links":{"self":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/posts\/180181","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=180181"}],"version-history":[{"count":0,"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/posts\/180181\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/media\/180182"}],"wp:attachment":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=180181"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=180181"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=180181"},{"taxonomy":"dealstore","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fdealstore&post=180181"},{"taxonomy":"offerexpiration","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fofferexpiration&post=180181"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}